I deeply love this idea of specialized LLMs for search. It's also extremely confusing to me how rough Google's entrance here is.
When I, a human, need an answer to anything moderately complex, it's unlikely that I get it on the first (pre-AI) round of google searching. Simple stuff, sure, but more likely I'll need to go 2-5 rounds. Maybe click a few links. Double-check my assumptions.
An LLM that can do that quickly seems like a slam dunk. I wonder what other problems benefit from that 10x-100x increase in context + 2-5 rounds with the LLM.
For example, a couple of days ago I described a problem with my refrigerator's water dispenser to Google Gemini, and it told me exactly how to fix it. I then went looking for a video and fixed the thing in under 15 minutes. The only way that Gemini could have been better is if it linked to a video itself.
Do you mean search into less-well-known topics? Or something else?
About 15 years ago I would sometimes spend hours on Google image search discovering childhood toys and filling in vague memories of locations or things. I tried this recently and it’s basically impossible. I actually get to the end of the search results in like 3 minutes and the quality is horrible now.
I think in a lot of ways Google peaked and is now on the decline into a profitable but much less relevant services company.
I guess someone who has used a search agent (or a dedicated subagent) can speak when I'd reach for a tool like this vs either just 1) a smaller general model or 2) a non-llm approach to the problem? Like it's interesting I'm just curious how a search agent compares to say a model with dedicated rag pipelines is that much different?
Mixedbread Search is a multimodal & multilingual search product, where you can upload any kind of data and make it searchable. Its powered by Wholembed [1] v3, a late interaction retrieval model.
I know everyone loves to hate on google but i find search overviews and asking gemini to search for things way faster than any alternative. I was curious about a development near me and asked literally that and gemini pulled court records in about 20 seconds
Anyway, back to this - it seems to be more like the AI equivalent of algolia than google
Long time user of your embedding models. I'm trying to understand how this works and if I can leverage it.
It sounds like this is a new layer on top of your existing storage layer? So to use this, would I need to give you all of my data first? Or is there a version that can be run on prem?
Looks good as I use something similar with the SearXNG MCP, but a shame this isn't an open weight model. There are some wrappers around SearXNG which seem to reduce the token counts returned thus making it easier for the calling model to understand, but a full dedicated model for search is nice. How does it compare with Perplexity, Gemini with search, and Parallel AI? Those are the cloud providers of search based models that I've seen so far.
31 comments
[ 3.5 ms ] story [ 31.0 ms ] threadWhen I, a human, need an answer to anything moderately complex, it's unlikely that I get it on the first (pre-AI) round of google searching. Simple stuff, sure, but more likely I'll need to go 2-5 rounds. Maybe click a few links. Double-check my assumptions.
An LLM that can do that quickly seems like a slam dunk. I wonder what other problems benefit from that 10x-100x increase in context + 2-5 rounds with the LLM.
For example, a couple of days ago I described a problem with my refrigerator's water dispenser to Google Gemini, and it told me exactly how to fix it. I then went looking for a video and fixed the thing in under 15 minutes. The only way that Gemini could have been better is if it linked to a video itself.
Do you mean search into less-well-known topics? Or something else?
I think in a lot of ways Google peaked and is now on the decline into a profitable but much less relevant services company.
[1]: https://www.mixedbread.com/blog/wholembed-v3
Anyway, back to this - it seems to be more like the AI equivalent of algolia than google
Bread-first search, is it?
It sounds like this is a new layer on top of your existing storage layer? So to use this, would I need to give you all of my data first? Or is there a version that can be run on prem?
Sadly its another software company.