159 comments

[ 0.20 ms ] story [ 6.6 ms ] thread
[dead]
Funny that the 2B loses to the 0.8B. Question about the benchmarks: BBH and JudgeBench are more reasoning, where you fall behind, but for zero-shot classification there are more relevant ones like Banking77 or CLINC150. Was there no temptation to pick something closer to where System 1 models are actually used?
That's some local hardware.
Was going to make the same observation. Cool to run locally - but renting seems the saner choice?

Would love to see what this run would cost from something like Verda. Just out of curiosity. I'm not going to be installing any sparks at home anytime soon.

Great project! How does this compare to asking qwen to reply with 1 token in terms of speed?
Positively identifiable as slop before even clicking, click and find 6 commits. Why was this even posted? Do you really expect to still be working on this even by next week?
First commit of six commit was 4 hours ago.
This type of project looks extremely useful. There was a lot of buzz around Jev, but having models that run locally and can be fine-tuned is extremely helpful.
I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier

I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes

The classifiers also run in <1ms, so they can be very fast and precise at the same time

But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)

For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data

"Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification"

Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?

If so, I'm curious whether you compared that approach (A) with:

B) Jev only, with a single output.

C) Jev with multiple outputs fed into a logistic classifier.

Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)

You are correct, I didn’t train the embeddings model

Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller

I haven’t compared different ways of sending requests to Jev

The data to train the classifiers comes from the datasets used to test them (not from Jev)

What do you use to determine that a particular task in a heterogeneous pile of tasks requires reasoning? The logistic classifier itself is too dumb to recognize the details of the problem that make it reasoning-sensitive (IIRC recognizing the “fiddliness” of a given problem requires a recognizer at least as complex as the problem itself.) And if you’re using the lightweight LLM for that, then you may as well skip the classifier and just use the LLM all the time, since that eval step is already going to be dominating your response time anyway.

My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.

I’m in the process of piecing together the different task/dataset-specific classifiers

Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach

For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time

I have not personally reviewed the benchmarks, but recall seeing a post where a Linear SVM with bag-of-words features outperformed Jev on many simple NLP classification tasks, like the ones you cite. You don't even need embeddings!
I've used LLM's to classify things for numerous projects and it's always really hard to beat linear classifier/decision tree over simple embeddings or even the sklearn HashingVectorizer. However, you DO need to trust your validated data, ground truths for this to work - but you should have these anyway to validate a Jev or similar solution.
> The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

You can already so that with classification models such as ModernBERT, at 0.4B.

Jev's value is its zero shot performance without having to fine-tune.

I am sure I am missing something obvious here, but why is that valuable? Like, what kinds of projects are there where you need to classify stuff but are unable to make a bespoke model targeting the specific problem?
For my org, it meant we could trial classifiers across various internal systems with little to no engineering effort. In one case we ended up building our own classifier instead of Jev, but in others we kept Jev because it was zero-effort for a great impact.
[delayed]
Because a lot of people employed as engineers can’t build classifiers. They’re able to glue libraries and tools together, build UIs, and write APIs, but don’t have the curiosity or creative problem solving to learn and master a new domain. Even when “master” is scoped to something like this.

On top, most EMs wouldn’t take a risk on an exploration of something “unknown” (to them) and couldn’t get buy-in from a PM.

I say this as an EM. Interview hundreds of people and, while, yes, some people don’t interview well, you might be shocked at the level of creative thinking. Even when “creative” is narrowly scoped to “this is a solved problem in a related domain.

I guess this makes sense. I am no business guy, but if my product/company was focused on some sort of classification problem, my naive intuition would be to focus on hiring guys that can do it, rather than try to make the problem easier for them. But perhaps at the end of the day this is still cheaper? It's the same reason why we use Postgres instead of hire database experts to build something special?
> if my product/company was focused on some sort of classification problem

It's more likely the product is focused on something else, but a classification model could come in handy...

> unable to make a bespoke model targeting the specific problem?

There is a fixed cost (and some maintenance) to e.g. fine tuning ModernBERT.

Maybe once you include all of that it might be a half-day to a day of engineering time to set everything up in a maintainable fashion.

For Jev, it takes all of 30 seconds of prompting. And it's not that much more expensive to deploy vs. a BERT model.

The biggest disadvantage of Jev is that it's a proprietary product and you need to send them your data. A bespoke solution makes much more sense in many scenarios.
Tons. For example, most web and mobile app developers won’t know where to start with making a bespoke model (and would likely have no interest in making one), but they will have lots of usecases for a classifier.
its a 1 hour project on claude code on ur local machine. lol...
Only for folks who already know what they’re doing.

What I said will only make sense if you take yourself out of your current context and think entirely from the perspective of someone who knows little-to-nothing about ML.

It’s the same mistake folks on HN made when Dropbox launched, drawing comparisons to rsync and other Unix tools as if they were somehow equivalent.

Anything where you don’t have a decent amount of training data.
You need to have enough high quality data to train with, knowledge how to do it, and developer time.

In practice, that's enough of a barrier to not even try the approach on a number of cases where it might potentially be useful.

I wouldn't be surprised if Jev turned out to be a "gateway drug" that validates approach on a use case, the team gathers experience and labeled data, and switches to an in house locally tuned model to minimize costs.

Exactly. And that labeled data could be collected by just recording what they feed jev and what the decision is.
You don't seem to understand what the point of Jev is when you say "you can fine tune it for your use case". Building your own classifiers for your specific business problems is the type of work we all used to do back in 2016 or so. It costs very much. Jev is a cheap and fast general purpose classifier.
My use case is simple classification for job ads. Things like, industry, work settings (remote, hybrid, onsite) and job type (full time, part time .. etc).

I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff

I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.

I suppose with larger model, this could be useful, but would require more ram and will be slower.

i've tried all the "open source" me too Jevs

they all suck

most of these dev clones need fine tuning on your use case, might as well fine tune modernBERT then. Jev generalizes well while being fast and cheap
Von 1.2 had a better Doom score :D

https://github.com/wfzyx/von

Oh wow, those Doom scores for Jeff are pretty terrible

The Von numbers have led me on a rabbit whole of getting a classifier to play Doom

I got it to average 22 kills (the max is 26) on that same scenario that Jeff and Von are testing on (it’s called Defend Center)

Now I’m having it play a more advanced scenario, and it’s doing about 45 kills (SOTA is ~59 kills)

It’s amazing what you can do with small classifiers if you can collect some data. These models I’m testing train on CPU in seconds (what takes the longest is running the game, doing test runs and collecting data), they are <1MB in size and do inference in <1ms on CPU

One day we’ll be using the same Kills/SOTA metric for models driving physical kill-bots.
Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
No need, as that functionality can run locally no problem.
the appeal would be if they can deliver it at a much higher performance and similar speed, which is plausible
BeRT and FLAN-T5 were used as classifiers 5-7 years ago, they were technically "frontier" for their time.
This is the exact comment I’ve been waiting for, what is the difference between classifiers and jev?
Jev is a classifier. The big thing about it is that it has high accuracy on domains it wasn't fine-tuned for, like an LLM, but with speed and cost comparable to traditional classifiers.
BERT need to be fine tuned for your use case, Jev generalizes. It’s a pretty big difference!
So what you're saying is it's artificial... general... intelligence? /s
FLAN-T5 generated text (Jev does not generate text), and BERT wasn't able to do tasks without fine-tuning.

Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.

Eg, "Rank these 5 things by increasing order of how big they are: car, cow, mouse, ant, building", the model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.

So it generates logits in a 255 token output space? ;)
logits assumed some form of softmax or logistic, which may not be the case
How does it know that you’re asking for “rank” instead of something else if it’s not generating text?
By reading, not generating, text.
It has to output “rank” - that requires generation unless I’m mistaken
no, 255 numbers come out at once all the time, the order is determined by the inputs
That’s not what I’m asking about - the text prompted for rank and it output “rank” in the response object. How did it do that? Like if I’d asked it to group similar items, how would it know how to structure that output and know that the key should be “groupings”
The output key is determined through code by the harness from the inputs, and the model generates 255 floats in order that follow that schema by reading the expected return type "rank" in the input. The harness then programmatically uses the floats of the 255 floats that are useful. The harness can assume that the correct float will be in the correct position as the model is trained to follow the schemas
Is Jev's architecture public somewhere? I'd love to read more about it.
The fact that its not narrow and stupid is what's different
If I had to guess, it's already built and is just waiting on Product's/Marketing's desk. How do you position this without looking like your roadmap is being determined by newcomers? Probably don't want to adopt the same verbiage+acronyms - but also can't be seen to be just sherlocking features.
I think Apple has demonstrated that shipping second has essentially no negative impact if your product is seen as higher quality.
There is zero stigma to shipping second. If anything, the labs’ customers are probably begging for them to add these features natively so they don’t have to deal with the hassle of adding another provider to their stack.
The confidence scores need to be good if we're gonna forego fast and cheap.
I don't think there would be any utility for that. Anything jev can do, a frontier model can also do. Just not as quickly or as cheaply.
I think these products (Jev and the inevitable offerings from Anthropic, OpenAI, etc) want to become more than end-user output machines. They'd benefit from being in the hotpath of other services. Not backgrounded generation but in-band, request-time work.

1M x $0.50 == 1B x $0.0005

To expand a bit for my current use cases. Inline routing of work to heavy task specific models, and prompt/context generation (user is asking something, what and how much should we prompt the expensive LLM with). Latency or time to first token does matter for some applications.
Probably at the frontier stage - you will only see it where Jev is better regardless of cost.

For everyone else who is conscious of cost, you're already seeing this being built into harnesses.

Almost certainly, you'll see versions of this from all the Chinese labs as fast as humanly possible.

If I had to guess, Cursor/Grok or Google/Antigravity will be the first major players to natively support something like this to drive down cost, as they're primarily the budget conscious choices.

I would be astounded if Anthropic leads the way on a cost reduction.

But for those of us that prefer open source and self hosting, JEV alternative LAYA will beat anything the frontier models package up.
Negative three years, give or take. Although recent Anthropic and OpenAI models no longer expose the capability. But for any open model you just tell it to respond with a single token "Y/N" and take the logit difference. If you want multiple distinct questions answered you just ask them independently and put the shared context first so it gets cached.

OpenAI and Anthropic don't want to give out logprobs these days but could trivially add a dedicated classification API to their existing models if there was enough demand.

I think the main differentiator offered by Jev is not the ability to answer questions, most models can be coerced into that function if they don’t already have a dedicated pipeline for it, rather it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.
> it’s the extreme speed of the evaluation, and very low cost that opens new possibilities.

It also returns confidence scores for all choices.

Granted, they are not stable. They fluctuate even when you reorder choices, but it still counts as an additional feature.

They’re useable if you’re trying to rank autosuggestions or something, not great at implying actual understanding
Do people really have zero awareness that Structured Outputs with a constrained schema has been a thing for a while now, and open weight models that give you logprobs can give you distributions per key?

Like what am I missing? I use Structured Outputs every day and this just seems like that with fewer steps?

Many people yes, but it’s probably for the best. Without fine-tuning on such style the results would be unreliable. It wouldn’t be the probability of the outcome of whatever you intend, but just the next token probability - which could alter if you add a space or punctuation to the prompt. Very flaky.
Pricing

The same reason I want diffusion language models to be mainstream

FWIW you’re absolutely correct on how most developers are unaware of this as they’re mostly operating on the API level and not aware of server side (vLLM/SGLang) capabilities.

The aspect that I like the most is the typesafe API that introduces new probabilistic concepts that are more sound than json schema and constrained decoding with quasi-confidence scores. Developers were asking LLMs to also emit confidences which made absolutely no sense whatsoever.

Classifiers (instead of LLMs) return results with confidence scores.

Name Entity Recognition (NER) is one example.

So many of them... https://huggingface.co/models?language=ner&sort=trending

Also used to block SSN and CC #s from logs, etc... as small and fast enough to do it. You don't want to call OpenAI GPT-6 and ask it to return your text with the SSN blanked out. I am sure people do though... (SSN is a bit simple, but all kinds of PPI in one model is more likely).

The nice thing about Jev is that people started taking about models that are not LLM text streams again.

> Classifiers (instead of LLMs) return results with confidence scores.

This is also bogus unless you are talking about Bayesian inference. No classifier can output CI for a single point estimate. In every ML theory textbooks worth their $, it's always stressed not to treat these sigmoid'ed or softmax'ed numbers as probabilities or confidence scores, there is no such thing as CI for point estimate.

I remember reading a paper about a model that used a parser bound to the end of a LLM where it squashed all logprobs for outputs that were parsing errors. I have not seen this used in the manner I thought I would. I thought it would have been entirely possible for a model to construct it's own grammar for how it would prefer to respond and then opt to generate tokens that matched (with a metacode to turn it off obviously).

That said, I think the advantage of Jev style approaches is not their capabilities, but rather the capabilities that they have for a much lower resource requirement.

Yeah, I'm with you.

Nowadays any LLM and any harness you use will just do this for you.

But there are helper libraries like Instructor that have been around since like gpt3, which abstract away retries and stuff to make this super easy.

Yeah I though llama.cpp had this during the very early days, if my memory serves me correctly.
Can we get a price comparisson ?

Edit: Running them for the masses.

I assume the cost is whatever you run the model on. Jev is already dirty cheap, $0.42 per million input tokens and output is free and the input cost is covering tokenization + API utilization
Well, I'm no expert, but this runs on Qwen3.5 and Gemma 4, stripped down, but they're pretty pricey.
Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
The main difficulty with fast (low-latency) inference, is not actually computation but loading parameters from memory. The problem with generating 1 token at a time isn't that that its expensive computationally (it is, but so is training), but that you need to stream your entire model from memory for every single token. So the strategy here is to process big batches of user requests, letting you share the memory loads across users. This means individual answers aren't that fast, but you can do lots at once.

My hypothesis for Jev is that they simply generate many answers independently in parallel from your prompt, and then discard the duplicates (or train to avoid duplication). In that way the entire batch is 1 user's prompt.

My guess is that they use an encoder-only model as the foundation and then do RLDC. Why? 1) it doesn’t need text generation 2) limited context window (40k last time I check)

Those are telltales of a Bert model.

What proportion of commercial LLM use is classification? I'm just wondering what happens to business AI spending/data centre usage when they realise they don't need full LLMs.
The American stock market probably loses 20% of its value in a few days.
Can you share your short positions so we can see how confident you are in these predictions?
Can you uhh elaborate on why?
I'd expect the vast majority at this point is coding. Classification is a thing but in my experience tends to run on light, cheap models, not the proprietary frontier ones.
part of agentic engineering is classification as well tho. reviewing/gating, what to read etc. don't need a full LLM
We've stopped upgrading the models of our classification workflows for >1 year at this point, meaning they're running acceptably on early/mid-2025 models. That said I believe there is a long tail of non-production-ized users who throw this into their everyday LLM chats.
In one of my client projects we've had to do forced upgrades of LLM models because of deprecations (I believe three or four of them over the course of 2 years). Each time our internal benchmarks have shown REDUCED performance after upgrading to "better" models.
My name jeff
No, he jefe, man.

I've been holding it in since I saw the title. I was surprised it wasn't all over repo!

Anything like this in the VLM side? Classification on images...
Isn't jev just a less nuanced classifier? What am I missing?
The zero-shot learning is what makes these different from a traditional classifier.
Zero-shot just means you give it zero examples. Jev lets you add examples, so it's one-shot/few-shot depending on how many you provide. After playing with it, it seems like once you wonder off their guide examples domains - you have to provide examples to get anything useful from it.

IMO you going to get better results by doing a small tune of a tiny model. Making training dataset for it with LLMs is easy, serving it is going to be cheaper.

I've always been more afraid of these types of models than LLMs. These are what enable mass surveillance at scale and autonomous real time combat drones. Now they are spreading and being optimized. Gg.
For those of a certain age - the fact that Askjev.com is still available astounds me.
I can’t believe ask.com threw in the towel in 2026. Like ride that LLM hype - how much better could the SEO for ask.com be? Do a pivot!
I am astounded you thought about this but didn’t buy it!
Blows my mind that Jeeves never came back as an AI model. Even if it's poor quality, it'd still be better than Google.
Awesome, I was just looking for a decision-making model that can be deployed locally, and here you are. Thanks a lot!
I think Jev is still the king and while I really appreciate this and other similar projects, Jev gives assurance and value that is hard to beat. Well.. this works locally which is always best, even if slower.
Everyone saying you could replace Jev or decision type models with an LLM with bolted schema output constraining are missing the point completely. Its about extreme speed and cost effectiveness with high quality, neither of which you are going to get with LLMs even with these KV-cache tricks
There is no free lunch. You cannot get high quality on open domain and have it be fast.
83.1 vs 83.0 on your panel against 70 vs 94 in someone's actual use case is the whole story with zero-shot classification. Any sense of what the 0.8B does on a phone NPU instead of an M4 Max? That's what decides if it's shippable on device.
Training 0.8B models at home with that latency is seriously impressive. What kind of hardware setup did you use for training?
Sorry, do we have actual clear implementation details for Jev? I keep seeing these "recreations" or "Do Jev at home" but do we have access to their architecture? I haven't even used the product, I just find it strange.
What they did was immediately obvious and trivial to replicate
They make it clear that their edge is in their synthetic data. Train your own jevs miss the point. And if you have that much data, you didnt need jev or these replacements in the first place.

HN's desire to pretend the data pipeline doesn't exist or isnt meaningful is silly.

(comment deleted)
If you're looking for a more impressive doom example, I put together laya-duum which uses open source micropython implementation of doom (duum) and freeware freedoom1.wad. It uses the standard jev api and will play through the first two levels to completion: https://github.com/Hadlock/laya-duum
Why is this more impressive?
It's not just shooting at static monsters in an empty room, this will go after monsters, make decisions about prioritizing health, changing goals, and finish the level, etc. It's not unique, it's just more impressive than the demo they're using.
Their demo could also change goals, e.g. avoid monsters instead of shooting them, by changing some variable. I'm sure that could also be used to change goals.
But it doesn’t. That is why that other person is claiming their demo is more impressive.

What you are saying is like me losing 13kg is as impressive as having the potential to do so if only I would stop eating ice cream and too many snacks every day.

You're confusing demo with technology.

The potential to lose weight, in this context, is the original post because there was potential but it wasn't shown

This commenter is the person who lost 13kg.

The technology is not different but the demonstration is more impressive

- Laya receives semantic snapshots, not the framebuffer.

What does tjis mean? Another generative model?

Laya is the model this uses. Like Jev, it's blind - it can't take an image as an input. So how can it play Doom, or Atari games, or do other visual tasks, as people have shown it to do? They write a bit of app-specific adapter code that transforms the game state into a structured piece of data (here, the "semantic snapshot"). And that is what these models take as input
laya sucks. it's a dead end and for some treat it as a peer of jev or even the qwen jevs. base model doesnt have enough world knowledge to generalize.
[flagged]
btw your comment was immediately marked as flagged/dead. Odd for a high quality comment. I vouched for it.

My main point of disagreement would be that I fundamentally see a different use for a jev sort of model (generalism is applealing), but if you're finetuning, a bert base is not bad.

Thanks for the vouching. I don't really understand what I did to anger the algorithms! But yeah, I think that that's the real sticking point. If the value prop of jev is "plug it in anywhere and it generally does smart things" you're absolutely right that Laya is Not A Thing. I've spent a bunch of time working on fine tuning models (i teach a class on it! https://earino.github.io/applied-deep-learning/ using the same ModernBERT stuff!) and so for me, I naturally went: "wait a second, so how good is Jev compared to fixing the problem domain."

Anyways, thanks for the vouching!

Great course, thanks for sharing, going through your slides and papers right now.
My guess is, it's because you sounded like an LLM in a few phrases there.
I think that's a pretty good guess. I move around different countries a lot so I also assume I can look like a bot farm from multiple source IPs.

However at this point I talk to LLMs more than anyone except probably my wife. As a multiple times immigrant, I can absolutely believe I'm adjusting my speech patterns to its vernacular.