Just curious, where has this term 'noul' come from for yes/no ansers?
/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
My bots are all named Jeeves lol. I have a CLI tool I use that connects up to a LLM I made and I call it Jeeves too ...so funny. I really didn't use Jeeves all that much I tended to use...I think it was called Web crawler pre-google era
Ask Jeeves hired hundreds of cheap liberal arts majors to classify data, some users thought Jeeves was real, the stock spiked when big companies hired Jeeves to automate support thinking it was a silver bullet, and the whole thing collapsed when a better model came along, and it degenerated into ripping off rubes with bottom of the barrel ads.
Doesn't the fact that it's general purpose warrant a new term? It's partly that it doesn't need to be trained, but it's also able to play games based on game state, I'd imagine it would be hard to train a classifier to do something like this because you'd need to represent a good distribution of all the states. The general purpose llm world understanding underneath it allows for this.
I've used it to do web research where it follows the most appropriate links, decides what to record in state, etc. I struggle to see how you could implement something with a classifier. That said, I have no idea how deep the technology is and it might be replaced with open source pretty quickly since its drafting of the frontier models and the open source models seem almost as good.
outside of ML, for normies i mean, classifier models are supposed to be general purpose. Like I don't think we say segmentation model to do X. In ML, we know of their constraints so we usually say classifier models trained to do X.
Not complaining that decision models are a better term though.
it's a generalistic classifier w/training needed tho. as far as I'm concerned that is somehwat novel and very practical. you get an ok classifier without working out training data and whatnot. sure, you can do classic ML experiments, but the closest that comes to mind is auto ml. not sure how far we got there, it's been a minute for me on this front.
so I don't agree on the premise that it's "just a classifier". it's not something earth shattering, but practical nonetheless
This isn't eating Jev's lunch. This is someone who doesn't understand the entire use case of Jev replacing it with something that doesn't handle it at all.
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.
You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).
That's a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
There are two parts in the data you supply to Jev for classification - the prompt describing your classification and the data. The data can be quite small - a simple chat message. And prompt part could be considerable since you need to describe your rubrics well.
With Jev you each time pay for your prompt, you can't cache it.
I mean, it sounds like it's only ideal for cases with significant system prompt overhead. I don't think Jev was built to have a large well described prompt setup. To me its more like a happy go lucky small label classification tool with important decisions left to stronger agentic models or yk humans.
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long.
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
Update: Jeeves took about 2 hours to moderate 394 data points and performed really well, it's not as good as Jev, but super close! In general, it's super cool that you can turn how strict you want with content moderation with these models.
just get an LLM to think and then force it to output a specific json with prefill post-think.
make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.
Cool engineering, but 17s p90 latency kind of defeats the point of a Jev-class model, which is supposed to be fast and cheap. Losing 10 points on MMLU along the way doesn't help.
88 comments
[ 1320 ms ] story [ 355 ms ] thread/a bit more digging and..
A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true.
I hate it :)
Like, yeah, you don't hallucinate, but only because you force the user to decide in the end.
And that's...bad?
I think, it's a bit much to call this "no hallucinations".
Technically true, but in practice you could still choose the wrong result or the probabilities can be off.
Ask Jeeves hired hundreds of cheap liberal arts majors to classify data, some users thought Jeeves was real, the stock spiked when big companies hired Jeeves to automate support thinking it was a silver bullet, and the whole thing collapsed when a better model came along, and it degenerated into ripping off rubes with bottom of the barrel ads.
But funny that jev is getting its lunch eaten apparently in under two weeks?
I've used it to do web research where it follows the most appropriate links, decides what to record in state, etc. I struggle to see how you could implement something with a classifier. That said, I have no idea how deep the technology is and it might be replaced with open source pretty quickly since its drafting of the frontier models and the open source models seem almost as good.
Yes, one general classifier would be very hard to train. However, you can create a sort of ensemble of classifiers, each trained in different tasks
I’m currently experimenting with this. So far I’ve combined classifiers for 13 different datasets, my target is 95 (the ones Laya used for training)
Not complaining that decision models are a better term though.
so I don't agree on the premise that it's "just a classifier". it's not something earth shattering, but practical nonetheless
I guess it's not really a benchmark but you could say if it can do it faster it sort of could be taken as one.
For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
That model scales very well with quantities of requests.
You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).
With Jev you each time pay for your prompt, you can't cache it.
So why didn't they show both??
It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine.
[1] https://tn1ck.com/blog/jevdit
just get an LLM to think and then force it to output a specific json with prefill post-think.
make sure to include good conditioning text in the prompt with examples of exactly what the output should be like. you don't want dissonance in the probabilities on the prefill.