1 comment

[ 0.22 ms ] story [ 7.2 ms ] thread
When you are measuring how intelligent model is, in fact you mostly measure how good it was following the plan, or question, based on your own intelligence.

Difference is subtle, but crucial. And it is so easy to be confused, while measuring small model performance, like various Flash variants or local models.

Smaller models became quite good at doing the work under the well specified goal, like fixing the well specified bug. But can it find the bug, without you giving it any clues?

In my benchmark the answer is mostly No. Being able to invent the questions - thats the real intelligence - and thats where the big models show true strength.