unlikely I think, they're likely doing this to garner some interest in their company but they seem pretty interested in revenue (judging by the companies they're working with)
I never thought i'd see the day they released a model, rather than a blog post. The Figure 3 demo being a screencap of chrome in localhost made me feel better about myself. Jokes aside, best western open weights model- very cool.
Seems like this is particularly good at instruction following, but not as strong at coding as others. It's always great to get more diversity of open weight models though! I'll need to test this out to see what its "personality" is like.
It's nice to see a strong long context open weights model that is multi-modal.
There are many applications that will benefit from the strength in audio here and until z.ai and co work in visual this could be very strong for general agentic applications, though I see there's a bit of weakness in the benches for areas that might make that less true.
Like all models need to slap it in your harness and do proper evals on the tasks you care about.
They also indicate they have a 276B A12B version, but it doesn't seem the weights are available. This might actually be able to fit in 128GB when quantized to 2 bits or so which makes it interesting.
Raised 2 billion dollars at a 12 billion valuation and debuts at 41 on the Artificial Analysis Intelligence Index, while KIMI and DeepSeek will release Fable-class models this week. What a joke.
I think your comments might be overly negative. Would you expect the first model from an organization to top the chart?
It's a process, I think they did pretty good. They have enough resources to continue improving on it
For a first model, and given it's open, I am gaining some faith in American Open research labs again...
I couldn't test it since it's not on openrouter or something, but even if it's only as good as GLM5.1 that's more than good enough first attempt, I think.
Perhaps a lot more labs will catch up to ballpark frontier esque level soon, I am all for more competition in any field.
Interested in the implied strategy - that training a bespoke model for what you need will make economic sense over using a mass-trained model. I wonder if that's true?
Interestingly, when opening this page, the first thought I had was not that the benchmarks should be high, but 'I really hope they did not benchmaxx'. I think a model with modest benchmark scores can have much better real world utility as opposed to the current frontiers that are RL'd into being robotic and rigid.
This supposedly is better than KimiK2.7, as much hype as GLM5.2 gets, I find myself using KimiK2.7 half of the time, so if the benchmark is true, then this can definitely go in the mix. My hope is that it might have strengths in some areas to beat all other open weight models.
I'm sure it's better than KimiK2.7 and GLM5.2. Benchmarks aren't the full picture. Despite GLM5.2 performing well on benches and supposedly near frontier, in reality it was nothing close to frontier in actual usage.
85 comments
[ 4.5 ms ] story [ 88.8 ms ] threadThinking Machines might be it.
What does "winning" mean to you?
Maybe for the multi modal?
There are many applications that will benefit from the strength in audio here and until z.ai and co work in visual this could be very strong for general agentic applications, though I see there's a bit of weakness in the benches for areas that might make that less true.
Like all models need to slap it in your harness and do proper evals on the tasks you care about.
https://artificialanalysis.ai/models/inkling
I couldn't test it since it's not on openrouter or something, but even if it's only as good as GLM5.1 that's more than good enough first attempt, I think.
Perhaps a lot more labs will catch up to ballpark frontier esque level soon, I am all for more competition in any field.
If you want to run locally, checkout https://github.com/danielhanchen/llama.cpp/tree/add-inkling https://unsloth.ai/docs/models/inkling https://huggingface.co/unsloth/inkling-GGUF https://huggingface.co/unsloth/inkling-NVFP4
This supposedly is better than KimiK2.7, as much hype as GLM5.2 gets, I find myself using KimiK2.7 half of the time, so if the benchmark is true, then this can definitely go in the mix. My hope is that it might have strengths in some areas to beat all other open weight models.
How can you tell?
I just looked at the benchmarks and was kinda disappointed that it seems to be between KimiK2.6 and KimiK2.7 on most of the benchmarks.
Do you refer to what it feels like to use the model? Or are there other benchmarks I haven't seen?