A company whose first demo was completely fraudulent announces that its model beats GPT-5.5, on its own benchmark? I’m gonna wait a little before I trust this.
This whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technical jargon...
Imagine how far community might have pushed if 2 past versions of 'morally superior' Anthropic and 'completely Open AI' open sourced their models for the community to build on top of them
I've always had mixed feelings about Cognition. Obviously they have some very, very smart people working there (I even know a few), and they do make real products. But at the same time, they've made suspicious marketing claims more than once and even been caught making outright fabricated ones; and while they certainly seem to have shaped up from that, I still find their claims to be in a sort of grey area where they seem to avoid unfavorable comparisons and lean on their own benchmarks. Certainly when I've tried their models they have not been nearly as useful as comparable versions of Claude, GLM, etc. -- though I haven't had a chance to try SWE-1.7 yet.
We need more models that optimize for coding and that can be cheaper than frontier models, like what SWE 1.7 and composer 2.5 are trying to do. I don't think there's an effort to make something GLM-5.2 level but focused only on coding.
Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot.
What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.
I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.
I'm looking forward to trying this out. I've been using SWE 1.6 quite a lot for grunt work alongside Opus for higher level planning and tricky stuff - a good combo.
As a (former) Windsurf user I'm pretty happy with the progress of the Cognition/Devin ecosystem after they took over Windsurf, now known as Devin Desktop.
The benchmark debate is fair, but I think the more interesting signal is how quickly coding models are becoming a category of their own rather than just smaller frontier models. More specialization, more competition on cost, and probably a lot more benchmark gaming along the way :)
I think it's a bit odd to show the API prices for competitors when that's not how most people pay for them. I do like that it's provisioned by Cerebras though. I think I'd have leant towards focusing on the TPS.
While I am skeptical of the results here, I am very excited for this new trend of making models faster. Running capable models at 1k TPS is more valuable for me than running better models at 30 TPS. I can only imagine the trend continues to move from "let's only make models smarter" to just incremental intelligence gains but with step improvements in speed.
Apparently 'free' on the $20/mo Devin plan (presumably within some quota still)
and that is "via Cerebras at 1000 TPS" according to the announcement
I live on Opus 4.8 High and their benchmark scores SWE-1.7 slightly higher ... if at all realistic that sounds like a great deal ... too good to be true?
Cognition... oh what a ride... We were customers when they acquired Windsurf, stopped offering customer support, raised prices, dismantled the brand, and raised prices again. We are not customers anymore. Benchmarks are not the only thing to worry about when you are using models.
I’ve unfortunately had to temper my excitement with Cognition’s models/products given the amount of unwarranted hype they created with Devin on first release, but hopefully this is good.
36 comments
[ 3.2 ms ] story [ 61.6 ms ] threadThis whole company seems to optimize for raising money and impressing VCs. Lying about their products, ignoring consumer market to target enterprise, bragging about how they work their employees like slaves, and writing these posts full of intimidating technical jargon...
Imagine how far community might have pushed if 2 past versions of 'morally superior' Anthropic and 'completely Open AI' open sourced their models for the community to build on top of them
What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.
I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.
1. https://cursor.com/blog/composer-2-5
As a (former) Windsurf user I'm pretty happy with the progress of the Cognition/Devin ecosystem after they took over Windsurf, now known as Devin Desktop.
https://x.com/theodormarcu/status/2074896486047834380
Apparently 'free' on the $20/mo Devin plan (presumably within some quota still)
and that is "via Cerebras at 1000 TPS" according to the announcement
I live on Opus 4.8 High and their benchmark scores SWE-1.7 slightly higher ... if at all realistic that sounds like a great deal ... too good to be true?
Time to support it in my agent IDE just like Cursor's...
But here, both Kimi 2.7 and its derivative SWE-1.7 are ahead of GLM 5.2. This tells me the benchmarks they use are cherry-picked.