Is there a reason the AI companies usually announce new products so close to each other. Like not just the same day but literally hours apart. GPT Live then an hour later Grok 4.5. As if they try to one up. I expect something new from Anhtropic as well today.
It seems to be extremely economical - 4x better reasoning efficiency compared to Opus while being priced at $2/$6. For comparison, GPT 5.4 is $2.5/$15, GPT 5.5/5.6 are $5/$30, Opus 4.8 is $5/$25, Fable is $10/$50.
Grok 4.5 is a huge step up from their next best model and now around the same performance as GLM 5.2, but it's not exactly at the frontier of the cost efficiency curve in our coding evaluations. That curve is defined by the 2 lighter GPT 5.6 models.
However, the fact that they finally have a strong post-training and RL setup bodes well for future releases. They certainly are not compute-constrained anymore.
Of the 3 models I tried, Grok did the best at making an iOS app I wanted for personal use (a bike computer with specific qualities). (Claude just gave up and did an HTML/CSS implementation but I insisted on native SwiftUI+Metal.) Grok definitely fumbles sometimes, but I have been surprised what it CAN intuit versus me having to micromanage it.
(I am not an iOS developer, so getting something specific that I needed in a few hours/days was really helpful instead of spending months/years learning the language, APIs, etc.)
(I am absolutely not "vibe-coding" Caddy btw, just tinkering with it for personal projects.)
I am very curious what your Claude thread looks like. I have never had Claude swap languages, in fact my experience is the opposite, sometimes it holds on too much when working on a large code base.
> Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.
Its remarkable how Anthropic is able to maintain their edge against all competition. Anyone have any idea what the secret sauce is that has Anthropic at the top of all leaderboards for the past few years?
Isn't this the same Twitter company that was supposed to go bankrupt a few years ago? Now it is somehow part of a Space company that has an AI division inside of it?
I think we are going to be waiting a long time for Twitter / X to go bankrupt as it was (erroneously) predicted a long time ago.
The solar system diagram doesn't work for me. When I click on the planets, it will center on them. When I click on the sun, nothing happens. When I click on a planet next, it goes to the sun.
So basically since US stopped OpenAI and Anthropic for 4 weeks, it allowed all other AI Labs to almost catch up.
GLM 5.2 caught up, Cognition RL'ed Kimi 2.7, Grok 4.5 is out, DeepSeek v4 GA is out in a few days...
What is the moat? and why should we pay for the expensive tokens today instead of just waiting a few months/weeks and getting AI for significantly cheaper?
I must say, I feel like companies spending Millions on Anthropic tokens are just negative capex'ing and wasting money, even OpenAI is barely ok pricing...
Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding, so this is an encouraging release.
What are you talking about gemini 3.5 flash beats fable at tool calling and is 5x faster... I think it's very competitive and what most normal people are using.
> Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments.
This is what the big money was for. Cursor is the first big player that had real-world data from real-world projects, before cc / codex were a thing.
> We used reinforcement learning on difficult problems in realistic environments spanning both software engineering and broader knowledge work. These environments teach the model to investigate problems, use tools, recover from mistakes, and verify results.
> Many of these problems had to be designed to be difficult enough that even frontier models fail at them. As models improve, existing tasks stop teaching them anything new, and problems that once required extensive reasoning become routine.
> We developed a distributed agent system to construct these environments at scale. Engineers specify a problem and how a solution is verified, and large groups of agents construct, test, and refine each environment.
This is where scale comes in. You use the previous gen model to prepare datasets for the next model iteration. The better the models, the better the data, the better the next models. (they also have a comparison with their composer2.5 training run, for people still thinking chinese models are "close to SotA"...)
Reports of xAIs demise (after giving a lot of compute to Anthropic) were slightly exaggerated, it seems.
> Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs
Props to them for including three benchmarks that actually seem to say something, instead of focusing on totally gamed benchmarks like regular SWE-Bench. That could mean this model is actually pretty close to the SOTA as the benchmarks indicate.
Most labs - including OpenAI and Anthropic, but also Google and Chinese labs - highlight their scores in benchmarks that have fixed, widely available answers. Those answers end up in the training data and so models can just regurgitate training data instead of actually doing the benchmark. As a result, most benchmarks often quoted are essentially meaningless for gauging model performance.
Terminal-Bench still publishes answers, but neither DeepSWE and SWE-Bench Pro do. Especially for DeepSWE it's been difficult for models to fake good results so far. SWE-Bench Pro does have weird outliers like good performance for e.g. the atrocious Muse Spark, but it also doesn't provide answers for the training data.
So either they're good, or they found a way to game DeepSWE. Given that the Cursor team previously published the well-received Composer 2.5 a good score here doesn't come out of nowhere, so this might hold up. Cursor has enormous amounts of training data to train good coding models with.
- Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency
- Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency.
- Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same tier as opus, occupying the lower end of that tier together with GLM 5.2.
The fact that it is more token efficient will itself lead it to be smarter since the context will be smaller for the same task. However, in opus models, you'll have built internal correlations like "if it did X it will usually do Y" which may not be true here, since grok 4.5 may have done X purely due to the smaller context size, but can't do Y cos it wasn't RLd on that pattern enough. So it will be a unique experience as far as opus tier models go.
Can someone breakdown to me how this makes any sort of economical sense? Spending billions and billions to have the 3rd best model while even the number 1 and 2 players already seem to struggle making a profit. What am I missing here? Not trying to go full Ed Zitron but this doesn’t make sense to me.
They only struggle to make a profit because of investments into their future, basically: training, aggressive hiring of AI researchers. Anthropic seems to have 90% margin on their API pricing and all enterprise customers have to use this.
And the reason they can do this is because they can create a $1Tr company in 5 years, so they know the investment will pay off.
Why Elon wants his own model so much is a good question with many possible answers, but if Cursor/xAI can produce a truly good model at competitive pricing I don't see why many people won't jump on it.
Stock market / investor driven products do not make economical sense, yet here we are; just because they (may) not make a profit, doesn't mean they don't generate value. My house doesn't make a profit, but it does appreciate in value over time. AMD famously didn't make a profit for years.
146 comments
[ 4.6 ms ] story [ 107 ms ] threadAnd by benchmarks (unless they gamed them), seems to be at around Opus 4.7 level, which is what Elon mentioned in https://x.com/elonmusk/status/2074911038286295049.
I guess the Cursor data was very useful.
However, the fact that they finally have a strong post-training and RL setup bodes well for future releases. They certainly are not compute-constrained anymore.
Data at https://gertlabs.com/rankings?mode=oneshot_coding
(I am not an iOS developer, so getting something specific that I needed in a few hours/days was really helpful instead of spending months/years learning the language, APIs, etc.) (I am absolutely not "vibe-coding" Caddy btw, just tinkering with it for personal projects.)
Notably:
> Grok 4.5 and Composer 2.5 are two different model weight classes, and we're excited to support both sizes and weights. Composer 2.5 will remain offered, and we will release new models of this size going forward.
I think we are going to be waiting a long time for Twitter / X to go bankrupt as it was (erroneously) predicted a long time ago.
I'll give this one a try with a grain of salt and lowering my levels of expectations
GLM 5.2 caught up, Cognition RL'ed Kimi 2.7, Grok 4.5 is out, DeepSeek v4 GA is out in a few days...
What is the moat? and why should we pay for the expensive tokens today instead of just waiting a few months/weeks and getting AI for significantly cheaper?
I must say, I feel like companies spending Millions on Anthropic tokens are just negative capex'ing and wasting money, even OpenAI is barely ok pricing...
Edit: Gemini 3.5 Pro. Expectations grow with each day it is not released.
terminal is nice but codex desktop app is very useful
> Training included trillions of tokens of Cursor data which capture a wide-range of user interactions with codebases and software tools. This dataset lets the model learn both from existing software as well as developer-agent interactions, capturing how developers work and how agents interact with their environments.
This is what the big money was for. Cursor is the first big player that had real-world data from real-world projects, before cc / codex were a thing.
> We used reinforcement learning on difficult problems in realistic environments spanning both software engineering and broader knowledge work. These environments teach the model to investigate problems, use tools, recover from mistakes, and verify results.
> Many of these problems had to be designed to be difficult enough that even frontier models fail at them. As models improve, existing tasks stop teaching them anything new, and problems that once required extensive reasoning become routine.
> We developed a distributed agent system to construct these environments at scale. Engineers specify a problem and how a solution is verified, and large groups of agents construct, test, and refine each environment.
This is where scale comes in. You use the previous gen model to prepare datasets for the next model iteration. The better the models, the better the data, the better the next models. (they also have a comparison with their composer2.5 training run, for people still thinking chinese models are "close to SotA"...)
Reports of xAIs demise (after giving a lot of compute to Anthropic) were slightly exaggerated, it seems.
> Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs
In the blog post, it is unclear whether Grok 4.5 is also a finetune on top of Kimi; they do imply it is also a finetune.
> Training included trillions of tokens of Cursor data… We used reinforcement learning on difficult problems
If xAI pivoted from a frontier base model company, to a finetuning company, it does mark a stark change to their relevance in the industry.
[0]: https://cursor.com/blog/composer-2-5
Most labs - including OpenAI and Anthropic, but also Google and Chinese labs - highlight their scores in benchmarks that have fixed, widely available answers. Those answers end up in the training data and so models can just regurgitate training data instead of actually doing the benchmark. As a result, most benchmarks often quoted are essentially meaningless for gauging model performance.
Terminal-Bench still publishes answers, but neither DeepSWE and SWE-Bench Pro do. Especially for DeepSWE it's been difficult for models to fake good results so far. SWE-Bench Pro does have weird outliers like good performance for e.g. the atrocious Muse Spark, but it also doesn't provide answers for the training data.
So either they're good, or they found a way to game DeepSWE. Given that the Cursor team previously published the well-received Composer 2.5 a good score here doesn't come out of nowhere, so this might hold up. Cursor has enormous amounts of training data to train good coding models with.
- Very fast, easily beats GPT 5.5/Opus 4.8/GLM 5.2 because of higher t/s (around 90?) and very high token efficiency
- Very good price, no contest vs GPT and Opus which are very overpriced if you pay API costs, and probably cheaper than GLM 5.2 when you take into account the token efficiency.
- Will take quite a while to get a feel for how smart it is, but it's definitely good, I'd say in the same tier as opus, occupying the lower end of that tier together with GLM 5.2.
The fact that it is more token efficient will itself lead it to be smarter since the context will be smaller for the same task. However, in opus models, you'll have built internal correlations like "if it did X it will usually do Y" which may not be true here, since grok 4.5 may have done X purely due to the smaller context size, but can't do Y cos it wasn't RLd on that pattern enough. So it will be a unique experience as far as opus tier models go.
And the reason they can do this is because they can create a $1Tr company in 5 years, so they know the investment will pay off.
Why Elon wants his own model so much is a good question with many possible answers, but if Cursor/xAI can produce a truly good model at competitive pricing I don't see why many people won't jump on it.
Did anthropic found their moat or we hit a Wall?