Sol 6 is in there? You may be on Enterprise where it didn't roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)
Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
I think they are hitting compute restrictions. And buying compute right now can be 3-4X. And the costs are increasing. If they train a larger model and demand is high, that’s a lot of compute for Codex subscriptions, which is a loss leader for them. Especially Pro 20X which they just nerfed to 10X.
>Another piece of evidence on the pile that the sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
It's not so much that they're hitting a plateau in capability, as we're saturating long horizon benchmarks and it's not greatly improving general usability. On the other hand, newer models have been amazing for people interested in 3d, graphics, video editing, etc. The difference between Opus 5.5/Astra and earlier models is night and day even if for many coding tasks they're not a revolution.
I agree that they're not hitting a plateau and I see it in my reserach. I had a math/code benchmark paper [1] at NeurIPS last year that is still unsaturated. At the time of writing the paper, the best model was o3, which was scoring 3-4%. By the time NeurIPS came around, GPT-5.2 was the latest model but it was getting similar scores to o3. The models were still in the flat part of the usual hockey stick curve. The newer models are getting into the steep part. I evaluated gpt-5.6-sol+codex a week or two ago and it got ~16%. Astra+codex got ~24%.
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
> sudden panic and desire to "slow down" is because they're hitting the plateau on capability
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
Is there anything that could happen that you wouldn't use as evidence that they are hitting a plateau?
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
I am a lawyer not a coder, but for me the new models make the exact dumb mistakes they did in 2023. Everytime I come here I feel I'm in an alternate reality.
I'm a software engineer and I'm in a similar boat. The models have definitely gotten better, but all the latest frontier models have been within the same order of magnitude of usefulness for many months now. Opus 4.5 was a huge boon in productivity, and I certainly write less code by hand than I did when it came out, but I can't say that my workflow has changed drastically for many months. Every model still takes some amount of babysitting to ensure it's doing the right thing, and they all make silly mistakes sometimes.
IANAL, but it sounds as though law hallucinations are the final frontier. Sorry.
But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.
Hence your alternative reality.
(All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)
Some version of this claim has been made for the past 4 years. There's a data cliff, there's no more compute to buy, the financials don't make sense and all of these orgs will be out of business by end of quarter.
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
>Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
I think we'll eventually hit an information theoretic type of wall with physical hardware and GPUs and need a similar AI breakthrough as well as the development refinement of logical/physical qubits in the quantum computing space with some analogue to the transformer architecture to continue accelerating. However, I think there must be many years of development and refinement that can take place before that paradigm shift to overcome the physical compute wall is necessary. This is just my theory, but I'm young enough that I'm expecting with the rate that we are advancing, I will see AI / LLM analogues developed and run on a quantum computer in my lifetime.
Yep. I've made the claim (and been wrong). I was convinced the data cliff was going to be a real problem. Now I feel like we are on the cusp of having Tony Stark's Jarvis at our fingertips.
Incredible how many times I read similar comments over the years, containing 'on the cusp' and 'what a time to be alive'. Indeed, what a time - not a single user-facing thing on the internet has improved since then, considering the power tool we got. The most used web services get drowned in generated stuff and so are the users
Not a single thing? In my house, we are using LLMs to:
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
> I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
Companies need to radically change to be able to take advantage. Most companies are afraid to do that and are letting their engineers serve as slow meat proxies, doing software development basically the same way as before.
The difference now is that they've hit the "good enough" point. LLMs are a tool, and that tool is useful but not incredibly valuable unto itself.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
This is how I feel about it. I've stopped looking at all the scores of new releases and just look at the price to see how much usage I can get in a month. Seems like I'm not the only one either, from comments above like
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
>The difference now is that they've hit the "good enough" point.
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
Right, that's exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation? Or is most of the economic value in the stuff that already exists, and can be performed nearly as well by qwen/deepseek/kimi/GLM/etc?
And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
>Right, that's exactly my point though. Those are absolutely valuable use cases. But are they useful enough to justify a trillion dollar valuation?
Replacing white collar work would be worth dozens of trillions of dollars. Software is not the only valuable job that can be done on a computer.
OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.
Chatgpt is used by a billion people every week. Their ads program hit $1B Annual Revenue Run Rate in 200 days. And Anthropic is growing so fast they're on pace to hit $100B in Annual Revenue.
>And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic and OpenAI are still growing enterprise usage heavily. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit about what qwen does. And specialized models often perform worse than generalized ones.
> ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
I didn't use the word LLM. I'm talking AI capability, you're focused on this or that current approach to AI. I think it's fair to assume that the approach will change as new ideas are learned, new and more hardware will be purchased and applied to the problem, and then capabilities will (for now) continue on their exponential curve, same as it has gone for the past several years.
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
tl;dr it changes the weights, it does not add new ones.
RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.
Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.
The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.
Hmm interesting idea. I’m pretty confident there is generalization and learning that occurs during RL that does make the model smarter. So I think the distinction doesn’t fully hold up.
Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.
For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.
I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.
This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.
Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.
Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?
Has it been exponential this whole time? I feel like GPT-4 was pretty dang good. Maybe it’s rose tinted glasses cause I could finally have a bot write my dockerfiles and bash scripts, which knocked my socks off
I think this is just a case of the bitter lesson that increasing compute just makes all these predictions meaningless. LLMs just keep going when everyone predicts them to fail constantly.
This is literally the plan, open weight models are something like 60% of token spend, and it will get worse. many companies now have model gateways where you can slot in cheaper models via cli for cheaper. we've been using glm 5.x and it's pretty close to SOTA frontier models.
it's also why there have been so many calls for regulation and slowdowns.
Exactly what I've been doing. I don't need the all-powerful GPT-6 Math Scoopa, or Opus T-1000, just to write react, svelte and C# for me; my local Qwen3.8 is more than capable, and I can switch to Deepseek and GLM on OpenRouter when I need speed. I just pop in to read the comments on HN for the latest drama and navel gazing, then I click the Hide button and move on. Couldn't give a wooden nickel what their latest and greatest models are capable of anymore, it's just PR buzz.
There is already tooling to automatically pick models within an organization. Eventually it could be as easy as flipping a switch in group policy that forces everyone to switch to the cheaper models.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
Pretty standard business to identify and compete on every axis (cost, speed, intelligence, etc). Often, nobody will be able to maximize every axis so you end up with a polyhedron derived from the axes where there’s a niche for everyone.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
Switching models is _very_ expensive in compute (you have to rerun everything from the beginning), and highly variable in cost. Cursor tried doing this for awhile, but inconsistent performance/usage means most users turned it off and pick models specifically.
These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.
I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
DeepSeek v4.1 Flash is fascinating and uneven. It's way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex/claude tui's and it's usable. but it seems to be "brilliant and yet stupid" in a way I can't quite put my finger on. I've got too much real work to get done to dig into it so until the big boys price me out of the market I'm back to my $100/month deal.
I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
I am using the API only with them for now. But what you describe is nothing compared to Qwen or Mimo. These models are more capable than Opus in general and at a fraction of Opus's cost.
What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?
If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.
I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
Probably one of the factors.
Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.
Luckily it's not a mistake as now we have access to
.
.
.
dots.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
I see. One of the biggest issues I had with 5 is that it constantly made mistakes calling tools and making API calls that even lesser models didn't struggle with. Mistakes were crazy high.
Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
They have also cut allowances for subscriptions in half. So even in the best case scenario it's about 2.5 times cheaper for Codex users. They just seem to have matched Claude Sonnet 5.5 *API pricing*, but from what I see online, it seems Claude Code now has a much more generous subscription allowance.
I love free market competition. We're getting insane advancements every day. I remember when llms used to cost an arm and a leg for decent intelligence
This is great. But maybe part of the motivation is that 6-Sol wasn't as good as initially advertised so they needed to tweak it. I felt a clear degradation in quality in some simple refactoring tasks vs 5.6-Sol.
Yes, obviously. They're both working to make it cheaper, faster, and better at different industries (3d animations, etc). The only direction they are slowing is raw intelligence.
I wish they'd list the environmental cost. My employer has an unlimited AI budget so I don't care about using Astra if it's just more profit for OpenAI. I care more if it actually uses 5x more energy.
I don't understand the point of this, why just now when it comes to llms. Why wasn't anyone enraged with the environmental costs of kids playing video games. I would not be surprised the environmental cost of that is an order of magnitude bigger than what llms have.
If you recall history past the last 5 minutes, you will remember that people have indeed been enraged with the environmental costs of things for a long time. Its just that AI seems to have induced a mass amnesia, and people tend to forget about what happened pre 2024.
There are movements against consumerism and the environmental impacts of industry in general. Greenpeace is over half a century old.
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
Because people find video games fun, though I suppose there's some vocal people that think of them as bad for society. In contrast the AI companies are promising a torment nexus future.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
Given that the number one cost of inference is memory and compute, and the incremental cost of each is energy, cost per inference is roughly proportional to energy consumption.
Yes I do. I've got a spare desktop that isn't too efficient (probably ~100W idle but annoyingly I've lost my power meter) so I don't leave it on even though I would like to use it as a server.
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
I want to energymaxx. Every home should have a nuclear generator for free limitless clean energy. Do not energysimp, we want prosperity for all we must energymaxx and invest heavily in solar/battery/nuclear.
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
I've only tried animating models in Astra-6, and I was quite impressed! It's rarely able to one-shot things perfectly, but it usually gets pretty close.
From the results of a lot of YouTubers in the space, I think Opus 5.5 is pretty competitive with Astra in 3D. It's slightly worse at spatial detail but better at aesthetics and little touches.
The poster above carries this tone in his posts where he is the divine one. I bet on a lot of stuff he ain’t got a clue what he’s talking about but hopes people like you don’t catch him out.
After all the hype, I’ve been kinda disappointed tbh. Modeling specific models are so much better (eg. Tripo3d). Astra still models some janky crap for me.
If you run out of sol medium with $100 you're doing something wrong. Astra destroys your usage, I get 1 day of usage with Astra, but 6 sol is almost unlimited and I only use xhigh.
> "You're holding it wrong." Is hardly a retort from a real paying customer having problems with their paid services.
It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.
Sides? Hate? This is all very emotional. Try to put the facts down plainly and see how ridiculous it is --
It's a product and if you're using it incorrectly, we can either
1. say so
2. pretend that you don't to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?
How is 2 better in any way for anyone involved? Why would you, as a paying customer, holding it wrong, want other people to keep that information from you?
Literally the prompting and task definition is main variable how LLM's performs. There are literally millions of examples of vibe coders and new AI adopters who run out of tokens since they don't know how the LLM's work.
yeah I use sol constantly and have done maybe $15 of spend in the past week. it's solid and cheaper. this is at least 4-5 investigations, prs, whatever per day.
It’s only nearly unlimited if you haven’t just used a banked reset. After a banked reset your weekly usage gets cut by about 80% (not the week you need to wait to get your normal limits back though). ChatGPT has given me a really good reason to cancel.
Can you elaborate? I've been getting great usage out of my $200/mo plan, and thought I'd try a reset (first time) which was expiring just for giggles. Am I going to get only 20% of it effectively?
I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)
I can’t say what will happen to you, but yes, that has been my experience. It is better to wait for your normal full limit to return, because if you use a banked reset you get only 1/5th of the tokens but you still have to wait the full week afterwards for it to reset. 20% would be fine if it didn’t also reset the date your normal reset fires.
You have a lot of control over compaction, both directly by changing compaction settings, and indirectly by how you structure your codebase/docs so agents use less tokens.
Context window is only 275k or something. And honestly compaction is not that bad in Codex. I often don't even notice I went through 5 compactions in a session.
Same for me, I started wondering if maybe workflows using compaction instead of clear + markdown memory would be more efficient. Writing a plan or tasks to a file often has the next session repeat part of the exploration, compaction seems to keep most relevant context.
Sounds like that's the problem then, 275k is a tiny context window. I regularly have sessions that go to 450k or even up to 700k for an unattended overnight Claude Opus session.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
Its really not tiny; you can't compare Claude to GPT, they have honestly diverged enough that as the other reply said, 256k GPT is about equal to 1M Claude. The compaction is slightly annoying, and you can turn it up to 1M as you said if you truly need everything in context, but otherwise it's perfectly serviceable
Absolutely not my experience. I've been a long time Claude user and at work people slowly started using Claude too. I canceled Claude for my personal account due to the Astra hype, but I'm back after a single month and the absolutely ridiculous context window is one of the many reasons.
I barely compact at work in a very complex monorepo (neither with Fable 5.1 nor Opus 5.5), and yet in my personal greenfield project Astra keeps compacting all the time, to the point of it being unusable.
Not sure I understand if this was meant as a slight against Claude? Or agreement?
These are often my best sessions - they're unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep & wake up to an entirely new application completed. Claude never uses compacting in my sessions.
I haven't used GPT as much as I should have, so I'm prepared to be incorrect & out of date. It just intuitively feels like I wouldn't get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek & GLM have 1 Million context windows now, so it "feels" strange for people to actually prefer the 275K window. But that's just my intuition.
neither, actually, just that unattended "oneshot" sessions are incredibly token inefficient
if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you're comparing apples to oranges
a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail
I don’t usually have a problem doing a complete task in that context size. OMP does make a lot of use of rewind which may be helping - basically forks itself and sends back a summary after a long tangent. Coding tasks use a Luna max agent.
I’ve also found compaction not to be a problem when it does happen.
Its in settings under Tools->Checkpoint/Rewind. I don't know why its not enabled by default and actually forgot I had to enable it. But its a great feature that can really stretch context.
If it's compacting every 5 mins, you're going to notice it in your cache miss ratio and your costs...
It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...
you can actually leverage 400k and 1M contexts in codex with very little code changes to the harness. note that excess context past the.. 250k or 400k mark (i don't remember) is charged at 2x the price.
Your tool calls (MCPs?) are very likely too wasteful. Apply some filtering logic on the offending tool’s output. Either a wrapper CLI, or just tell codex how to filter.
> Cache doesn't help you much when you are compacting every 5 minutes...
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
iirc you can still turn the compaction limit up in codex, though they don't make it easy. It costs way more when you use "large context" though, more than the ~256k that codex allows by default. You can also use the large context via the api directly
what. I use Astra xhigh, sometimes max, never ran out of tokens on the 100$ thing. I'm using pi though which is by definition harder better faster stronger than claude code/codex.
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
Subagents will inherit the context window at the point in which they are spawned, but it sounds like you're more referring to orchestrating/conducting/managing multiple agents?
Both... even subagents inheriting the context window doesn't cost a huge amount if the context window was never that large, but yes orchestrating/conducting/managing multiple agents is even better though higher thought cost (but the newer claude agents are really good at this in my experience, part of why I am using Claude a lot lately despite the models being more expensive that ChatGPT's for the same performance when taken alone).
Whenever I see my main agent do a compaction, that to me is a clear sign I didn't have it delegate bounded tasks enough.
The backdrop being deepseek offering 1% (I remember it was ~1% when 4-pro first came out early this year - 4-pro is now removed) / 2% (current for 4.1-flash).
People's volume and approach varies. I'm a happy customer and I use my entire double max subscription on planning and analysis and have other models doing all my implementation work because I would burn through my subscription in a day or less. It's difficult to calculate, but I'm something like 10-20 billion token per week consumer and I can't use a US-based model to do this volume of implementation work.
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.
708 comments
[ 4.2 ms ] story [ 104 ms ] threadThese moves all make sense when you take into account the enterprise market.
https://news.ycombinator.com/item?id=49889873
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
[1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] https://oeis.org/A000530
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
People were talking about plateau for years already.
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
But coding-wise, models keep getting better and cheaper. You can train for code correctness in a way you can't train for legal correctness, and you can test your code in an agentic loop in a way you can't test a legal opinion.
Hence your alternative reality.
(All that said, 2023 was GPT-4 territory. GPT-4o wasn't released until 2024. No matter what question you're asking, I struggle to believe you wouldn't notice the difference between GPT-4 and the current frontier model set. You can download and run any number of sub-27B local models that will be better than GPT-4. The pace of change in this field really has been insane.)
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
do you think it will be exponential forever?
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
I think it's fully possible that it continues being exponential for decades like Moore's law did (and still is depending on exactly what you measure)
What a time to be alive.
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
Replacing white collar work would be worth dozens of trillions of dollars. Software is not the only valuable job that can be done on a computer.
OpenAI and Anthropic already have what it takes right now to become trillion dollar companies even if the above doesn't materialize.
Chatgpt is used by a billion people every week. Their ads program hit $1B Annual Revenue Run Rate in 200 days. And Anthropic is growing so fast they're on pace to hit $100B in Annual Revenue.
>And beyond that - how long until those individual Astra capabilities are distilled into separate Qwen-27b size models, with harnesses and scaffolds specifically designed to support that functionality?
How long until...you could say that about the capabilities of past models but OpenAI still dwarf everyone else in consumer usage, and Anthropic and OpenAI are still growing enterprise usage heavily. In the end, neither the billion+ users of gpt or the enterprise customers are going to give a shit about what qwen does. And specialized models often perform worse than generalized ones.
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
What distinction are you drawing?
RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.
Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.
The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.
For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.
I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.
This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.
Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.
Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?
it's also why there have been so many calls for regulation and slowdowns.
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.
The speed I'm having to update that document has not gone unnoticed.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring for some.
Astra is a pretty impressive model. Excited to try this.
Yeah? Show me the big movements against computer gaming.
The differences with AI are: 1) we are starting off (mid 2020s) from a baseline point of already being in a hopelessly shitty situation, past the 1.5C warming target; and 2) Electronics, chips, data centers etc were already a thing for a long time, but industry took _decades_ to ramp up production to pre-AI levels, and these things are used everywhere for a huge number of things. Now we're consuming electronics/data centers/water/power at an unheard-of rate, and for a single purpose (AI) with questionable benefits, besides the private interests of a handful of people.
I'd be curious as to how much of internet infrastructure is dedicated to gaming though.
Even then people do care about the power consumption of non-AI things. Look at the energy label on your TV or tumble drier for example.
But this is not that, the same gpus you play games with are used to run llms. How was energy consumation by gpu not a topic before llms?
> I don't think video games consume nearly as much power. A PS5's power consumption is apparently around 200W. That's not enough to run even one GPU, let alone the armada it presumably takes to run Astra.
Just Steam has 200 million monthly active users. Add Steam, PS, Xbox, and whole other devices having gpus and I'm pretty sure you at least 10x the energy consumption of all ai companies.
I dunno what you're not getting but a GPU to run games is like 200-500W. A GPU cluster to run Astra is probably more like 10kW.
Also gamers tend not to spin up dozens of other machines to also game for them.
Laptops use very minimal power - you don't need to worry about them. If they didn't their battery life would suck.
Huge misstep releasing it.
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
Why is the little dog barking? Where be the big dog?
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
It's very appropriate in the cases when you're holding it wrong. The fact that you're paying doesn't mean that you can't make mistakes or waste resources.
If this is how you want to get people on your side, I can understand why the entire country/human race are against these companies.
It's a product and if you're using it incorrectly, we can either
1. say so
2. pretend that you don't to get/keep you on "our side"? or not say is because you're skeptical or hate it? (how does that last bit even follow logically?
How is 2 better in any way for anyone involved? Why would you, as a paying customer, holding it wrong, want other people to keep that information from you?
I overused Astra in order to drain my weekly, figuring I'd have the reset. (not wastefully, I did get more work done)
I create a lot, but I can make a full month with Astra on the current Pro plan. What are you doing to spend that much?
1 day is kind of generous, it probably lasts like 12 hours of running non stop. In my testing 6 Astra uses about 7x as much as 6.1 Sol
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
https://x.com/thsottiaux/status/2089082893804896524
There's at least a forum thread about it here:
https://community.openai.com/t/why-does-codex-report-a-258-4...
I barely compact at work in a very complex monorepo (neither with Fable 5.1 nor Opus 5.5), and yet in my personal greenfield project Astra keeps compacting all the time, to the point of it being unusable.
yes, exactly
These are often my best sessions - they're unattended overnight, because by then we have the specification figured out, and I can just leave Claude to build out the rest, making good choices if it does find gaps in the spec. I regularly go to sleep & wake up to an entirely new application completed. Claude never uses compacting in my sessions.
I haven't used GPT as much as I should have, so I'm prepared to be incorrect & out of date. It just intuitively feels like I wouldn't get the same from a 275K context window - maybe it uses lots of subagents? Even Deepseek & GLM have 1 Million context windows now, so it "feels" strange for people to actually prefer the 275K window. But that's just my intuition.
if you talk about them (in which you lean on an LLM as a sort-of independent employee) and conservative, chunk-based usage (in which you use the LLM as more of an extension of yourself), you're comparing apples to oranges
a predefined spec obviously reduces that gap but how much is highly dependent on the level of detail
I’ve also found compaction not to be a problem when it does happen.
It also presumably means it's regularly not able to get everything it wants to have to make decisions in context, which means it's going to perform poorly...
~/.codex/config.toml
> model = "gpt-6.1-sol" > model_context_window = 700000 > model_auto_compact_token_limit = 630000
It's crazy on Codex. I sometimes get just 2-3 turns before it compacts. It has forced me to use persistent project documentation for everything. Maybe that's not a bad thing but unless it reads all the documentation after every compaction (and uses half its cache), it goes off the rails. By comparison, Opus 5.5 is a breath of fresh air. It takes FAR longer to hit the cache limit and that means it keeps useful information in working memory far longer. I think this alone has resulted in a massive productivity and efficiency increase for me.
The longer your chat gets, the slower and more expensive it gets.
Subagents are expensive but they scale way closer to O(n) than O(n^2).
Have some agents make bug reports/feature requests/roadmaps (linear is very AI friendly), others coordinate, others work on grinding out an individual ticket.
If there is a good ticket-level description, it's a waste of time IMO to have a main agent do it, that should be an agent with fresh context that will do it better faster (the shorter the context, the better models are at using the context they're given).
Whenever I see my main agent do a compaction, that to me is a clear sign I didn't have it delegate bounded tasks enough.
Clearly an OpenAI employee.
Also, a lot of this work is verification to ensure that AI generated code does what is intended and is safe to merge and deploy. That verification work is critical and uses a lot of tokens.
Guess not?