175 comments

[ 6.8 ms ] story [ 104 ms ] thread
This might explain the outtages today, as folks race, misguidedly or not.-
Claude Opus and Fable are so bad compared to GPT-5.6-Sol it's ridiculous, their desktop client is worse, and the value is worse because OpenAI has been spamming discounts. They better get their shit together at Anthropic!
Sure OpenAI has open source tools, and equivalent or better models, and lower prices, and they get along better with other agentic tools, but Anthropic has much better marketing and I guess that's what matters.
This is just the cycle. If the OpenAI offerings are better and cheaper, people will shift there and then those discounts that made it such a deal will evaporate. Anthropic or someone else will entice people back with the discounts.

The whole thing is getting ridiculous.

I saw this announcement when it came out and completely forgot about it. That explains why Claude Code felt so surprisingly generous these past few months.
The promotion is ending; the limits are reverting to pre-promotion levels.

"From May 13, 2026 through August 19, 2026, your weekly usage limit in Claude Code is 50% higher."

If they remove that extra 50%, I'm moving my $200 to Codex. Good thing we will know tomorrow, just a day before renewal.

"Your subscription will auto renew on Aug 20, 2026."

But I doubt they will remove it, just like they didn't remove Fable from the subscription. There's simply too much competition.

God I hope so, I'm rinsing my cc limits across two max accounts with ease atm
prediction: extension of promotions because of the progressive move to codex.
Anthropic seems cooked right now in terms of compute
They've been really struggling all of 2026. All of these "limited time" promos, just to have less usage than codex, and the shenanigans around "peak hour" reduction earlier in the year.
To be fair, their growth is ridiculously insane. I don't know how else that can be described
It’s like we’re hot-spotting, but across providers rather than across infrastructure!
it looks cooked to you, because growth numbers are cooking (to them)

(nobody uses claude anymore, it's too over-subscribed)

Flavor of the month LLM gets inundated and swamped with users.

Sideline LLMs have free compute and offer cheap prices to draw in crowd. Becomes flavor of the month LLM.

Back to step one.

It should be pretty clear by now that token prices are predominately a function of available compute.

Meanwhile my $100/mo subscription allows me to spank Claude all day doing the work of 4 coders and never hit any limits.
Looking like this will be the last month with Anthropic. Between the outages and just overall crap utility of Opus/Fable lately...
Maybe with lower weekly quotas the load will be lower and we'll have fewer outages? A guy can dream.
They'll all be getting around to this sooner or later, its just too expensive
I get incredible utility out of $200 per month

It's hard to put a number on it, but even accounting for all the time in meetings, talking with stakeholders and developers etc I'm over 10% more productive overall. I earn substantially more than $2000/month, the ROI is there

It's only expensive compared to the currently very strong offering from OpenAI. Or other models - my hobby projects are all on DeepSeek

What you’re saying is that they should be charging $2000 or more per month then.

Perhaps that’s true, but it’s a price tag that’s a lot harder to swallow for many orgs than $200/mo, and would require some hard justification for how your increased productivity contributes to the business bottom line.

I’ll agree that you can probably do that with hard numbers. I am skeptical that most $200/mo users could.

I think the rapid improvements in local models will make a $2,000/mo price point completely impossible to sustain. This feels like a race to the bottom.
yes and no, anthropic and openai are losing money on people who max out their sub, but openai has a lot more room to play with with much cheaper models to serve (by all signs we have from actual api/task pricing)
I haven’t had any issues with Fable or Opus code-wise, but the way it has started writing recently has almost become incomprehensible.

At the end of a long session it will start saying stuff like “There are smoke tests on the foundation-gates that are left for the cutting seam checks on these domains, which is genuinely your decision”

Never before the past couple months have I ever not been able to understand wtf it’s even saying lol

I'm seeing a lot of that too.

I was working on a project recently that required the attribution of a data source, and it added to the "licenses" page of the project something like "We use <blah> and per their terms we owe an acknowledgement of attribution to you, the user."

So in addition to the stuff it says in a session there's gobbledy gook that it prints out in copy as well.

The writing really is inscrutable at times.
Same. especially Opus. It is unbearable.
Hahahaha I wasn’t sure if I was getting dumb or something. Mine finishes the tasks and then outputs paragraphs of that nonsense. But the code is great.
My subscription expires tomorrow. Had subscribed ~March, eventually ended up paying for the $200 tier. However, there isn't enough of a moat for me to keep spending $200 if the models aren't being helpful anymore.

The unnecessarily jargon filled language was already pretty bad, but the latest Opus/Fable also just didn't seem to go far enough when asked to look into something, often leading me in circles.

Yeah it feels like they tune or 'optimize' the models after release and they get less useful to me as time goes by. I'm most productive the first few days after a release. Noticed this since opus 4.6.
I'm getting to that point as well. Has anyone noticed that sometimes claude just seems way dumber on a random day? Like they're doing some internal tweaking? Also, I'm 100% convinced that it's verbose writing style is to generate tokens to make you hit your limit quicker. I have so many rules and it still talks in this bizarrely verbose, cryptic style where it's sometimes borderline impossible to understand what it's saying.
I just cancelled my Max 5x plan. I was already mad at the limits, now Anthropic says it will get worse ? Nah, thanks.
Well the subscriptions are kind of coming to an end anyway because they can't support the subsidies. I've been building www.freepi.ai which is totally free (ad+training) supported inference. It's using a PI harness but I have a webfrontend coming soon. I'd love your feedback!
I wonder if this promotion was because of the summer holidays.
Will Anthropic be the Netscape or Yahoo of our time ?
This is ridiculous neither Netscape nor Yahoo made much money (source: I worked at Netscape). Anthropic is one of the largest fastest growing businesses of all time.
I wonder if they schedule these promotions around training runs?
I forgot this was in effect. I’ve been hitting my weekly rate limits two to three days in because of needing to steer Slopus 5 with Fable. Absurd that they’re being cut further in the face of steep competition from open weights models.
I got a max account during this promotion and it's been very fun but I'm a little burned out and am kind of looking forward to going back and tinkering with game engines without the help of an LLM for code gen (will still use it for documentation questions but I can do that with the free tier)
I too have the same experience with LLM generated code. I just get burnt out reading it and all the verbosity it has. I eventually gave up reading through it and keep it as a archive of previous work, while handwriting the code. If I get stuck somewhere, I asked the agent to generate a snippet based on my archived code and what i need to achieve and manually copy paste it.

I wish inline completions would have gotten better but it seems no effort is being put into it and it is frozen in time right now. Thankfully, the stuff i use this code for is simple and small scale. I can’t even begin to imagine how actual programmers are feeling about this.

Will they, though? I could see Anthropic extending it since OpenAI cut the pricing of Sol by 50% for the time being.
Wasn't that a limited time promotion from OpenRouter? I didn't see anything from OpenAI themselves.
Only API or also sub?
Only on openrouter for some reason. Not OpenAI Subs/API.
Yet another scramble to use up weekly quota before an Anthropic cliff. This isn't fun.
Quite a number of the smartest people I know got scalped to work at OpenAI. I don’t know anyone who went to Anthropic. Something I think which is under emphasized is that OpenAI has the human capital in addition to the financial capital advantage over Anthropic.

Over the long term I think OpenAI will produce the better experience when it comes to model quality, harness quality, and availability. I have been using codex the past few months and never looked back.

I suspect the net result will be that more Pro $100 subscribers upgrade to $200/mo. If you explicitly do simpler stuff with Opus, you generally have enough Fable time for most programming and planning tasks.
$100 Claude and $100 ChatGPT Pro is the best value for $200/month. They can see each others' mistakes.
Not to mention a lot of people who already have more than one max20 sub getting additional ones.
As we speak, Claude is down. I was already complaining about how extremely slow and dumb it has gotten since Opus 5. I guess this will be the last nail in the coffin for Claude. At least for me.
I actually use claude these days mostly for mission critical tasks and code reviews. Everything else I use Sol. I have $20/month plans on both. For non-coding I also use Gemini on a $20/plan. So far I haven't had any issues but we'll see how things go. I have Qwen3.8-27B installed locally but until I upgrade my mac it's not for day to day stuff.
they never made money on Claude code, it was always a way to get used recommend API access to your company
I think the difference between the Anthropic token maximization approach (vibe code all the things!) and OpenAI's focus on efficiency, terseness and token reduction are going to be the defining features of who wins the long-term race.

My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements. And as we see from some of the Chinese models, they can quickly distill frontier models with the tax of being slower and more token-guzzling, while retaining most of the quality. The real differentiators are becoming speed and efficiency, which translate to cost and user velocity more than incremental capability improvements.

It's great that frontier models can solve complex math equations, but bread and butter LLM usage (where the money is made) has already shifted from "I need the best always" to "what solves my day-to-day problems quickly and consistently". Fable usage as a percentage is flat-lining. We are already at the point where output quality is negligible. What wins going forward is cost, speed, consistency and the compounding effects of "softer" improvements to the harness.

Oh no, I use the top shelf models every day and I really think they have a lot of room to improve in pretty much every regard. I suspect Fable usage is flatlining because it's not that good comparitively and way too expensive.
I really would like to see OpenAI’s focus on efficiency but everytime I use Codex, it wastes tokens like there’s no tomorrow, hitting week limit in a day, where I’m able to use Claude just fine. Maybe it’s based on the codebase, I don’t know, but I have better results with Claude than Codex.
I had the opposite experience. Claude models and its harness feel like they are set to eat tokens for everything, especially if it’s ultracode effort. I have seen it spawn 6 agents and eat my 4-hour quota right in front of me. Codex is on point and follows instructions well even with max effort. I like my analogy of Claude being garrulous and Codex being laconic. If I see more limit hits in same sitting session, I have to rearrange my workflows. my experiments with qwen and deepseek have been good, cant wait to try glm and other models.
Requesting ultracode is basically asking for maximum token usage. If you want to limit consumption, use medium or high.
Here is an excerpt from the system prompt for UltraCode (Same for Fable,Opus,Sonnet):

"Ultracode. When a system-reminder confirms ultracode is on, that opt-in is standing: author and run a workflow for every substantive task by default. The goal is the most exhaustive, correct answer you can produce — token cost is not a constraint."

It might be your prompting.

My boss uses opus and gets good results when I use it always burns tokens. The other way is also true too I get great results with sol/terra but my boss does not.

Maybe. But if Claude is fine with my prompts while Codex uses them as an excuse to burn tokens. I have nothing against the model, they are more or less on pair, it’s just that Claude is giving me more value for same money.
If Claude is working for you then stick with it! Happy you’re getting good value from Claude.
One of my challenges is that I often end up burning tokens trying to build the right solution rather than building a solution that is good enough and deferring the right choices until later. I have been actively working to change my expectations for building with AI to accept worse solutions to get things going rather than trying to solve everything at launch.

This has been a recurring problem for me - as a security engineer catasrophization is a fundamental skill to finding vulnerabilities in complex systems, but makes me to conservative when building. For some of my leaders they are much better at saying 'good enough, ship it, and fix it later'.

At least with OpenAI you can use a third party harness like Pi.
Are you kidding me? OpenAI weekly limits feel _infinite_. I run out of Anthropic limits reliably halfway through the week
Codex is fabulous at work, where token use is near limitless and ultra-thorough tool use is welcome. Go ahead and fire off searches for look-alike terms on my 96-core cloud instance.

By contrast, Claude Code's bias to make assumptions of reasonableness about underlying systems has proven to be immensely frustrating over the last month or two, both personally and at work. I've wasted days on "that was my mistake. I've been reporting numbers on the old architecture because I hadn't enabled the new one in the config" both at work and home. It's immensely frustrating.

But here we are. Wrestling with energetic idiots in model form, wrangled by over-specific harnesses that struggle to stay off of deranged side-quests.

What a time to be alive!

> who wins the long-term race

Isn’t the race between Chinese open-weight models and the others more decisive for the future?

I think the race is actually between locally optimized rigs with extremely strict context management, and the rest.
Those rigs are running one of the open Chinese models though, right?
You've articulated what I found unsettling about Boris' pov that "coding is a solved problem". I listened to a few of his talks and was instinctively off-put by that sentiment. I figured a fellow programmer would understand and speak on the nuances.

Granted it did make me think about my biases and to lean into more future facing inevitabilities. But you've nailed it, for Boris and Anthropic, they are betting that coding is a solved problem in the sense that any person can one-shot any random idea and the output will be in some abstract sense "good". And then at what cost and toward what end?

On the one hand it feels true. On the other I ask - what good software have anthropic, or anyone else, produced that was fully vibed?

As someone who does near 100% of my coding via LLM these days, i still find that for anything complex i am still looking at and thinking in terms of code. Im still quality checking and steering at some interval via code. And im still not sure how or whether i can replicate that level of thought without still dealing in code at times.

You'll have to define "good" before you can expect quality responses.
In this last week I've started reviewing code more and in just a short time, I've found 3 fairly simple things that were introduced by Claude that were not wrong per se, however, they were very inefficient and didn't address the root of the issue. I still think we're in a place where the output looks good as long as you don't look under the hood or keep it scoped to small, vibe-coded projects. Once you get beyond that it can fall apart. Anthropic must have a large codebase by now though. Yet I haven't seen much released from them about how they actually work day to day on development.
Claude used to be 10x more efficient. I think there are no barriers of entry between one coding agent and the other. Therefore, I expect them to reverse as soon as people switch. At least this is what I will do.
There is no evidence regarding distillation. It is impossible to distill a model in just a month which was the gap between fable and Kimi k3. Anthropic wouldn't even keep up with the load. It is just another example of American exceptionalism.
(comment deleted)
Not to sound like a mark, I try to not get attached to any of these providers.

I've jumped between copilot, claude, gemini and chatgpt since the start of the year. chatgpt wasn't even worth looking at early this year.

Anthropic has the smarter models for sure, and seems to be default in corporate. However, the amount of budget you get with GPT as a user is much better, the harness feels more polished, and the models are faster. They are also much nicer to work with, I can just read the output for the most part. With claude I get pages of text and need to skim to find where the actual information i need to care about lies. So much more cognitive overhead.

Sol is smart enough for anything I've thrown at it, it's not one-shotting like fable, but I'm more willing to actually go back and forth with it, and it's likely producing better output to keep a human in the loop rather than trying to solve the world independently and making multiple incorrect assumptions.

I think GPT sees the market changing and is correctly repositioning themselves. Anthropic is down the wrong road, and if they don't correct course quickly I'm sure many of those enterprise contracts will start pivoting.

> I think GPT sees the market changing and is correctly repositioning themselves

I'm not sure how they'll survive their creditors tbh

I'm ok with that. When the.major AI players implode under debt, it will create a massive proliferation of folks who go on to start new businesses unencumbered by the debt and bad decisions by the current leaders.

The collapse of the massive overvaluation and circular investments will be the undoing of several major tech investors and companies that have a serious contribution to the things in tech we don't like. It will be bad for the economy, but it will also weaken the abhorrent control that those companies have over regulation in the United States.

It's gonna be a rough time, but it's pretty clearly necessary.

The only big problem is that if that collapse happens under the current administration in the US, they will be utterly incompetent to respond to a real domestic crisis, since they have been virtually unable to do anything without creating more problems.

> and it's likely producing better output to keep a human in the loop rather than trying to solve the world independently and making multiple incorrect assumptions.

I was just discussing this with a co-worker yesterday. I would really like a model (or harness?) that worked with me instead of for me. Walk me through its choices and decisions, let me correct it and guide it along. I would be way more confident in it's output, I would be more familiar with the changes that are being made, and it would make reviewing the final code way easier since I was making the decisions along side it. I'm sure it would also reduce the "brainrot" we're all going to experience the more we hand work to these models.

Those exist and they're called orchestration skill frameworks like Superpowers, GetShitDone, etc.
does "Superpowers" exist outside of the Claude ecosystem? I feel like I've become hooked on this workflow - honestly I could take or leave the models ... it's the workflows and the way that they essentially create tightly focused loops over multiple sessions that I've becoming fairly dependent on.
Yes, Superpowers is just a set of (agent agnostic) skills that you can use in any harness.
the "grill" skill from matt pocock is quite nice for this
Aside from the skills mentioned, I recently found that asking the model to create a simple HTML presentation to walk you through it's proposed design can do wonders. Fine tune your prompt to your liking (language style, what to include, what not, etc). Then, read that thing thoroughly, and keep asking questions and iterating on the plan until you fully understand what's about to be built, and are happy with it.
Seems like it wouldn't be hard for Anthropic to tweak a few prompts or RL pipelines to tune for terseness and token-efficiency if that stays as something that consumers want.

I find it unlikely that there's some fundamental property of OpenAI's models' "personality" or style which Anthropic (or any other serious AI firm) wouldn't be able to match if they wanted to.

> My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements.

As much as I wish you were right, everything about software economics for the last 30+ years has favored _less efficient software_. Traditional hardware has been optimized for traditional software for decades and we still see bloated software win consistently. LLM hardware is at the start of its cycle, with abundant low-hanging fruit to conquer--I would expect the pro-bloat dynamics to weigh even more heavily in the LLM space than in the traditional software space.

If the problem isn't that hard, I've been using Cursor's Auto or Grok 4.5 (not 4.6, it's too slow).

They're both pretty damn competent and more important, Extremely Fast! I find the speed more useful than trying to be a hundred percent complete on every task. The big intelligent models screw up all the time as well, but I have to wait twenty minutes to three hours to find out.

GPT 5.6 is especially tenacious and seems to want to solve every bug in edge case 1000% all the time. Sometimes that's what you need, but a lot of times you're just trying to move fast and figure out what the product is.

LLMs are already of sufficient capability that if they were super cheap and fast and made no progress on intelligence, we will still make huge leaps in our ability to build systems and solve problems.

I'm just happy there is capital behind the efficiency path because I am confident we can get positive value on that path.

They are both on borrowed time. The only thing keeping them afloat business wise is that current hardware costs are prohibitive which keeps local models artificially constrained. The next era will be open weight / local model for the masses and SOTA for the big players.
If your prognosis is true I'm not sure it's fully positive. Inference without batching is just massive less efficient, even if local hardware can be more energy efficient per operation and can skip out on cooling.

Open weights, obviously beneficial. Local compute when not necessary, seems like it'd be significantly worse for the environment?

To more fully flesh out what my thoughts on this:

I think because more devs cant run things locally, we are not getting the "open source gains" that you would get by having the network effect of more eyes on the work. Theres still little nuggets of gold in there though (eg antirez's DS4 project). But once more people can get their hands on hardware, I think the flood gates will open. And once that happens, I dont necessarily think people will run locally. But that will enable fungible service providers which we already see, but at a much larger scale.

To a lesser degree, open weights also leads to an eventuality of these two companies not being able to sustain their capital investments. Not to mention their pricing models are unsustainable as they currently stand. So they are being eroded from the outside-in, while also trying to fight the internal pressure to continue to raise prices/margins.

EDIT: I actually think the harness is more of a future for these big players. And theres real value in providing a high quality harness. So maybe they move to SAAS for claude code in the future. Who knows

The 'tokenmaxxing' and the extreme gambling fiesta with Anthropic's latest Claude Code slot machine engine called 'Fable 5' which has given many subsidies and free spins of the wheel and cheaper tokens are all incompatible with the desires of their future owner: Wall Street.

If Wall Street sees a single outage or a tiny drop in usage, they won't be happy and will pressure Anthropic to take away the free tokens.

Better to reduce the limits now rather than to wait until Wall St. tells them to just to avoid a stock punishment.

Buy a subscription, get something different every month.
To anyone like me that switches between Codex and CC based on limits, but prefers CC to Codex, PI has been really great to migrate over. It can customize itself very easily, I was expecting to have to clone the repo or whatever, but no you can just tell it "add a manual mode where you present all changes as diffs in VSCode, letting me edit them or save to accept" and it will do it.

I wish I could use my Anthropic sub with it but I heard you get banned, but at least you can use it with any other subscription or model.