184 comments

[ 0.17 ms ] story [ 57.7 ms ] thread
I seem to recall Anthropic going on record saying that they don't do anything to model performance to stretch their compute capacity. I've anecdotally noticed massive peaks and troughs in performance week to week (albeit with Opus, not Fable).

I wonder what their official explanation for this behavior is.

Last time they were called out, it was a regression in Claude code itself.

At least that's their explanation. Either way, it wasn't a good look for "vibecoding" but it got brushed over.

When something is new, its capabilities feel incredible. Over time, those same capabilities become mundane, and you start to notice the flaws.

(Now, if TFA is actually measuring reasoning tokens, that's quite different! It's not entirely obvious to me how he is measuring.)

I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.

What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?

And, maybe it's one thing to defraud normies like us (although class action lawsuits exist), but I don't think major enterprises or the US military would take too kindly to it.

Are you telling me that companies might defraud people for millions and billions of dollars and pay fines that are 1000% less than their profits?" My goodness, you must live on a hell planet.

Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.

And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.

The question isn't whether the penalty would be less than their profit, it's whether the penalty would be less than whatever they make by secretly downgrading the models (or whatever it is you suspect), which remember also causes consumers to get less value out of the product and more likely to cancel.

The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.

is it? it's still the same model, they can claim the quantization down to q4 still retains 98% of the performance therefore it's fine.

nothing on the fine print tells you what the weights are, you're just getting Fable 5, whatever that is

It's called hedonic adaptation.

> What is actually stopping these model companies

You can say this about any company in the world, selling anything.

It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't.

But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs.

>What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?

As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.

Not true. I can read what Fable output with ease but when it sprout Claudish like Opus 5, I know they are doing something to the model. Yes, you can immediate know the claudish language if you work with opus long enough
They are deploying optimizations weekly (if not daily) with various AB tests. They don't manipulate model performance, but they do actively perform tests.
You're right to push back, and one honest caveat -- they could just be lying.
Your caveat isn’t just a side note, it’s worse than that, they have incentives that go against your best interests!
Their exact phrasing IIRC was that they "never intentionally degrade" their models.

This still leaves an absurd amount of wiggle room for arguments like "oh no, our evals show that this quantization has no detectable effect on performance (in the eval distribution) therefore running the quant doesn't degrade quality"

How do you measure thinking tokens? They don't send those back to the client.
They tell you how many tokens are used, however, right? Otherwise you couldn't see your own token consumption.
Obviously. The standard pattern is that model X is basically AGI and wins all benchmarks, followed the next day by Y and Z, which both win all benchmarks, too.

Then weeks later people find out that they have been duped and complain that the models have been quantized or employ worse inference.

Buy decent coffee instead of your $200 subscription and sidestep all the scams.

Well sorry, still have to get decent AI somewhere. Productivity without AI is about 5x less. I am not comfortable with paying Chinese companies, and no Western companies provide subscription-based pricing for open models.
Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.
Reminds me of how slot machine users swear the odds have changed on a machine.

also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order

The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.

Wtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.
Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.

But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets.

At that point I might as well do it myself
Well yeah. If for some task you find it easier to just do it yourself then you should. But you can improve the situation even if you can't entirely automate it by automating parts of the verification thereby making it easier to human-do larger verifications when X>1. But in many cases even that is not possible.

The progress however is such that the number of tasks that you can do with >p% automated and X=1 keeps increasing. So many times just waiting works. Of course, here also it changes from field to field. There are some tasks at which AI hasn't even gotten started, others where it has already peaked, others where it's increasing slowly, and others where it's increasing fast.

I am skeptical of this as well but slot machines are programmable and the house can change the odds.
How do you create repeatable tests in a non-deterministic system? Every time you send the same prompt you get a different answer.
I strongly believe that the real Fable is the one we had for a few days in June. Then they nerfed the model a bit after the government pulled it off the market. What we have now is something less, but still good
I also believe this. Fable post-ban was never the same. At the least, whatever system prompt munging or pre/post filtering they did to strengthen the guardrails nerfed it.
question is if they ever let the general public access borderline AGI
It is clear by now to me that Anthropic is constantly trying to find a kind of “auto” degradation perhaps to save money on work it thinks does not require high reasoning. I always use max reasoning and I can clearly see differences between the models when they release and after 3-4 weeks. I think they give a kind of intelligence boost also for new accounts.
Just yesterday I was thinking about gpt-5.6-luna. I made it my default model in Hermes during its fist week of launch. It was just as good as 5.5 which was my previous default. But over the last 2 or 3 weeks I've seen how dumb it is now. I have to be very explicit with it.

For example, I used to be able to prompt "Check the system logs on <server> for...." and it would just figure it out. Yesterday I asked "Did <service> on <server> complete the overnight job" and all it said was "that service is not installed on my host"

I had to tell it to ssh into the server and run journlctl to check it

Anecdotal, I know, but they all seem to be less capable with time.

(comment deleted)
Same exact experience. I worked with both Fable and Sol foe the last two months, daily for several hours, and got used to the very bright, quick thinking, proactive even.

As of last 2 weeks or so both models are nearly on par with DeepSeek4.1 now, which I also use a lot. They're still better, but that difference is not as pronounced as before and, importantly, the frustration level is now on par.

Whatever they're doing will surely drive people to less advanced but predictable, self hosted open models. I sure would rather use DS4.1 with Qwen/GLM in adversarial mode than deal with this b/s I pay significant amount of money.

Me and my friends have been contemplating on getting an Ultra M5 256 and splitting the cost. PI harness is so good now that this is really a viable alternative.

I'm fairly sure it's just luck of the draw if you get put onto a quant'd model or not. I've seen luna xhigh change intelligence fairly drastically on a day to day basis.
Really this is the base problem. You have zero idea where and how your prompt is being executed.

If for example AWS sells you a 2xLarge server there may be some variability in performance but it's going to be averaged out very well.

When it comes to AI services executing your model there is absolutely no information on what and with what settings your model is being executed. Hell, you have no idea if it even is the model you're paying for. Add that models are not deterministic so variability can be pretty large.

This leads to a common set of dynamics that induce cheating behavior in humans. For example, is there a mix of different hardware. Does lessor hardware use different settings? How do you know xhigh is what your prompt ran under. Anthropic has a proven history of running your prompt silently under different models.

This is a huge mess that needs and will be regulated or sued heavily over. Hell, with as many people out there that hate AI it might be easier than one thinks to have a state sue the providers on this and elicit a huge amount of discovery.

& the nice thing about Hermes (since its open source) is you can be reasonably sure that behavior change is coming from the model and not the harness. (probably)
Same experience with Sonnet on low effort. It used be when I used a "table_name/id" format to reference a db record, it knew exactly how to find it using connected mcp tools. Today it failed 4/4 times (I tried the exact same prompt in 4 separate sessions and each time it replied "I don't have access to [...]"). On medium effort it got it right the first time.
FWIW, logged-out ChatGPT claims to be Luna on high. (Unless it's variable for some reason.)
The smart takeaway is not skepticism or snark, but understanding that once the new datacenter buildout starts coming online, cheap and widespread access to even the current frontier models (without strict thinking limits) will blow the economy wide open.

(ie, even a pause in AI training isn't going to stop the train where AI flips the economy upside down, we've barely even seen the impact of the current frontier)

Anthropic is straight up scamming its users at this point.
The question I have is this only happening for a subset of users working in specific areas, such as AI or distributed systems (https://news.ycombinator.com/item?id=48742153), or is this across the board? I am working on distributed systems. Today Fable is mostly unusable. It resembles Opus, so I went looking to see if anyone else is having issues. Sure enough.
I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Just an hour ago I had Fable correctly identify an unused method that could be deleted. I then immediately get a diff for an exact duplicate method, and then Fable outputting, "I accidentally duplicated <method> instead of deleting it. Removing both copies now."

The remaining morning complaints that makes it feel like something's off is that it will do a lot of "thinking" for simple things that previously took very little time. And it got very lost and completely mixed up DE-91M predicate names and implementations. Just absolute disaster code that I had over the past months come to generally expect it to do without issue.

Glad I carefully review everything. I think what I need is reliability and consistency. But it feels like picking a model from the list doesn't guarantee that: that the models' "brain" is open on the table and they're screwing with it.

Maybe they are jealous of Navier Stokes and try the Hodge conjecture with 80% of total compute at the expense of their customers.
I would rather wait in a queue than be routed to a degraded model. And if they _have_ to degrade the models, then I wish they would fucking tell us. Instead, it's "I have a strong feeling".

That we have to guess at this is by far the worst part of the AI era. It feels like a dark cloud over my productivity. It makes my body tense for the entire day when it happens. Not healthy.

They have repeatedly said they do not ever intentionally reduce model quality and do not degrade in this way, and that a model version number is always the same.

But, of course, OP is an empirical claim to the contrary, and I'd be curious to see if anyone (who's been capturing data over these timeframes) can replicate the same results and if Anthropic has any comment.

Every official statement I've seen around this is careful to say that they "don't intentionally reduce model quality", which leaves plenty of room for "we adjusted some knobs and our evals show performance is materially the same".

However, I also agree that I haven't seen any robust data from someone tracking it daily/weekly. The handful of sites purporting to do this aren't even running it enough times to hit stat sig.

I kind of feel "reduce the amount of thinking tokens produced" would fall under degrading model quality.

In any case, I am willing to believe it's possible something degraded, but so far I have not seen any empirical evidence of it since the previous incident with the inference and harness bugs. I lean towards Anthropic probably not intentionally doing anything like this without disclosing it beforehand.

The issue here is you have to think like a lawyer trying to weasel out of making an empirical statement.

For example "We didn't change any settings, but when GPU use gets high the run time of a prompt is lessened. But you must remember this is always in effect so nothing changed at all. This happens occasionally on random prompts some of the time, and when it's busy it happens all of the time".

In someones eye this would fit the letter of the law but not the spirit of the law that you hold.

See March 26, 2026 incident. Model not degraded, but harness changed to strip out past thinking tokens when a session went out of cache to save money and ease capacity constraints (affected API users too), resulting in bad degradation.
Sounds like me without coffee.
The Claude models definitely felt more susceptible to moods, like you could leave them for a few hours, come back and it suddenly was unable to do things which it was doing just earlier, which tellingly is never an experience I've had with an open model.

Honestly I lost patience with Anthropic both clearly messing around with things like this and their agitation over regulation. They aren't good actors, and quite why so many blindly trust them with their company crown jewels is a mystery.

If you follow reddit forums for claude code, its common to see people, on the same day, claiming that Opus/Fable is especially smart today, and especially dumb today.

I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

If you have a bank of rigorously tested benchmarks that you run every few days, with enough trials to know what your standard deviation is, and you are getting significant trends over time with those, that would be interesting.

But "I have a feeling" and "Seems like" really isn't a reliable signal at all, humans just can't handle perceiving these things reliably. On top of that changes in your work environment can easily pollute LLMs and change quality of results. Are things getting added to your memory or claude.md files that you don't realize? Is your project growing in size and thus claude is performing worse as more context is needed to work with it? etc etc

> and human perception is absolutely horrible at evaluating trends like this

The need to have a measure of competence for your fellow man is, most likely, a pre-human skill, probably with a dedicated bit of neurons for it. I think the problem is that those instincts were co-evolved with our fellow man, and, as you say, don't apply at all to a more non-deterministic system that, fundamentally, lacks some logic faculties that even small children have (simple riddle modifications, car wash question, etc).

> I think people are still not used to non deterministic tools like this, and human perception is absolutely horrible at evaluating trends like this no matter how smart, clever, and experienced you are.

The implication is that humans are unreliable and shouldn't be trusted.

Or humans have certain shorthands when they complain on reddit, but their diagnoses are accurate for the specific context? If my AI does something stupid, am I not allowed to call it out? A NS-solving AI is still capable of not satisfying the abstract thing called the user experience. People have intelligent thoughts without compiling to lean.

OK, you say. Then let's get an aggregate benchmark for "intelligence". That doesn't prove that AI didn't flounder a specific use case that the user requested.

Classic moves: Humans are unreliable, converge to some "objective" benchmark that necessarily will quotient out the special cases, etc. Wonder how we'll be solving these issues in the AGI era - well, if you have an AGI that just replicates itself, dominates everybody because it's a machine and humans are soft fleshy creatures, and agrees with itself, fine. But part of the beauty of human experience is the messy part, and providing value is in the messy part.

In other industries of chance we have regulators that ensure compliance and that the providers aren't cheating.

At the end of the day the highest quality of benchmark tells you nothing if the man behind the curtain is constantly changing variables on you. You have no idea if you're really testing the same thing at all. So when you run your test at the top level on their system you're seeing lets say a 30% difference in quality most of the time, you have no idea if you should really only see a 5% difference in quality if you were running a local model with stable settings.

It’s load shedding. They’re reducing consumption for capacity balancing at your expense. Whenever there are rate limiting storms Claude gets dumber. They also shift capacity for new releases, and Claude gets dumber leading up to it.

Self run infrastructure won’t have this cost but you have to manage the capacity and rollouts yourself, at which point it’s more obvious what’s happening, but the effects will be the same. The not knowing makes it harder, but also harder to plan your own work around.

Correct, but they should explicitly announce this ahead of time.
A general rule of corporate behavior unless they are forced to under duress.

If this is duress of competition or at gunpoint of regulators is up for the population to decide.

Luckily for us open weight models exist. Until regulatory capture anyway.
Unfortunately there's just no demand for that. People have suspected they do this for years, but they get shouted down for even suggesting it.

I find it odd, because it's not just plausible, it's understandable that they would need to dynamically throttle these things or degrade people to a different model. I would be fine with that if there was transparency. I'd think a business building any type of logic on AI decision making would like to know if they are actually using Opus and not Fable. But apparently the idea that these companies don't have infinite resources is conspiratorial fear mongering to most AI boosters.

> I have no hard data but I have a strong feeling this morning that something's wrong with Fable 5 compared to Friday evening.

Fable is effectively worse than Opus 4.6 now. They severely messed with the model.

Gemini Chat is constantly throwing, "Pro is in high demand right now, a different model was used for this generation," too.

I'm thinking they're all running out of physical resources. It's the DotCom bubble all over again; rollout of the physical infrastructure that's necessary to keep all of the pie-in-the-sky promises will not happen on the timescales that investors can work with, and they will panic when they realize this.

They want transparency from everyone else but not for them ... you don't say.
This is like shared clouds back in the day where if someone is using the CPU more it impacts you, just pool every one to the same service. There should be an SLA but for the intelligence of these models, otherwise, you are sold fable but with the intelligence of a table.
This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.

Tests their intelligence, not their diligence.

Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.

I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.
So in 5 years will they lose a suit for intentionally deceiving users? Or is something baked into the ToS by now that allows them to adjust things like this?
I do not know a single senior developer who likes Claude anymore. I do not use their API (Sonnet, Haiku, Opus) anymore and am sending my money to offshore companies such as z.ai (GLM) and QWEN.

The American companies have become extremely deceptive and scammy. I hate Anthropic and OpenAI and can't wait to have a decent GPU at home to use at least Opus or a fable-like open-weight model. This is the current dream of every developer. But NVIDIA is not going to let that happen anytime soon, so maybe China can come up with a GPU that destroys NVIDIA. I pray.

I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool ergonomics changes, etc.; I'd be shocked if they didn't A/B every release. Less thinking as measured by tokens isn't necessarily bad if you can get the same results by making it think about the "right" things or structure. They obviously screw up sometimes, and I've always been suspicious with hidden tokens, but I haven't found evidence quality intentionally degrades over time.
Same. With some 500 hours of usage in just my project at home, across both the $200 Claude and Codex subscriptions, I have not once encountered a situation where I would have attributed unsatisfactory results to a degradation in the model.

I've seen bugs in the harnesses, sure, but never anything in the actual model where I could have said with any certainty that it's not just regular variation or me having a bad day myself.

No idea where people get the confidence from to make such claims every other week.

These analyses are much better than these Twitter charts.

I don't think anyone is reading the details for the Twitter post because it was not an actual benchmark. They did a post-hoc analysis of their logs from day to day.

Their random collection of prompts for each day is not a benchmark.

The site you linked is a much better example of a real benchmark being repeated over time.

The Office of Weights and Measures exists because, long before any of us were born, in 1836, companies were up to shady shit and consumers were paying for inconsistent products. I.E. Being scammed.

AI companies should be subject to the OWM like any other company that sells a product that varies in weight. Perhaps when a sane administration is re-elected; one that can read history books and comprehend why our regulations exist in the first place. Or have even a semblance of respect for its citizenry.

THIS EXACTLY.

The only regulation that we need right now is the model that's on tap

Anthropic terms of service:

> 12. General terms

> Changes to the Services. Our Services are novel and will change. We may sometimes add or remove features, increase or decrease capacity limits, offer new Services, or stop offering certain Services.

> Unless we specifically agree otherwise in a separate agreement with you, we reserve the right to modify, suspend, or discontinue the Services or your access to the Services, in whole or in part, at any time without notice to you. Although we will strive to provide you with reasonable advance notice if we stop offering a Service, there may be urgent situations—such as preventing abuse, responding to legal requirements, or addressing security and operability issues—where providing advance notice is not feasible. We will not be liable for any change to or any suspension or discontinuation of the Services or your access to them.

You're not buying a gallon of milk or a pound of flour. You're buying hosted software that the host reserves the right to modify.

Companies can say whatever they want it doesn't mean we'll agree it's okay or legal.
I doubt that would change the perception. Every model release is followed by accusations of nerfing.

There are several projects that repeat benchmarks on published models. None has ever found significant fluctations

Here's one example https://marginlab.ai/trackers/claude-code/

Fluctuations of a few percentage points are to be expected and should not surprise anyone who knows how LLMs work.

This Twitter analysis of Fable 5 is not that at all. They analyzed their coding sessions and blamed all of the fluctuations on Fable changing. They then compared to ARC-AGI-2 questions as the benchmark for thinking tokens and tried to stir up anger that coding turns don't produce as many thinking tokens as the ARC-AGI-2 problems.

It would change my perception but only if there were a competent and stringent administration in place. I didn't used to have to wonder if the ground beef I was buying was actually 1lb because there were inspections and repercussions, but stuff is kind of chronically underweight these days.

A properly run OWM enables you to stop wondering if you're being ripped off and that's what AI needs because I think it's incredibly easy to just assume we're being ripped off because these companies are all built on a foundation of wonton theft. (Not that I really care about that — I think all information should be free, but still.)

I thought METR was supposed to fill this role?

... But what exactly is the "weight" metric you have in mind?

we could even just repurpose the same office, "weights and measures" is oddly relevant
Another benefit of open weight models although providers could still run another model in the background.
I stopped using Fable long time ago. It's worse than Sonnet. Opus is not much better.

This cycle of new model running at full quantisation and then nerfed few days / weeks after premiere should be called out. Anthropic should also drop the adaptive reasoning scam.

If I pay for Fable, I should get full, not nerfed model at honest pricing.

Regulators should investigate them.

OpenAI is no different. Astra has basically the same problem.

The ROI just isn't there. It feels like Fable is in the same place Opus was early last year; at best marginal improvement that's barely noticeable over the lower model, for 10x the cost.
It's not really 10x the cost though, with the low cost of cached read it's more like maybe 1.2x the cost.
Makes sense, no? Test time compute is something you can vary, so it makes sense that you start covertly reducing it once the model has already made it's splash.
Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.

I have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.
I’m not sure why this point keeps coming up — if your service/product is so popular that it’s capacity-constrained, then the answer is to raise prices, not degrade service, because the demand should be inelastic.
This really depends where the load shedding point is.

A very small raise in prices may cause a very large loss in customers that you risk never getting back.

For example if customers figure out that the Chinese models are just as good, they are gone because they are so much cheaper.

Exactly, it’s not a good business to be in if they’re capacity-constrained and can’t raise prices.
Raising prices also has second order effects, like consumer and business expectations around how widespread the tech can be. Valuations depend on it being reasonably affordable to roll out on a much more massive scale than today. If people get the impression that it seems too limited to very rich people (200 is affordable for a North American / Western European professional), the impression about the trajectory will change.
The Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code.

Fable seems to be following the same enshittification arc of other Anthropic models.

Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.

> Making your product worse over time rarely plays well with customers.

On the other hand, New Coke was a resounding success. Well, it, itself wasn't, but in the aftermath, Coke outsold Pepsi 2:1.

True. New coke is a good example. So was unity licensing. But these even feel slow motion compared to the Anthropic rise and self immolation.
I’m so behind on this topic but I find it interesting how quickly things change. I feel like just yesterday I way hearing how anthropic is far and away better than OAI, and now this.

I have no way to judge myself. I don’t even use them. But it’s interesting to follow by just reading stories and comments

This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason
There must be some benefit if all the providers are doing it independently.

GPT5.6-Sol on Max thinking just became regarded as of a few days ago.

The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).

The cycle repeats.

(comment deleted)
(comment deleted)
Again, I’m out of my element here, but isn’t the entire industry dependent on “new better releases frequently”? If so, and if no one has made any meaningful breakthrough, might they all pursue this kind of deception just to stay afloat/“competitive”/relevant?

Thanks for your insight

Kinda. Off the top of my head, DeepSeek and their thinking model was pretty new and interesting. Multi input models are also newish (combined input of text, image, video, audio, etc). Then there's Jev, a recently release that has a lot of people talking. It isn't really an LLM, but also is one.

Sam Altman believes he can train a model entirely on synthetic data, which he admits would not have human world knowledge but is interesting none the less, which likely led to their mathematical models.

Overall models have become cheaper to run and smarter per token.

> might they all pursue this kind of deception

They might but multiple competitors engaging in ongoing deception as an intentional corporate strategy isn't required to explain what we're seeing. It's entirely possible to get the same clearly unethical outcome without any employees knowingly participating in an explicitly unethical plan of record.

Instead it happens without overt coordination when individuals and groups within an org each pursue their local metrics and incentives. In isolation, no individual action seems obviously unethical on its own. They just look like 'optimizing performance', 'maintaining ASP or ARPU targets' or 'achieving operating margin', etc. Customers are still getting deceived and receiving less for their money than they think. The difference is most of the people involved in enabling it get to not feel bad about themselves.

See Shepard tone. Similarly model releases could be engineered to appear that they’re always getting better by slowly degrading and upgrading at the right time. That plus hitting some benchmarks and making a lot of noise around that.
Astra is also useless and completely ignoring instructions at random intervals.

We are being A/B tested on and there is nothing you can do about it.

We are being A/B tested on and there is nothing you can do about it.

Oh, yes there is. DeepSeek 4.1 Flash on max thinking can simply be dropped into Claude Code. Close your eyes as the chain-of-thought traffic scrolls by and you can easily fool yourself into thinking you're still running Opus, in terms of both cognition and throughput.

To be fair, matching Opus's throughput costs about as much as a new car, but cars suck nowadays and you didn't want a new one anyway, right...? Failing that, rent a cloud server, one that you control.

They obviously test various quants and other serving cost saving strategies. Models like Fable are probably trillions parameters with hundreds of billions active MoE. They probably try to squeeze and quant each piece until people notice.
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Not unless your competitors do the same, or else you will only be perceived as falling behind others.

Yes that makes sense. In my hypothetical, the industry frontier is stagnating, meaning no one is making big breakthroughs, so they all resort to this.

If one lab makes a breakthrough, can the other labs just distill to bear parity anyways and then set a new baseline industry wide.

I’m quite ignorant on this topic, so if any of this sounds moronic, forgive me

So are they making big breakthroughs, or are they stagnating?
...releasing a new model that’s marginally if at all better than the original...

This isn't what we see in benchmarks.

Yep, that's what they've been doing for a long while now. Also the amount of tokens you get per sub varies drastically from month to month. Needs to be regulated.
> For an industry that’s stagnant in progress

Yes, the AI technology is known primarily for how stagant it is.

Yes, I freely admitted I was entertaining a pure hypothetical I pulled out of my butt.

I have no idea, just had a thought and put it out there

You can serve Fable from a cloud vendor (like AWS, Azure). They have frozen versions of the models, so likely this should not be an issue?

I would do a test to verify my suspicions.

Sounds like a good smoke test.

I’m actually so far removed from this tech that I couldn’t run such a test myself lol

In my experience, the API versions are as good as ever; it's the subscriptions that are severely degraded.
> to create a perceived improvement when in reality there isn’t really one?

This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.

but there is a gap between benchmarks and user feel.

Opus 5 came out with better benchmark results than Fable, but it really did not feel better to use at all.

This is a good point I hadn’t considered, thank you.

Is there any training variable here? For example, can a model released in October perform better on the same benchmarks vs its predecessor released in July just by virtue of training on newer data that was made available on those 3 months?

Sorry if it’s a dumb question, I don’t really know much about the topic.

Also, in a world where there are several models competing with each other for public perception of which is best, that seems like an extremely bad move.
Overfitting to benchmarks. And puff, you have the exact same effect.
The public frontier is not the frontier. Actual frontier models are too expensive to serve to public, and also risk distillation by competitors.
Using Dieselgate trick is another way for benchmaxxing.
Nah its because they cache and preprocess requests by dumb models and send them too often to another dumb models instead of the top tier model.
The Shepard tone of "progress"
This would only provides a benefit if we're approaching some sort of theoretical limit of how good LLMs can be with the current approaches and data.

Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so all working together?

So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.

There doesn't need to be an explicit conspiracy. It only needs all the US frontier labs to be facing the same economic pressure (logarithmic improvement / $). It's a pretty obvious strategy - it's not like the large labs can magic up huge volumes of extra compute as demand comes online; there are almost certainly tweaking model performance to occupy the compute available and margin/cash burn targets.

That was one of the main points of the movie 'A Beautiful Mind' - that actors can coordinate without any explicit communication.

No, what they are doing is trying to optimize inference to increase margins which leads to degradations. Model deployment is not like websites, you can continuously tune performance based on usage, new memory optimizations, etc.
> For an industry that’s stagnant in progress

Surely you're not talking about the AI industry. Astra was released less than 3 weeks ago, and Fable-level models became public only 6 months ago. The rate of change is dizzying.

And yet they have only improved marginally in my use cases since around Opus 4.5.

The harnesses have improved somewhat, but the code produced on large or legacy code bases is still very average and I still see similar mistakes made that I saw back a year ago (although less now that harnesses have become better at steering).

For my use cases, we are definitely on the flatter part of the curve at the moment.

This is wild to me, but to each their own. Mythos-class stuff is insanely better at nearly everything than Opus 4.5 was in my experience.
Same experience here, anything frontier human knowledge wise, same if not a regression. For human understanding and emotional intelligence, for many tasks regressiin is so bad that many near anchient llama era models now beat frontier anthropic/oai models. Notable exceptions to capability rot seem to be qwen models, and previously deepseek but the latest gen of models has started showing the same rot. General writing quality is down significantly accross the board, often it is outright ass. For example, I didnt mind reading 4.5's outout, but opus 5 makes me goddamn near violent, its fucking insufferable.
It’s a hypothetical statement that seems to have confused a lot of people.

I’m not saying it is stagnant. I’m saying for a hypothetical industry that was (maybe that fits AI, maybe not, I have zero authority to say myself)…

>For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

Are you really saying AI is a stagnant industry?

You couldn’t make it through 4 whole sentences?
Uhh, they're expressing incredulity, not a lack of reading comprehension.

Unlike you, however...

There was a coding horror story I read some years ago where a developer bragged that he improved performance by artificially increasing iterations on some critical path in an app and then lowering the iterations occasionally while bragging to management about squeezing out more performance.

Kind of reminds me of that, but with more smoke and mirrors

One persons horror story is another persons career roadmap!
AI just solved a millennium problem two weeks ago. "The pace is insane. And there is no reason to be this fast." to quote Terence Tao word by word.

HN: Well, must be a stagnant industry...

I literally said I have no idea if it’s stagnant. My entire comment is a hypothetical. Perhaps you didn’t have the patience to read all 4 sentences?
> to create a perceived improvement

In addition to the dozens of opaque model parameters and hardware variables that can nerf or buff model intelligence, speed and profit, there's also the very real possibility that models aren't just training on benchmarks but could be evaluating if they are being benchmarked in real-time and applying more resources adaptively. 'Driver optimizations' that detected benchmarks in real-time were deployed in the first 'GPU Wars'.

> I have no idea is the actual frontier is stagnating.

Like a lot of complex, rapidly evolving tech, the truth is it's probably rapidly accelerating on some measures for a few and stagnating on many others for most - hence the divergence in user reports. It's depends on how you use it, for what problems, how rigorously you assess the output and whether you happen to be on a server bank, RAM pool or shard at this moment which hasn't yet been sufficiently 'cost optimized' by the margin algorithms. They don't call them load balancers anymore. They're Margin Balancers.

That definitely isn't what has been happening; Fable was much better than anything seen before it.

Could be what happens next, though.

Check gpt I think they recently started taking the same route