I've always wondered why everyone flocks to SV's latest darling company. Have we not learned from our history of glorifying these SV darlings that turn hostile?
Most of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
A mixture of opportunistic edge-seeking, FUD, FOMO, novelty-seeking, the need to impress shareholders, the tendency of salespeople to believe other salespeople are telling the truth, the ever-present need to stay in front of relentless commodification, and pragmatic curiosity.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
Some of us want to get work done and don’t feel the need to either glorify or hypothesize about what might happen.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
A model is just another thing to plugin to a harness. I don't give it much more thought than that. If developers are still caught up on Claude Code, or Codex that's just not a long term thing. It's best to develop workflows locally and in the cloud with open harnesses. I know this will be the future because that's how it worked on every other system that developers use.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
That’s not true. Before AI, I have been using Jetbrains IDEs as far as I can remember. Also have been using MacBooks for work since my first job. You don’t have to generalise everything. If a particular specialised tool is good at its job just use it instead of re-inventing the wheel
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
Just cancelled my Claude subscription and used Ox Alpha Free through OpenCode (so that's one).
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
I am because Claude has figured out I am a biochemist and therefore even asking Fable what the weather is gets me bumped back to Opus, sometimes even Opus 4.8 instead of 5. Kimi? GLM 5.3? DeepSeek? No such problem.
I can literally open a new chat with just "Hello" and it gets bumped.
Memories break Fable 5 for me as well in chatbot. I ask Opus a lot of sec related stuff and now if I even type “hello” in chat it gets insta-downgraded to Opus 5.
I wish Fable were as good as you make it sound. A plan created by Fable is good, but in my case, it always contain issues caught only when it's reviewed again (whether by itself, Opus, Sol etc.). That's (almost) not different from plans created by Sol, GLM 5.3 etc. The one thing where it's genuinely better is the front-end, but then again it's far from perfect, it just needs less iterations.
>> but in my case, it always contain issues caught only when it's reviewed again
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
I think highly complex work and “professional” work are basically completely orthogonal. You can have highly complex work you do as an amateur, where AI can be very useful. For example working through a difficult mathematical problem, building or contributing to an operating systems, or researching a highly technical topic for a hobby project such as microscopy, chip design or lithography.
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
Would I, a non developer be able to utilize or benefit from Fable? I have been using Sonnet to update and make changes to a piece of open source software that was basically abandoned. It found and fixed a handful of security issues. I am curious if I am missing out on anything. If I change from Sonnet to Fable mid-stream will I lose any history? The one challenge I have run into is turn-by-turn session history. It constantly has to go back and re-learn what it did in our sessions a week or two ago. It accidentally deleted many of its own files so I convinced it to store everything in a .claude sub-directory in my fork of the other persons code so that I can keep it backed up when the claude sandbox has issues. Claude seems to want to talk me out of changing models which I found odd.
I'm just doing this as a hobby for fun and adding features to a program I use. I am trying to make the most of the boost they gave me for August.
Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.
I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.
Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
It's kind of funny you describe it that way. I'm literally doing the exact opposite, which is why I have both.
I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.
It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.
Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.
'Loops' are the hackiest thing ever invented I don't think they're good for anything.
I think a researched plan is much better, the agent will follow the plan.
You can allow for an 'iterative' strategy for solving a problem, with guardrails so that it doesn't try to many times.
Yes - you can use Codex as the 'manager AI' but I don't see much benefit - it has a much shorter context window. Codex is better 'at the front' where it matters.
That said - it's 'auto-compaction' is quite good.
I don't believe the hype over OP5 lagging that much, it's fine as a working AI.
Other than for truly automated tasks, I don't think there's much that an AI can 'work a few days on'.
That form of 'loop for days until it works' produces nightmare code and architecture.
Fun for experiments and learning but not for code u want to keep.
In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex math prompt. Most of the time you get one or two if you reduce context and question complexity. Most of my colleagues are cancelling their subscriptions because of that, and just using Sol instead, which is virtually unlimited on the Pro tier.
Hmm. Would you mind sharing an example of a complex math prompt that you would use? Because I found that even Sonnet can solve fairly complex math problems fairly easily if you give it the right tools, so I'd like to give it a shot myself if you'd like.
> the only task that requires that level of intelligence is frontier scientific research
Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.
They've put themselves in a corner. Fable was too good and they gave it away with the $20 plan. It had to be a big step from Opus 4.8 to show progress, and Opus 4.8 is GREAT at coding in many different domains.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
Might as well not be - I routinely get rate limited in a single review session.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
Just started reviewing a feature PR and it would hit the 4 hour limit on max effort. On high as soon as I do some clarification turns it would also hit the limit.
I still havent used it much, and they crippled Opus for what feels like no reason. Some suspect that Opus knowing theres a better model degrades it some. Shame.
Yeah for me Fable works great on the $100 plan. The problem with Fable is you will hit the weekly limit. The $200 plan does not give you 4x or even 2x the weekly limit. For me i only get about 1.3x weekly quota, so I don’t think it’s worth it.
And they've hobble Fable and Opus so hard with their safety guardrails, I ask innocuous questions and tasks and they get flagged so often I gave up on it. I can get all the work I need done in GPT5.5 or 5.6 without the hassle.
> It tends to bail occasionally for me once auth-related code comes into play.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
I hit it once when I asked a question about whether butterflies remember anything from their time as caterpillars. I've never hit it for coding, but I also don't really do much related to security.
They want to be able to zoom in and zoom out as needed while analyzing aggregated user requests to better understand the different kinds of ways (well-resourced) actors use to distill their most capable models.
> "The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives."
> "Some attacks only become visible across multiple requests. Best-of-N jailbreaking, for example, sends hundreds of slight variations of a prompt in the hope that one will work. Larger patterns of misuse, such as state-sponsored espionage or data extortion campaigns, only surface when our safeguards classifiers can zoom out across many requests. Detecting these threats requires temporarily retaining prompts and outputs so they can be analyzed together, rather than one at a time."
I kind of still chill out on Opus 4.6 too. 4.8 is good too. I go between them. Opus 4.8 is a little smarter some of the time. Their use of language is both very different from Opus 5. Some of the time I have a hard time believing Opus 5 is even related to Opus 4.6 and 4.8.
I'm going to be sad when they retire 4.6. It's not my daily driver but it's still my go-to when the other models are being stupid in one way or another (either being too verbose, or lecturing me about how what I'm asking for is evil and bad).
That seems about right. 4.8 is like in between 4.6 and 5 in terms of capability and language and they are all pretty close honestly. I just default to using Opus 5 for a coding agent that I don't interact with and I like driving with Opus 4.6 or Fable. Fable thinks too much though. Fable is like that engineer on your team that will over-engineer the shit out of something if you let them. Fable was like, "Here are 34 yaks, which shall we shave first <hands rubbing together>" and I was like can't we just... write the script first and then decide of any of these poor yaks need shaving?
Also on a dark reader? It says at the bottom, but gets dimmed out pretty seriously with dark reading. It's a 7 day average business spend, relative to June 1st (2025 presumably) indexed at 100.
The other thing I forgot to mention about Opus 5 is that at least out of the gate, it seemed very intent on spinning up agents and obliterating my token budget. It was noticeably more token hungry than 4.8. It would make sense for them to intend this behavior.
*This is likely because of your thinking level. The difference between max and ultracode is primarily that the latter is max with a bunch of agents.*
I do not know why people even use Max or Ultra levels of thinking effort. Most of the time, i found its better to just run any of the frontier level models, with high or medium. Their capability beyond those levels is often diminishing returns, in exchange for double or quadrupling the costs (or usage).
And the bonus is that often at those levels, they tend to spin up less sub-agents that just eat away at tokens/usage, like its water in the desert.
A Max or Ultra is something that really only belongs if medium or high can not solve a very nasty bug or issue. Even in planning its often too much.
I'm using company Claude for work with no choice what provider I use.
I did a quickie test with Fable when they were OMG TRY IT NOW, didn't see much of a difference from Opus 4.8. Except i ran into one of their stupid security guardrails. Yes I'm doing security on the software my employer owns, silly.
Then I didn't notice the US unbanned Fable after banning it, or that they extended the "trial" Fable period where it worked on fixed price plans. Not in time to bother trying again.
After that Opus 5 showed up. Language changed, okay, I don't care much. However it seems to create busywork for itself. Spawning subagents on tasks that don't do that with 4.8. Overall taking longer.
So I'm back on pinned 4.8 and doing actual work. Main problem - for Anthropic - is it's good enough.
This whole article seems sort of based on a misunderstanding. Anthropic is _actively discouraging_ people from using Fable at this point. It's not a failure, they're steering people to Opus 5 intentionally.
Yeah I mean if I run out of tokens every couple of hours and have to pause my work or shell out more money I’ll switch to other tools that don’t have this problem. Though they turned this down a bit it seems, I can work with Fable reasonably now and I enjoy it actually. I think they were just testing out how much they can raise the cost without users leaving when having the best model. I guess not much after all!
Hasn't it always been the premise that intelligence would get cheaper? To me, on the enterprise side, it seems like firms are finally getting the memo that, whether you are locked into the Ant/OAI ecosystem or not, you don't need the smartest, most expensive model to do every single task. This is a good thing for overall adoption. Whether that trickles down into regular user behavior, especially with subscription pricing, remains to be seen; even though I intellectually know I don't need Sol for a simple refactor, I am sometimes hesitant to downgrade, as it's hard to accept using something positioned, even implicitly, as 'worse'. Remembering that the smaller models tend to be faster is what usually puts me over the edge.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
> whether you are locked into the Ant/OAI ecosystem or not
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Agreed 100% for the consumer case: an empty chatbox is just about the least sticky surface I could ever imagine. I saw a mobile interstitial ad for Kimi recently whose hook was basically "Tired of paying for expensive ChatGPT? Download the Kimi app, it's the same thing but cheaper". I myself bounce between token subscriptions like no one's business and use Pi/OMP for maximum model flexibility when coding (and it's a few env variables or lines of (TO|YA)ML|JSON to switch providers in Codex, Grok Build, CC). I even self-host and try to use OpenWebUI + CLIProxyAPI when I can for all my chats.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
Anthropic’s issue is churn because of the peak verbosity vomit coming out of Opus 5/Mythos/Fable. What the hell did they train it on. The sane model is still Opus 4.6.
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
Commented elsewhere, I still use Opus 4.6 because it is the only model that feels decent to interact with. 4.8 is decent and some times smarter but you can see it trending towards Opus 5 levels of nonsense.
As a small background, I have a local server and I've been trying out different models with different inference engines, quants, configurations etc... I'm also using Opus and Sol at work consistently. I've used AI since the first wave, first as a toy, then as a highly specific tool, last 6+ months as the primary LoC generator.
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use it still *feels* much better. I think benchmaxxing is ruining the value of these benchmarks, if they ever had any. The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. Probably the last month of my Claude subscription.
> When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace.
I've wondered if this is part of why we don't see the reasoning traces for Anthropic's models before -- Open models might just be accurately surfacing how the sausage is made.
I'm assuming it's definitely part of the equation, but considering that I'm getting more tps but still waiting a lot more time for code to come out I'd assume it's not a 1:1 comparison. Plus I'm running quants, maybe with full precision it's better.
In my opinion, the big issue with Fable is that Claude Code cannot use it properly. I know, that sounds weird, but I've had Fable run down the wrong lane (and never stop) or give up and claim that something was impossible so many times (until I pointed at a GitHub repo that solves the "impossible" issue).
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
Transformer models are quickly becoming a commodity, and I suspect in time we'll all be running them locally. Even now, you can run something pretty useful on a 16 GB graphics card, and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
Well nobody thought you could write software that helps tightly coordinate processes happening simultaneously from millions upon millions of nodes on nearly every corner of the world, but here we are. When industry expands further off earth, we will need more complex and intricate software to coordinate its movements, why wouldn’t our systems become more powerful. If you are hopeful for humanity than you must expect the scale of industrial necessity to only ever increase alongside the imagination and capacity of its people’s.
My theory is that it's because the Bluetooth protocol is named in honor of the Viking king Harald "Bluetooth" Gormsson, who conquered several disparate peoples into a single kingdom, so as an act of solidarity toward those peoples, the Bluetooth protocol is written in a way where a devices are as likely to rebel against consolidation as they are to pair together.
>and I suspect a decade into the future, entry level hardware will be running better models than high-end graphics cards can run now, as entry-level hardware gets better and models get more efficient.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
> The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription.
I had not considered that as a possibility. It is somewhat dark and unlikely, imho, but a real possibility.
I am more inclined to believe that the fab capacity will grow over time and "commodity" compute will be available to us all again.
However, I can also imagine going back to the 60s era "hyper verticalized" mainframes, in which case, the frontier labs might not just be producing models, but chips and an ecosystem around themselves and their suppliers/customers.
I had an internship at HP in the early 2000's, and my computer was a terminal that connected to an X server in a data center. Doing everything as a service is not a new idea, and suppliers have been trying to push it for decades.
The only time it really succeeds is when regulations force it. I can store files on my own computer or a home NAS just fine, but if I start a medical practice, I have pretty much no hope of being HIPAA compliant without signing up for a data hosting service. The same goes for tax preparation, banking, and several other fields.
It may very well be the case in China, Europe, or Australia that regulatory restrictions force users into models-as-a-service, but in the US, regulations censoring models, even if they frame that censorship as a safety measure, won't pass constitutional muster.
But the Chinese models exist, are competitive, and have locally useful quantized versions.
If you're actually willing to pick up old DC inference hardware on the cheap you can already do a lot at home. The new hardware not so much admittedly but if the Chinese models keep getting better and running on less hardware... I don't see a reason to be quite so pessimistic.
This won't happen so long as the bandwidth bottleneck is a thing. Local models are just several orders less efficient at scale, and this limitation is inherent to the architecture and won't go away unless both hardware and model structure change enormously. There are really only two usecases I can see for local models going forward:
- 10-1000 person groups such as corporations where you can amortize serious hardware with parallel use
This is really it for me as well. At the heights of complexity the AI can do magical things. But really a lot of the time I just want it to do mundane things right. And currently it just cannot. It writes garbage text, consistently ignores something you have told it, makes mistakes a human makes once but the AI remains uncorrectable.
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
I've spent many years doing work in compliance for airlines. Hundreds of pages of documentation to describe the different rules and regulations which must be followed in specific scenarios. We'd convert those documents into a programmatic definition (rule engine) to alert when rules are at risk of not being followed.
This work was fraught with bugs, a large portion of which came down to disagreement of what was coded v/s what was written. Even if you had airlines sign off sentence by sentence exactly what you wrote down in English, that's too open to interpretation.
People don't appreciate how the same sentence can be read five different ways by different people (or the same person on different days). We had to structure our documents to be closer to pseudo-code than to English to get any meaningful consensus on the definition.
I’m not sure if the underlying data is counting subscription use for Fable, which is where a lot of people are using it because token pricing is very expensive. I wouldn’t be surprised if this was counting enterprise token usage only. As rich as enterprise customers are, they’re not exactly willing to double the cost of SWE salaries on tokens.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing code base. It’s the closest thing I’ve seen to nearly one-shooting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
Yeah as soon as CFOs realized AI was racing to become one of the most expensive line items along with salaries and AWS bills, they started cracking down on the most expensive ones.
Very much agree on Fable. Over the past month or so it has shown to be the only Anthropic model that can understand a largish dbt model codebase. Opus 5 gets almost everything wrong. (High reasoning on both)
More importantly other way cheaper models can do what Opus 5 does. So you can pay for Claude to use Fable 5 exclusively for harder stuff and planning, then get the same value you'd otherwise get from switching back to Opus by using other $20 plans for day-to-day coding tasks.
Perhaps nerfing the cap out of your best model for press attention and hosting valuable features like thought traces isn’t such a great business model?
GLM series has made it very practical to self host. If the new update for Deepseek flash holds up, I think it would be silly for some companies to not self host.
My company still hasn’t been able to deploy wide access to Fable because it’s not available on a ZDR basis. This wasn’t mentioned in the article but I imagine this factor is not irrelevant.
It’s far from clear to me that this is directly connected to the no-ZDR requirement. I’m heavily involved in this stuff with my company and I’ve never heard that the lack of ZDR fable is a Trump admin thing.
I get where it's coming from but I think it is also very funny if someone actually believes their claims (except as a CYA strategy)
The AI companies have already shown how little they care about other peoples intellectual property and they won't care about their customers either if it stands between them and the promises they made to their investors.
Did you think that these things are just a handshake deal? ZDRs are backed by contracts negotiated against extremely capable and well-resourced counterparties. Being found to violate them would bankrupt even OpenAI and Anthropic many, many times over.
679 comments
[ 0.24 ms ] story [ 16.7 ms ] threadMost of the people pushing this are just hoping that they can cash out before the hype pops and financial gravity crashes the party. Sam Altman recently claiming that the singularity is here is so stupid on its face he should just be treated as what he is, a huckster.
None of this stuff ever made any sense on what it was being sold initially. It was always insulting that the media and business leaders tried to argue that the tech could replace entire call centers or vast swaths of entire industries.
People keep arguing, but it will or it has based on extrapolating certain, reasonable use cases. Klarna has shut up about replacing call centers with bots because Markov chains with memory only can do so much.
TBH I don't think any of that is unique to the IT industry's relationship with the Valley. Other technology-driven industries have a similar worship-ish relationship with a few rarified businesses. But the culture of the IPO exit accelerates all the most short-term motivations to do anything.
I’m happy with Claude. If they become (bigger) jerks, I’ll switch to something else. I don’t ha e the energy to praise Anthropic today and I won’t have the energy to demonize them tomorrow. The emotional investment people have for/against these companies feels like celebrating or being offended by the weather.
Nobody has complained and seems like for every use case we have Opus is more than powerful enough, especially with Opus 5
I find it funny how OpenAI got caught lacking for a very brief window, but it turned out to be a very critical turning point.
Like a guy that that's at the top of their game the entire year, and the one day they have the flu, the CEO does a surprise performance review.
Sure there are Microslop and Oracle db users but most of the world we live in is Postgres and Linux. That's why I think most companies will run llm's like that.
> good at its job just use it instead of re-inventing the wheel
Exactly why is everyone reinventing a harness every month. There will be Microslop / Oracle harnesses and there will be 1 or 2 open source ones that win.
In my view, from the testing I did since Thursday, it's better than Fable. I had just finished a rather large task that Fable completed, including a /review and an /ultrareview.
Ox Alpha found bugs that Fable and Opus missed, and it continued to build things like a pro.
It does have issues with availability - but it's on a free promo right now. That also means that I don't know how much it would have cost if I had to pay API prices for it, which might not be cheaper than the subsidized small company / consumer usage, but for large companies paying API prices for either offering, the difference will be considerable.
https://openrouter.ai/stealth/ox-alpha
I can literally open a new chat with just "Hello" and it gets bumped.
I would compare it to a extremely high end $15k PC, or an expensive pro-grade video camera, or a freight train, or a …
I would say at least 95% of the global population will not encounter a situation once in their life where it would be actually useful/warranted.
Yes but those issues will be much less severe with Fable-written plans than those written by lesser models. I know this because my workflows at both my regular job and my startup involve multi-step agent reviews via codified adversarial review skills. Fable as a reviewer will frequently find blocker-level issues with plans written by GPT 5.6 Sol, and sometimes with Opus 5. The opposite almost never happens. In fact I cannot remember the last time it happened.
It’s more reliable and makes less dumb errors than Opus.
It still messes up, of course. But for my working style, I definitely prefer it.
That said, Opus 5 is broken. Use 4.8 or another vendor for the build agent.
And you can have extremely simple work that you nevertheless have to do as a professional. For example, drafting a routine customer email, summarizing a meeting, formatting a report, filling in standard documentation, or making a trivial code change.
So I don't think the $15k workstation / professional camera analogy really holds. Those are specialized tools whose capabilities are mostly useful within a fairly narrow domain. A general-purpose AI can be useful across thousands of completely unrelated tasks, including one-off problems encountered by ordinary people.
I'm just doing this as a hobby for fun and adding features to a program I use. I am trying to make the most of the boost they gave me for August.
I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.
I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic
Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.
Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.
And - trick - give it a little cli so it can run Codex if you have them both.
Let it do codex to do the bulk of the work, get a OP5 sub-agent to audit the work of the codex worker.
Just let Op5 manage and have 'specific oversight.
You can run for 2 days on 1 context window in the manager, the advantage is that it will stick to a broad plan.
I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.
It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.
Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.
I think a researched plan is much better, the agent will follow the plan.
You can allow for an 'iterative' strategy for solving a problem, with guardrails so that it doesn't try to many times.
Yes - you can use Codex as the 'manager AI' but I don't see much benefit - it has a much shorter context window. Codex is better 'at the front' where it matters.
That said - it's 'auto-compaction' is quite good.
I don't believe the hype over OP5 lagging that much, it's fine as a working AI.
Other than for truly automated tasks, I don't think there's much that an AI can 'work a few days on'.
That form of 'loop for days until it works' produces nightmare code and architecture.
Fun for experiments and learning but not for code u want to keep.
They can definitely be powerful.
Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.
But they're getting killed on token cost. They have to get people paying more for tokens. So then they put Fable in the $200 plan and release Opus 5. I'm suspicious of Opus 5. It is mostly worse than 4.8. It _seems_ like they nerfed it to create more distance between it and Fable.
So what have most of us done? Stayed on Opus 4.8. The statistics bear this out. 4.8 still dominates.
Now they're stuck. If they take 4.8 away, everyone will riot. If they make Opus 5.x better than 4.8, they disincentivize everyone from moving to Fable and most importantly, paying more.
Really, all they can do is take the L for now and just let 4.8 be the apex of the $20 pro plan for the foreseeable future while they work like hell to make Fable THAT much better that it earns the $200 to $infinity that they really want everyone to pay.
I've honestly stopped using CC and moved to codex. Sol has it's warts but I've never once hit limits on a 100€ and I get a similar level of performance for what I'm doing.
I wouldn't mind bumping Fable to 200$ plan if it was actually better but between the insane caps, reverting to opus/sonnet randomly and having similar perf as OAI - I'm done with it.
Next step is to put 100$ into open router and try some western hosted open models with Pi when OAI starts pulling up prices.
I just kept a $20 plan going for use on my phone.
I don’t doubt people are hitting it… shrugs
2nd time was literally me being lazy and telling it to commit and open a PR.
Very strange.
Nowadays Codex handles the bulk of the implementation and Fable/Opus on the planning.
Not sure if Anthropic patched it, but early on its release the web UI Fable guardrails will trip if you mention you're a biologist.
I have the same experience. I found that if I add to the prompt: “but it’s not security related, we only focus on the auth library design”, it “bypasses” the safeguard.
Used to straight out bail to Opus on this, now it's "Honing" and "Pondering" for 10 minutes on it.
https://status.claude.com/incidents/vgz5psbjmt1h
That would be absolutely insane... I don't think it's doing that?
https://readysolutions.ai/blog/2026-06-10-claude-fable-5-sil...
> "The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives."
From: https://www.anthropic.com/news/claude-fable-5-mythos-5
> "Some attacks only become visible across multiple requests. Best-of-N jailbreaking, for example, sends hundreds of slight variations of a prompt in the hope that one will work. Larger patterns of misuse, such as state-sponsored espionage or data extortion campaigns, only surface when our safeguards classifiers can zoom out across many requests. Detecting these threats requires temporarily retaining prompts and outputs so they can be analyzed together, rather than one at a time."
From: https://support.claude.com/en/articles/15425996-data-retenti...
(yes, it was a very lazy prompt I could easily have googled, but that makes the refusal even more bewildering)
Where can I find these stats?
I do not know why people even use Max or Ultra levels of thinking effort. Most of the time, i found its better to just run any of the frontier level models, with high or medium. Their capability beyond those levels is often diminishing returns, in exchange for double or quadrupling the costs (or usage).
And the bonus is that often at those levels, they tend to spin up less sub-agents that just eat away at tokens/usage, like its water in the desert.
A Max or Ultra is something that really only belongs if medium or high can not solve a very nasty bug or issue. Even in planning its often too much.
I did a quickie test with Fable when they were OMG TRY IT NOW, didn't see much of a difference from Opus 4.8. Except i ran into one of their stupid security guardrails. Yes I'm doing security on the software my employer owns, silly.
Then I didn't notice the US unbanned Fable after banning it, or that they extended the "trial" Fable period where it worked on fixed price plans. Not in time to bother trying again.
After that Opus 5 showed up. Language changed, okay, I don't care much. However it seems to create busywork for itself. Spawning subagents on tasks that don't do that with 4.8. Overall taking longer.
So I'm back on pinned 4.8 and doing actual work. Main problem - for Anthropic - is it's good enough.
Anthropic in particular is much more compute-constrained than OpenAI and SpaceXAI and has relied on partnerships to provide inference. This reality factors into their pricing and usage limits (they started 'adjusting' the 5-hour limits during peak hours, and it certainly wasn't an upward adjustment). Accordingly, this is presumably what Anthropic wants, given they develop and release the lower-end models, suggest users use them in various nudges within their product, position the bigger/more expensive models as "For the most complex tasks" in their UIs, and so on.
I think the problem (for Ant/OAI) is that there is no sensible lockin or moat. LLMs are essentially interchangeable and stuff like a harness doesn't offer enough value on its own for someone to be locked into using one of them.
Now with the onslaught of the Chinese models that offer almost the same quality for much less money they have a very serious problem on how to proceed. Investors now might be looking through rose tinted glasses but their patience has its limits.
Enterprise is a whole different ballgame IMO with countless technical, compliance, and employee adoption considerations that add friction to switching. It's also where both Anthropic and (as of last week) OpenAI get the bulk of their revenue, and, incidentally, the venue where US Government regulations on Chinese models would have the most impact.
What if:
- Opus 4.6 was the last Opus generation that got a lot of use by Anthropic's own employees
- After that they primarily used Mythos internally
- 4.7, 4.8 and 5 were RLAIFd by Mythos "teachers"
- Hence why 4.6 is the last Opus gen who doesn't report back like a robot wanting to cover every potential hole another AI system would've spotted and criticized
- Hence why coding style in Opus 5 also gets criticized, not only behavior in CC
Worth noting that Fable (ie, Mythos) is actually nice to interact with.
[1] https://github.com/zachahn/vomit
This is the first time I've felt, and I use the word *felt* since I don't have a suite of benchmarks or any sort of material approach towards comparing models, that Opus has declined in quality compared to before. Primarily I think its powers of deduction and understanding, even on xhigh, have become much worse. Before, being vague and providing a simple prompt would be enough, it could deduce and expand the details it needed, plus ask you clarifying questions, now this is no longer the case. A concrete, personal example, for a personal project, I've asked it to setup ssl over local IP. I didn't go into too much detail in the prompt as there are many approaches it could take and I didn't care too much to choose. It did horrible. The first thing it did was say the best lightweight approach is to add a reverse proxy. I'm like ok, makes sense. Then after asking it to proceed, it went and added a bunch of config to my golang service and didn't even setup a reverse proxy even when it said that is the way to go. It even said it didn't set it up lol. Then after I told it to do so it failed building the config in a way it was asked of it (support LAN IP and tailscale IP). Etc etc...
When Fable came out it was huge, the benchmarks told the story, and the story mostly matched the experience. It felt, again, intentionally saying felt, like it was miles ahead. Now benchmarks say that there are many models that are close, but in actual use it still *feels* much better. I think benchmaxxing is ruining the value of these benchmarks, if they ever had any. The price for Fable is definitely too much for any personal use now that it's no longer included in the subscription, and GLM 5.2, Deepseek Flash and Qwen 3.8 served locally or via cloud provide a lot, requiring a bit more babysitting though. Considering the price of Fable, my 5k USD Epyc server would pay itself off in less than a year if I used Fable or Opus in the same manner so at least for me the decision seems easy. Probably the last month of my Claude subscription.
I've wondered if this is part of why we don't see the reasoning traces for Anthropic's models before -- Open models might just be accurately surfacing how the sausage is made.
I switched from Opus 4.7/4.8 to test Kimi K3 a few weeks back and the test hasn't finished; it's my daily driver now.
Given their general behavior and preference toward social engineering to scare the shit out of normal people...this seems fitting.
But a while ago, I had access to a Fable harness that just never gives up. And that verifies itself. It burned $100 in API tokens in 15 minutes ... but it succeeded for all the prompts where Fable + Claude had failed.
And I believe that's a real issue for Anthropic. Fable+Claude is not too expensive thanks to the subscription, but Claude severely nerfs Fable. To save money, I guess. Fable API + Custom Harness is a different class, it's so much better. But API tokens are so expensive, you're cheaper off hiring a freelancer.
It doesn't mean hosted frontier models wont exist, they'll just be rare. It's no different than any other commodity market, for example most cars are cheap commodity models, with rare individuals buying expensive luxury cars and businesses buying expensive trucks and specialized equipment.
My theory is that it's because the Bluetooth protocol is named in honor of the Viking king Harald "Bluetooth" Gormsson, who conquered several disparate peoples into a single kingdom, so as an act of solidarity toward those peoples, the Bluetooth protocol is written in a way where a devices are as likely to rebel against consolidation as they are to pair together.
I doubt it. The play seems to be: lock what was once commodity compute up into datacenters depriving us regular folk of it, then sell it back to us on subscription. Even if my #NeverSubscribe movement succeeds, all that misdirected hardware [into datacenters] won't likely be practical for home use.
I had not considered that as a possibility. It is somewhat dark and unlikely, imho, but a real possibility.
I am more inclined to believe that the fab capacity will grow over time and "commodity" compute will be available to us all again.
However, I can also imagine going back to the 60s era "hyper verticalized" mainframes, in which case, the frontier labs might not just be producing models, but chips and an ecosystem around themselves and their suppliers/customers.
The only time it really succeeds is when regulations force it. I can store files on my own computer or a home NAS just fine, but if I start a medical practice, I have pretty much no hope of being HIPAA compliant without signing up for a data hosting service. The same goes for tax preparation, banking, and several other fields.
It may very well be the case in China, Europe, or Australia that regulatory restrictions force users into models-as-a-service, but in the US, regulations censoring models, even if they frame that censorship as a safety measure, won't pass constitutional muster.
If you're actually willing to pick up old DC inference hardware on the cheap you can already do a lot at home. The new hardware not so much admittedly but if the Chinese models keep getting better and running on less hardware... I don't see a reason to be quite so pessimistic.
- 10-1000 person groups such as corporations where you can amortize serious hardware with parallel use
- Porn.
I don't want Shakespeare, I want Bob the builder.
Half of my work is telling claude how to behave. I'm pretty certain they have enough _data_ to realize people do the same thing time and time again.
Check this comment of mine for a better explanation of this: https://news.ycombinator.com/item?id=49413353
But people are not as alike as you think. I doubt I share your unique preferences.
That said, I don't spend much time telling it how to behave. Are you sure you're not fighting the default system prompt?
Dijkstra in the Foolishness of Natural Language Programming
[...] the "naturalness" with which we use our native tongues boils down to the ease with which we can use them for making statements the nonsense of which is not obvious. It may be illuminating to try to imagine what would have happened if, right from the start our native tongue would have been the only vehicle for the input into and the output from our information processing equipment. My considered guess is that history would, in a sense, have repeated itself, and that computer science would consist mainly of the indeed black art how to bootstrap from there to a sufficiently well-defined formal system.
Imagine if only we had languages at our fingertips whose explicit purpose was to precisely and unambiguously tell a machine what to do!
https://www.cs.utexas.edu/~EWD/transcriptions/EWD06xx/EWD667...
But that would require people to think! And the marketing is they don't need to do that...
Last consumer (ish!) product that required people to learn something to use it was PalmOS with Grafitti, wasn't it?
This work was fraught with bugs, a large portion of which came down to disagreement of what was coded v/s what was written. Even if you had airlines sign off sentence by sentence exactly what you wrote down in English, that's too open to interpretation.
People don't appreciate how the same sentence can be read five different ways by different people (or the same person on different days). We had to structure our documents to be closer to pseudo-code than to English to get any meaningful consensus on the definition.
Either way, a model used to solve the top 10% of problems that people use AI to solve for, being used 10% of the time … seems like it’s in a decent place.
I find Fable indispensable, and measurably better than alternative models, for complex feature development in an existing code base. It’s the closest thing I’ve seen to nearly one-shooting features. Even still, I only use it for the hardest features and Opus 5 does a good enough job on the rest.
This is in the article.
The AI companies have already shown how little they care about other peoples intellectual property and they won't care about their customers either if it stands between them and the promises they made to their investors.
Zero Data Retention, for the uninitiated readers in this thread.
The no-ZDR is clearly to permit surveillance. I would be shocked if NSA wasn't all up in these SOTA model providers' systems.
Never subscribe!
It's to train better models. The three letter agencies don't need to spell it out in a ToS, they just access it if they want.