866 comments

[ 0.27 ms ] story [ 57.2 ms ] thread
This time it came with a usage reset
Has anyone been able to get anything substantial done with Fable in the first place? I more or less had totally given up on using it since the alignment checks were so sensitive that it pretty much always threw me back to Opus.
(comment deleted)
Nope. Always failed within 2-3 prompts. The most basic REST service you can imagine. Cookies are signed, that's crypto, banned. Completely useless model.
In my experience, the guards are less strict than they were at first. When Fable came out, it dropped back to Opus 4.8 for about 50% of my prompts. Now it's maybe 20%.
I successfully used it to re-slop Claude slop from 1+ year ago. Really puts things into perspective.
Yes, I used it to do a big push of my self-maintaining project All I need next is to push it through enough cycles to trust it (I have high confidence based on results so far) and then I can just setup a cron to do routine maintenance on my project
Bit of a discount if you're using caching:

> same input and output prices, with cache reads at a quarter of the cost

This should impact any long-running agent since subsequent calls can benefit from cached reads for previous transcripts.

And yet, despite this, the quota limits went down by 17%.
In my opinion, this is a bit disingenuous.

They were _temporarily_ increased in May by 50% [1]. They continued to extend them through July and August (admittedly, their messaging around this has just been a complete mess and they frequently pushed the deadline back as it approached).

So, now they are giving you a 25% quota increase compared to where things originally stood in May.

So, let me ask you this: assuming you knew that the 50% quota increase was temporary all along, would you then have complained about Anthropic restoring things back to the original limit?

[1] https://www.anthropic.com/news/higher-limits-spacex

On the contrary, you and Anthropic are being disingenuous by pretending that a usage reduction is actually an increase. Especially when the 20x max plan isn't actually anywhere near 20x, as people have recently realized.
Yes, some people will complain about anything (and everything) related to AI. And relentlessly push the most negative interpretation of any datum.
Yeah but haiku 5 when?
In my mind Sonnet 5 is haiku 5.
Haiku would have to be a banger, with a significant price drop, to make any sense.

It's currently priced 33% above Gemini 3.7 Flash, and several multiples of 5.6 Luna.

I think the signal from Anthropic is pretty clear between Haiku not getting an update in a year and the Sonnet issues this year. They don't care about low intelligence models. You should go elsewhere.

That's what we've done, migrated workflows away from Haiku and Sonnet. I actually think this is not a crazy position because these lower models have so much competition from Grok, OpenAI, DeepSeek, and about 20 other labs with really solid models in the Haiku to Sonnet range. So what is the point of Anthropic competing in these spaces where everything is going towards zero cost?

Exactly, that’s the low margin part of the market. They don’t care about it.
(comment deleted)
Isn't this basically the phenomenon known as low-end disruption/upmarket migration? Which is often viewed as an unhealthy sign for a company that recognizes it can't keep up with efficiency of a new market entrant but unwilling or unable to make changes to their business model to stay competitive.

The trap being that it's a rational decision at the beginning to focus on the most profitable lines of business with highest margins. But the disruptor then captures that value to finance innovations to move up into higher margin tiers, forcing another retreat.

The cycle can repeat until the once dominant firm is relegated to a tiny niche with no growth prospects or until fixed costs exceed dwindling revenues thus eventually resulting in insolvency.

It's not always a bad strategy but I think a pre-IPO company that's only a fear years old would generally prefer to grow the size of their potential market vs. shrink it preemptively.

https://online.hbs.edu/blog/post/low-end-disruption

What's the point in paying them for Haiku-class models? You can run those on your own graphics card.
Is there a polymarket for Haiku 5 vs Gemini 3.5 pro?
Tbh with that price , not even willing to try . What are the benefits for a regular coding agent ? I barely have any errors already with 4.8 level , eg grok 4.6 , gpt 5.6 sol/terra behind router . Why do I need to pay so much money for this ? Any reason ?
Do you only ever tackle problems you've never dealt with before or something?

When I discuss something new with an agent I want to feel like it genuinely gets what I mean, which has only started feeling true with fable 5 for me.

The biggest change is the price cut of course.
"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%."

Glad to see this!

The big issue they face right now is that vastly cheaper open models are proving capable for more and more uses at cents on the dollar.

This is the right direction, but they aren't going to get there fast enough.

They will list, investors who don't know anything about tech will buy, the world will realise that China just put out a model that is good enough at a fraction of the price, they will crater.

US enterprise customers aren't going to convert to overseas models, average consumers might though.
At 1/10th the price they will. Claude is way overpriced for most peoples needs.
This is just cache reads. In real usage it costs 15% more than Fable 5 -- all for marginal gains.

https://artificialanalysis.ai/

Cache reads dominate in modern workflows (coding CLIs and modern web clients such as ChatGPT Work and Claude Cowork (web)).
Output tokens are 5x more expensive than input tokens, so I'm not sure "dominate" is entirely correct.
do you think we'll go a full year without a new haiku lol
"Content provenance" seems to be activated with this model.
"with cache reads at a quarter of the cost"

OK, I think that's what they meant when they suggested reduced extra promo usage will not sting this much.

I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big differences? I did notice Opus maybe making more mistakes repeatedly but I don't have hard numbers on this. I hope Fable 5.1 brings noticeable improvements. I am giving it a go now on my 20x Max plan on a problem that Opus 5 has struggled for more than week now and has made very slow progress with regular regressions on the way.
My impression is that Opus 5 can be very impressive if you don't care about maintenance, novel-length comments, and really having any input in general. But otherwise it's borderline-to-totally unusable. It seems tailor-made to not have a human in the loop.
If you're doing something cutting edge like math or formally verifying algorithms, Opus 5 is a steaming pile of shit compared to Fable 5 and Sol 4.6, it makes countless stupid mistakes and is essentially incapable of completing the task without extreme hand-holding.
I don't know how I feel when all the documentations are written by AI for humans.

AI to AI doc share: sure, do what you please.

AI to human: please make it legible and flowly.

example, "Every thinking block records which model produced it, and it's preserved in one direction only: Claude Fable 5.1 reads earlier models' thinking blocks, and no earlier model reads Claude Fable 5.1's." is a very Claude-isk way of writing. Choppy, long, and lacking flow.

Unless these people start offering free, unlimited inference for a cautionary period so we can test the new model without an up-front (re-)investment, I am not touching this load-bearing pile of neuralese spew with a ten thousand token pole.-
Are you going to post the same comment on every AI article on HN? How is this useful to the discussion?
This, is the first time I have not only made this comment, but opined on this issue at all, actually.-
(The funny thing is that the comment itself was highly resonating, until the PR - or Claude! - came in, after a lag).-
You tend to take a rather aggressive approach when commenting in here.
“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.”

I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better.

I don’t think Anthropic realizes that humans have a token limit too and it can be exhausting to read Claude’s output. Prose density is not the same thing as succinctness.

Amen. I would trade some stupidity (say ten points on any benchmark) in exchange for a version of Opus or a similar model that actually gave me direct, concise answers.
You should try setting claude code to opus 4.6. With the style instructions I set in my user CLAUDE.md it does exactly that. It's like night and day: Opus 5 gave me a page and a half of word-vomit, yet the exact same task and prompt with 4.6 and I got maybe 100-150 words total, entirely readable.
x2 on opus 4.6. still works great, and it's fast
I've developed a habit of adding into my prompts "please keep your response concise and succinct" or "I'm trying to cram, please only provide the minimum level of technical detail necessary to understand this topic"

I find it helps immensely but it'd be nice if I didn't have to do that.

why so many people add 'please' when asking machine to do something? Was there actually research that when you SCREAM or curse it follows your instructions better?
Probably because polite people are already in the habit of saying please when typing out requests in chat. We're not consciously thinking about it, regardless of whether a human or machine is on the other side.
I consciously write it to keep the habit. I've seen some people sometimes talk to others in the same way they talk to their dogs or horses.
I think about removing please/thanks, but then I accidentally add them back in during some edit/rewrite of the prompt... It's just how I'm used to asking for things
I'm polite to LLMs. It's not for the models it's for myself. If I start being rude to models then I might accidentally start being rude to other people as well.
add to your system prompt?
It only sticks to the instruction for maybe 3-4 turns. This is why when Anthropic released "concise output style" feature in claude code, it basically spams the model's context with "be concise" system reminders every other turn.
ya, but now i have this system that ive build that works with any ai agents, putting boundaries, gates every single time and it generate memories from the runs so it can inject them as needed.
just try kaplira, it will help u with that.
i tried using claude codes output style option to do something like this and it worked for like three prompts and then it was back to normal lol
Same, currently on a mix of Kimi Vivace (K3), GLM Max (5.3 and 5.3 Flash) and OpenAI Max (Sol and Terra mostly).

I will say that Kimi feels nice but slow, GLM feels faster but has limited tokens (even off-peak) and OpenAI is nice and fast but has limited context (258k shows up in Codex, really).

Neither of them are perfect, but I prefer their type of prose across the board to what Opus 5 and Fable 5 kept outputting. I'll probably check out Anthropic again in a year, but for now I need a break from its brand of slop. Oh also all of the other ones allow usage in OpenCode with their subscription plans.

One thing I've noticed and HATE, is that when you increase thinking-effort, that seemingly increases response-length. Meaning that X.High is longer than High, which is longer than Medium, etc.

Which is kind of the inverse of how people work; a really smart person can condense difficult ideas into simple[r] terms. Whereas people who struggle speak a lot but say very little.

High/X.High do seem to deliver better quality results, but it sometimes feels like needle-in-haystack extracting that from the word vomit.

I hate this too, I had to switch to Codex, because the skill to force Claude Code not to think too much about very, very basic things no longer worked
It's so bad I've made myself a Pi extension that rewrites responses in side by side view using models on Cerebras (insanely fast tps)
It's so bad I made my own chat client for Claude, so I can attach steering prompts in conversation. They are applied just at the end, before the last LLM response, then removed and response kept.
The higher the effort the more things Claude checks, and it's eager to tell you about all of them

See, this insight it had early on looked like a red hering for a while, but then turned out to be load-bearing. And that's not just a difference in semantics, it changed the whole conclusion (spoiler: it didn't). And Claude is very eager to tell you about this exciting journey

I just go over the comments with Gemini 3.1 Pro at the end which has a much more normal "voice" and it doesn't lose nuance as a cheap model would. I don't care so much about what Claude writes during the debugging as I just do all the cleanup at the end instead of at every commit.
“I have made this letter longer only because I didn’t have the time to make it shorter.” - Blaise Pascal
> when you increase thinking-effort, that seemingly increases response-length

I use the /sss writing style - synthetic, short and simple - and it helps a lot.

Just remember Charles Dickens was paid by the word too
Today, Opus talked about "rotation slabs" in relation to logging. (and not log rotation). I didn't even bother asking what that was supposed to mean and switched over to Sonnet.
> In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks

that feels like they just blocked words like load-bearing but can't actually fix the real problem. The insane word slop density and run on sentences was the real reason it became annoying to work with claude, colored with way too many analogies and pointless linguistic comparisons.

"Humans have a token limit too" - that's so good and it explains so much of the fatigue that myself and colleagues/peers have about Claude in particular.
I think it's not just token limits - I think it's because it's so _dense_.

You get a week of research and debugging and testing compressed into a few pages. Even if it's explained well, it's just so much information. And since it's AI, I'm constantly second guessing "is that really true?" and it's exhausting.

I just can't stand how often Claude says something like "And the honest part? It's..."

Like, were the other parts not honest? I don't understand how Anthropic let it get like this, it's been such a clear regression

i think they took a huge bet that speaking like a ted talk was going to be a vast popular differentiator in their offering, i don't think they anticipated that people were going to make fun of it, that it could become a meme..that it could get in the way of getting stuff done and result in cancellations.

it's downright exhausting to read claude, the language style was a regression imo.

If you ask it why it uses the term honest so much it'll tell you it was actually trained not to. lol
It can't truthfully answer "why" questions, only infer them in a way that aligns with its training for conversational engagement.
Of course, that's how LLMs work. It's still amusing, at least to me.
Geminis is "it really is". The Notebook podcasters use it _constantly_.
Sometimes Opus 5 (high/xhigh) feels like I'm dealing with the programmer equivalent of Zeno of Elea.

Every time, without fail, it would get me 90% of the way there and then leave a small note, exception, or deferral. When instructed to address that, Opus would somehow take nearly the same amount of time as the first 90%. And then it would finish with yet another deferral. Repeat ad infinitum.

You can sometimes get around it using the `goal` directive provided you are not subject to the constraints of mortality.

They got that from Anime seasons. Every prompt has yet another cliffhanger to keep you hooked. But the Season II story arc where Claude-chan fights the NsPasteboard boss battle on the journey to the UIViewMainController, I thought that was pretty intense. I guess I just gotta keep watching my terminal to see what happens to the main character input - rooting for him to survive the next season, but you know they always kill off the good input characters early.
Me too. And it does it so often, that I've added a stop hook that detects "honest*" in its response and forces it to regenerate without the banned word.
Same here. I still have access until my account churns but Anthropic has huge issues comparative to everyone else with token / usage burn down. K3 Swarm also delivers better results than Fable at a fraction of utilization. The Pro plan is definitely not worth it anymore and if I do want to burn some money I can always just leverage the API. But Anthropic went from simply amazing last year to a dumpster fire in less than 6 months for my use cases, anyway.
If you ask any model to write as tables to enumerate points, and BDD for logical flows, it’s like 50x less strain on you
> sentences run longer and there are fewer paragraph breaks.

Gotta fit in the watermarking.

You can change CC's output style (https://code.claude.com/docs/en/output-styles). You can also put style notes in your global claude.md. I've instructed claude to treat me like I have adhd, get to the point, and be succinct, ... More or less eliminates the problematic prose.

I took time to figure this out after Fable spat out "...then stays purely as cascade-debugging provenance rather than load-bearing arbitration."

That sentence is fine. As a long-time HN reader, HN is full of this kind of performative erudition and I’m already used to it. Fable probably learned from the worst parts of HN.
My biggest frustration with Anthropic with Opus being too verbose is that they tried to put this on users. It’s pretty clear that Anthropic employees don’t use the day-to-day models that everybody else use. They have access to the next tier model so they don’t see the problems that everybody else is dealing with.
Yep, they have no clue what their users are complaining about since all they use all day is mythos max preview.
I don't think they care about humans... ... ...
> In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.”

This sentence reads like Claude wrote it. Perhaps it did, or perhaps Claude has learned to write like the folks who work at Anthropic?

(Had I edited this, I would have said that a colon is not the right separator here. The second clause does not _explain_ the first, per se, bur instead expands upon it. Consider instead: "In some cases, however, its prose is denser than Claude Fable 5's, with longer sentences and fewer paragraph breaks.")

I just used it to do a review of a smallish (10k SLOC) codebase that Fable 5 / Opus 5 largely built, cost like $2 and caught some good stuff, but more importantly, it communicated very directly and was pretty light on bizarre metaphors. No "let me read the source before opining" type verbiage launched at me. Honestly night and day for me vs before.
I wasted a lot of tokens last month asking "Please explain the meaning of this sentence in plain language"
Kimi K3 is the best explainer. I switch to that for code review and insights.
Forgive me this long letter, I hadn't the time to make it short. —Pascal
Just the other way I was thinking that if I asked "What does Lamborghini do?" the only correct way to answer is a single sentence "Which Lamborghini are you referring to?".

But LLMs will fail at this question: they will tell you about Lamborghini's latest car and mix some history in it. Just try.

Which is the wrong answer anyway, because there's at least two major companies called Lamborghini, one making cars, one making agricultural equipment and at least one famous person (Elettra) with that family name.

This very simple test/question makes me realize how much do I hate LLMs in a sense: while I agree that the answer it gives is the most plausible for 90% of the users, it's ultimately both wrong and long. And that 90% compounds.

But there's no "correct" answer in my eyes than "who are you referring to?". Possibly without listing all the possible Lamborghinis.

This is ... unnecessarily pedantic. Anybody in my social universe who asked me that question would undoubtedly expect "they make cars".

If you're picking nits, why not focus on the word "do" and (wrongly) expect an answer like "Lamborghini (either of the two main companies of that name) does not 'do' anything - the companies employ humans who 'do' things. Lamborghini is a legal entity established to allow humans to 'do' things, such as make cars, or agricultural equipment."

Shared context is a thing. Reducing every conversation to first principles is not always required. Get a grip.

This is not a nit. This is a real problem in the technology.

It assumes the average and plausible answer token by token.

And this tendency shows in every single field it's applied to.

At the end of the day I want *correct* answers, to the point.

Instead LLMs, no matter if it's version 3500, are bound to producing average results: slop.

I switched to using Codex for the last two weeks, and while the prose has been better, there have been a lot more technical oversights. I'm now having fable review codex commits and it finds deep issues. I'v also done the reverse where opus/fable do the work and then I have codex revise all of the prose prior to reading anything myself. This has also been effective; I'm not sure which is the better approach.
Yes! I have my .md's have

"If you respond with more than 3 paragraphs, give me a TLDR"

"Do not assume I know all technical jargon, please explain things plainly"

> Prose density is not the same thing as succinctness

Can't agree with you more. I review 2-3 PRs a day from my team of eight data engineers. Most of my team members use Claude to write SQL, dbt and Python code. Some of them use Claude a lot, some less so. I can easily tell when I review the code that is mostly Claude generated vs. the one that is not. In dbt models where we have a lot of biz logic in intermediate layers, that's where I really have a difficult time following Claude-generated comments. So much jargon copied over from other adjacent dbt models (yet inconsistently), and the prose is super choppy (for the lack of better word).

After reading a looooong sentence/comment line, I still can't figure out what it really means. Had to always re-read the line 2-3 times (sometimes, more) to sort of understand. Reading code, however, is so much easier and usually, I just skip to reading the code and then come back to the comments. :D

It still talks the same claudish, but now it's indeed denser. I'm not quite sure what step up they're talking about.
You know you can just... change this right? What a weird thing to cancel over. I have in my global CLAUDE.md something to the effect of I have ADHD and give me succinct responses with headings and lists etc, works great.
Did we get thought traces back? If no, it's useless.
Lol we got literally the opposite:

> *Fewer progress updates during long tool runs.*

> The model writes less user-facing text between tool calls, especially at higher effort. Set thinking.display to "updates" (beta) to receive the progress updates it does write, and remove any prompt line that tells it to hold findings for the final response.

I love when I make a request or ask a question, and Claude Code immediately queues up a dozen tool calls to edit files instead of explaining what it wants to do or why, despite repeated constant reminding that I will not approve tool calls without context and reasoning
Ive been through this a lot myself, and thats why im building this tool called kaplira, it is an control layer for ai coding agents. it does not allow ai agents to touch files outside a determined scope, it learns as u code, learns from past mistakes, regressions, it injects cirurgical memories into the ai agents. u can download it for free at kaplira.com
At least half the changes are just anti-distillation strategies...
I really don't think they can stop it, only make it somewhat more expensive. As long as the model need to make tool calls on the user's computer, the user can record the trajectory and use it to reinforce another model to follow the same trajectory.
"Democratization" through AI means that everyone has to pay a monthly Anthropic tax and only a small secret guild gets access to the real model.

Jane Street is a partner? How sad indeed. Anthropic could front run them because they leak all the data.

Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from other models verbatim (since it can see the decrypted version). I get that in their eyes it's an "exploit" but still kinda disappointing that they patched this
To be fair, I assume they want to hide that not from their customers, but adversaries who use the way Claude models think and reason to refine their own models.
I have a hard time believing whatever prompts get Claude to reason can stay relevant secret sauce for long anyways. It’s not hard to A/B test something that gets you close enough, and it’s not Ike anthropic has uncovered the global optima of reasoning prompts.
I don't really want the models I use learning from Claude at this point. Open weight models of similar scale are available now too, so I expect this "distillation"/"stealing" chatter to wind down.
Maybe, but that's sort of begging the question that those open weight models aren't significantly trained using "distillation"[0]

[0] not technically distillation. https://thomasdullien.github.io/posts/2026-06-15-rl-economic...

distillation is a minor piece of training data, you have to have a good foundation for it to be helpful, and even if you have good traces, you need a good RL reward scheme at the point it is used (very challenging)
I think this whole distillation argument is between fully overblown and bogus.

In any case, highly misunderstood.

That's their motivation, for sure. But it's also unambiguously making their product worse and harder to use legitimately. Which pushes customers further towards use of open weight models which don't have these restrictions. I don't think this is a fight they're going to win.

It's also hard to have sympathy for them - they want to protect their IP, sure. But their IP was built on a corpus of dubious legal provenance. And even if the courts decide their training data are legal, most of the authors of the data would disagree. There was no consent given.

I think LLM's are great - don't get me wrong. I'm glad they were built the way they were, because it's unlocking an amazing new world. But I just don't have sympathy for the "I stole this and now it's mine so you can't steal it" argument behind concealing reasoning traces.

These draconian "Preserved Thinking" measures they're taking are going to be an absolute pain in the ass. This alone is enough for me to move our API use off their platform entirely. It's a HUGE breaking change that they're trying to dampen by having it not affecting current customers until "in the future", see: https://platform.claude.com/docs/en/build-with-claude/preser...

You're no longer allowed to edit the context anywhere! The whole context is to become append-only, says Anthropic. No more editing the system prompt as the conversation progresses, no more dynamic loading of custom tool calling formats. Everything has to go through their built-in tools API and you aren't allowed to mess with anything in the context if it has any thinking blocks following it. This is the most intrusive "model DRM" we've seen so far!

Interesting, I liked to experiment with a second model "simplifying" and summarizing the previous messages and continue.

Needless to say, it improved output on following messages by whatever metric I cared for.

Not sure why would they prevent it.

I give you a chain of messages, what do you care for what the origin is?

> No more editing the system prompt as the conversation progresses, no more dynamic loading of custom tool calling formats.

Hm, aiui you can support both of these via mid-conversation system turns https://platform.claude.com/docs/en/build-with-claude/mid-co... - and in general you'd want to to preserve the cache and recency of the instruction anyways rather than frankensteining an off-distribution transcript. Not sure though.

I really don't see how that is an option, as if appending to the system prompt was ever enough to override previous instructions. Their example isn't very confidence inspiring either:

"The user switched the workspace to read-only mode. Do not write files until told otherwise."

Great! Now we just have to trust that the model never misinterprets any of the system prompt, which has always been so reliable before. Instead of your meticulously crafted prompt, it will now be some junk like this:

"The workspace is in write mode. The user switched the workspace to read-only mode. Do not write files until told otherwise. The workspace is now in write mode again. Wait, back to read-only!"

And who knows how this integrates with their context summarisation that we will be FORCED to use. How does it summarize multiple user + assistant/thinking blocks without messing up the system "appends"? If all it did was append the mid-convo system messages right under the original system prompt then they'd be ripe for all the same distillation "vulnerabilities" as before. I guess we'll never know!

You seem intent on being mad, so maybe this isn't helpful, but context summaries are optional.
(comment deleted)
You think moving will help you? OpenAI is going to do exactly the same thing soon. Unless you're moving off the frontier entirely, that is.
How do they still purport to champion alignment and explainablity of AI if the reasoning traces are going to be hidden.
Making clear the scale of distillation they’re combating.
It's only theft when people pay Anthropic for inference in other to improve their own datasets. It's not theft when Anthropic grabbed basically all ebooks and web content on the internet to build their own dataset, without paying anything to anyone
Aren't the "thinking" chains always just reconstructed anyways? It would be like using a debugger that just looks at the source code rather than the actual binary.
So basically, Anthropic can charge you for tokens you don't even see. "Trust me bro, you really did use $5,000 worth of tokens to generate that pelican". We need AI consumer rights, urgently.
I think that it would be sufficient to realize that no one is forced to use a specific AI model or model harness that has anti-consumer features built in.

People should just walk away if they see something like that. This is where the actual, practical, consumer rights start. Not with the regulation. But with the customers being determined to stand up for themselves. And not just fold.

There are plenty of good enough models that are open weight and whose use comes with almost no strings attached. And plenty of great harnesses such as various flavours of Pi.

Agreed, and that's exactly why I don't use Claude anymore and haven't for months now. But I'm not most people. Most people don't vote with their wallet, unfortunately.
Anthropic:

AI models should be explainable so that we can ensure and verify alignment. Responsible AI 101.

Also Anthropic:

No not like that.

Don't care unless it is priced in as other models.
Sadly still not available for Pro subscription. At least they reset everyone's limits.
I didn't see it anywhere on their announcements, but when I restarted Claude (on a Claude Max account) I see the model is now Fable 5.1
That feels bad, my weekly limit was going to reset today. (I wonder if mostly everyone's reset day is today as well...)
Mine reset yesterday but I won't complain since I profited from the last two resets that were on Friday :D
So bullshit safeguards are still there.
> Denser prose in places.

Really? Interesting choice. Pretty much every CLAUDE.md file I have starts with something about Hemingway, terseness and treating every word you use like you're carving it on your own back, but different strokes for different folks. I suppose I haven't heard from anyone who enjoys how wordy Claude is because they aren't done writing their post yet.

I yearn for a model that can churn through claude text and write sensible text. so far gemini is pretty good at that, even in the low variant
"Explain in simple terms" works.