I've had good luck getting it to debug (and patch) a tricky WebRTC issue that had all the other models stumped. Sorry it didn't work on your problem, I guess?
Terrible title. Should be "Fable's guard rails are way too sensitive", which I don't think you can really blame Anthropic for. They likely had to whack them way up so it would block whatever trivial stuff got demoed to the government.
I would expect them to dial down the sensitivity in a few months when nobody is looking.
This post can essentially be distilled down to: yes, Fable's classifier (which is meant to downgrade cybersecurity, biology, or jailbreak attempts to Opus 4.8) is definitely overly sensitive to the point of uselessness.
e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, despite only being very marginally, tangentially, somewhat related to biology.
Do we think that someone at Anthropic, OpenAI, the government... has access to SOTA models without censorship? "How do I build an effective weapon?", "How do I effectively control the masses?"...
It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerous when you think about who will be gaining access to it first. I do believe the goal is to pull away from the rest of humanity in a near trans-humanistic state. Are we ready for that and how do we counter it?
I have only really used Fable as a final pass on something. A "Take a look at everything we did so far, and make sure we didn't forget something" kind of review prompt.
But it is a huge waste of money for most coding tasks. Opus is still overkill most of the time, too.
I asked it a question about indoor carbon dioxide levels (wholly innocuous question), which it flagged as involving biology, therefore downgraded to Opus.
It's a pretty good strategy if they're hoping to fail as a business, I guess.
This honestly just reads as “this model failed exactly where the company said it would but I’m very special and deserve special treatment rather than the same overactive guardrails I and everyone else were told we would get.”
The honest way to say this is that Fable is not useful for bio-related work. The author is working on processing RNA sequences and similar biology tasks, and Fable's classifier has a hair trigger on those tasks.
Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work.
I guess it’s useful for making inane one shot games and websites, lol.
I'm curious what the state of alignment research is. My gut says this is basically impossible. People have different moral frameworks. Each individual probably has an inconsistent moral framework. Even granting perfect consistency, applying these typically requires some knowledge of reality. And these LLM / harness combos are turing complete.
So you don't know what it should do, you may not even know what you would do, you don't necessarily know what's happening, and can't predict what will happen. How do you align that?
Seems like these overly sensitive filters are responding to this difficulty.
For anyone using these models for anything remotely sensitive, keep in mind that Anthropic says [0]:
> We retain inputs and outputs for up to 2 years and trust and safety classification scores for up to 7 years if your chat is flagged by our automated trust and safety systems as violating our Usage Policy.
And, since those automated systems apparently have a ludicrous false-positive rate, you should assume that your inputs and outputs are being retained for 2 years even if you are doing nothing that any reasonable person would consider to be problematic.
Oh, and they'll train on that data [1]:
> We will use your chats and coding sessions (including to improve our models) if:
>You choose to allow us to use your chats and coding sessions to improve Claude, learn more here
> Your conversations are flagged for safety review (in which case we may use or analyze them to improve our ability to detect and enforce our Usage Policy, including training models for use by our Safeguards team, consistent with Anthropic’s safety mission)
It appears that the usual controls (including for businesses) to prevent Anthropic from training on your data will not apply.
53 comments
[ 3.0 ms ] story [ 59.6 ms ] threadI would expect them to dial down the sensitivity in a few months when nobody is looking.
Typo second paragraph, 4th line. I think you meant "what"
e.g. a colleague asked Fable to help create an simple app to help calculate the statistics for phase II and III trials. (Ignoring that such things already exist) it passed his request down to Opus, despite only being very marginally, tangentially, somewhat related to biology.
So yeah, if I can't ask about nicotine withdrawal, then I think almost anything biology related is going to get downgraded...
It's very concerning that we get the nerfed models but you know that somewhere, people with a lot of resources have access to the raw, uncensored, probably more powerful models. The sprint toward AGI looks even more dangerous when you think about who will be gaining access to it first. I do believe the goal is to pull away from the rest of humanity in a near trans-humanistic state. Are we ready for that and how do we counter it?
But it is a huge waste of money for most coding tasks. Opus is still overkill most of the time, too.
It's a pretty good strategy if they're hoping to fail as a business, I guess.
For the future of AI, we need to look elsewhere.
From my experience, the model itself is very useful when it isn't refusing any of your prompts.
I'm a bioinformatician
So you don't know what it should do, you may not even know what you would do, you don't necessarily know what's happening, and can't predict what will happen. How do you align that?
Seems like these overly sensitive filters are responding to this difficulty.
However when it's happy to do the task, its relatively fantastic.
> We retain inputs and outputs for up to 2 years and trust and safety classification scores for up to 7 years if your chat is flagged by our automated trust and safety systems as violating our Usage Policy.
And, since those automated systems apparently have a ludicrous false-positive rate, you should assume that your inputs and outputs are being retained for 2 years even if you are doing nothing that any reasonable person would consider to be problematic.
Oh, and they'll train on that data [1]:
> We will use your chats and coding sessions (including to improve our models) if:
>You choose to allow us to use your chats and coding sessions to improve Claude, learn more here
> Your conversations are flagged for safety review (in which case we may use or analyze them to improve our ability to detect and enforce our Usage Policy, including training models for use by our Safeguards team, consistent with Anthropic’s safety mission)
It appears that the usual controls (including for businesses) to prevent Anthropic from training on your data will not apply.
[0] https://privacy.claude.com/en/articles/7996866-how-long-do-y...
[1] https://privacy.claude.com/en/articles/10023580-is-my-data-u...