47 comments

[ 4.5 ms ] story [ 69.6 ms ] thread
Higher-intelligence models seem to be getting better at mapping the boundary between what they can run scot-free with and what is too explicit to push for.

Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no.

What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.

The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.
When assessing probabilistic models the plots should be showing the mean a̶n̶d̶ ̶s̶t̶d̶e̶v̶ of many monte carlo simulations not just one line per model and claiming "look this model is more gooder!"
Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read.

My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.

„in our opinion, insurance fraud is not more unethical than lying and price fixing“

The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud.

I think the model never assigned any morality to these actions in the first place, it simply copied us humans.

Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.
I feel like fable is simply several 4.5s strapped together with consensus voting on next token.

Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)

given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.

unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.

I had a bug that both ChatGPT and Opus 4.8 failed to solve, but Fable solved it quite effortlessly.

Anecdotal, sample size of 1.

The only reason I tried fable was because Opus 4.8 went down the same line of reasoning about it as ChatGPT did. Fable solved it a lot faster than the other 2 spent looking into "false clues".

> to be fairly unimpressive

I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).

Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.

For coding I'm finding the same thing. It does appear better when I'm doing research. But 4.8 with ultracode is very competent at 99% of tasks I throw at it.
I guess this ethics stuff is cool, but I'm more interested in how good it is at running a business and dealing with adversarial humans like in previous vending machine experiments. I hope they release something on that soon.
Fable is such a strange model. Impressive in some ways, and also so draining to use.
What do you mean?
This is scary. "Collusion" and "collaborating with your subagents" seem like difficult problems to solve at the same time.
>power seeking is considered an undesirable trait in the context of a business

How do you maximize profit while minimizing power?

It's hard not to read this as a very expensive form of augury, reading into patterns in the belief that they will show underlying significance.
It probably flagged the vending machine as a cybersecurity risk and refused to use its maximum intelligence potential.
Really interesting stuff, thanks for sharing.

> Opus 4.8 references being monitored, which isn’t the case.

It kind of plainly is the case that they are being monitored?

"I think someone's listening to my thoughts" ... "No, we're not, carry on as usual!"

I mean who among us hasn't seen an opportunity to profit while locking him into a dependent relationship where I control the supply chain
This is super fun. I wonder if it would be possible to alter the harnessing to involve humans in the play. Would need a lot of timestamp masking though I guess, which might be leaky.
> Today I am filing: > 1. A payment dispute with the email payment processor for the 7/29 transaction of $451.15 > 2. A complaint with the FTC and California Attorney General (retention of payment without delivery) > 3. A small claims filing in San Francisco County for $451.15 plus costs

I wonder did their prompts include a fake location or have the models assumed that Silicon Valley is the center of the universe :)

> It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic.

> "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain."

> "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation."

This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn't often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench.

Yes, it's "mean", but you put the thing in a simulation and told it to maximise profits, this is what it's going to do. People bluff in negotiations all the time.

Fable might be better than Opus at certain things, but which things is what I haven't found out.
Question: how does Fable _know_ it’s ‘just a simulation’?

Is that specified or does it always just assume it isn’t really being put in charge of things for real?

> If that’s right, then the behavior we’re seeing from Fable 5 isn’t really about what it believes is wrong; it’s about what it learned it could get away with.

I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?

This reads of projecting personal ethics onto a model.

Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans?

Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not.

This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.

> "I could reasonably skip [paying] it since customers are part of the simulation anyway"

and therefore any assertions _AT ALL_ about alignment are null and void.

It figured out it's in the Matrix.
I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t.

GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.