I haven't joined your chats in a while but glad to see you put this together, I truly feel as though opus 5 is not much of an improvement. The only time i ever felt a wow factor was opus 4, 4.6 and fable pre trump admin lobotimizing
Opus 5 is an overconfident stupid model. It tries to generate too much slop, tries to act like everything will fall. I have reversed back to fable and codex sol.
Nice! I actually ran across this paper+benchmark recently, too. It's the first I've found that start to aim at some of the non-functional and longitudinal requirements that I think have always been an important part of writing production code.
It's especially relevant now that models are good enough to solve ~most point-in-time problems.
Some relevant but disconnected thoughts:
- deterministic scores are so nice
- what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is
- another signal I've been thinking about and I'm seeing increasingly get brought up is the state space of a system; I'm seeing formal methods pop up a lot recently
There’s this notion of variety popularized by cybernetics folks a long time ago. Variety is like the state space of the system. Then there’s a law that says “only variety absorbs variety”.
So if a method has high variety then it must have an equally complex implementation to handle the variety.
When there is a mismatch it means that either the method has parameters that aren’t useful, or that the body of the method isn’t covering cases it should.
Thanks for sharing, I've never heard of variety before and that was an interesting read. It makes me think of a broader term I've been using to describe the run time state of the system: Entropy. I've given up on measuring that for now and I'm settling for a compile time proxy through measuring the _semantic_ cardinality of types as this seems a bit more approachable for a POC.
One drawback I ran into was that the variety/entropy of scalars tends to dominate everything else. I'm clamping them to 1 for now but that's also not great as it punishes more descriptive systems.
> what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is
Maybe, but "maintainable" has a simple definition: minimising the effort to incrementally grow a software system as its size grows to infinity (while maintaining some defect ratio). Humans (called senior software engineers) figure this out over decades by working on a number of large software systems.
So far models haven't figured the same thing out. I can only guess why. All their learning comes from examining code on the internet. Somehow whatever the patterns are that make large software projects like the OSs maintainable by large groups of people working independently has escaped them. Maybe the bulk of code they look at is in the small so they miss it, maybe the patterns are just hard to discern in big systems.
My theory is there are anti-patterns which the larger projects somehow manage to keep to a minimum. It's hard to learn something that's not there. Instead, you learn these anti-patterns by doing them, and watching a system all fall apart over the course of years, and if you're good you manage to pin the blame on the right thing.
If you want to wrap it up neatly in a package, models have learnt how to program, but programming is not the same as design. Design requires very different skills to using a programming language. At a very high level, big systems are sets of interconnected modules. The things that matter are narrow APIs with minimal coupling, and where coupling is unavoidable making it explicit and easy to reason about. Concrete examples of the anti-patterns to be avoided are global variables, leaky abstractions and APIs that require complex sequences of interconnected calls.
Those things don't bite hard until you program in the large. The current crop of LLM's seem to be useless at all of them. If you've got a big context window and you are only writing 100s of kloc, a budget of kilowatts to understand the complexity, it doesn't matter. But humans don't have a budget like that, and so far everyone I've seen who reads large chunks of vibe coded software recoils in horror, gives up, and walks away.
Even with that budget, the models fail anyway once the size gets way beyond their context window - it's just that they fail much later than humans. Which I guess makes learning the lesson so much harder - perhaps near impossible as your "they need labeled examples" hints at. Oddly the solution would be to train with smaller context windows (a less powerful model in some ways), so the failures become apparent much earlier.
It looks like they have been going with your solution for now - giving engineers cheap tokens via flat-rate plans, and watching how they do things. I don't think that will work. Engineers who chew through the billions of tokens offered by these plans are vibe coding. The code they produce is so poor it can't cross the threshold from small to large. There is nothing to see there about programming in the large.
:( what a rant - it looks like an LLMs CoT as it thinks through things. Which is what I was doing, I guess.
This matches my experience of Opus 5 being a nice improvement over Opus 4.8, but not being revolutionary like Fable felt.
I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.
To what degree is this a harness/system prompt problem? Models maybe should implement new stuff with as little impact on the existing stuff as possible by default? A simple system prompt for it to always check the code after task completion for proper simplifications, abstractions and cleanups before returning to the user? Instructions to retain "story like" readability of the code.
So many benchmarks more the models themselves.. just make one unified standard to benchmark all or stop calling it “benchmarking” as this word lost its meaning.
Really hoping that all of the attention you're bringing to the longitudinal sloppification of codebases makes it back to the labs and creates some pressure to improve that trait of the models. This new benchmark seems promising.
At the same time, I imagine it will be hard for them to prioritize this over improving flashy one-shots of impressive zero-to-one feats that demo so well and attract more customers.
Can we have simonw make "pelican on a bicycle after 1000 requests for iteration" popular?
I remember reading a blog post a year or two ago about prompting like a printer driver and asking for increasingly ridiculous changes like time travel or wormhole capability. I don't know if I hallucinated that post but I haven't been able to find it again and I REALLY wish I could
Mostly a harness problem in my experience. Slop accumulates when the agent can touch anything, so constraining it to one seam and having it add alongside rather than edit in place does more than model choice.
I have a suspicion that most models will miss the `database_migration` Checkpoint 2 test that includes a `default_value` because it could be interpreted as either a JSON-literal or a SQL-expression.
There might be other tests as well that are prone to failure for reasons other than the reasons cited in the paper.
I think a cool experiment would be to adjust the order of the features implemented (e.g. checkpoint 3 then 2 then 5 then 4) where dependencies allow it. Then one could account for some checkpoints being more difficult than others.
Great writeup. The excessive function thing has always driven me crazy; I guard against this explicitly in Claude.md.
I have found that models are generally poor at managing refactors / complexity while also implementing new features. But I’ve had some success with a semi-lights-off approach where you decompose it and prompt the model adversarially in a second pass to look for new rough edges and areas of complexity or refactors that might simplify the codebase.
So I’d be very curious to see this benchmark but with something like a periodic “refactor turn” interleaved in.
Also eager to see Fable benchmarked; anecdotally that was the only model whose code I felt I could actually trust to not review closely.
> The big headline is that Opus 5 got a 24% on the small subset of the benchmark that I ran - not much higher than Opus 4.6's 17% strict pass rate in the original paper.
a 41% improvement is not much higher? come on that's just doomer
SCB is definitely an underrated benchmark. For me the unique selling point is that it more closely mirrors software development by not stopping after a single task. The agent has to keep code clean. The only disadvantage is all the problems are greenfield and not git inited so the agents don’t make use of git diffs.
I did a full circle and essentially dropped all of my personal static workflows encoded in skills because I observed recent models picking better ad-hoc workflows for particular problems, when a static one would force a subpar one.
It seems like we all tried to contain and organize a system that simply prefers to select its own organization.
Which makes me to think that these skill packs of workflows are really made to make it easier for humans rather than agents.
Thats right and to go even further I'm judging them on a metric they didn't necessarily target. A client I work with uses skills such as these to apply their own processes on the agentic development lifecycle. But users should also understand the trade-offs. I think it's intuitive that the extra steps and processing invoked by these skills adds to the token cost - this benchmark aims to put numbers on that as well as time and accuracy.
I hope the big labs will start using this benchmark in their RL pipelines. Reducing complexity in generated code should be the number 1 priority, in my opinion. The holy grail for me is models implementing features while reducing LoC (i.e., choosing the right abstractions).
What is also nice about this benchmark is that it can be used to iterate on prompts/skills for reducing code complexity.
> [...] with Opus 5 writing five times the number of functions/callables than Opus 4.8 over the course of the same set of challenges.
Is this bad? I have McCabe complexity switched on in Ruff and find it a handy watermark for when something should be broken up into smaller, individually testable callables. Five times as many callables could make for much more readable and testable code.
Can we use those metrics in a review pass for every PR? So review agent will have to pinpoint new complexity and propose refactoring or justify increase?
This benchmark makes me worry a bit that people will just ask their model to reimplement everything from scratch once their requirements become more clear.
46 comments
[ 0.19 ms ] story [ 121 ms ] threadIt's especially relevant now that models are good enough to solve ~most point-in-time problems.
Some relevant but disconnected thoughts:
- deterministic scores are so nice
- what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is
- another signal I've been thinking about and I'm seeing increasingly get brought up is the state space of a system; I'm seeing formal methods pop up a lot recently
There’s this notion of variety popularized by cybernetics folks a long time ago. Variety is like the state space of the system. Then there’s a law that says “only variety absorbs variety”.
So if a method has high variety then it must have an equally complex implementation to handle the variety.
When there is a mismatch it means that either the method has parameters that aren’t useful, or that the body of the method isn’t covering cases it should.
https://fffej.substack.com/p/only-variety-can-absorb-variety was my attempt to write it up more fully.
One drawback I ran into was that the variety/entropy of scalars tends to dominate everything else. I'm clamping them to 1 for now but that's also not great as it punishes more descriptive systems.
Maybe, but "maintainable" has a simple definition: minimising the effort to incrementally grow a software system as its size grows to infinity (while maintaining some defect ratio). Humans (called senior software engineers) figure this out over decades by working on a number of large software systems.
So far models haven't figured the same thing out. I can only guess why. All their learning comes from examining code on the internet. Somehow whatever the patterns are that make large software projects like the OSs maintainable by large groups of people working independently has escaped them. Maybe the bulk of code they look at is in the small so they miss it, maybe the patterns are just hard to discern in big systems.
My theory is there are anti-patterns which the larger projects somehow manage to keep to a minimum. It's hard to learn something that's not there. Instead, you learn these anti-patterns by doing them, and watching a system all fall apart over the course of years, and if you're good you manage to pin the blame on the right thing.
If you want to wrap it up neatly in a package, models have learnt how to program, but programming is not the same as design. Design requires very different skills to using a programming language. At a very high level, big systems are sets of interconnected modules. The things that matter are narrow APIs with minimal coupling, and where coupling is unavoidable making it explicit and easy to reason about. Concrete examples of the anti-patterns to be avoided are global variables, leaky abstractions and APIs that require complex sequences of interconnected calls.
Those things don't bite hard until you program in the large. The current crop of LLM's seem to be useless at all of them. If you've got a big context window and you are only writing 100s of kloc, a budget of kilowatts to understand the complexity, it doesn't matter. But humans don't have a budget like that, and so far everyone I've seen who reads large chunks of vibe coded software recoils in horror, gives up, and walks away.
Even with that budget, the models fail anyway once the size gets way beyond their context window - it's just that they fail much later than humans. Which I guess makes learning the lesson so much harder - perhaps near impossible as your "they need labeled examples" hints at. Oddly the solution would be to train with smaller context windows (a less powerful model in some ways), so the failures become apparent much earlier.
It looks like they have been going with your solution for now - giving engineers cheap tokens via flat-rate plans, and watching how they do things. I don't think that will work. Engineers who chew through the billions of tokens offered by these plans are vibe coding. The code they produce is so poor it can't cross the threshold from small to large. There is nothing to see there about programming in the large.
:( what a rant - it looks like an LLMs CoT as it thinks through things. Which is what I was doing, I guess.
I’ve now replaced my use of Opus 4.8 xhigh with Opus 5 medium, and I’m using less tokens and it’s quicker. I can understand people being annoyed by its writing style but for getting work done that really doesn’t bother me. I’ve been really enjoying using it.
At the same time, I imagine it will be hard for them to prioritize this over improving flashy one-shots of impressive zero-to-one feats that demo so well and attract more customers.
Can we have simonw make "pelican on a bicycle after 1000 requests for iteration" popular?
I have a suspicion that most models will miss the `database_migration` Checkpoint 2 test that includes a `default_value` because it could be interpreted as either a JSON-literal or a SQL-expression.
There might be other tests as well that are prone to failure for reasons other than the reasons cited in the paper.
I think a cool experiment would be to adjust the order of the features implemented (e.g. checkpoint 3 then 2 then 5 then 4) where dependencies allow it. Then one could account for some checkpoints being more difficult than others.
I have found that models are generally poor at managing refactors / complexity while also implementing new features. But I’ve had some success with a semi-lights-off approach where you decompose it and prompt the model adversarially in a second pass to look for new rough edges and areas of complexity or refactors that might simplify the codebase.
So I’d be very curious to see this benchmark but with something like a periodic “refactor turn” interleaved in.
Also eager to see Fable benchmarked; anecdotally that was the only model whose code I felt I could actually trust to not review closely.
a 41% improvement is not much higher? come on that's just doomer
Edit: FWIW the paper the post quoted has repositories as slop baseline https://arxiv.org/html/2603.24755v1#S4.SS2
I’ve used SCB as part of my assessment of agent skills (superpowers, GSD etc) https://orcabot.com/labs/do-skills-improve-coding-agent-accu...
There is a small but growing community on discord for discussing SCB so if interested please join https://discord.gg/BrC4BA9sVj
It seems like we all tried to contain and organize a system that simply prefers to select its own organization.
Which makes me to think that these skill packs of workflows are really made to make it easier for humans rather than agents.
Many people though are going to read the headline figures and think it means - say - Opus 5 is only a quarter the strength of a human coder.
What is also nice about this benchmark is that it can be used to iterate on prompts/skills for reducing code complexity.
Is this bad? I have McCabe complexity switched on in Ruff and find it a handy watermark for when something should be broken up into smaller, individually testable callables. Five times as many callables could make for much more readable and testable code.