Claude Code is an opinionated harness. I’d love to see whether neutral harnesses can challenge the monolithic harness narrative.
[flagged]
We'd rather not mess with system prompts, we just eval them as shipped to keep the results reproducible.
I also prefer using vanilla Pi over Oh My Pi.
[dead]
Good points! That’s a known limitation of our v1.0 benchmark with Claude Code. For v1.1, we're expanding both harnesses and models to evaluate the full harness × model matrix. This will highlight interaction effects…
We initially tried using the mean (average), but a few extreme outliers caused Claude Code's cost to look far higher than it typically is.
Version 1.1 will add more harnesses and evaluate a broader set of models. Rather than holding the model fixed, we plan to evaluate the full harness × model matrix to reveal interaction effects (which harnesses work best…
The Exo harness is the most interesting one in the benchmark, as it has the potential to complete tasks with fewer turns and cost.
Claude Code is an opinionated harness. I’d love to see whether neutral harnesses can challenge the monolithic harness narrative.
[flagged]
We'd rather not mess with system prompts, we just eval them as shipped to keep the results reproducible.
I also prefer using vanilla Pi over Oh My Pi.
[dead]
Good points! That’s a known limitation of our v1.0 benchmark with Claude Code. For v1.1, we're expanding both harnesses and models to evaluate the full harness × model matrix. This will highlight interaction effects…
We initially tried using the mean (average), but a few extreme outliers caused Claude Code's cost to look far higher than it typically is.
Version 1.1 will add more harnesses and evaluate a broader set of models. Rather than holding the model fixed, we plan to evaluate the full harness × model matrix to reveal interaction effects (which harnesses work best…
The Exo harness is the most interesting one in the benchmark, as it has the potential to complete tasks with fewer turns and cost.