2 comments

[ 2.6 ms ] story [ 12.6 ms ] thread
I'd be super interested in knowing how other models perform by default and knowing the exact details of the prompts as well and seeing how two agents with different contexts or isolated contexts perform against one another, one as an implementer and another one as a validator.

This actually has benchmarking potential. It's especially interesting now that some of the models start having knowledge about themselves in the training dataset.

is there any published code or prompts regarding this for the either the prompts or the evals?

*also trying it against several different delegation topologies would be super interesting but it seems like none of the harnesses by default offer control over how many levels of delegations and which parts of the task each instance or each agent is gonna focus on