Multi-Agent Step Race Benchmark: LLM Collaboration and Deception Under Pressure (github.com) 7 points by zone411 1y ago ↗ HN
2 comments
[ 1.9 ms ] story [ 16.2 ms ] thread