MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025 (arxiv.org) 3 points by janandonly 1mo ago ↗ HN
[–] dima853 1mo ago ↗ How do you handle partial success in the benchmark? If an agent picks the wrong tool mid-chain but fixes it with backtracking, does it still count as a full success?
1 comment
[ 0.91 ms ] story [ 12.1 ms ] thread