Evaluating performance and efficiency of the GitHub Copilot agentic harness (github.blog) 3 points by mariuz 2mo ago ↗ HN
[–] brammertottens 2mo ago ↗ It's an interesting post, but i'm a bit skeptical on their decision to report the best run for each agent, and not just the mean over the 5 runs. We have seen this as well in running benchmarks, that variance within one setup can be pretty big.
1 comment
[ 3.7 ms ] story [ 13.5 ms ] thread