HumanEval is saturated: new coding LLM benchmark released (bigcode-bench.github.io) 1 points by eitanturok 2y ago ↗ HN
0 comments
[ 4.0 ms ] story [ 12.1 ms ] threadNo comments yet.