APIEval-20: A Benchmark for Black-Box API Test Suite Generation (huggingface.co) 5 points by AkshatVirmani 5mo ago ↗ HN
[–] akshay_93 5mo ago ↗ like that the scoring bias is toward bug detection & not test generation only. generating lots of tests with AI is easy but that doesn't necessarily mean they're good
3 comments
[ 3.7 ms ] story [ 22.6 ms ] thread