4 comments

[ 3.3 ms ] story [ 17.5 ms ] thread
Cool launch - let me try this with our in house setup!
Kind of wild that we're finally getting benchmarks for AI-generated API testing. Feels like the equivalent of SWE-bench, but for finding actual bugs instead of writing code.