HWE Bench: A new unbounded Benchmark for LLMs (GPT 5.5 is on top) (hwebench.com) 6 points by fesens 4mo ago ↗ HN
[–] fesens 4mo ago ↗ Current benchmarks have ceilings, usually 100%. This benchmark aims to be a long lasting, high correlation with the ability to solve real world problems and follow complex instructions, and unbounded (meaning it can always go higher).
3 comments
[ 3.1 ms ] story [ 23.8 ms ] thread