1 comment

[ 70.3 ms ] story [ 63.7 ms ] thread
>On a composite of all three dimensions, a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009)

note this was GPT-5.2