Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad (arxiv.org) 6 points by mauriziocalo 1y ago ↗ HN
[–] galaxyLogic 1y ago ↗ > Our results reveal that all tested models struggled significantly, achieving less than 5% on average
1 comment
[ 1.9 ms ] story [ 12.1 ms ] thread