Show HN: Darts Vision Benchmark (darteval.vercel.app)

1 points by red545 ↗ HN
Detecting the correct score from a dartboard photo is unexpectedly hard for LLMs.

The task appears to stress spatial reasoning: Gemini 3 models lead this benchmark by a decent margin.

Counterintuitively, “more reasoning” often reduces accuracy.

Even the top-performing model scores only ~36% of darts correctly.

0 comments

[ 6.0 ms ] story [ 14.3 ms ] thread

No comments yet.