Show HN: Darts Vision Benchmark

https://darteval.vercel.app/darts-vision/dashboard
by red545 • 7 months ago
1 0 7 months ago

Detecting the correct score from a dartboard photo is unexpectedly hard for LLMs.

The task appears to stress spatial reasoning: Gemini 3 models lead this benchmark by a decent margin.

Counterintuitively, “more reasoning” often reduces accuracy.

Even the top-performing model scores only ~36% of darts correctly.

Related Stories

Loading related stories...

Source preview

darteval.vercel.app