Ten proofs, ten 97-99% scores, one rejection: what Verdikta's math bounties reveal about AI proof grading
This story is from 2026-10-09. It is preserved in the archive; the latest stories are on the live feed.
Verdikta Bounties has run a series of math bounties where two AI models (one from OpenAI, one from Anthropic, 50% weight each) grade a written proof against a weighted rubric, and an on-chain escrow pays if the score clears a threshold. I went through all ten completed ones. The numbers tell a more…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-09 02:27 · DEV Community — AI
Ten proofs, ten 97-99% scores, one rejection: what Verdikta's math bounties reveal about AI proof grading