AINewsnow

Ten proofs, ten 97-99% scores, one rejection: what Verdikta's math bounties reveal about AI proof grading

This story is from 2026-10-09. It is preserved in the archive; the latest stories are on the live feed.

Verdikta Bounties has run a series of math bounties where two AI models (one from OpenAI, one from Anthropic, 50% weight each) grade a written proof against a weighted rubric, and an on-chain escrow pays if the score clears a threshold. I went through all ten completed ones. The numbers tell a more…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-09 02:27 · DEV Community — AI
    Ten proofs, ten 97-99% scores, one rejection: what Verdikta's math bounties reveal about AI proof grading

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Claude Pro vs ChatGPT Plus vs Copilot Premium: which one would you choose for this use case? — r/ChatGPTPro
  3. What to know about Mistral's ML4 as it bets on EU sovereignty in the US-China open-weight AI race — Euronews Next
  4. Hot take but AI mode is by far the most useful AI out of chatGPT/Claude/Gemini — r/GeminiAI
  5. What is AI model distillation, and why is it so hard to stop? — Scientific American
  6. SpaceX looks to raise $40bn to buy Nvidia chips — Financial Times AI
  7. Rogue AI or human error? The real story behind the OpenAI-Hugging Face incident — Scientific American
  8. OpenAI will watermark ChatGPT outputs by default—but only in the EU — Ars Technica AI

Get the daily brief of stories like this at 6:30 every morning →