GoBench: Evaluating LLMs on the game of Go [R]
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
GoBench evaluates LLMs on 9x9 Go games against a ladder of KataGo opponents, from random to superhuman. It measures general reasoning ability, strongly correlates with ARC-AGI 2 (r=0.83 correlation), and remains highly unsaturated. GPT-6 Astra max achieves 2500 Elo, much lower than the best KataGo,…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-16 18:54 · r/MachineLearning
GoBench: Evaluating LLMs on the game of Go [R]