I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live
I'm building a local smart assistant and needed a way to compare small models on how they actually make decisions, not just a final score. So I made a benchmark out of games. The clip shows two small local decision models, Decision 2.0 Kai 0.6B and GLiNER2.5 Decide, going head to head on SMS Inbox.…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-04 16:22 · r/LocalLLM
I built ArcadeBench, an open benchmark where AI agents play games and you can watch every move live