A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.
A while ago, I posted here getting OSS-120B and GLM-4.6 playing full games of Civilization V . Since then, models have moved pretty far, and we wanted a better understanding about models' capabilities playing the game. Introducing the controlled version of CivBench on newer models: The controlled v…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-10-05 23:38 · r/LocalLLM
A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. - 2026-10-05 23:16 · r/LocalLLaMA
A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.
More stories
- Introducing GLM 5.3 on Amazon Bedrock — AWS Machine Learning Blog
- can i run qwen flash next with these specs, or am i out of luck? — r/LocalLLM
- Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
- The Story of Qwen: Alibaba's AI Models From 7B to 2.4T — MarkTechPost
- Aleph Alpha releases open-weight Kolibri with 1M context — TestingCatalog AI News
- Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen — r/LocalLLaMA
- One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui
- My frontier class agent fact-checks my local AI before I grade it. How do you grade your Agents and LLMs? — r/AI_Agents
Get the daily brief of stories like this at 6:30 every morning →