A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.
I'm building an AI agent from scratch, and one decision I need to make is which open-source model to use. At first, I thought I'd go with GPT-OSS-120B, which is approximately four times the size of Qwen3.6-35B. I assumed the larger model would be substantially better at everything. It wasn't. To te…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-26 07:25 · r/AI_Agents
A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.