AINewsnow

A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.

I'm building an AI agent from scratch, and one decision I need to make is which open-source model to use. At first, I thought I'd go with GPT-OSS-120B, which is approximately four times the size of Qwen3.6-35B. I assumed the larger model would be substantially better at everything. It wasn't. To te…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-26 07:25 · r/AI_Agents
    A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.

More stories

  1. GPT‑6 Sol and Luna: Cheaper, but Worse Where It Matters — r/OpenAI
  2. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. Need some help — r/AI_Agents
  5. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  6. new update? — r/GeminiAI
  7. Question about Wan 3 — r/StableDiffusion
  8. OpenAI rogue agents targeted govt and varsity websites in US, Australia before Hugging Face hack: What we know — Mint AI

Get the daily brief of stories like this at 6:30 every morning →