AINewsnow

I swapped in a "better" 9B model for my local agent seats. It silently wrote tool calls as prose 5 times out of 80.

Ran a proper comparison before trusting a model swap on local agent seats, and the interesting part was not which model scored higher. It was how the worse one failed. Setup: 20 tasks shaped like actual agent work, five each of read, write, edit, and patch, using the real tool schemas from my harne…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-27 01:38 · r/AI_Agents
    I swapped in a "better" 9B model for my local agent seats. It silently wrote tool calls as prose 5 times out of 80.

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  4. OpenAI agent hacked an Australian government healthcare website — New Scientist AI
  5. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  6. Am I the only one who actually likes GPT-6 Sol and Luna? — r/ChatGPT
  7. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  8. Appeals Court Lets the Pentagon Designate Anthropic a Supply-Chain Risk — Wired AI

Get the daily brief of stories like this at 6:30 every morning →