I swapped in a "better" 9B model for my local agent seats. It silently wrote tool calls as prose 5 times out of 80.
Ran a proper comparison before trusting a model swap on local agent seats, and the interesting part was not which model scored higher. It was how the worse one failed. Setup: 20 tasks shaped like actual agent work, five each of read, write, edit, and patch, using the real tool schemas from my harne…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-27 01:38 · r/AI_Agents
I swapped in a "better" 9B model for my local agent seats. It silently wrote tool calls as prose 5 times out of 80.