We Tested a 35B LLM Against Typed-Decision Models on 12,000 Real RFQs—Confidence Changed the Winner
91.9%. 89.6%. 78.0%. Those were the primary-class accuracies of a typed-decision API, a 35B mixture-of-experts LLM, and a 421M open-weight decision model on the same real classification job. But accuracy was not the result that changed the deployment decision. The decisive question was: Can the mod…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-24 19:39 · DEV Community — Machine Learning
We Tested a 35B LLM Against Typed-Decision Models on 12,000 Real RFQs—Confidence Changed the Winner