A 3B model that beats a 7B: failure-driven orchestration on uncontaminated knowledge
How a fully sovereign QA stack — local Wikipedia index, distilled Qwen2.5-3B reader, and crutches built only from measured failures — went from 33% to 52% on facts no LLM can have memorized, and matched a zero-shot 7B more than twice its size. TL;DR Configuration Post-cutoff-150 accuracy 3B naked (…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-01 19:53 · DEV Community — Machine Learning
A 3B model that beats a 7B: failure-driven orchestration on uncontaminated knowledge