I Ran a 125B Model on Three RTX 3090s at 80 Tokens/s by Teaching llama.cpp Which Experts Matter
This story is from 2026-09-26. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: Qwen3.8-Flash-Next is a 125B mixture-of-experts model: on paper it beats the Qwen3.8-27B I run in production everywhere, by +16.5 points on agentic coding. It does not fit in the 72 GB of VRAM my three RTX 3090s have, and stock llama.cpp, spilling a quarter of the experts into system RAM, ru…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-26 16:02 · DEV Community — AI
I Ran a 125B Model on Three RTX 3090s at 80 Tokens/s by Teaching llama.cpp Which Experts Matter