Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
This story is from 2026-10-03. It is preserved in the archive; the latest stories are on the live feed.
I run local models on one home machine: an RTX 5060 with 8 GB , a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. When NVIDIA published Qwen3.8-Flash-Next in NVFP4 (133 GB on disk), the obvious answer was "it doesn't fit". It does now: it runs as a normal chat mode…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-03 11:46 · DEV Community — AI
Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe