AINewsnow

Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe

This story is from 2026-10-03. It is preserved in the archive; the latest stories are on the live feed.

I run local models on one home machine: an RTX 5060 with 8 GB , a Core Ultra 5 225F, 31 GiB of RAM and a Gen5 NVMe drive used only for model files. When NVIDIA published Qwen3.8-Flash-Next in NVFP4 (133 GB on disk), the obvious answer was "it doesn't fit". It does now: it runs as a normal chat mode…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-03 11:46 · DEV Community — AI
    Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU. — r/LocalLLaMA
  3. Griffin, the first Human Interaction Model to pass video Turing Test it's already #1 on NVIDIA's benchmark for full-duplex AI video - 44% of people thought it was a real person while other systems are at ~3% — r/singularity
  4. NVIDIA Vera CPU Is Coming to CoreWeave: Pack In More Agents — CoreWeave Blog
  5. What It Takes to Bring Up a Multi-Rack NVIDIA Vera Rubin NVL72 Cluster — CoreWeave Blog
  6. What Comes Next: Operating and Evolving the Production AI Factory — CoreWeave Blog
  7. Why AI Factories Need Proof Before Production — CoreWeave Blog
  8. Liquid-Cooled Switching Doubles AI Network Bandwidth Per Rack — CoreWeave Blog

Get the daily brief of stories like this at 6:30 every morning →