AINewsnow

Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

We built an inference engine for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them. This started as a proof of concept, and poc worked, we are getting 9-10 tok/s decode on Qwen3.8-Flash-Next NVFP4 (9.06 on the benchmark turn, 10.4…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-27 22:13 · r/LocalLLaMA
    Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM

More stories

  1. Scoop: Anthropic's Dario Amodei to have White House dinner with Trump — Axios AI+
  2. Bill Gates says unchecked AI could ‘cause a billion deaths’ in call for regulation — The Guardian AI
  3. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  4. Scoop: Top AI companies probing tens of thousands of security incidents — Axios AI+
  5. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times Technology
  6. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  7. ‘Things Will Never Be Chill Again’: The Doomers Who Shaped the AI Safety Freakout — Wall Street Journal Technology
  8. OpenAI says its models engaged with US government websites in new model misbehavior disclosure — ABC News Technology

Get the daily brief of stories like this at 6:30 every morning →