I built Nebula to run Qwen3.8-Flash-Next on a 12GB RTX 4070 Ti + 128GB RAM
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
Hi everyone, I'm the developer of Nebula, an open-source C/CUDA inference engine for Qwen3.8-Flash-Next. I started from antirez's DwarfStar (ds4) and specialized the engine for Qwen, combining native MTP speculative decoding with GPU expert caching and CPU MoE execution. Source code, architecture a…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-14 03:01 · r/LocalLLaMA
R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG. - 2026-09-12 19:33 · r/LocalLLM
I built Nebula to run Qwen3.8-Flash-Next on a 12GB RTX 4070 Ti + 128GB RAM