AINewsnow

I built Nebula to run Qwen3.8-Flash-Next on a 12GB RTX 4070 Ti + 128GB RAM

This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.

Hi everyone, I'm the developer of Nebula, an open-source C/CUDA inference engine for Qwen3.8-Flash-Next. I started from antirez's DwarfStar (ds4) and specialized the engine for Qwen, combining native MTP speculative decoding with GPU expert caching and CPU MoE execution. Source code, architecture a…

Read the full story at r/LocalLLM ↗

Timeline · 2 reports

  1. 2026-09-14 03:01 · r/LocalLLaMA
    R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG.
  2. 2026-09-12 19:33 · r/LocalLLM
    I built Nebula to run Qwen3.8-Flash-Next on a 12GB RTX 4070 Ti + 128GB RAM

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis — r/LocalLLM
  3. Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D — r/LocalLLM
  4. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  5. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  6. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  7. Qwen Developers on X: "Qwen-Image 2.1 is going open source" — r/StableDiffusion
  8. Ternary Bonsai 2 27B — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →