Qwen3.8-Flash-Next (180B MoE) on one DGX Spark: 43.9 tok/s. The same machine gave 11.2 tok/s with Qwen3.8-27B dense model.
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
Before this test, I ran Qwen3.8-27B NVFP4 on the same DGX Spark. It gave 11.2 tokens each second. Qwen3.8-Flash-Next has 180 billion parameters. It uses 7.31 billion parameters for each token. It also has a built-in draft head for speculative decoding. The result is 43.9 tokens each second. That is…
Read the full story at r/LocalLLM ↗