[DGX Spark] Qwen 3.8 27B (FP8) at ~32tok/s generation
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
I finally got a good, usable setup for the new dense Qwen3.8-27B model in FP8 on the DGX Spark. Z-Lab released a new drafter, [Qwen3.8-27B-DFlash2]( https://huggingface.co/z-lab/Qwen3.8-27B-DFlash2 ), together with the new DFlash 2 speculative-decoding approach. With this setup, Qwen3.8-27B has bee…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 06:50 · r/LocalLLM
[DGX Spark] Qwen 3.8 27B (FP8) at ~32tok/s generation