Qwen3.8-27B at ~39 tok/s chat and ~117 tok/s context replay on one GB10 (DGX Spark / ASUS Ascent GX10)
This story is from 2026-08-24. It is preserved in the archive; the latest stories are on the live feed.
TL;DR Hardware: one NVIDIA GB10 — DGX Spark / ASUS Ascent GX10. Model: Qwen3.8-27B NVFP4 with a DFlash2 W4A16 drafter. Result: ~39 tok/s in ordinary chat and ~117 tok/s when lookup can reuse the prompt. Long-context improvement: cold TTFT at 50k fell from 29.94 to 23.46 seconds. Trade-off: chat-onl…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-24 19:43 · r/LocalLLM
Qwen3.8-27B at ~39 tok/s chat and ~117 tok/s context replay on one GB10 (DGX Spark / ASUS Ascent GX10)