DFlash2 speeds Qwen 3.8 27B up to 4 times
This story is from 2026-08-19. It is preserved in the archive; the latest stories are on the live feed.
llama.cpp pr #27342 adds dflash2, so i rented an rtx 6000 and ran the same four prompts through four decoding setups on qwen3.8 27B median results over the four tasks: baseline 47.4 tok/s mtp 114.7 tok/s dflash 99.3 tok/s dflash2 140.6. tok/s so on average 3x for dflash2 though i have to point out…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-08-19 18:14 · r/LocalLLM
DFlash2 speeds Qwen 3.8 27B up to 4 times - 2026-08-19 18:10 · r/LocalLLaMA
DFlash2 speeds Qwen 3.8 27B up to 4 times