Getting Qwen3.8-27B with decent speed on my 4080 with 16Gb card
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
I saw that Q2 is actually very good and produce real good results and I also saw how dflash2 make its running at generating >60 t/s with a 120k context lenght. And I like what its doing!! Heres how to set it up (ai wrote this) DFlash2 speculative decoding on 16GB VRAM — 1.72x faster (setup guide) D…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-26 01:41 · r/LocalLLaMA
Getting Qwen3.8-27B with decent speed on my 4080 with 16Gb card