This draft model is OP on 16 GB cards for Qwen 3.8 27b
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
https://huggingface.co/HermiHg/Qwen3.8-27B-DFlash2-Q2_K_S-MIX-GGUF I used this draft model with https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with the IQ3_XXS with 128k context and I saw it averaging about 60 tokens per second tg speed on the 16 GB RX 9070 XT. This is much better than…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-12 14:57 · r/LocalLLaMA
This draft model is OP on 16 GB cards for Qwen 3.8 27b