Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.
I've got Qwen3.8-Flash-next running on RTX 3090, Ryzen 9 3950X, a PCIe 3.0 motherboard, and 64GB DDR RAM from 2020. IQ4_XS weights, full kvarn5 context, vision on GPU, experts in host RAM, n-grams on disk. MTP works but actually slows decode down even with 80% draft acceptance, as expected since ev…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-28 15:40 · r/LocalLLaMA
Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)