6-7tk/s on qwen3.8-next-flash Q2_0 at 128k context on 8gb vram and 32gb ram worth it?
i got this thing badtheorylabs/Tinfield-1-Mini-GGUF running on my 8gb vram GPU and 32gb ddr5 ram, it is like a post trained qwen3.8 flash next, at q2 quant, really really bad KL btw (0.340 against q8) So it has like 34.7 GiB of total weights in experts and then around a ~35 gb ngram table, so all t…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-26 16:14 · r/LocalLLM
6-7tk/s on qwen3.8-next-flash Q2_0 at 128k context on 8gb vram and 32gb ram worth it?