Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
I was using UD-Q3_K_XL until now with more than 140000 context. Quality wise it's very good, very few erroneous tool calls. Then I saw many others here reporting good results with IQ3_XXS, so I gave it a try. The downside is prompt processing speed went down from 700-800 tk/s to 400 tk/s. Quality d…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-08-29 12:50 · r/LocalLLaMA
Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp) - 2026-08-27 19:45 · r/LocalLLaMA
Over 200k context on 16GB VRAM with Qwen 3.8 27B UD-IQ3_XXS