Qwen3.8-27B @ 100K context on an RTX 4080 16GB — ExLlamaV3 MTP results
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Been playing around with Qwen3.8-27B on my 4080 and figured I'd post the numbers since this turned out better than I expected. The goal was to see how much of the model/context I could squeeze into 16GB while keeping generation speed decent, and then see whether MTP was actually worth the extra VRA…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-14 10:16 · r/LocalLLM
Qwen3.8-27B @ 100K context on an RTX 4080 16GB — ExLlamaV3 MTP results