Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public
I've been working on getting Qwen3.8-27B running with a genuine 131,072-token context and KV cache , Vision , and MTP-3 speculative decoding on a single RTX 5080 16GB . The project has moved on quite a bit since my original 128K/Vision validation, so I thought it was worth posting an updated summar…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-22 00:26 · r/huggingface
Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public - 2026-09-22 00:24 · r/LocalLLM
Qwen3.8-27B at true 128K + Vision on a single RTX 5080 16GB — NInfer v1.3, MTP-3, ~3.95 BPW, model now public