Qwen3.8-27B NVFP4 5090 via NInfer on Windows/WSL2 - 205 tok/s coding, ~9k prefill, 192K ctx
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
Setup GPU: 1× RTX 5090 32GB (stock) OS: Windows 11 + WSL2 (Ubuntu) — engine runs in WSL, served on 127.0.0.1:8081 , reached from Windows fine Model: Qwen3.8-27B Huihui abliterated , NVFP4 (~19.7 GB weights) Engine: NInfer KV: fp8, 192K context (196,608), vision on Spec: MTP, 3 draft tokens, --lm-he…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-17 16:17 · r/LocalLLM
Qwen3.8-27B NVFP4 5090 via NInfer on Windows/WSL2 - 205 tok/s coding, ~9k prefill, 192K ctx