Swift-1.5 Qwen3.8-27B on a single RTX 5090: all-NVFP4 + DFlash2, full 262k context, ~160 tok/s decode
We're serving a single-GPU LLM box (RTX 5090 32 GB) and built this artifact: ukisai's Swift-1.5 Qwen3.8-27B (RL+OPD post-training, agentic/coding focus) converted to all-NVFP4 (W4A4 gs16) + z-lab DFlash2 drafter, for the NInfer engine (v3). Artifact (public, SHA-pinned): Qwen3.8-27B-swift15-nvfp4fu…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-25 03:16 · r/LocalLLM
Swift-1.5 Qwen3.8-27B on a single RTX 5090: all-NVFP4 + DFlash2, full 262k context, ~160 tok/s decode