(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
Hey! I forked NInfer (a from-scratch C++20/CUDA inference engine for Qwen models) and added two things: tensor-parallel across two GPUs, and YaRN ×4 rope scaling. Together they let Qwen3.8-27B NVFP4 run a 1,048,576-token context on two consumer 5090s — 27.4 GB per card, no NVLink. Numbers (single s…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-29 18:24 · r/LocalLLaMA
(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s