Qwen3.8-27B on 2x RTX 5070 Ti now matches an RTX 5090 in decode on the same weights (tensor-parallel NInfer fork)
I've been using a local Qwen3.8-27B as a sub-agent for Opus for a while now. For everyday coding it's more than good enough, and it saves a lot of tokens. For the harder, more novel stuff Opus is still needed. When I built my workstation I went with two RTX 5070 Ti instead of a 5090 to stay on budg…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-28 13:04 · r/LocalLLM
Qwen3.8-27B on 2x RTX 5070 Ti now matches an RTX 5090 in decode on the same weights (tensor-parallel NInfer fork)