Qwen3.8-27B on 2× RTX 5070 Ti (16 GB, no P2P) with a tensor-parallel NInfer fork vs an RTX 5090, same weights: 94-96% plain decode, 79-87% MTP3
My NInfer fork (C++/CUDA Qwen engine; upstream targets one RTX 5090) runs Qwen3.8-27B tensor-parallel on two 16 GB cards, too big for either alone. Benchmarked with upstream's own tools/bench suite against its published 5090 runs. Same weights (the official NVFP4 artifact) on both: 94-96% of the 50…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-25 08:42 · r/LocalLLM
Qwen3.8-27B on 2× RTX 5070 Ti (16 GB, no P2P) with a tensor-parallel NInfer fork vs an RTX 5090, same weights: 94-96% plain decode, 79-87% MTP3
More stories
- Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080 — r/LocalLLaMA
- I hate it when they do this! — Matt Wolfe
- Qwen Image 2.1's editing capabilities are mind-blowing! Generating Character Design Sheets without any LoRAs — r/StableDiffusion
- Qwen-Image 2.1 LoRA testing — r/StableDiffusion
- I added Qwen-Image 2.1 + LoRA support to TensorSharp (GGUF, local inference) — r/LocalLLaMA
- You are going to love this one, working on a 3D pose editor tool for qwen image edit. Amazing Qwen-Image 2.1 🤩! — r/StableDiffusion
- Qwen 2.1 Might Be Just TOO Good at Face Swap... [Free Workflow] — r/StableDiffusion
- Character Design Sheet V2.0 Update: A Practical Approach to Character Sheet Generation. — r/StableDiffusion
Get the daily brief of stories like this at 6:30 every morning →