AINewsnow

Qwen3.8-27B on 2× RTX 5070 Ti (16 GB, no P2P) with a tensor-parallel NInfer fork vs an RTX 5090, same weights: 94-96% plain decode, 79-87% MTP3

My NInfer fork (C++/CUDA Qwen engine; upstream targets one RTX 5090) runs Qwen3.8-27B tensor-parallel on two 16 GB cards, too big for either alone. Benchmarked with upstream's own tools/bench suite against its published 5090 runs. Same weights (the official NVFP4 artifact) on both: 94-96% of the 50…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-25 08:42 · r/LocalLLM
    Qwen3.8-27B on 2× RTX 5070 Ti (16 GB, no P2P) with a tensor-parallel NInfer fork vs an RTX 5090, same weights: 94-96% plain decode, 79-87% MTP3

More stories

  1. Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080 — r/LocalLLaMA
  2. I hate it when they do this! — Matt Wolfe
  3. Qwen Image 2.1's editing capabilities are mind-blowing! Generating Character Design Sheets without any LoRAs — r/StableDiffusion
  4. Qwen-Image 2.1 LoRA testing — r/StableDiffusion
  5. I added Qwen-Image 2.1 + LoRA support to TensorSharp (GGUF, local inference) — r/LocalLLaMA
  6. You are going to love this one, working on a 3D pose editor tool for qwen image edit. Amazing Qwen-Image 2.1 🤩! — r/StableDiffusion
  7. Qwen 2.1 Might Be Just TOO Good at Face Swap... [Free Workflow] — r/StableDiffusion
  8. Character Design Sheet V2.0 Update: A Practical Approach to Character Sheet Generation. — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →