Qwen 3.8 27B with Asymmetric Dual GPUs (RTX 3080 20GB Mod + RTX 3070): Pipeline vs Tensor Parallelism, Oculink
I spent the last couple of nights tinkering with an asymmetric multi-GPU setup to see how far I could push context length on Qwen 3.8 27B without compromising quality. I wanted to share my findings, benchmarks, and the rabbit holes I fell into along the way. Disclaimer : I used Gemini to help write…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-08 08:33 · r/LocalLLM
Qwen 3.8 27B with Asymmetric Dual GPUs (RTX 3080 20GB Mod + RTX 3070): Pipeline vs Tensor Parallelism, Oculink