AINewsnow

Qwen 3.8 27B with Asymmetric Dual GPUs (RTX 3080 20GB Mod + RTX 3070): Pipeline vs Tensor Parallelism, Oculink

I spent the last couple of nights tinkering with an asymmetric multi-GPU setup to see how far I could push context length on Qwen 3.8 27B without compromising quality. I wanted to share my findings, benchmarks, and the rabbit holes I fell into along the way. Disclaimer : I used Gemini to help write…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-08 08:33 · r/LocalLLM
    Qwen 3.8 27B with Asymmetric Dual GPUs (RTX 3080 20GB Mod + RTX 3070): Pipeline vs Tensor Parallelism, Oculink

More stories

  1. Nano Banana 2.1 is rolling out now. — r/GeminiAI
  2. Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance — r/LocalLLM
  3. Opus 5.5 vs. GPT-6 Astra vs. DeepSeek V4.1 Flash vs. Gemini 3.8 Flash ✈️ — r/ClaudeAI
  4. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA
  5. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Gemini 4 is coming today (or tmr depending on your time zone.) — r/GeminiAI
  7. What's up, Docsy? Google’s docs project joins the Linux Foundation as AI agents become readers — The New Stack AI
  8. Gemini 4 Argon — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →