Qwen3.8-27B on One RTX 3090 vs Two: +20% Decode, +14% Cold Prefill, and 3x on Cached Prompts
This story is from 2026-09-23. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: Qwen3.8-27B — the dense 27.8B that dropped on August 14 — fits on a single RTX 3090 at W4A16 and decodes at ~125–155 tokens/sec . Splitting it across two cards with tensor parallelism buys roughly 20% more decode (noisy: +13% in one run, +22% in the next) and 14% faster prefill on average on…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-23 03:27 · DEV Community — AI
Qwen3.8-27B on One RTX 3090 vs Two: +20% Decode, +14% Cold Prefill, and 3x on Cached Prompts