38 t/s on an RTX 3060 for Qwen3.8 27B (and 56 t/s for Qwen3.6 35B-A3B)
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
I see posts for cards like 40 series and 50 series but unfortunately im still stuck with a 3060. Tried to push this humble 3060 to its limits hosting Qwen 3.8 27B (quantized of course). Got it from 22 to ~40 t/s and the 35B MoE to 56 (just used HumanEval), on both Ubuntu headless and WSL2. Still ca…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-09 15:08 · r/LocalLLM
38 t/s on an RTX 3060 for Qwen3.8 27B (and 56 t/s for Qwen3.6 35B-A3B)