Cherry picked Qwen3.8 27B token generation speed with 4x RTX A3000 12GB
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I'm sure this is a fast generation but prefilling crawling probably due to small batch size tensor split across 4 gpus with that MTP configuration.
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-11 12:32 · r/LocalLLM
Cherry picked Qwen3.8 27B token generation speed with 4x RTX A3000 12GB