Users Share Qwen 3.8 27B Token Speeds on Apple Silicon and RTX GPUs
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
Reddit users report Qwen 3.8 27B speeds: 17-20 tok/s on an M5 Pro MacBook Pro via LM Studio, and 74 tok/s on an RTX 5080 with IQ4_XS-Smaller quantization and MTP enabled. Another user asks about RTX Pro 4000 SFF performance.
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-08-24 03:20 · r/LocalLLM
Qwen 3.8-27B on RTX 5080 - 2026-08-22 12:16 · r/LocalLLM
Any of you running Qwen 3.8 27B on an RTX Pro 4000 SFF Blackwell? - 2026-08-22 04:48 · r/LocalLLM
people running Qwen 3.8 27B on apple silicon… whats your best token generation speed and how did you attain it?