Squeezing more performance out of Qwen 3.8 27B
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
What's up, everyone? Fellow local LLM-er here, trying to perfect my development environment. I am using the Macbook Pro, M5 Max with 128 GB unified memory. I have been running the OpenAI server with: mlx_lm.server \ --model mlx-community/Qwen3.8-27B-8bit \ --host 0.0.0.0 \ --port 8080 \ --max-token…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-13 16:49 · r/LocalLLM
Squeezing more performance out of Qwen 3.8 27B