M5 Max users: what models are you using & what tk/s are you getting?
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
I was using antirez’s ds4 for a while and getting around 20 tk/s, which worked for my purposes. But I know there have been big advancements between Qwen, the DS4 vision model, and GLM. I’m not sure how the quants affect performance, so what’s the best thing to run right now & how fast is it? submit…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-06 01:32 · r/LocalLLaMA
M5 Max users: what models are you using & what tk/s are you getting?