DeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps)
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
I liked DeepSeek V4.1 Flash as an agent model but at 16 t/s on ds4 it was painful to sit through a real turn. It seemed some redditors and m3 ultra owners appreciated my glm 53 flash optimizations, so I forked antirez/ds4 for V4.1 Flash to see how much I learned optimizing GLM for the M3 Ultra coul…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-15 11:51 · r/LocalLLaMA
DeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps)