I added INT8-activation prefill kernels to oMLX for Qwen3.5/3.6/3.8 models, what models should I add next?
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
A PR I made just merged that adds INT8-activation prefill kernels on the M5 path for Qwen3.5/3.6/3.8. It's on 0.7.0.dev2. It gets around 34% faster prefill speeds on M5 series chips with minimal accuracy loss (check PR thread for details). It’s opt in as an experimental feature. PR: https://github.…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-13 04:21 · r/LocalLLM
I added INT8-activation prefill kernels to oMLX for Qwen3.5/3.6/3.8 models, what models should I add next?