Splash on a 40-core M5 Max: +20% decode by tuning the kernels for your own chip
FYI the engine's default kernel rules were measured on smaller chips (16/20-core M5s and a 32-core M4 Max), so a 40-core M5 Max runs guesses. Splash's repo includes a developer tool, “make tune-kernels” that tests every available way of running each quantised matrix-multiply on your hardware. On my…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-25 13:13 · r/LocalLLaMA
Splash on a 40-core M5 Max: +20% decode by tuning the kernels for your own chip