MacBook M5 and AMD Strix Halo sharing large models
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
Running a GLM 5.3 321B shared on a MacBook M5 max and a Strix Halo, both with 128 GB ram. TB4 connection. Found out that llama.cpp's default tensor split puts about half the layers on the slower box, so the pair was actually slower than one Mac. Loaded the model Mac-heavy instead: IQ1_S went 188 to…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-11 09:23 · r/LocalLLM
MacBook M5 and AMD Strix Halo sharing large models