AINewsnow

MacBook M5 and AMD Strix Halo sharing large models

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

Running a GLM 5.3 321B shared on a MacBook M5 max and a Strix Halo, both with 128 GB ram. TB4 connection. Found out that llama.cpp's default tensor split puts about half the layers on the slower box, so the pair was actually slower than one Mac. Loaded the model Mac-heavy instead: IQ1_S went 188 to…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-11 09:23 · r/LocalLLM
    MacBook M5 and AMD Strix Halo sharing large models

More stories

  1. Release b11003 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM
  3. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  4. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/LocalLLM
  5. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  6. Digit-logits-based classifier with llama.cpp — r/LocalLLaMA
  7. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  8. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →