For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
I've been making a lot of comments about optimal setup for Strix Halo (gfx1151) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the s…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-08 00:43 · r/LocalLLaMA
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput