Strix Halo 128GB local LLM tuning: 96GB UMA made a ~100GB Qwen3 235B model go from timeout to 17 tok/s -> scripts + benchmarks
I’ve been benchmarking and tuning a Bosgame M5 / AMD Ryzen AI MAX+ 395 with 128GB RAM and Radeon 8060S (GFX1151) for large local GGUF models under Linux. I put the results, scripts, benchmark data and graphs into a public repo here: GitHub repo: strix-halo-128gb-llm-guide There is also a frozen v1.…
Read the full story at r/LocalLLM ↗