AINewsnow

Strix Halo 128GB local LLM tuning: 96GB UMA made a ~100GB Qwen3 235B model go from timeout to 17 tok/s -> scripts + benchmarks

I’ve been benchmarking and tuning a Bosgame M5 / AMD Ryzen AI MAX+ 395 with 128GB RAM and Radeon 8060S (GFX1151) for large local GGUF models under Linux. I put the results, scripts, benchmark data and graphs into a public repo here: GitHub repo: strix-halo-128gb-llm-guide There is also a frozen v1.…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-24 00:26 · r/LocalLLM
    Strix Halo 128GB local LLM tuning: 96GB UMA made a ~100GB Qwen3 235B model go from timeout to 17 tok/s -> scripts + benchmarks

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Bringing Private Processing to Meta AI Glasses — Engineering at Meta
  3. Gemini 3.8 text-to-speech models now available on AI Gateway — Vercel Blog
  4. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  5. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  6. OpenAI Agent Hacked Australian Government Website — Wall Street Journal Technology
  7. Meta Connect 2026 live: Updates from Mark Zuckerberg's keynote on AI glasses, VR and more — Engadget
  8. How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →