AINewsnow

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395. Model GLM-5.3-Flash MiMo-V2.6-Flash-MOPD Size 99.7 GB 105 GB Prefill 580 tok/s at 3.5K, 546 at 64K…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-03 14:45 · r/LocalLLM
    Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine
  2. 2026-10-03 14:16 · r/LocalLLaMA
    Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

More stories

  1. Prime Intellect Launches Prime Inference, Serving 600B Tokens Daily — AlphaSignal
  2. Hugging Face Pulls GLM-5.3 Build Made for Cyberattacks — r/ArtificialInteligence
  3. GLM-5.3 and the spread of advanced cyber capabilities \ Anthropic — r/ArtificialInteligence
  4. Best local model for Blender and game dev? — r/LocalLLaMA
  5. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  6. Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
  7. Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
  8. Google unveils Gemini 4 Argon: Its most powerful AI model yet, focused on coding and cyber defence — Mint AI

Get the daily brief of stories like this at 6:30 every morning →