AINewsnow

I got Mimo 2.6-Flash-RL running at 30-45 tok/s on the Strix Halo 128GB 2TB

https://preview.redd.it/yr2p7xsydcrh1.png?width=2455&format=png&auto=webp&s=e421690b54307775ece00731c300850f5e70455a A lot of this came from looking at how projects like Halogen and CIRU handled Qwen 3.8 Flash Next. I leaned on Claude for the automated benchmarking and the math I definitely couldn'…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-23 22:00 · r/LocalLLM
    I got Mimo 2.6-Flash-RL running at 30-45 tok/s on the Strix Halo 128GB 2TB

More stories

  1. Help? — r/GeminiAI
  2. I built a desktop app where Claude Code designs machines, writes their firmware, and simulates both together — r/ClaudeAI
  3. 5070 ti + 64 ram Qwen 3.8 27b — r/LocalLLM
  4. Qwen 3.8 27B on one 5090: 22 GB VRAM, 175k context — a coding driver, not a Claude replacement — r/LocalLLM
  5. Qwen 3.8 with Claude is amazing — r/LocalLLM
  6. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  7. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  8. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI

Get the daily brief of stories like this at 6:30 every morning →