AINewsnow

Qwen3.8-Flash-Next 125B at 17–26 tok/s on a 16 GB GPU + 32 GB RAM

I’ve been working on running Qwen3.8-Flash-Next on relatively modest hardware: RTX 5060 Ti 16 GB 32 GB DDR5 NVMe SSD The model is ~76 GB, so the experts are split across VRAM, RAM and NVMe. Current real-workload speeds: Code: 26.4 tok/s Agent: 21.3 tok/s Reasoning: 21.8 tok/s 21K context: 17.7 tok/…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 18:26 · r/LocalLLM
    Qwen3.8-Flash-Next 125B at 17–26 tok/s on a 16 GB GPU + 32 GB RAM

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
  3. Google rolls out Gemini 4 Argon to trusted cyber defenders through Fairwind and says it is participating in the US government's voluntary pre-release process (Madison Mills/Axios) — Techmeme
  4. Introducing dots — OpenAI News
  5. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  6. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  7. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  8. OpenAI pauses AI model training after another agent bypasses network restrictions — InfoWorld AI

Get the daily brief of stories like this at 6:30 every morning →