AINewsnow

[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed. Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-29 08:40 · r/LocalLLaMA
    [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. The Future Is for Everyone: Muse for Small Business — Meta Newsroom
  4. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  5. OpenAI Scraps Debut of AI Model as It Sets New Guardrails — Bloomberg AI
  6. Introducing Claude Sonnet 5.5 on AWS — AWS Machine Learning Blog
  7. AMD will acquire Fei-Fei Li's World Labs for $8.2 billion — TechCrunch AI
  8. Meta launches enterprise AI business seeking to cash in on vast spending — Financial Times AI

Get the daily brief of stories like this at 6:30 every morning →