AINewsnow

Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME

I got Qwen3.8-Flash-Next-GSQ-RCO-Abliterated running at IQ3_S with just 12GB VRAM and 32GB system RAM, achieving 20-30 tok/sec decode (Q2 achieves 39-45 tok/s) & 300 to ~90 thousand tok/sec prefill @ 131k context, on a custom fork of Strata. This fork has tonnes of architectural changes, all are ve…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-09 14:34 · r/LocalLLaMA
    Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME

More stories

  1. Introducing GPT-6 in ChatGPT with Intelligent UI — OpenAI YouTube
  2. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  3. Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude — The Guardian AI
  4. Philadelphia police receive false homicide tip from Anthropic AI model — The Hill Technology
  5. Anthropic launches free AI security scans for open-source projects — The Verge AI
  6. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  7. Sophos cuts threat investigation time by 96% with OpenAI Daybreak — OpenAI News
  8. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog

Get the daily brief of stories like this at 6:30 every morning →