AINewsnow

Qwen3.8 Flash Next GSQ-RCO IQ3_XXS running Strata on 2x Intel B70s

Flash-Next GSQ-RCO IQ3_XXS via Strata running on 3975wx Threadripper with 256gb DDR4 with 131k context. No speculative decoding. 4096 chunk prefill 1,563.4 tok/s and decode up to 58 tok/s. Speed is great, prefill could probably be tuned even higher and MTP/Dflash should be tested eventually but the…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-05 18:02 · r/LocalLLM
    Qwen3.8 Flash Next GSQ-RCO IQ3_XXS running Strata on 2x Intel B70s

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  3. Introducing Mistral Large 4 — Mistral AI News
  4. Sharing AI progress in mathematics — OpenAI News
  5. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  6. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  7. Introducing Playground: Create and play custom games — Google AI Blog
  8. Surface RTX Spark Dev Box is available for preorder for $5,999 — The Verge AI

Get the daily brief of stories like this at 6:30 every morning →