AINewsnow

~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.

Coverage of "~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5." from 2 sources, with a live timeline of who reported what and when.

Read the full story at r/huggingface ↗

Timeline · 2 reports

  1. 2026-10-05 19:39 · r/LocalLLM
    Qwen3.8-Flash-Next Q4 vs Qwen3.8 27B Q5 on single R9700 (32GB) + 64GB RAM: 2x128k context, almost similar performance
  2. 2026-10-05 11:11 · r/huggingface
    ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  3. Introducing Mistral Large 4 — Mistral AI News
  4. Sharing AI progress in mathematics — OpenAI News
  5. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  6. Everything announced at Microsoft's Windows and Surface event — Engadget
  7. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  8. Introducing Playground: Create and play custom games — Google AI Blog

Get the daily brief of stories like this at 6:30 every morning →