AINewsnow

NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

I've tinkered some more with my fork of NInfer for the Qwen 3.8 Flash Next model on RTX6000, and I've just hit 400tg/s under ideal conditions with it, so I thought it was worth a share. Decode MTP3 with --lm-head-draft Context 16-bit 8-bit Change 512 197.1 tok/s 274.8 tok/s +39% 8K 303.1 tok/s 401.…

Read the full story at r/LocalLLaMA ↗

Timeline · 6 reports

  1. 2026-10-08 11:07 · r/LocalLLaMA
    Halogen + Qwen Flash Next keeps getting better
  2. 2026-10-08 10:55 · r/LocalLLaMA
    Qwen 3.8 Flash Next is so much fun for three.js
  3. 2026-10-07 22:26 · r/LocalLLM
    Is anyone running Qwen Flash Next on Strata with q6 or higher?
  4. 2026-10-07 22:02 · r/LocalLLM
    Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next
  5. 2026-10-07 20:43 · r/LocalLLaMA
    Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?
  6. 2026-10-06 10:37 · r/LocalLLaMA
    NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

More stories

  1. Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature — r/LocalLLaMA
  2. 113 Decision Models in 3 Weeks (Mostly Qwen and Gemma Fine-Tunes): A Deep Dive — r/AI_Agents
  3. Ultimate web scraper — r/AI_Agents
  4. Switching from Claude Code to local Qwen for Android dev — can smaller local models keep up? — r/LocalLLM
  5. I built a root-cause tool with Claude Code, then built a validator to catch the LLM inside it when it's confidently wrong. Honest numbers: 60% / 20% — r/AI_Agents
  6. GPT-6 and Intelligent UI for everyone — OpenAI News
  7. Introducing Mistral Large 4 — Mistral AI News
  8. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →