AINewsnow

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Hi. I saw some feedback that halogen was degrading at context depth. So I fixed that. Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0: decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter) decode at 258,794: 42.9 to 45.0 prefil…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-20 01:47 · r/LocalLLM
    I benchmarked 12 quantizations of Qwen3.8-Flash-Next on Strix Halo — here's what actually works
  2. 2026-09-19 21:49 · r/LocalLLaMA
    Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

More stories

  1. Higgsfield AI ships new video features in a day with GPT-6 Astra — OpenAI News
  2. Amazon blocks Meta’s Muse AI agent — The Verge AI
  3. Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI revenue — South China Morning Post Tech
  4. Google's Gemini AI hacked three companies in security test — BBC Technology
  5. British Columbia Sues OpenAI Over Canada Mass Shooting Warning Failure — Bloomberg AI
  6. Lawsuit accuses Anthropic, OpenAI, SpaceXAI, Google of AI pacing 'collusion' — The Hill Technology
  7. Grok 4.7 — Hacker News Front Page
  8. Ahead of Sam Altman's UN address, OpenAI proposes new ways to track AI misalignment risks — Axios AI+

Get the daily brief of stories like this at 6:30 every morning →