AINewsnow

I got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming

This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.

Bit of a side project I wanted to share. The metrics are 17.5tks decode, 360tks prompt processing based testing against my normal ai usage. I tested a couple of new things others haven’t done (at least that I’ve seen). Setup a carousel buffer for streaming in experts for prompt processing which got…

Read the full story at r/LocalLLaMA ↗

Timeline · 11 reports

  1. 2026-10-07 22:26 · r/LocalLLM
    Is anyone running Qwen Flash Next on Strata with q6 or higher?
  2. 2026-10-07 22:02 · r/LocalLLM
    Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next
  3. 2026-10-07 20:43 · r/LocalLLaMA
    Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?
  4. 2026-10-06 14:58 · r/LocalLLaMA
    Qwen3.8-Flash-Next on Strata
  5. 2026-10-06 10:37 · r/LocalLLaMA
    NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s
  6. 2026-10-05 17:38 · r/LocalLLM
    Strata: Qwen 3.8 Flash next in loop
  7. 2026-10-05 12:28 · r/LocalLLM
    Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance
  8. 2026-10-05 10:50 · r/LocalLLaMA
    MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?
  9. 2026-10-05 07:27 · r/LocalLLM
    4x 3080 20GB (modded, alibaba) + Strata (Qwen Flash Next 125B IQ3_XXS) = 105 t/s generation 5000 prompt processing on 150k context
  10. 2026-10-05 07:21 · r/LocalLLM
    Upgrade existing PC to run Qwen 3.8 Flash Next via Strata or swap to a Strix Halo?
  11. 2026-10-05 06:44 · r/LocalLLaMA
    I got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming

More stories

  1. What to know about Mistral's ML4 as it bets on EU sovereignty in the US-China open-weight AI race — Euronews Next
  2. Upgrade existing PC to run Qwen 3.8 Flash Next via Strata or swap to a Strix Halo? — r/LocalLLM
  3. ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5. — r/huggingface
  4. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA
  5. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM
  6. Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
  7. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  8. FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks — arXiv cs.LG

Get the daily brief of stories like this at 6:30 every morning →