AINewsnow

Got Qwen Flash Next Q4 running on my Mac Mini M5 64GB with ssd streaming

This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.

Bit of a side project I wanted to share. The metrics are 17.5tks decode, 360tks prompt processing based testing against my normal ai usage. I tested a couple of new things others haven’t done (at least that I’ve seen). Setup a carousel buffer for streaming in experts for prompt processing which got…

Read the full story at r/LocalLLM ↗

Timeline · 12 reports

  1. 2026-10-07 22:26 · r/LocalLLM
    Is anyone running Qwen Flash Next on Strata with q6 or higher?
  2. 2026-10-07 22:02 · r/LocalLLM
    Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next
  3. 2026-10-07 20:43 · r/LocalLLaMA
    Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?
  4. 2026-10-06 14:58 · r/LocalLLaMA
    Qwen3.8-Flash-Next on Strata
  5. 2026-10-06 10:37 · r/LocalLLaMA
    NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s
  6. 2026-10-05 17:38 · r/LocalLLM
    Strata: Qwen 3.8 Flash next in loop
  7. 2026-10-05 12:28 · r/LocalLLM
    Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance
  8. 2026-10-05 10:50 · r/LocalLLaMA
    MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas?
  9. 2026-10-05 07:27 · r/LocalLLM
    4x 3080 20GB (modded, alibaba) + Strata (Qwen Flash Next 125B IQ3_XXS) = 105 t/s generation 5000 prompt processing on 150k context
  10. 2026-10-05 07:21 · r/LocalLLM
    Upgrade existing PC to run Qwen 3.8 Flash Next via Strata or swap to a Strix Halo?
  11. 2026-10-05 06:44 · r/LocalLLaMA
    I got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming
  12. 2026-10-05 06:35 · r/LocalLLM
    Got Qwen Flash Next Q4 running on my Mac Mini M5 64GB with ssd streaming

More stories

  1. What to know about Mistral's ML4 as it bets on EU sovereignty in the US-China open-weight AI race — Euronews Next
  2. ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5. — r/huggingface
  3. Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance — r/LocalLLM
  4. Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open — r/LocalLLaMA
  5. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM
  6. Qwen Image 2.1 Uncensored MCP — r/StableDiffusion
  7. A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well. — r/LocalLLaMA
  8. FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks — arXiv cs.LG

Get the daily brief of stories like this at 6:30 every morning →