AINewsnow

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier post , I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around ~23–27 tok/s and slowed down as context grew. Today I'm releasing Slipstream : a c…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-01 20:21 · r/LocalLLaMA
    Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

More stories

  1. Qwen Flash Next MTP work restarted — r/LocalLLaMA
  2. RTX 5070 TI and 32GB RAM DDR4, WHAT IS THE BEST VIDEO MODEL I COULD USE? — r/comfyui
  3. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  4. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  5. FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users. — r/LocalLLaMA
  6. The Rise of Overfit Inference Engines — r/LocalLLaMA
  7. Flash next rig born from mining parts. — r/LocalLLaMA
  8. Imma just say it, Strata absolutely clowned llama.cpp — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →