Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant
I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier post , I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around ~23–27 tok/s and slowed down as context grew. Today I'm releasing Slipstream : a c…
Read the full story at r/LocalLLaMA ↗