200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)
Strata runs Qwen3.8-Flash-Next on gaming PCs, but with 32 GB of RAM its low-RAM mode only works on one GPU. Split the model across two cards and the experts that don't fit in VRAM get read from the SSD. My 4060 Ti sat next to the 5080 doing nothing. So I forked it. The RAM copy of the experts now w…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-03 14:50 · r/LocalLLM
200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)