Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?
I was actually pretty happy with my Qwen3.8-27b setup, and I'd been tinkering with Ninfer to have a version that was "fast but maybe a bit stupid" and the speed was nice to have as a backup. But I was curious how the Flash-Next version might work, after I learned it didn't need to all fit in VRAM t…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-24 21:30 · r/LocalLLaMA
Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?