Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken
I was somewhat disappointed with the performance of Qwen 3.8 Flash Next on my single RTX 5090 using llama.cpp. One issue is that llama.cpp still has no Expert Caching implemented for MoE models. There are various PRs and discussions (see here , for example), but nothing is merged yet. Then I stumbl…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-20 18:16 · r/LocalLLM
I ran claude vs Qwen 3.8 Flash Next - 2026-09-19 22:55 · r/LocalLLaMA
Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken