Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15…
Read the full story at r/LocalLLaMA ↗
Timeline · 5 reports
- 2026-10-04 20:23 · r/LocalLLM
Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram - 2026-10-03 20:15 · r/LocalLLM
Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD - 2026-10-03 20:06 · r/LocalLLaMA
Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD - 2026-10-03 08:33 · r/ArtificialInteligence
Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM - 2026-10-03 06:49 · r/LocalLLaMA
Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4