Qwen3.8-Flash-Next IQ1_S on a single 5070 (12GB VRAM)
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
Guys, if you have low VRAM, you should start with a small quant first to verify that everything works correctly. command line: .\bin\Release\llama-server.exe -m J:\llm\models\Qwen3.8-Flash-Next-UD-IQ1_S-00001-of-00003.gguf --parallel 1 -c 10000 results: 2.08.059.754 I slot print_timing: id 0 | task…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-29 14:47 · r/LocalLLM
Qwen3.8-Flash-Next IQ1_S on a single 5070 (12GB VRAM)