Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram
mrdefaultuser@team-green : ~ $ /srv/ai/llama.cpp/build/bin/llama-server \ -m /srv/ai/models/qwen/Qwen3.8-Flash-Next-Q8_0/Qwen3.8-Flash-Next-Q8_0-00001-of-00002.gguf \ -c 262144 \ -np 1 \ -ngl 64 \ --cpu-moe \ --lazy-mode on \ --load-mode auto \ --agent \ --tools all \ --host 0.0.0.0 \ --port 8080 E…
Read the full story at r/LocalLLM ↗
Timeline · 5 reports
- 2026-10-06 05:39 · r/LocalLLaMA
unsloth/Qwen3.8-Flash-Next-GGUF is being updated - 2026-10-06 01:47 · r/LocalLLaMA
Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra - 2026-10-05 23:48 · r/LocalLLM
RX 7600 (8 GB) on Linux: Qwen3.8-Flash-Next (~125B) at 24 tok/s with Strata, Qwen3.6-35B-A3B at 32 tok/s with llama.cpp + MTP. Numbers and how-to - 2026-10-05 14:40 · r/LocalLLM
Uncensored models: what they are and how they work, with a file-by-file hash check of one uncensored Qwen3.8-Flash-Next upload against Qwen's original - 2026-10-04 20:23 · r/LocalLLM
Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram