Qwen3.8-Flash-Next drew a self-portrait site from one prompt on an RTX 3060, then signed it “Claude”
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
I gave a local model one prompt and walked away. No system prompt, default sampling. Setup: Qwen3.8-Flash-Next UD-IQ3_XXS, 82 GB file on a machine with 62 GB RAM and a 12 GB 3060, llama.cpp with the built-in MTP head. It doesn’t fit in memory; it works because the hot experts stay in page cache and…
Read the full story at r/LocalLLM ↗
Timeline · 4 reports
- 2026-09-06 19:15 · r/LocalLLM
CPU only, 64 GB DDR5: Qwen3.8-Flash-Next UD-Q3_K_XL - 2026-09-05 19:53 · r/LocalLLM
Qwen3.8-Flash-Next workcase part 2: after the self-portrait, a painting and a map sheet, both unattended on a 3060 - 2026-09-05 12:52 · r/LocalLLaMA
Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal? - 2026-09-05 08:08 · r/LocalLLM
Qwen3.8-Flash-Next drew a self-portrait site from one prompt on an RTX 3060, then signed it “Claude”