Dual 3090 Qwen 3.8 Flash Test
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Hardware: i9-12900K, 128GB RAM (DDR4), 2x RTX 3090 24GB. I tested Qwen3.8 Flash-Next UD-Q4_K_XL vs UD-IQ4_XS locally. Q4_K_XL needed ~27 expert layers on CPU and topped out around 8.5–8.8 tok/s. It survived a ~96K agent context, but my first serious repo/tool task took ~30 minutes and never produce…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-01 02:53 · r/LocalLLM
Dual 3090 Qwen 3.8 Flash Test