Qwen3.8-Flash-Next Q6 on MSI MEG Z790 ACE and 6 consumer GPUs (RTX 3090): ~90 tok/s shallow, ~40 tok/s at 80k context
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
Hi, a quick datapoint for anyone interested in what Qwen3.8-Flash-Next UD-Q6_K_XL (unsloth) can do on a consumer multi-GPU rig. Maybe there is even more possible but this is the current status :) Hardware: MSI MEG Z790 ACE, i9-13900K, 96 GB DDR5@4400MHz, 1× RTX 4090 + 5× RTX 3090 = 144 GB VRAM . 40…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-09-12 16:52 · r/LocalLLaMA
2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s - 2026-09-11 23:11 · r/LocalLLM
Qwen3.8-Flash-Next Q6 on MSI MEG Z790 ACE and 6 consumer GPUs (RTX 3090): ~90 tok/s shallow, ~40 tok/s at 80k context