2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
What I have: - CPU: EPYC 7551 (32c/64T, Zen 1) - Board: Supermicro H11SSL-i (SP3), Rev 2.0 - RAM: 128 GB DDR4-2133 (all 8 channels full) - GPU: 2x RTX 3090 (48 GB total, PCIe 3.0) - 1500 W PSU What I run: - Qwen3-Flash-Next (177B total / ~6B active MoE, IQ4_XS) on Ilama.cpp. Experts live in system…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-12 16:52 · r/LocalLLaMA
2×RTX 3090 + EPYC box running qwen3.8-flash-next at ~38 tok/s