Dense 27B vs MoE IQ2_XS on one RTX 4090: synthetic bench plus 3 graded real tickets
Rig: RTX 4090 24 GB (450 W), i9-14900K, 32 GB RAM, Linux, desktop on the iGPU. A: Qwen3.8-27B dense on NInfer. 262K context, rk4v4-e8 KV, MTP, 1 slot. B: Qwen3.8-Flash-Next IQ2_XS on Strata 0.1.39. 262K context, q4_0 KV in VRAM, experts in host RAM, MTP. Both models: thinking budget 4096, same prom…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-05 03:52 · r/LocalLLM
Dense 27B vs MoE IQ2_XS on one RTX 4090: synthetic bench plus 3 graded real tickets