Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions
UPDATE: My usage refers to the NVFP4 fork of strata: https://github.com/sergqwer/strata-nvfp4 with qwen3.8-flash-next-uncensored-nvfp4 I've been daily driving a number of 27b models over the past couple of weeks (notably the swift 1.5 dflash2 variant). My 27b experience has been reasonable, but I f…
Read the full story at r/LocalLLM ↗
Timeline · 7 reports
- 2026-10-11 08:23 · r/LocalLLaMA
What setup do you use, which harness? and how do you customize it to run qwen 3.8 flash next swift with a high context length and a reasonable amount of tok/sec? - 2026-10-11 02:20 · r/LocalLLM
Qwen3.8-Flash-Next on an RTX 5090: 128GB RAM: IQ3_XXS vs IQ3_S vs UD-IQ4_XS — Speed and Accuracy with Strata - 2026-10-10 19:21 · r/LocalLLaMA
Benefits of using bigger models than Qwen 3.8 flash next? - 2026-10-10 00:14 · r/LocalLLaMA
Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix - 2026-10-09 18:41 · r/LocalLLaMA
Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact) - 2026-10-09 12:05 · r/LocalLLaMA
rtx 4090 + huawei atlas duo for qwen flash next ? - 2026-10-09 07:05 · r/LocalLLM
Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions