Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?
5080 16GB, 64GB RAM. Currently on 3.6 35B-A3B (Unsloth quant, fits fully in VRAM) and happy with it. Anyone tried Flash-Next at Q2/Q3 on a similar setup? Worth the switch, or should I stay put? Asking before committing to a 30GB+ download.
Read the full story at r/LocalLLaMA ↗
Timeline · 6 reports
- 2026-10-09 07:05 · r/LocalLLM
Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions - 2026-10-08 11:07 · r/LocalLLaMA
Halogen + Qwen Flash Next keeps getting better - 2026-10-08 10:55 · r/LocalLLaMA
Qwen 3.8 Flash Next is so much fun for three.js - 2026-10-07 22:26 · r/LocalLLM
Is anyone running Qwen Flash Next on Strata with q6 or higher? - 2026-10-07 22:02 · r/LocalLLM
Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next - 2026-10-07 20:43 · r/LocalLLaMA
Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?