AINewsnow

Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?

5080 16GB, 64GB RAM. Currently on 3.6 35B-A3B (Unsloth quant, fits fully in VRAM) and happy with it. Anyone tried Flash-Next at Q2/Q3 on a similar setup? Worth the switch, or should I stay put? Asking before committing to a 30GB+ download.

Read the full story at r/LocalLLaMA ↗

Timeline · 6 reports

  1. 2026-10-09 07:05 · r/LocalLLM
    Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions
  2. 2026-10-08 11:07 · r/LocalLLaMA
    Halogen + Qwen Flash Next keeps getting better
  3. 2026-10-08 10:55 · r/LocalLLaMA
    Qwen 3.8 Flash Next is so much fun for three.js
  4. 2026-10-07 22:26 · r/LocalLLM
    Is anyone running Qwen Flash Next on Strata with q6 or higher?
  5. 2026-10-07 22:02 · r/LocalLLM
    Water cooling 2x RTX 6000 Pro Workstations; temps down 50% at max load, 500W max under Qwen 3.8 Flash Next
  6. 2026-10-07 20:43 · r/LocalLLaMA
    Qwen 3.8 Flash-Next on 64GB RAM + 16GB VRAM, worth it over a fully-loaded 3.6 35B-A3B?

More stories

  1. What to run overnight and such when llm idle? — r/LocalLLM
  2. Local Qwen3.8-Flash on a DGX Spark vs Claude Opus 5.5: 21 graded tasks, 3 harnesses. Matches Opus on everyday coding, 70% on hard tasks, and the harness settings matter more than you'd think — r/LocalLLM
  3. 113 Decision Models in 3 Weeks (Mostly Qwen and Gemma Fine-Tunes): A Deep Dive — r/AI_Agents
  4. Ultimate web scraper — r/AI_Agents
  5. Switching from Claude Code to local Qwen for Android dev — can smaller local models keep up? — r/LocalLLM
  6. I built a root-cause tool with Claude Code, then built a validator to catch the LLM inside it when it's confidently wrong. Honest numbers: 60% / 20% — r/AI_Agents
  7. What to know about Mistral's ML4 as it bets on EU sovereignty in the US-China open-weight AI race — Euronews Next
  8. GPT-6 and Intelligent UI for everyone — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →