AINewsnow

Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system

Hey everyone. So ive recently got my Qwen 3.8 27b stack running and tuned using ninfer. This is my daily driver, and will continue to be until 4.0 replaces it. I'm happy with it, but i'd still like to see what Qwen 3.8 Flash Next can do and see how that performs compared to ninfer + Qwen 3.8 27b. W…

Read the full story at r/LocalLLM ↗

Timeline · 3 reports

  1. 2026-09-29 10:21 · r/LocalLLM
    Qwen flash next on 12+16gb vram, and 32gb ram viable?
  2. 2026-09-28 07:49 · r/LocalLLaMA
    If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.
  3. 2026-09-27 20:18 · r/LocalLLM
    Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system

More stories

  1. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  2. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  3. Qwen 3.8 is a workhorse — r/LocalLLaMA
  4. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy — r/LocalLLaMA
  6. Layer Extract & Layer Remove Loras For Qwen Image 2.1 — r/StableDiffusion
  7. I built Slopus, a free, open-source desktop app for generating and editing AI videos locally (Minimax H3) — r/StableDiffusion
  8. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →