AINewsnow

Best current Qwen Flash Next Q4-ish? + worth using?

Im running a 5090 and 64gb of ram, so im limited on what I can run. I have currently been able to fit the following - Atomic Q4_k_m 4.27bpw @ 31 layers offload Swift IQ4_xs @ 32 layers offload. Im about to try the Unsloth IQ4_xs as well. I could get a "bigger" (non IQ) quant for atomic because its…

Read the full story at r/LocalLLaMA ↗

Timeline · 6 reports

  1. 2026-09-28 07:49 · r/LocalLLaMA
    If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.
  2. 2026-09-27 20:18 · r/LocalLLM
    Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system
  3. 2026-09-26 21:11 · r/huggingface
    Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle.
  4. 2026-09-26 19:54 · r/LocalLLaMA
    Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much?
  5. 2026-09-26 13:02 · r/LocalLLM
    We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it.
  6. 2026-09-26 08:52 · r/LocalLLaMA
    Best current Qwen Flash Next Q4-ish? + worth using?

More stories

  1. Qwen 3.8 is a workhorse — r/LocalLLaMA
  2. 85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri — r/LocalLLaMA
  3. Another "Harness matters" post (codex cli > pi and opencode) — r/LocalLLaMA
  4. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  5. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy — r/LocalLLaMA
  7. Which provider actually wins on pure affordability right now for gemma qwen gpt oss and deepseek under one roof — r/AI_Agents
  8. When is the next generation of "B tier" models releasing? — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →