AINewsnow

Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much?

I wish Qwen also released dataset and method to fully train a model ourselves but it is what it is. However, I come here with my stupid question because someone can answer it better. And will the model still be an over thinker of faster inference will make up for that.

Read the full story at r/LocalLLaMA ↗

Timeline · 5 reports

  1. 2026-09-29 10:21 · r/LocalLLM
    Qwen flash next on 12+16gb vram, and 32gb ram viable?
  2. 2026-09-28 07:49 · r/LocalLLaMA
    If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.
  3. 2026-09-27 20:18 · r/LocalLLM
    Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system
  4. 2026-09-26 21:11 · r/huggingface
    Abbiamo inserito 100 informazioni nella tabella engrammatica di Qwen 3.8 Flash Next e abbiamo creato un sito web per illustrarle.
  5. 2026-09-26 19:54 · r/LocalLLaMA
    Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much?

More stories

  1. Qwen 3.8 is a workhorse — r/LocalLLaMA
  2. 85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri — r/LocalLLaMA
  3. Another "Harness matters" post (codex cli > pi and opencode) — r/LocalLLaMA
  4. A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling. — r/artificial
  5. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  6. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  7. Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy — r/LocalLLaMA
  8. Which provider actually wins on pure affordability right now for gemma qwen gpt oss and deepseek under one roof — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →