AINewsnow

Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D

MTP-3 restored prompt caching. Our deeper speculative mode cleared the cache for correctness. Using ordinary state backups at depth 3 let follow-ups reuse it. Better CPU kernels unpack IQ4_XS weights once for several speculative tokens, reducing repeated work. Lower-memory GPU attention processes a…

Read the full story at r/LocalLLM ↗

Timeline · 4 reports

  1. 2026-09-20 03:53 · r/LocalLLM
    Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3)
  2. 2026-09-20 03:44 · r/LocalLLaMA
    Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3)
  3. 2026-09-19 19:44 · r/LocalLLaMA
    Finally got qwen 3.8 q4 k_M 27b runnin slick 40tkps load at 240k inf 4qnl for stable agentic coding on a 7900xtx. I have been finally able to give hermes a task and come back to results. Built a panel to manage my inference servers from hermes.
  4. 2026-09-18 22:38 · r/LocalLLM
    Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  3. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/StableDiffusion
  4. Qwen 3.8 27B running on a single RTX 5090 researches and creates a full animation using only code. — r/artificial
  5. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  6. Qwen Image 2.1 Examples — r/StableDiffusion
  7. Ternary Bonsai 2 27B — r/LocalLLaMA
  8. Get rid of GPT images using QWEN IMAGE 2.1. 🔥👍 — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →