AINewsnow

Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected)

TL;DR — Qwen3.8-27B is a hybrid model (Gated DeltaNet linear attention + full attention). Unsloth Dynamic GGUFs of it (Q4_K_XL, Q6_K_XL) work fine for everyday chat, but whenever I let it think deeply on a long task, it never converges: no closing token, endless tail-looping, or the engine just die…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-06 05:28 · r/LocalLLM
    Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected)

More stories

  1. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  2. LLM Inference Dashboard — r/LocalLLaMA
  3. Photon Announces $4.5M Seed Round to Help Developers Build AI Agents for iMessage and WhatsApp — AI Insider
  4. Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp — r/LocalLLaMA
  5. Overclocking DDR5 For Faster MoE Prefill and Decode — r/LocalLLaMA
  6. llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more — r/LocalLLaMA
  7. Llama.cpp + WebGPU = agants.html — r/AI_Agents
  8. SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →