Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected)
TL;DR — Qwen3.8-27B is a hybrid model (Gated DeltaNet linear attention + full attention). Unsloth Dynamic GGUFs of it (Q4_K_XL, Q6_K_XL) work fine for everyday chat, but whenever I let it think deeply on a long task, it never converges: no closing token, endless tail-looping, or the engine just die…
Read the full story at r/LocalLLM ↗