AINewsnow

Qwen Code 400 "failed to parse grammar" against llama.cpp? Here's the fix

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

If you run Qwen Code against a local llama.cpp server, you may have woken up one morning to every request dying with: API Error: 400 Failed to initialize samplers: failed to parse grammar Before any token comes back. The model never even starts. And here's the annoying part. The same model works fi…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-02 17:05 · DEV Community — AI
    Qwen Code 400 "failed to parse grammar" against llama.cpp? Here's the fix

More stories

  1. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  2. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  3. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  4. My Version of Jev running locally, playing doom. — r/LocalLLM
  5. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  6. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA
  7. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  8. Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →