OK I give up, how do I turn down reasoning in vLLM / Pi
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
I've been getting great results from Qwen3.8 in llama.cpp by passing in this on startup: --reasoning-effort low Now I keep seeing posts about how much better vLLM is, and it is, like 2x the tokens/sec, but I cannot figure out how to set reasoning-effort to low. I mean, I can when I call it directly…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-06 17:30 · r/LocalLLM
OK I give up, how do I turn down reasoning in vLLM / Pi