Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s
Oh My Pi + vLLM on two 3090s. Average wait per turn went from 28s to 7s, mostly from changing omp settings: explicit effort level on every role (unset ones defaulted to xhigh) thinking_token_budget of 7500 maxTokens 8k → 32k (file writes were getting cut off) tool output over 10 KB goes to a file m…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-19 01:25 · r/LocalLLaMA
Tuning Qwen 3.8 27B and OMP as a coding agent on 2× 3090s