UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-16 13:58 · r/LocalLLaMA
MiniCPM5-2B vs. Spark-X2.5-4B / -64% thinking, x1.5 speed while keeping the accuracy of xhigh - 2026-09-14 15:57 · r/LocalLLaMA
UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh