AINewsnow

UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh

This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.

Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results…

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-09-16 13:58 · r/LocalLLaMA
    MiniCPM5-2B vs. Spark-X2.5-4B / -64% thinking, x1.5 speed while keeping the accuracy of xhigh
  2. 2026-09-14 15:57 · r/LocalLLaMA
    UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh

More stories

  1. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  2. Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis — r/LocalLLM
  3. US government website used Chinese model the FBI called "malicious" — Ars Technica AI
  4. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  5. Qwen Developers on X: "Qwen-Image 2.1 is going open source" — r/StableDiffusion
  6. Ternary Bonsai 2 27B — r/LocalLLaMA
  7. 10 hours left fo Qwen Image 2.1 Public Open Source Release — r/comfyui
  8. M2 Mac ultra128gb Qwen flash next — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →