MiniCPM5-2B vs. Spark-X2.5-4B / -64% thinking, x1.5 speed while keeping the accuracy of xhigh
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
We gave two small LLMs a real customer-service task. One got 100/100. The other stopped with the job half-done. I wanted to test something more meaningful than benchmark scores: can a small model actually complete a multi-step agent task without handing the work back to the human? The task: A custo…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-16 13:58 · r/LocalLLaMA
MiniCPM5-2B vs. Spark-X2.5-4B / -64% thinking, x1.5 speed while keeping the accuracy of xhigh