I Ran the Same 40 Prompts Through Qwen2.5 and Qwen3. Here's the Script and Results.
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
Why I Didn't Just Trust the Benchmarks Qwen3 's published benchmarks look like a clean win over Qwen2.5 — real gains on MMLU-Pro, MATH, and coding tasks, plus a much larger training set (roughly 36 trillion tokens versus Qwen2.5's 18 trillion) and support for far more languages. On paper, swapping…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-08-26 08:38 · DEV Community — AI
I Ran the Same 40 Prompts Through Qwen2.5 and Qwen3. Here's the Script and Results.