Qwen 3.6 vs 3.5: Same 37 tok/s on RTX 4070, +43% on Frontend Generation
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
The first number I saw on Qwen3.6-35B-A3B was 12 tok/s . I almost hit publish on "Qwen regressed at generation speed" and moved on. The 3.5 baseline on the same RTX 4070 was 34.6 tok/s. A new generation running at a third of the old one would have been a hell of a headline. It was also completely w…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-02 13:00 · DEV Community — AI
Qwen 3.6 vs 3.5: Same 37 tok/s on RTX 4070, +43% on Frontend Generation