Tokens per second told me nothing about agent performance
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
I benchmarked a dozen free AI APIs on generation throughput, then plugged the same models into an agent harness and gave them real tasks. The rankings barely overlapped. Two examples from the same platform: Gemma 4 31B — 50.9 tok/s, second fastest I measured. Inside the agent: hung until the 900-se…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-02 09:26 · DEV Community — AI
Tokens per second told me nothing about agent performance