Spent a week optimizing throughput, then realized I was measuring the wrong thing
Chased tokens per second for days. Then I looked at wall clock across a whole multi-turn session instead of per request and the picture was completely different. Most of the time wasn't generation at all. Has anyone else had the thing you were optimizing turn out to be the small half of the problem…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-24 00:12 · r/learnmachinelearning
Spent a week optimizing throughput, then realized I was measuring the wrong thing