AINewsnow

Spent a week optimizing throughput, then realized I was measuring the wrong thing

Chased tokens per second for days. Then I looked at wall clock across a whole multi-turn session instead of per request and the picture was completely different. Most of the time wasn't generation at all. Has anyone else had the thing you were optimizing turn out to be the small half of the problem…

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-09-24 00:12 · r/learnmachinelearning
    Spent a week optimizing throughput, then realized I was measuring the wrong thing

More stories

  1. GPT-6 Sol and Luna now available on AI Gateway — Vercel Blog
  2. Bringing Private Processing to Meta AI Glasses — Engineering at Meta
  3. Gemini 3.8 text-to-speech models now available on AI Gateway — Vercel Blog
  4. Sam Altman’s remarks at the United Nations Security Council — OpenAI News
  5. Alibaba unveils new AI chip to challenge NVIDIA, plans Qwen models with up to 10 trillion parameters — Mint AI
  6. OpenAI Agent Hacked Australian Government Website — Wall Street Journal Technology
  7. Meta Connect 2026 live: Updates from Mark Zuckerberg's keynote on AI glasses, VR and more — Engadget
  8. How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →