AINewsnow

Long Responses API runs are costing us more from context than the output itself

I’ve been looking closer at the token usage from one of our longer-running workflows using the Responses API and a big portion of the tokens on the expensive runs aren’t coming from what the model generates. The first couple steps are pretty normal but the workflow uses a few tools and keeps going…

Read the full story at r/OpenAI ↗

Timeline · 1 report

  1. 2026-10-07 16:22 · r/OpenAI
    Long Responses API runs are costing us more from context than the output itself

More stories

  1. Introducing Mistral Large 4 — Mistral AI News
  2. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  3. GPT-6 and Intelligent UI for everyone — OpenAI News
  4. Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China — Wired AI
  5. Sharing AI progress in mathematics — OpenAI News
  6. NVIDIA, Microsoft Kick Off a New Beginning for Windows PCs With RTX Spark and AI Agents — NVIDIA Blog
  7. Surface RTX Spark Dev Box is available for preorder for $5,999 — The Verge AI
  8. Introducing Playground: Create and play custom games — Google AI Blog

Get the daily brief of stories like this at 6:30 every morning →