AINewsnow

Routing by task difficulty: the numbers that changed how our AI company spends on models

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

Until recently we spent on language models the way most teams do. Pick the strongest model, make it the default, move on to the next fire. Then we instrumented production traffic and looked at where the money actually went. One frontier model, gpt-4o, was carrying 77 percent of our production calls…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-09-03 22:31 · DEV Community — AI
    Routing by task difficulty: the numbers that changed how our AI company spends on models

More stories

  1. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  2. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  3. Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History — New York Times Technology
  4. Anthropic mulls new AI model ahead of IPO to counter OpenAI's GPT-6 Astra, says report: What we know — Mint AI
  5. what's the state of the art recipe for running Qwen3.8-Flash-Next with a pair of 3090s and a ton of system RAM rn? — r/LocalLLaMA
  6. I built an iOS app with Claude code to break out of my usual chord habits and unlock new progressions. — r/ClaudeAI
  7. I ran Claude code and Codex in parallel for 15 days. Here's what I found. — r/AI_Agents
  8. ChatGPT for Word is now available — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →