Routing by task difficulty: the numbers that changed how our AI company spends on models
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
Until recently we spent on language models the way most teams do. Pick the strongest model, make it the default, move on to the next fire. Then we instrumented production traffic and looked at where the money actually went. One frontier model, gpt-4o, was carrying 77 percent of our production calls…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-03 22:31 · DEV Community — AI
Routing by task difficulty: the numbers that changed how our AI company spends on models