Prompt Caching: Why cache_control Writes But Never Reads
This story is from 2026-09-12. It is preserved in the archive; the latest stories are on the live feed.
I turned on prompt caching for an agent loop that resends a 12K-token system prompt on every turn. Obvious win, right? Input tokens are the whole bill in a tool loop. The bill went up. Not a little. Roughly a quarter. And nothing in the logs looked wrong. Every request returned 200. Latency was the…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-12 16:50 · DEV Community — AI
Prompt Caching: Why cache_control Writes But Never Reads