Why the second question in a chat is 80–96% cheaper than the first
If you run a local model behind an API and your client sends a multi-turn conversation, you're re-prefilling the entire history on every turn. Turn 10 re-processes turns 1 through 9. On an edge box that's the biggest single waste I know of, and it's what the naive setup does by default. So keep the…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-10-03 00:10 · r/machinelearningnews
Why the second question in a chat is 80–96% cheaper than the first