AINewsnow

Why the second question in a chat is 80–96% cheaper than the first

If you run a local model behind an API and your client sends a multi-turn conversation, you're re-prefilling the entire history on every turn. Turn 10 re-processes turns 1 through 9. On an edge box that's the biggest single waste I know of, and it's what the naive setup does by default. So keep the…

Read the full story at r/machinelearningnews ↗

Timeline · 1 report

  1. 2026-10-03 00:10 · r/machinelearningnews
    Why the second question in a chat is 80–96% cheaper than the first

More stories

  1. NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
  2. Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
  3. Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
  4. Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
  5. Google announces Gemini 4 Argon AI model, but you can't use it yet — Ars Technica AI
  6. The latest AI news we announced in September 2026 — Google Gemini Blog
  7. Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
  8. OpenAI Fires Researchers for Allegedly Sharing Information with AI Safety Group — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →