AINewsnow

Your LLM bill is 80% hidden thinking tokens. One parameter fixes it.

This story is from 2026-10-11. It is preserved in the archive; the latest stories are on the live feed.

I asked three cheap, popular models the same simple question: "Write a 150-word explanation of how HTTP caching headers work, for a junior developer." Each one answered in about 150 words. Each one billed me for 800 to 2,500 output tokens . The difference is reasoning tokens: "thinking" the model d…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-11 07:37 · DEV Community — AI
    Your LLM bill is 80% hidden thinking tokens. One parameter fixes it.

More stories

  1. Microsoft unveils Microsoft-Decision-1, a fast decision-scoring model trained on Qwen3.5-9B, and says it will soon rebase it on MAI, OpenAI, and other models (Achint Srivastava/Command Line) — Techmeme
  2. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  3. Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude — The Guardian AI
  4. Philadelphia police receive false homicide tip from Anthropic AI model — The Hill Technology
  5. Make a game using GPT-6 with Intelligent UI — OpenAI YouTube
  6. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  7. Sophos cuts threat investigation time by 96% with OpenAI Daybreak — OpenAI News
  8. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog

Get the daily brief of stories like this at 6:30 every morning →