AINewsnow

I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.

This story is from 2026-10-11. It is preserved in the archive; the latest stories are on the live feed.

This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none , it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-11 11:54 · DEV Community — Machine Learning
    I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.

More stories

  1. Daily Driving Qwen 3.8 Flash-Next MoE (NVFP4) on RTX 5090 + 128GB RAM — Telemetry & Impressions — r/LocalLLM
  2. Set up your dot in the ChatGPT mobile app — OpenAI YouTube
  3. Asana cuts model costs 76x in browser tests with GPT-6.1 Sol — OpenAI News
  4. Make a game using GPT-6 with Intelligent UI — OpenAI YouTube
  5. OpenAI caught Russians and Iranians using ChatGPT for influence campaigns — NPR Technology
  6. How Oracle Uses ChatGPT Work to Transform Recruitment — OpenAI YouTube
  7. How Oracle turns days of work into minutes with ChatGPT and Codex — OpenAI News
  8. Introducing GPT-6 in ChatGPT with Intelligent UI — OpenAI YouTube

Get the daily brief of stories like this at 6:30 every morning →