I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.
This story is from 2026-10-11. It is preserved in the archive; the latest stories are on the live feed.
This is a submission for the Kaggle Benchmarking Challenge I gave gpt-5.4-mini a logic puzzle: seven people, seven days, ten clues, "Who gives the talk on Friday?" With reasoning effort set to none , it replied: Cleo FINAL ANSWER: Cleo 18 output tokens. $0.00024. Wrong. (The answer is Fay.) It gave…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-11 11:54 · DEV Community — Machine Learning
I Turned the Reasoning Dial to 'High' on 4 Models. It Fixed One Thing and Billed Me for Everything.