AINewsnow

My benchmark scored GPT-5.4 mini 1.00 by silently skipping the 3 questions it failed

This is a submission for the Kaggle Benchmarking Challenge GPT-5.4 mini got a perfect 1.00 on my benchmark. Then, in a separate repeats notebook, I asked it the same three questions five times each with the cache off, and it failed all 15. The model didn't change. My benchmark had four bugs. Two ma…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-09 08:48 · DEV Community — Machine Learning
    My benchmark scored GPT-5.4 mini 1.00 by silently skipping the 3 questions it failed

More stories

  1. GPT-6 and Intelligent UI for everyone — OpenAI News
  2. Sharing AI progress in mathematics — OpenAI News
  3. Claude Pro vs ChatGPT Plus vs Copilot Premium: which one would you choose for this use case? — r/ChatGPTPro
  4. ChatGPT for Teens Is an ‘Unacceptable Risk,’ Watchdog Group Says — CNET AI
  5. Hot take but AI mode is by far the most useful AI out of chatGPT/Claude/Gemini — r/GeminiAI
  6. OpenAI publishes 722 mathematical proofs & manuscripts — r/singularity
  7. How Oracle Uses ChatGPT Work to Transform Recruitment — OpenAI YouTube
  8. How Oracle turns days of work into minutes with ChatGPT and Codex — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →