Same Score, Different Habits: What Frontier Models Did Inside VORTEX
This is a submission for the Kaggle Benchmarking Challenge Static benchmarks have a shelf life. Once a question is public, it can end up in a training set, and the score starts measuring memory. I wanted tasks where memory cannot help. So I generated them with code, from a fixed seed, and checked e…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-10 13:18 · DEV Community — Machine Learning
Same Score, Different Habits: What Frontier Models Did Inside VORTEX