Evidence Boundary: when an AI benchmark measures its own grading rules
This is a submission for the Kaggle Benchmarking Challenge In the first Evidence Boundary experiment, Gemini 3.7 Flash scored below Gemini 3.1 Flash-Lite Preview. Removing a single enclosing Markdown fence from responses reversed their order, without changing either model's answers. That was the mo…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-01 09:35 · DEV Community — Machine Learning
Evidence Boundary: when an AI benchmark measures its own grading rules