Breaking the Subword Boundary: Probing Architectural Vulnerabilities in Modern LLMs (ASTRAL-Bench on Kaggle)
This is a submission for the Kaggle Benchmarking Challenge When we look at standard LLM leaderboards, everything seems solved. Models boast 90%+ on MMLU and 85%+ on GSM8K. We are told these models possess advanced reasoning, can act as autonomous enterprise agents, and are ready to manage multi-mil…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-08 16:27 · DEV Community — Machine Learning
Breaking the Subword Boundary: Probing Architectural Vulnerabilities in Modern LLMs (ASTRAL-Bench on Kaggle)