Can LLMs Actually Audit Code, or Just Fix Commas? A 12-Task Security & Jailbreak Benchmark
This is a submission for the Kaggle Benchmarking Challenge . Public leaderboards love telling us how well models solve LeetCode problems or pass high school exams. But in real-world software engineering, syntactically valid code that runs without errors is often the most dangerous code in productio…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-01 19:05 · DEV Community — Machine Learning
Can LLMs Actually Audit Code, or Just Fix Commas? A 12-Task Security & Jailbreak Benchmark