I collected every single LLM coding benchmark, and computed their Intelligence Density
This story is from 2026-08-30. It is preserved in the archive; the latest stories are on the live feed.
The intelligence in my context is an aggregate index, I called the Agentic Coding Index , across most relevant agentic coding benchmarks: SWE-bench Pro, DeepSWE v1.1, Terminal-Bench (v4, v3, v2.1), Code Arena Elo, and LiveCodeBench v6. Intelligence/Parameter=Scale x (Agentic Index / Norm) ^ (Super_…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-30 22:20 · r/LocalLLaMA
I collected every single LLM coding benchmark, and computed their Intelligence Density