The biggest improvement in my skill evaluation came from a skill that was never invoked
This story is from 2026-09-20. It is preserved in the archive; the latest stories are on the live feed.
Last week I measured three agent skills twice over: once with Claude Code's built-in claude plugin eval , and once with a runner I maintain. Same skill text, same task prompts, same rubrics. I expected the interesting part to be which tool scored higher. It wasn't. The interesting part was a result…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-20 17:59 · DEV Community — AI
The biggest improvement in my skill evaluation came from a skill that was never invoked