Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
We build an automated security scanner for Solana programs. To know whether it works, we keep a small benchmark: 19 programs with a planted vulnerability each, 7 written to be deliberately clean, and a written description of every planted flaw so a judge can check whether the scanner's findings act…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-02 13:25 · DEV Community — AI
Our benchmark was leaking the answers to the model. The numbers looked fine the whole time.