How do you stop evals from becoming a cheat sheet for the prompt?
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
We kept tuning a prompt against the same small golden set until every check passed. But then the paraphrases failed in ways the score never predicted... Most synthetic cases shared one template, near duplicates leaked across train and holdout and the scorer rewarded memorized formatting more than i…
Read the full story at r/PromptEngineering ↗
Timeline · 1 report
- 2026-09-14 20:07 · r/PromptEngineering
How do you stop evals from becoming a cheat sheet for the prompt?