Use failed runs to generate prompt edits, then judge the edits on fresh tasks
A trace that shows a model fighting an instruction is a good reason to edit the prompt. It isn't yet evidence that the edited prompt will work better elsewhere. Reef Infra's GEPA recipe connects those two jobs. It keeps an archive of prompt candidates, selects a parent, and reflects on that parent'…
Read the full story at r/PromptEngineering ↗
Timeline · 1 report
- 2026-09-20 07:50 · r/PromptEngineering
Use failed runs to generate prompt edits, then judge the edits on fresh tasks