AINewsnow

How do you stop evals from becoming a cheat sheet for the prompt?

This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.

We kept tuning a prompt against the same small golden set until every check passed. But then the paraphrases failed in ways the score never predicted... Most synthetic cases shared one template, near duplicates leaked across train and holdout and the scorer rewarded memorized formatting more than i…

Read the full story at r/PromptEngineering ↗

Timeline · 1 report

  1. 2026-09-14 20:07 · r/PromptEngineering
    How do you stop evals from becoming a cheat sheet for the prompt?

More stories

  1. Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
  2. Anthropic adds support for the AGENTS.md instructions spec to Claude Code; OpenAI contributed AGENTS.md to the Agentic AI Foundation last year (Thomas Claburn/The Register) — Techmeme
  3. Alibaba ships Qwen3.8-Omni-Flash to watch, listen and call tools — r/LocalLLM
  4. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  5. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  6. Introducing Astra for Law — OpenAI News
  7. Newsom signs executive order to explore new AI rules, consider ‘kill switch’ — Politico Technology
  8. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology

Get the daily brief of stories like this at 6:30 every morning →