I built EvalSeal v2.2.0: reproducibility receipts for LLM and agent evals
I’ve been working on EvalSeal, an open-source tool for making LLM eval results easier to trust. The problem I kept running into: A score changed, but I couldn’t always tell whether the model improved, the judge changed, the prompt shifted, or the eval itself was noisy. EvalSeal now runs eval cases…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-28 07:59 · r/AI_Agents
I built EvalSeal v2.2.0: reproducibility receipts for LLM and agent evals