I built EvalSeal: reproducibility receipts for LLM evals
I’ve been working on EvalSeal, a small open-source tool for making LLM eval results more trustworthy. The idea is simple: a single eval score is not enough. EvalSeal runs each case multiple times, measures flip rates, captures provenance, and seals the result into a tamper-evident ledger. A finding…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-18 08:06 · r/AI_Agents
I built EvalSeal: reproducibility receipts for LLM evals