The A/B that measured nothing: three ways my agent experiment was invalid
This story is from 2026-08-25. It is preserved in the archive; the latest stories are on the live feed.
I set out to test whether an MCP tool should return a rendered sentence or raw SI numbers. The experiment was invalid three times, for three unrelated reasons: the tools had never been called (a gateway tool-policy gap), the two arms were byte-identical at the model (the gateway drops a tool's text…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-08-25 20:16 · DEV Community — AI
The A/B that measured nothing: three ways my agent experiment was invalid