How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
I’m the author of a new preprint on repeated-query auditing of LLM brand recommendations, and the founder of Rankfor.AI. The practical question: how many times should we repeat a prompt before comparing results? The paper applies generalizability theory: estimate variance components from a pilot, t…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-09-04 06:53 · r/MachineLearning
How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R]