Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation
arXiv:2609.35824v1 Announce Type: new Abstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CL
Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation