Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation
This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.35824v1 Announce Type: new Abstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-30 04:00 · arXiv cs.CL
Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation