Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the mem…
Read the full story at arXiv cs.CL ↗