Identifying Introspection From the Inside
arXiv:2610.07186v1 Announce Type: new Abstract: Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.CL
Identifying Introspection From the Inside