Probe Generalization as Subspace Selection for OOD Deception Detection
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.02893v1 Announce Type: new Abstract: Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out decept…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv cs.CL
Probe Generalization as Subspace Selection for OOD Deception Detection