When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations
arXiv:2609.22337v1 Announce Type: new Abstract: Linear probing is the standard instrument for detecting social biases in the hidden representations of large language models. Yet reported probe accuracies come almost exclusively from \emph{counterfactual} evaluations in which every input carries an…
Read the full story at arXiv stat.ML ↗