The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.02959v1 Announce Type: cross Abstract: What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the mode…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv stat.ML
The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors