LLMs carry a judgment signal you can read with a linear probe (and why steering it fails)
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
LLMs carry a judgment signal you can read with a linear probe (and why steering it fails) LLM hidden states encode reward-related information: roughly, whether the current trajectory is heading toward a correct answer. We reproduced this on our own hardware, added the controls we felt were missing,…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-07 03:32 · DEV Community — Machine Learning
LLMs carry a judgment signal you can read with a linear probe (and why steering it fails)