Training on probes: What's going on
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
TL;DR If you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes dishonesty, the model will change its internal representation to evade the probe. Duh. If you train against the probe after all other training, it wor…
Read the full story at Alignment Forum ↗
Timeline · 1 report
- 2026-09-08 17:06 · Alignment Forum
Training on probes: What's going on