Distillation for Incrimination and Distillation for Capabilities
arXiv:2610.11012v1 Announce Type: new Abstract: Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.AI
Distillation for Incrimination and Distillation for Capabilities