Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
arXiv:2610.00320v1 Announce Type: new Abstract: Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, sugges…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.CL
Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning