Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
arXiv:2610.07023v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learnin…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.AI
Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment