Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione
arXiv:2609.25049v1 Announce Type: new Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying d…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-23 04:00 · arXiv cs.CL
Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione