Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs
arXiv:2610.09033v1 Announce Type: new Abstract: Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level var…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-08 04:00 · arXiv cs.CL
Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs