Psychological methods reveal major weaknesses in AI security testing
This story is from 2026-08-22. It is preserved in the archive; the latest stories are on the live feed.
Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to day. The study also off…
Read the full story at The Decoder ↗
Timeline · 1 report
- 2026-08-22 07:00 · The Decoder
Psychological methods reveal major weaknesses in AI security testing