CoT controllability evals seem very under-elicited
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
The CoTControl eval asks reasoning models to follow formatting constraints in their chain-of-thought (e.g. write in all lowercase, avoid a specific word) while solving questions. Models seem to mostly be pretty bad at this: recent models score between 0-30% with the exception of Mythos Preview [1]…
Read the full story at Alignment Forum ↗
Timeline · 1 report
- 2026-09-11 17:12 · Alignment Forum
CoT controllability evals seem very under-elicited