The Hidden Problem With AI Safety Tests: When Training Changes How Models Answer, Not What They Answer
AI safety scores can lie: new research finds RLVR training warps how models respond to tests, not just what they do, producing false alarms or missed.