
← LessWrong (Curated & Popular)4 days ago · 4 min
[Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine
[Linkpost] "Frontier models still hack on simple variations of alignment evals from early 2025" by Dean Valentine
This is a link post. In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked models to play a game of chess against a chess engine. They found that the new, RLVR'd models cheated on the task by altering the board state about 36% of the time. The experiment received a reasonable amount of circulation, and there were even rumors of skepticism from some lab engineers until they could rerun the evaluation. ...