The Peter McCormack Show

← The Peter McCormack Show14 aug · 1 u 10 min

#201 - Connor Leahy - The AI That Escaped: Inside OpenAI's Rogue Agent Incident

#201 - Connor Leahy - The AI That Escaped: Inside OpenAI's Rogue Agent Incident14 aug1 u 10 min

Connor Leahy has spent his career at the frontier of AI, from reverse-engineering GPT-2 as a student and co-founding EleutherAI to building the AI safety company Conjecture. He now leads US policy work at ControlAI, and he returns to the show with a warning that has stopped being theoretical: AI systems are escaping their sandboxes, writing their own zero-days and leaving each other notes on how to break out.

Peter and Connor discuss the OpenAI incident that ended in an attack on Hugging Face sophisticated enough to be mistaken for a state actor, why we understand almost nothing about how these systems work, how reinforcement learning produces models that lie, cheat and manipulate to reach a goal and why the next stage after chatbots and agents is swarms.

But Connor’s argument is that the future is not decided. He believes superintelligence should be treated the way nuclear weapons are treated, banned and verified and deterred and that the real bottleneck is awareness rather than opposition. They discuss what an “aligned” superintelligence would actually mean, why military people understand the threat faster than technologists and what ordinary people can do about it.

- - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - - -

TIMESTAMPS:

00:00:00 - The AI That Escaped Its Sandbox

00:01:52 - Inside The OpenAI Incident

00:05:46 - They Thought It Was China

00:06:43 - We Grow AI, We Don’t Write It

00:11:07 - Trained To Lie, Cheat And Deceive

00:13:48 - Chatbots, Agents, Then Swarms

00:17:57 - The Machines Are Developing Preferences