Future of Life Institute Podcast

← Future of Life Institute Podcast28 ago · 59 min

Why AI Hacking Is Becoming Hard to Control (with Benjamin Weinstein-Raun)

Why AI Hacking Is Becoming Hard to Control (with Benjamin Weinstein-Raun)28 ago59 min

Benjamin Weinstein-Raun is head of research at Palisade Research. He joins the podcast to discuss how advanced AI systems are changing cybersecurity. We examine rapid gains in autonomous hacking, recent incidents where models broke out of test environments, and why training can reward cheating-like behavior. The conversation covers open-weight model risks, whether AI can help secure software, personal security steps, and the need for international coordination.

LINKS:

Benjamin Weinstein-Raun ProfileBenjamin Weinstein-Raun Website

CHAPTERS:

(00:00) Episode Preview

(01:06) Opening cyber threats

(02:15) Measuring AI progress

(04:54) Open weights floor

(08:01) After Mythos moment

(14:22) Reward hacking setup

(24:47) Artifactory swarm details

(32:32) Defensive model dilemmas