
← Into AI Safety9 jul · 1 u 29 min
Pretraining Safety w/ Ethan Roland
What if the safest AI models weren't built by adding guardrails after training, but by shaping what gets learned in the first place? Ethan Roland, senior alignment researcher at AE Studio and first author on an ICML 2026 spotlight paper, joins Jacob to talk about gradient routing, a technique that routes dangerous capabilities into isolated parts of a model's architecture where they can be locked or removed entirely. They get into the absorption effect, KYC-style access control frameworks, and what it would actually take for frontier labs to adopt this kind of work before it's needed rather than after.
Chapters
(00:00) - Introduction
(06:39) - Inside AE Studio
(15:26) - China & the Alignment vs. Controllability Framing
(18:23) - Data Filtering & Gradient Routing (Aside)
(30:39) - Mixture of Experts Explained (Aside)
(36:25) - Why Pre-Training Interventions Are Rare
(42:43) - Ethan's Theory of Change
(56:17) - Access Control Governance and KYC (Aside)
(01:04:47) - The Researcher's Role in Policy Advocacy
(01:11:38) - Speed Round