Linear Digressions

← Linear Digressions7 sept · 32 min

Constitutional AI

Constitutional AI7 sept32 min

How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to be both useful and safe by giving them an explicit set of principles to reason from. It's a fascinating look at how alignment research is evolving beyond simple human feedback — and what it means to give an AI something like a conscience.

Links:

Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022)

https://arxiv.org/abs/2212.08073

Claude's Constitution

https://www.anthropic.com/constitution

Anthropic, "Teaching Claude Why" (2026)

https://www.anthropic.com/research/teaching-claude-why