Interconnects

← Interconnects16 Jun · 57 min

Frontier post-training recipe review with Finbarr Timbers

Frontier post-training recipe review with Finbarr Timbers16 Jun57 min

As I’ve been recapping fundamentals of post-training to wrap up my RLHF / Post-training book I knew I needed to get Finbarr Timbers back on the podcast to talk about the state of play. Over the last few months we’ve had many discussions on what we’d need to do to take an Olmo-style recipe to the frontier, supported by Finbarr’s extensive reading of recent model technical reports.

To prepare for this, I put together a summary slide deck on the key post-training recipes historically — the path from InstructGPT to today — and today — the key open frontier models. This deck is summarized below as the technical summary, but we do spend 20-35 minutes on it in the podcast, so watching on YouTube is likely the best experience for this one.

I previously interviewed Finbarr in December of 2024, shortly after the release of o1 and Tülu 3 (and before he joined Ai2) on the “We are so back” era of RL.

Chapters:

* 00:00 Introduction & Olmo reflections

* 06:28 Post-train recipes review (history)

* 23:00 2026’s model recipes (MiMo Flash, DeepSeek V4, GLM 5, Kimi K2.6, etc.)

* 39:05 Open-ended post-training discussions

* 48:22 Career advice in the LLM race

Listen on Apple Podcasts, Spotify, and where ever you get your podcasts. For other Interconnects interviews, go here.

For more educational post-training videos, see the course I’m putting together.

Technical Summary