Geek News Central Podcast

← Geek News Central Podcast10. Juli · 46 Min.

AI Distillation: How Frontier Models Teach Each Other #1870

AI Distillation: How Frontier Models Teach Each Other #187010. Juli46 Min.

In this episode, Ray Cochrane breaks down AI distillation, the teacher-student technique frontier labs now lean on to train smaller, cheaper models. He also covers GPT-5.6’s government-vetted rollout, Claude Sonnet 5 landing on AWS, Maryland’s two-year data center pause, and Microsoft’s climbing carbon numbers. Finally, he wraps with Apple’s $30 billion Broadcom deal, Meta’s tamper-proof recording light, Michigan’s parasite outbreak, and a simulation that erased a super El Niño.

– Want to start a podcast? Its easy to get started! Sign-up at Blubrry

– Thinking of buying a Starlink? Use my link to support the show.

Subscribe to the Newsletter.

Email Ray if you want to get in touch!

Like and Follow Geek News Central’s Facebook Page.

Support my Show Sponsor: Best Godaddy Promo Codes

Get 1Password

Full Summary

Cochrane opens with a quick personal update. Longer days have him outdoors, including a float trip on the Sandy River at Dabney State Park, where he found clearer water, clay-like sand, and easy footing. Next week brings both a move and a trip home, so he is stocking up on Trader Joe’s “Power Berries” and IKEA bags at his mom’s request. Then he turns to the lead story.

AI Distillation Explained: How Frontier Models Teach Each Other

Cochrane’s featured story comes from Hugging Face engineer Sergio Paniego. Distillation is teacher-student training for AI: a capable model generates the training signal, and a smaller student learns to match it. The classic off-policy version compresses giant models into cheap students, either through soft labels or piles of worked answers. Google’s Gemma models and DeepSeek’s R1-Distill line were built exactly this way.