
← AI Bites: The Academic Series14 ago · 22 min
EP 56 | CS224N: Open Frontiers in NLP & The Smart Scaling Era
<p>Welcome to the grand finale of CS224N! In our final episode, featuring insights from Professor Yejin Choi, we tackle the biggest open frontier in AI: what happens when high-quality web data runs out? We explore why the era of brute-force scaling is officially over, and how small language models (1.5B–7B parameters) are using "smart scaling" to out-reason 100B+ giants.</p><p><strong>Key Topics:</strong></p><ul><li><p><strong>The End of Brute-Force Scaling:</strong> Why internet text is the "fossil fuel" of AI, and why Ilya Sutskever says the future belongs to "smart scaling" rather than massive compute budgets.</p></li><li><p><strong>Prolonged RL (ProRL):</strong> How fixing "entropy collapse" with dynamic decoupled clipping allows a tiny 1.5B model (Nemotron-Reasoning-1.5B) to outperform DeepSeek-R1-7B.</p></li><li><p><strong>Prismatic Synthesis:</strong> Using model gradients as "reasoning fingerprints" and the G-Vendi score to generate hyper-diverse synthetic datasets with zero human labels.</p></li><li><p><strong>RLP & Front-Loading Reasoning:</strong> Why baking reasoning directly into the pre-training phase creates structural, compounding advantages that late-stage SFT simply cannot replicate.</p></li><li><p><strong>Unconventional Collaboration:</strong> How open-science initiatives like OpenThoughts3 prove that community-driven collaboration can beat closed-lab pipelines.</p></li></ul><p><strong>Note:</strong> This is an AI-generated discussion created using Google's NotebookLM, based on publicly available Stanford University course material (specifically CS224N) and personal study notes from my learning journey.</p>