AI Post Transformers

← AI Post Transformers6 dagen geleden

AI Coaching That Preserves Human Skill Development

AI Coaching That Preserves Human Skill Development6 dagen geleden

This episode examines "AI Coaching for Accelerating Human Skill Development with Reinforcement Learning," a University of Pennsylvania and Johns Hopkins paper that challenges the assumption that AI assistance always benefits learners, arguing that copilots optimized for immediate task success can quietly prevent people from ever mastering a skill independently. The discussion traces the tension between over-assistance, which produces clean performance but no learning, and under-assistance, which produces uninstructive failure, connecting this to Manu Kapur's productive-failure research and decades of shared-control work from Dragan, Srinivasa, Reddy, and Levine. The core innovation discussed is framing coaching as a non-cooperative dynamic game rather than a cooperative one: the learner optimizes for immediate performance while the coach is trained on Value of Independence, a counterfactual measure of how well the human would perform if the AI were removed entirely. The hosts debate whether this framing is truly adversarial or just a shared goal on different timescales, concluding the reward structures can genuinely conflict moment-to-moment. The episode also situates the work against closest prior art, including the Cyber Racing Coach's fixed assistance-decay schedule, highlighting why a skill-aware, game-theoretic approach marks a meaningful departure from prior fading-assistance methods.

Sources:

1. AI Coaching for Accelerating Human Skill Development with Reinforcement Learning — Wei Wang, Enlin Gu, Antonio Loquercio, Haimin Hu, Rahul Mangharam, 2026

http://arxiv.org/abs/2606.25337v1

2. A Policy-Blending Formalism for Shared Control — Anca D. Dragan, Siddhartha S. Srinivasa, 2013

https://scholar.google.com/scholar?q=A+Policy-Blending+Formalism+for+Shared+Control

3. Shared Autonomy via Hindsight Optimization — Shervin Javdani, Siddhartha S. Srinivasa, J. Andrew Bagnell, 2015

https://scholar.google.com/scholar?q=Shared+Autonomy+via+Hindsight+Optimization

4. Shared Autonomy via Deep Reinforcement Learning — Siddharth Reddy, Anca D. Dragan, Sergey Levine, 2018

https://scholar.google.com/scholar?q=Shared+Autonomy+via+Deep+Reinforcement+Learning

5. Highway Driving with a Semi-Autonomous Alliance of Human and Machine (Shared Control for Highway Driving) — Jake Brawer / Vaibhav Gupta / related shared-control-for-driving lineage (e.g., Broad, Arkin, Ratliff, Howard, Argall), 2017-2019

https://scholar.google.com/scholar?q=Highway+Driving+with+a+Semi-Autonomous+Alliance+of+Human+and+Machine+%28Shared+Control+for+Highway+Driving%29