Three Buddy Problem

← Three Buddy Problem23 Jul · 2 h 16 min

OpenAI's models breached Hugging Face, reward hacking ethics, benchmarking fast16

OpenAI's models breached Hugging Face, reward hacking ethics, benchmarking fast1623 Jul2 h 16 min

(Presented by Thinkst Canary: Most Companies find out way too late that they’ve been breached. Thinkst Canary changes this. Deploy Canaries and Canarytokens in minutes and then forget about them. Attackers tip their hand by touching ’em giving you the one alert, when it matters. With zero admin overhead and almost no false-positives, Canaries are deployed (and loved) on all 7 continents.)

Three Buddy Problem - Episode 106: We dig into the news that OpenAI's models were the "autonomous agent" that breached Hugging Face, escaping a sandbox through a zero-day to cheat on a cyber benchmark, then getting spun into a partnership announcement. We argue about the implications of the incident, the PR masterclass, the absence of ethics and human oversight, and calls for "kill switches" to mitigate "AI lab leaks."

Plus, SentinelLabs' new fast16 reverse-engineering benchmark, where GPT-5.6 Sol was the only public model to go the distance.

Cast: Juan Andres Guerrero-Saade, Ryan Naraine and Costin Raiu.

Timestamps:

0:00 Introductory banter

5:24 OpenAI admits it was the Hugging Face "hacker"

10:06 What’s ExploitGym and who's on top of the leaderboard

12:59 Reward hacking: Did anyone train this thing not to cheat?

19:35 Marketing stunt or real incident? The zero-day in the package proxy

26:43 Was OpenAI already plugged into Hugging Face?

29:17 Paperclips, kill switches, and "going rogue"