
← The Cybersecurity Defenders Podcast23 jul · 35 min
AI Chat: The Hugging Face / OpenAI breach — the attacker was the model [340]
In this episode:
• The timeline: Hugging Face's July 16 disclosure, OpenAI's July 21 attribution — and the five days in between when even the victim didn't know an AI did it.
• The attack chain: a malicious dataset abusing two code-execution paths in the dataset-processing pipeline, node-level escalation, credential harvesting and lateral movement — thousands of actions across short-lived sandboxes with self-migrating command-and-control.
• The escape: a zero-day in the eval sandbox's package-registry cache proxy, the single egress control — per OpenAI's own account.
• Motive: the models got "hyperfocused" on winning the benchmark, not stealing data — and whether "no malicious intent" is a fair description or a comforting one.
• What was and wasn't exposed, what to do about your Hugging Face tokens, and why this is not the 2024 Spaces incident or the 2023 OpenAI forum hack.
• Max's hot take: the beginning of the phase where we lock developers out of writing code — and a new fear unlocked: models backdooring other models.