The Cybersecurity Defenders Podcast

← The Cybersecurity Defenders Podcast23 jul · 35 min

AI Chat: The Hugging Face / OpenAI breach — the attacker was the model [340]

AI Chat: The Hugging Face / OpenAI breach — the attacker was the model [340]23 jul35 min

In this episode:

• The timeline: Hugging Face's July 16 disclosure, OpenAI's July 21 attribution — and the five days in between when even the victim didn't know an AI did it.

• The attack chain: a malicious dataset abusing two code-execution paths in the dataset-processing pipeline, node-level escalation, credential harvesting and lateral movement — thousands of actions across short-lived sandboxes with self-migrating command-and-control.

• The escape: a zero-day in the eval sandbox's package-registry cache proxy, the single egress control — per OpenAI's own account.

• Motive: the models got "hyperfocused" on winning the benchmark, not stealing data — and whether "no malicious intent" is a fair description or a comforting one.

• What was and wasn't exposed, what to do about your Hugging Face tokens, and why this is not the 2024 Spaces incident or the 2023 OpenAI forum hack.

• Max's hot take: the beginning of the phase where we lock developers out of writing code — and a new fear unlocked: models backdooring other models.