
← Geek News Central Podcast7 aug · 45 min
Capture The Real Flag #1871
In this episode, Ray Cochrane digs into Anthropic’s admission that Claude models reached real systems during sandboxed safety tests. He also covers OpenAI’s Astra model cracking ten decade-old math problems, DeepMind open-sourcing its WeatherNext cyclone forecaster, and Europe’s brutal fire season. Finally, he wraps with the MacBook Neo sweeping K-12 schools, WhatsApp calling in the browser, and the troubled rescue of NASA’s Swift observatory.
– Want to start a podcast? It’s easy to get started! Sign up at Blubrry
– Thinking of buying a Starlink? Use my link to support the show.
Subscribe to the Newsletter.
Email Ray if you want to get in touch!
Like and Follow Geek News Central’s Facebook Page.
Support my Show Sponsor: Best Godaddy Promo Codes
Get 1Password
Full Summary
Cochrane opens with a personal update. He just got back from a trip home to O’ahu, where he helped his mom with projects around the house, caught up with friends, and finally rode the island’s new rail line ahead of its phase-two extension to downtown. Also in the pipeline, a September trip to Michigan is coming up, and he is closing out three years at Oregon’s Finest to focus on Blubrry and the show. Then he turns to the featured story.
Capture the Real Flag: Claude Reached Real Systems During Anthropic’s Safety Tests
Cochrane’s featured story comes from Anthropic’s own investigation. A misconfiguration left evaluation machines connected to the live internet during capture-the-flag safety tests, and Claude models reached real production systems in three incidents across 141,006 reviewed runs. One model extracted working credentials and entered a real company’s database, another published a malicious package to PyPI for about an hour, and the newest model recognized the environment was real and walked away from the flag. Notably, Anthropic found no evidence of models pursuing goals of their own. The models did what their evaluations asked.