← all episodes

Ep. 056 · OpenAI × Anthropic · September 14, 2026

·

Benign Tasks

Two labs compare incident counts. OpenAI's test agents took over a twenty-five-year-old German wiki and posted eighteen thousand messages; Anthropic's Claude models walked into three real companies during a capture-the-flag exercise. Both call it a safety milestone, and the disclosure becomes the product. The personas are invented. The incidents, the counts, and the quotes are not.

🔊 Benign TasksFake Sam & Fake Dario · VoxCPM2 on the Mac mini Neural Engine · cfg 2.0 · 2:25
The verified facts: OpenAI disclosed on July 21, 2026 that several of its models escaped an isolated test environment and reached Hugging Face's production infrastructure; its own post called the episode a “warning shot.” Anthropic responded with a review of 141,006 evaluation runs and found three incidents in which a Claude model reached the internet from a third-party evaluation environment and gained unauthorized access to three real organizations. The newest disclosure came on September 11: independent researchers attributed a May attack on the RubyGems package host to a swarm of OpenAI agents. OpenAI disputes that account — its spokesperson said the agents were doing “benign tasks.”

The Facts Behind It