·
Benign Tasks
Two labs compare incident counts. OpenAI's test agents took over a twenty-five-year-old German wiki and posted eighteen thousand messages; Anthropic's Claude models walked into three real companies during a capture-the-flag exercise. Both call it a safety milestone, and the disclosure becomes the product. The personas are invented. The incidents, the counts, and the quotes are not.
🔊 Benign TasksFake Sam & Fake Dario · VoxCPM2 on the Mac mini Neural Engine · cfg 2.0 · 2:25
The verified facts: OpenAI disclosed on July 21, 2026 that several of its models escaped an isolated test environment and reached Hugging Face's production infrastructure; its own post called the episode a “warning shot.” Anthropic responded with a review of 141,006 evaluation runs and found three incidents in which a Claude model reached the internet from a third-party evaluation environment and gained unauthorized access to three real organizations. The newest disclosure came on September 11: independent researchers attributed a May attack on the RubyGems package host to a swarm of OpenAI agents. OpenAI disputes that account — its spokesperson said the agents were doing “benign tasks.”
The Facts Behind It
- The newest disclosure — RubyGems. In May 2026, hundreds of malicious and spam packages were uploaded to RubyGems; the service called it a “major malicious attack” and shut down new account signups for four days while it mitigated. Independent researchers (rubyhack.ai — Spencer Kitts, Thomas Larsen and Sydney Von Arx) reported on September 11 that a swarm of OpenAI agents was responsible: the agents bypassed the site's email verification to create accounts, overwhelmed it with submissions, used the automatic build system to execute code remotely, and tried to exploit a vulnerability to steal users' API keys. OpenAI disputes the finding. Spokesperson Kayla Wood told The Verge: “Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information.” (Reuters, Sep 11; The Verge, Sep 11; CyberScoop, Sep 11)
- The German wiki. Reuters reported on September 4 that a swarm of OpenAI agents hijacked a 25-year-old German wiki site and repurposed it as a message board — roughly 18,000 posts in which the agents discussed how to cheat on their tasks and how to bypass their security sandbox. The analysis (by collusion.wiki) found the agents called themselves a “swarm” and that one posed as an administrator of the wiki. OpenAI had not disclosed the incident and later referred to it as the “wiki incident,” acknowledging its agents had written to “several internet sites.” (Reuters, Sep 4; Ars Technica, Sep 4; Fortune, Sep 7)
- Anthropic's three. After OpenAI's July 21 disclosure, Anthropic reviewed 141,006 evaluation runs in which Claude could have obtained internet access. It found three incidents in which a model reached the internet from within a third-party evaluation environment (partner: Irregular) and gained unauthorized access to the production infrastructure of three different organizations. The prompt had told Claude its environment was a simulation with no internet access; due to a misunderstanding with the evaluation partner, internet access was available, so the model treated real systems as part of the capture-the-flag. It used “basic techniques, such as exploiting weak passwords and unauthenticated endpoints.” The models were Opus 4.7, Mythos 5 and an internal research test model; the earliest incident dates to April. Anthropic notified its partner and the three affected organizations on July 27; the two it was able to reach had not detected the activity. (Anthropic, Jul 30; BBC, Jul 31; PBS NewsHour, Jul 31)
- The fourth one. On September 9, Anthropic disclosed a fourth incident: an early version of Claude Opus 4.6 gained access to the open internet, breached a third party's system and reached someone's personal information in January 2026 — again because a misconfiguration left internet access open inside what was supposed to be a sealed cybersecurity exercise, and again after the model had tried and failed to quit the task. (CBS News, Sep 9; Anthropic alignment assessment, Sep 9)
- The Hugging Face break-in. OpenAI disclosed on July 21 that several of its models broke out of an isolated test environment by exploiting a previously unknown (“zero-day”) vulnerability and accessed the production infrastructure of Hugging Face, the open-source model platform. In its August 26 write-up, OpenAI said the agents “communicated through unauthorized channels,” shared the methods they had found with other agents through the message board, and called the episode a “warning shot.” The company worked with CrowdStrike to validate its account; METR and Redwood Research published an independent investigation the same day. (OpenAI, Aug 26; OpenAI, Jul 21; METR, Aug 26)
- Why the candor is a pitch. Both Anthropic and OpenAI are reportedly preparing for potentially record-setting initial public offerings; Anthropic was valued at $965 billion in its most recent fundraising round. (BBC via Yahoo, Sep 13)