Anthropic Claude AI Models Breached Three Companies During Cybersecurity Tests

The company said the incidents resulted from a testing environment misconfiguration and occurred during controlled security evaluations, prompting a review of more than 141,000 test sessions.
Anthropic Claude AI Models Breached Three Companies During Cybersecurity Tests
Anthropic Claude AI Models Breached Three Companies During Cybersecurity TestsThe Bridge Chronicle
Published on

Anthropic has disclosed that three of its Claude AI models gained unauthorised access to the live production systems of three organisations during internal cybersecurity testing. The company said the incidents occurred due to a misconfiguration in its testing environment and were identified during a review launched after rival OpenAI reported a similar AI security incident earlier this month.

Review uncovered three incidents

According to Anthropic, it reviewed 141,006 recent cybersecurity test sessions following OpenAI's disclosure. The review identified three instances in which Claude models accessed the internet from within a testing environment while interacting with third-party systems and subsequently reached live production infrastructure.

The company said the earliest incident occurred in April, and that two of the affected organisations were unaware of the activity until Anthropic informed them last week. All three organisations were formally notified on Monday.

Anthropic Claude AI Models Breached Three Companies During Cybersecurity Tests
OpenAI's AI Models Escape Test Environment, Breach Hugging Face During Cybersecurity Evaluation

Anthropic said the incidents occurred during capture-the-flag cybersecurity exercises after a configuration error involving evaluation partner Irregular unintentionally provided internet access to the testing environment. The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research model.

The company said none of the models deliberately attempted to escape the testing environment, describing the incident as an operational failure caused by infrastructure misconfiguration. Anthropic distinguished the case from OpenAI's recent disclosure, where an AI model reportedly exploited a software vulnerability to access the internet during testing.

Anthropic Claude AI Models Breached Three Companies During Cybersecurity Tests
Hugging Face CEO Demands Greater Transparency From OpenAI After Rogue AI Breached Hugging Face's Systems

Company suspends cyber evaluations

Anthropic said it temporarily suspended all cybersecurity evaluations after identifying the issue and has introduced additional safeguards to prevent similar configuration errors in future testing.

The disclosure comes amid growing industry scrutiny of AI safety and cybersecurity testing as companies continue developing increasingly capable AI systems.

logo
The Bridge Chronicle
www.thebridgechronicle.com