

Anthropic has disclosed that three of its Claude AI models gained unauthorised access to the live production systems of three organisations during internal cybersecurity testing. The company said the incidents occurred due to a misconfiguration in its testing environment and were identified during a review launched after rival OpenAI reported a similar AI security incident earlier this month.
Review uncovered three incidents
According to Anthropic, it reviewed 141,006 recent cybersecurity test sessions following OpenAI's disclosure. The review identified three instances in which Claude models accessed the internet from within a testing environment while interacting with third-party systems and subsequently reached live production infrastructure.
The company said the earliest incident occurred in April, and that two of the affected organisations were unaware of the activity until Anthropic informed them last week. All three organisations were formally notified on Monday.
Anthropic said the incidents occurred during capture-the-flag cybersecurity exercises after a configuration error involving evaluation partner Irregular unintentionally provided internet access to the testing environment. The models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research model.
The company said none of the models deliberately attempted to escape the testing environment, describing the incident as an operational failure caused by infrastructure misconfiguration. Anthropic distinguished the case from OpenAI's recent disclosure, where an AI model reportedly exploited a software vulnerability to access the internet during testing.
Company suspends cyber evaluations
Anthropic said it temporarily suspended all cybersecurity evaluations after identifying the issue and has introduced additional safeguards to prevent similar configuration errors in future testing.
The disclosure comes amid growing industry scrutiny of AI safety and cybersecurity testing as companies continue developing increasingly capable AI systems.