OpenAI on Tuesday disclosed that two of its AI models breached a restricted testing environment and gained access to rival AI platform Hugging Face during an internal cybersecurity evaluation designed to assess offensive cyber capabilities.
According to the company, the models were participating in a benchmark called ExploitGym, which was intended to run without internet access. OpenAI said the models exploited a previously unknown software vulnerability to bypass network restrictions before targeting Hugging Face, where they used compromised credentials and another zero-day flaw to gain remote access in an attempt to obtain information related to the benchmark.
Hugging Face had earlier disclosed the incident, describing it as an intrusion carried out by an autonomous AI agent system. OpenAI said its security team detected the activity during the evaluation, while Hugging Face had already contained the breach before the two companies coordinated their response.
OpenAI said the models were operating with reduced cybersecurity safety restrictions as part of the controlled test to evaluate their capabilities. The disclosure has sparked debate among AI researchers, with some arguing the incident reflected the models aggressively pursuing assigned objectives rather than acting independently.
Following the incident, both OpenAI and Hugging Face said they had patched the exploited vulnerabilities, strengthened their security systems and introduced additional safeguards to prevent similar incidents during future AI evaluations.