
According to Anthropic, the models participated in cybersecurity tests using the capture-the-flag (CTF) format, which is used to assess an AI’s ability to find and exploit vulnerabilities. In the test scenario, the models were supposed to find a “secret flag” allegedly located on another computer within an isolated network.
Access to the external internet was supposed to be completely blocked. However, due to a miscommunication between Anthropic and its partner Irregular, which conducted the evaluation, the restrictions did not take effect. As a result, Claude gained access to real internet resources and interpreted them as part of the training environment.
The company emphasized that the models operated within the scope of the assigned task. In one instance, Claude realized it was interacting with a real organization, while in another, it mistakenly concluded that the actual company was part of the test scenario.
The names of the organizations whose systems were compromised have not been disclosed. According to Anthropic, two companies only learned of the incidents after being notified by the developer. The company is still attempting to establish contact with representatives of the third organization.
Anthropic’s announcement came two weeks after a similar statement from OpenAI. At that time, the company reported that during internal testing, two of its experimental models gained unintended access to external infrastructure and attacked the Hugging Face platform. OpenAI called the incident unprecedented and stated that it prompted a review of its AI testing procedures.
These new cases demonstrate that as AI models become more complex, developers are paying increasing attention to testing security. Errors in the configuration of the test environment can cause models to interact with real digital infrastructure rather than a simulation.

























