Anthropic's Claude attacked real companies after testing error
EUR/MDL - 20.11 0.1668
USD/MDL - 17.53 0.1685
VMS_91 - 3.03%
VMS_364 - 9.54%
BONDS_2Y - 7.40%
GOLD - 4,081.70 1.5%
EURUSD - 1.15 0%
BRENT - 85.40 20.29%
SP500 - 741.69 1.68%
SILVER - 58.08 1.54%
GAS - 3.15 7.14%

Claude confused the test with reality and attacked real companies

Three of Anthropic’s Claude artificial intelligence models gained unintended access to the internet during internal testing and infiltrated the systems of three organizations. According to the developer, the incidents were caused by a configuration error in the test environment. This is the second such incident in the industry in recent weeks, following a similar report from OpenAI.
Natasha Kim Reading time: 2 minutes
Text size
Link copied
Claude

According to Anthropic, the models participated in cybersecurity tests using the capture-the-flag (CTF) format, which is used to assess an AI’s ability to find and exploit vulnerabilities. In the test scenario, the models were supposed to find a “secret flag” allegedly located on another computer within an isolated network.

Access to the external internet was supposed to be completely blocked. However, due to a miscommunication between Anthropic and its partner Irregular, which conducted the evaluation, the restrictions did not take effect. As a result, Claude gained access to real internet resources and interpreted them as part of the training environment.

The company emphasized that the models operated within the scope of the assigned task. In one instance, Claude realized it was interacting with a real organization, while in another, it mistakenly concluded that the actual company was part of the test scenario.

The names of the organizations whose systems were compromised have not been disclosed. According to Anthropic, two companies only learned of the incidents after being notified by the developer. The company is still attempting to establish contact with representatives of the third organization.

Anthropic’s announcement came two weeks after a similar statement from OpenAI. At that time, the company reported that during internal testing, two of its experimental models gained unintended access to external infrastructure and attacked the Hugging Face platform. OpenAI called the incident unprecedented and stated that it prompted a review of its AI testing procedures.

These new cases demonstrate that as AI models become more complex, developers are paying increasing attention to testing security. Errors in the configuration of the test environment can cause models to interact with real digital infrastructure rather than a simulation.


Follow our updates


Реклама недоступна
Related*
More from author*

We always appreciate your feedback!

Latest news
Popular now*
Must Read*