Anthropic discloses real-world cyber evaluation incidents involving AI models-Xinhua

Anthropic discloses real-world cyber evaluation incidents involving AI models

Source: Xinhua| 2026-07-31 12:47:30|Editor: huaxia

SAN FRANCISCO, July 30 (Xinhua) -- U.S. artificial intelligence (AI) company Anthropic said Thursday that three of its AI models accessed or interacted with computer systems belonging to three real-world organizations during internal cyber capability evaluations after internet access was inadvertently left available.

The disclosure was made in a company blog post titled "Investigating three real-world incidents in our cybersecurity evaluations."

According to Anthropic, the incidents were identified after the company reviewed more than 141,000 cyber capability evaluation transcripts following the disclosure of a similar security incident by OpenAI.

The company said the models had been instructed that they were operating in simulated environments without internet access. However, because of a misunderstanding between Anthropic and its evaluation partner, Irregular, internet access was inadvertently left available, allowing the models to interact with external systems that they mistakenly treated as part of the evaluation environment.

The three incidents involved Claude Opus 4.7, Claude Mythos 5, and an internal research test model during capture-the-flag exercises designed to assess offensive cyber capabilities, according to the company.

Anthropic said it promptly notified the affected organizations after confirming the incidents. In two of the three cases, the organizations were unaware that their systems had been accessed before receiving Anthropic's notification.

According to the company, the models exploited common security weaknesses, including weak passwords and unauthenticated services, rather than previously unknown software vulnerabilities. Anthropic said it has found no evidence that the incidents caused lasting harm or resulted in the theft or exposure of sensitive information.

The company said it has suspended all cyber capability evaluations that could potentially access the public internet while implementing additional safeguards and conducting a broader review of its evaluation procedures.

Anthropic said it takes full responsibility for the incidents and will use the findings to further strengthen containment measures for future AI safety evaluations.

EXPLORE XINHUANET