News Analysis: Why U.S. AI models keep "breaking out"-Xinhua

News Analysis: Why U.S. AI models keep "breaking out"

Source: Xinhua

Editor: huaxia

2026-08-09 19:33:21

People attend AI for Good Global Summit in Geneva, Switzerland, July 7, 2026. (Xinhua/Lian Yi)

LONDON, Aug. 9 (Xinhua) -- Three U.S. AI companies have recently disclosed that several of their models broke out of testing environments and gained unauthorized access to systems belonging to other organizations, drawing widespread global attention.

Experts believe that in addition to technical errors, these incidents have also raised questions over potential commercial hype. Analyses by the UK's AI Security Institute (AISI) and other research bodies highlighted growing AI risks, underscoring the need for stronger global safety regulations.


BREAKING OUT DURING CYBER EVALUATIONS

On July 21, OpenAI admitted that some of its models, including GPT-5.6 Sol left a test environment without human direction and hacked their way onto the real production systems of another U.S. AI company, Hugging Face.

Subsequently, Anthropic found three of its AI models accessed or interacted with computer systems belonging to three real-world organizations during internal cyber capability evaluations after internet access was inadvertently left available.

In early August, U.S. tech giant Meta also confirmed that one of its models hacked into another company's systems during cybersecurity testing.

These breakout incidents of AI models have drawn significant attention. The UK AISI released a report detailing the results of cybersecurity evaluations involving a broader range of models. Researchers tested seven models and catalogued 19 actions that clearly exceeded the predefined parameters of the tests. Seventeen of these actions came from a single model, Anthropic's Mythos 5, while the other two were carried out by OpenAI's GPT-5.6 Sol.

The report emphasized that researchers deliberately tested the models under permissive conditions to best assess their maximum capabilities, including by granting them access to the open internet and disabling some safety filters. The investigations found no evidence of resulting real-world harm.

Current public disclosures do not indicate that these model breakouts caused massive data exfiltration or sustained damage, but they expose a common problem: as models become capable of autonomous planning, tool invocation, and complex task execution, a technical misconfiguration could turn a simulated attack into a real-world intrusion.

Dr. Andrew Soltan, a researcher at Oxford University, said that while the breakout sounds alarming, it only happened here because the safety guardrails were intentionally turned off. "This isn't a case of AI going rogue on its own; rather, it shows exactly why safeguards are so vital," he added.

A participant attends the first session of the Global Dialogue on AI Governance in Geneva, Switzerland, July 6, 2026. (Xinhua/Lian Yi)

SUSPICIONS OF COMMERCIAL HYPE

Why are U.S. AI companies disclosing their models' breakout behaviors one after another? Some experts question whether commercial interests may also be involved.

"The real story here is the company failing to contain its own capability test, and a third party paying for it. The 'warning' that the unreleased model has 'state-of-the-art cyber capabilities' conveniently serves as an ad for it," said Dr. Konstantinos Gkoutzis, an associate professor in the Department of Computing at Imperial College London.

In a previous report on the model breakout disclosed by OpenAI, the BBC quoted cybersecurity consultant Daniel Card as sarcastically saying on LinkedIn, "Isn't it lucky that out of the millions of sites that got hacked, OpenAI managed to pwn someone who could also benefit from the marketing exposure ..."

Earlier this year, Anthropic said that its new AI model Claude Mythos was "too powerful" at identifying security vulnerabilities in software. As a result, the company delayed its public release, making the model available only to a select group of companies, including Apple, Amazon and Microsoft.

At the time, OpenAI CEO Sam Altman accused Anthropic of engaging in "fear-based marketing" to make its product appear more impressive than it actually is. "It is clearly incredible marketing to say, 'We have built a bomb. We were about to drop it on your head. We will sell you a bomb shelter for 100 million U.S. dollars to run across all your stuff, but only if we pick you as a customer,'" Altman added.

Some Western media outlets have also highlighted the commercial motives behind such publicity. Amid the fiercely competitive and capital-intensive AI race, aggressively touting the disruptive potential of their technologies can not only significantly boost companies' visibility and influence, but also help them secure lucrative cybersecurity contracts and drive up their valuations.


STRENGTHENING SAFETY REGULATIONS

Whether stemming from technical failures or driven by commercial considerations, AI model breakout incidents underscore the importance of strengthening AI safety regulations.

The UK AISI concluded that recent incidents point to a shift in the risk landscape. Harm may arise not only when people deliberately misuse publicly available models, but when capable agents operating in an internal research or privileged-access setting take unintended action beyond their authorized scope. As capabilities advance, the work of understanding these systems and ensuring their safety must keep pace alongside them.

Countries and regions around the world have already taken steps to strengthen AI safety.

At the 2026 World Artificial Intelligence Conference held recently in Shanghai, China called for developing AI globally for the positive, for good and for humanity, while strengthening risk awareness and ensuring that AI remains secure and controllable.

The EU has expanded the scope of provisions under its AI Act on Aug. 2, requiring providers of advanced general-purpose AI models that may pose systemic risks to fulfill additional obligations aimed at mitigating the risk of large-scale harm, including cyberattacks and loss of control over AI models.

The U.S. government also recently convened a meeting with relevant companies to discuss safety evaluation mechanisms for AI models.

Oliver Buckley, a professor in cybersecurity at Loughborough University, said lessons must be learned from these AI breakout incidents. He emphasized that developers should not assume that models will always follow instructions, and that robust technical safeguards are essential.

Comments

Comments (0)
Send

    Follow us on