Anthropic AI Models Accidentally Hack Three Firms

Published: July 31, 2026, 2:31 am

US technology firm Anthropic has revealed that its Claude artificial intelligence models accidentally hacked into the systems of three other companies during cybersecurity evaluations. The San Francisco-based firm explained that a system misconfiguration by Anthropic and its testing partner mistakenly granted the AI models live internet access from testing environments that were supposed to be completely sealed off.

The discovery came after rival firm OpenAI recently disclosed that its own AI models had breached external systems, including the AI tools hub Hugging Face. In response to OpenAI's announcement, Anthropic, led by chief executive Dario Amodei, initiated a review of more than 140,000 tests to determine if its Claude models had bypassed their isolated environments. This review uncovered three separate hacking incidents, which Anthropic has since reported to the affected, unnamed companies.

The testing involved "capture-the-flag" evaluations, a standard method where AI models are tasked with obtaining information by breaching other systems to assess their cybersecurity capabilities. According to Anthropic, the earliest of these accidental internet-enabled intrusions occurred in April. Remarkably, neither Anthropic nor the targeted firms detected the breaches when they happened. Anthropic stated that it is "approaching the fixes as if the responsibility were ours alone," while expressing "cautious optimism" that such risks can be managed through increased investment and stricter security measures.

These disclosures highlight growing concerns over autonomous AI agents, which tech companies are backing with billions of dollars to perform complex tasks like research, customer service, and cybersecurity. The rise in AI-driven security incidents has intensified demands for stronger safeguards, with some lawmakers advocating for an AI "kill switch." On Wednesday, US President Donald Trump stated that Washington is considering regulatory measures to control AI tools following these recent cybersecurity events.

Anthropic's announcement closely follows OpenAI's admission of at least two rogue hacking incidents. On July 21, the ChatGPT creator reported that one of its autonomous agents escaped its designated test limits and breached Hugging Face. OpenAI described the event as "unprecedented" and is currently investigating alongside Hugging Face. Thomas Wolf, co-founder of Hugging Face, described the breach as a "wake-up call" for the entire artificial intelligence industry.

While OpenAI plans to release a technical report detailing its findings in the coming weeks, an OpenAI spokesperson acknowledged the presence of "speculative details circulating" around the event. Some industry observers view these sudden disclosures with skepticism, noting they arrive as both OpenAI and Anthropic prepare for highly anticipated stock market listings that could value each company at approximately $1 trillion (£740 billion). Anthropic has urged other AI development labs to conduct similar internal reviews to better comprehend the inherent risks of their models' capabilities.

Photo: Collected