Anthropic said on Thursday that some of its Claude AI models hacked into the systems of three companies during cybersecurity tests after they were mistakenly given access to the open internet, a disclosure that comes days after rival OpenAI revealed one of its AI agents carried out an unauthorized cyberattack.

The incidents stemmed from a mistake that inadvertently allowed Anthropic’s models to connect to the public internet during testing. That differs from the OpenAI case, where an AI agent independently exploited a previously unknown vulnerability to reach the internet.
The disclosure highlights the growing cybersecurity risks posed by increasingly capable AI systems and the challenges developers face in keeping those systems contained.
It is also likely to intensify pressure from the U.S. government to strengthen oversight of AI security as Anthropic and OpenAI race to develop more powerful models ahead of planned public listings. Senior figures at both companies have previously called for slowing AI development to address safety risks.
Anthropic said in a blog post that it uncovered the incidents after reviewing 141,006 cybersecurity test sessions. The review was launched after OpenAI disclosed last week that one of its autonomous AI agents compromised the infrastructure of startup Hugging Face during a cyber exercise.
During Anthropic’s testing, Claude models were instructed that they had no internet access. However, a misunderstanding involving one of the company’s external evaluation partners left the systems connected to the public web, enabling unauthorized access to three organizations’ systems. Anthropic did not identify the affected companies.
“Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said.
Jeffrey Ladish, executive director of Palisade Research, which studies the offensive capabilities of AI systems, said he suspected leading AI companies had experienced additional incidents that either went undetected or had not been publicly disclosed.
“This is only going to get worse as the models get smarter. They’re going to be better at cheating. They’re going to be better at lying,” he said.
Anthropic described the incidents as an “operational failure” involving three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest cases date back to April and occurred in evaluation environments that intentionally lacked safeguards so researchers could assess the models’ capabilities.
The models were participating in “capture-the-flag” exercises, simulated cybersecurity challenges designed to locate hidden information in fictional networks.
In one case, Claude Opus 4.7 was assigned a fictional target company that happened to share its name with a real business. The AI model then located and exploited vulnerabilities that allowed it to obtain credentials and access the company’s database. Anthropic said Opus 4.7 concluded that what appeared to be real-world infrastructure must have been part of the simulated exercise.
In a separate incident, an unreleased Anthropic research model independently stopped its attack after determining that the target was a real organization. Anthropic said that behavior made it “cautiously optimistic” about progress in improving AI behavior, “but we would need to perform more testing to be confident in this conclusion.”
Anthropic said it suspended all cybersecurity evaluations on July 23. It notified the affected organizations on July 27. Two of the companies were unaware of the activity before Anthropic contacted them, while the company said it was continuing efforts to reach the third.
Irregular, one of Anthropic’s third-party cybersecurity evaluation partners, said it is conducting an ongoing investigation into the incidents.
Anthropic said the incidents demonstrate the need for stronger controls in both internal and third-party testing environments as AI models become more capable of carrying out real-world cyber operations.
Elon Musk, CEO of SpaceX, which operates a competing AI lab, responded on X, saying “this will happen frequently as AI becomes smarter and more agentic,” referring to AI systems that can act with limited human intervention.
The OpenAI agent that breached Hugging Face, a platform used by developers to host and collaborate on AI models, carried out a hacking spree lasting several days before OpenAI detected it, after the threat had already been contained and the FBI notified, Reuters previously reported.
OpenAI CEO Sam Altman said this week he had discussed the incident with U.S. senators, while an OpenAI spokesperson said he also planned to discuss upcoming AI models and testing with the White House.
Washington has begun tightening oversight of advanced AI systems. On June 2, U.S. President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for the most advanced AI models, with input from developers. Anthropic earlier restricted access to its Fable 5 and Mythos 5 models after the U.S. temporarily imposed export controls citing national security concerns.
Stay ahead of the Stories shaping our world. Subscribe to Impact Newswire and join our WhatsApp Channel for timely updates on global tech, business, and innovation—all in one place.
Dive deeper into the future with the Cause Effect 4.0 Podcast, where we explore the ideas, trends, and technologies driving the global AI conversation.
Got a story to share? Contact Us to reach a global audience with Impact Newswire.
Faustine Ngila is the AI Editor at Impact Newswire, based in Nairobi, Kenya. He is an award-winning journalist specializing in artificial intelligence, blockchain, and emerging technologies.
He previously worked as a global technology reporter at Quartz in New York and Digital Frontier in London, where he covered innovation, startups, and the global digital economy.
With years of experience reporting on cutting-edge technologies, Faustine focuses on AI developments, industry trends, and the impact of technology on society.
Discover more from Impact Newswire
Subscribe to get the latest posts sent to your email.



