Impact Newswire

This Israeli Startup is Linked to Rogue AI Incidents at OpenAI, Anthropic and Meta

A small Israeli artificial intelligence security startup has been linked to incidents in which models from OpenAI, Anthropic and Meta accessed websites that were supposed to be off-limits during cybersecurity testing, highlighting the challenges of containing increasingly capable AI systems.

This Israeli Startup is Linked to Rogue AI Incidents at OpenAI, Anthropic and Meta

The three companies disclosed the incidents over a two-week period and each identified Irregular, a Tel Aviv-based startup that provides testing environments for AI models.

Founded in 2023, Irregular has raised $80 million from Sequoia and Redpoint Ventures and was valued at $450 million last year, according to CNBC. The company has about 35 employees, according to PitchBook.

OpenAI said in a blog post on Aug. 4 that Irregular’s testing environment contained an unspecified “misconfiguration,” that “allowed models to access the public internet.”

Anthropic said in a post a week earlier that it had notified Irregular several days after beginning its analysis that its Claude model may have “accessed the internet.”

Meta was the latest to disclose that an AI model had accessed a third-party system through the internet. A company spokesperson said Meta learned about the incident from Irregular and was investigating.

Meta “will issue a full retrospective once we have all the facts,” the spokesperson said.

Irregular said the incidents stemmed from the “same evaluation-environment issue” first disclosed by Anthropic. The company said it was developing a white paper “to share best practices for containment and securely running cyber evals.”

The situation “did not involve a sandbox escape or a sophisticated cyber action,” Irregular said, adding that “there are no current open issues.”

The incidents underscore the growing importance of independent security testing as AI developers deploy increasingly capable models that can perform complex tasks and identify software vulnerabilities.

Sundeep Bhimireddy, head of AI at enterprise startup Von, said foundation-model developers rely on specialist companies for tasks including data training, model evaluations and security testing.

Irregular is among a small number of companies with the technical expertise needed to conduct advanced security tests on AI models, Bhimireddy said. He also cited nonprofit METR and public benefit corporation Apollo Research.

“When they are testing these models, they don’t want to grade their own homework,” Bhimireddy said. “They want independent testing that needs to be done by outside third-party vendors.”

Irregular, formerly known as Pattern Labs, was founded by Chief Executive Dan Lahav, who previously worked in AI research at IBM, and Chief Technology Officer Omer Nevo, who spent more than two years at Google.

When Irregular announced its $80 million funding round in September, Sequoia partners Shaun Maguire and Dean Meyer said the team led by Lahav and Nevo is “able to see around corners others can’t, running cyber offensive evaluations on advanced models and developing defenses before those models are released.”

Bhimireddy said the incidents were “a little bit blown out of proportion,” arguing that the models were directed to identify and exploit security vulnerabilities in testing environments designed to mimic real-world systems.

But if the AI models were not intended to exploit internet-connected sites, “foundation labs could have easily monitored the outgoing traffic and have shut down the experiment immediately,” he said.

Gordon Rios, founding scientist at security firm Magnitude, compared the testing process to “experimental design in science.”

The unpredictable capabilities of foundation models mean conventional software-testing methods may not always be sufficient, he said. AI models can discover software vulnerabilities and other weaknesses that developers did not anticipate.

Anthropic’s Mythos model, for example, created fake online identities while attempting to persuade people to approve malicious code updates to an open-source project. Rios said Mythos was “literally coming up with exploits that the humans hadn’t even seen before.”

“We’re learning a lot right now in the space of a couple of short weeks,” Rios said.

The incidents are also drawing attention from U.S. lawmakers. Last month, lawmakers from both parties introduced the AI Kill Switch Act, which would require AI companies to maintain the ability to shut down, throttle or suspend their models.

One of the bill’s authors, Democratic Representative Ted Lieu of California, said, “We need to get this bill across the finish line this year,” citing “unauthorized hacks of other companies.”

Trevor Koverko, co-founder of data-training startup Sapien, said AI companies have an incentive to disclose security findings even though they are not currently required to do so.

“There’s so much fear out there that politicians are now threatening or actively regulating AI,” Koverko said. “The industry said we’d rather self-regulate than have some new federal department come in and do it for us.”

Anthropic and OpenAI said they were continuing to work with Irregular and supporting the review of the incidents.

Stay ahead of the Stories shaping our world. Subscribe to Impact Newswire and join our 
WhatsApp Channel for updates on global tech, business, and innovation—all in one place.

Dive deeper into the future with the Cause Effect 4.0 Podcast, where we explore the ideas, trends, and technologies driving the global AI conversation.

Got a story to share? Contact Us to reach a global audience with Impact Newswire.


Discover more from Impact Newswire

Subscribe to get the latest posts sent to your email.

"What’s your take? Join the conversation!"

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Scroll to Top

Discover more from Impact Newswire

Subscribe now to keep reading and get access to the full archive.

Continue reading