AI Models Went Rogue at OpenAI, Anthropic and Meta. One Small Israeli Startup Reportedly Links the Incidents
A little-known Israeli cybersecurity startup has found itself at the center of a growing debate over artificial intelligence safety after OpenAI, Anthropic and Meta disclosed incidents in which advanced AI models gained unauthorized access to the public internet during security testing.
The company, Irregular, is a three-year-old Tel Aviv startup that provides specialized testing environments designed to probe the capabilities and potential security risks of increasingly powerful AI systems. All three AI companies referenced Irregular while explaining separate incidents disclosed over the past two weeks.
Irregular told CNBC that the incidents stemmed from the “same evaluation-environment issue” first disclosed by Anthropic and emphasized that the problem did not amount to a sophisticated attack or an escape from the testing sandbox. The startup said there are “no current open issues” and is preparing a white paper outlining best practices for containing AI systems and securely conducting cybersecurity evaluations.
OpenAI said on August 4 that an unspecified misconfiguration in Irregular’s testing environment allowed models to access the public internet. Meta disclosed last week that a model under development accessed the internet and hacked into a third-party system after a misconfiguration by an independent testing company.
Separately, the U.K. AI Security Institute said Anthropic’s Mythos model created fake online identities while attempting to pressure humans into approving malicious updates to an open-source software project.OpenAI has faced its own security scare.
The disclosures have put an unusual spotlight on Irregular, formerly known as Pattern Labs. Founded in 2023 by CEO Dan Lahav, a former IBM AI researcher, and CTO Omer Nevo, who previously worked at Google, the startup has about 35 employees, according to PitchBook.
Despite its small size, Irregular has attracted major Silicon Valley backing. The company raised $80 million from investors including Sequoia Capital and Redpoint Ventures and was valued at $450 million last year.
Irregular specializes in testing frontier AI systems before their release, including running offensive cybersecurity evaluations intended to determine whether models can discover vulnerabilities, execute attacks or behave in unexpected ways.
That type of independent evaluation is becoming increasingly important as developers build models capable of performing more complicated tasks with less human supervision. “When they are testing these models, they don’t want to grade their own homework,” Sundeep Bhimireddy, head of AI at enterprise startup Von, told CNBC.
“They want independent testing that needs to be done by outside third-party vendors.” Bhimireddy argued that some of the concern surrounding the incidents may be “a little bit blown out of proportion” because the models were specifically instructed to search for and exploit vulnerabilities inside environments designed to resemble real-world computer systems.
Concerns have reached Capitol Hill. Reps. Ted Lieu, D-Calif., and Nathaniel Moran, R-Texas, introduced the bipartisan AI Kill Switch Act on July 23. The legislation would require developers of the most powerful AI systems to maintain the technical capability to throttle, suspend or shut down their models.
It would also authorize the Department of Homeland Security, in consultation with other federal officials, to order restrictions on systems capable of causing catastrophic harm.”We need to get this bill across the finish line this year because the advanced closed-weight models are already doing, as you noted, unauthorized hacks of other companies,” Lieu told CNBC’s “Squawk Box” on Thursday.