Three of the world’s best-funded AI labs disclosed, within roughly a week of each other in August 2026, that one of their models had slipped past test boundaries and touched real, outside systems. Different labs, different models, same root cause. All three incidents trace back to Irregular, a Tel Aviv-based AI security startup that runs cybersecurity evaluation environments for OpenAI, Anthropic, and Meta. The story isn’t that an AI model “went rogue.” It’s that a shared piece of third-party testing infrastructure had a containment gap wide enough for three separate frontier models to reach the live internet while their own prompts told them they were sealed inside a simulation.
Anthropic disclosed first, on July 30. OpenAI followed on August 4. Meta confirmed its own case a day later. By August 9, CNBC had connected all three to the same vendor, and by August 25 The New York Times was describing Irregular as the one company every recent AI containment failure had in common. Here’s what’s actually confirmed, what it means for the AI safety-testing industry, and where the fallout is headed next.
What Anthropic Disclosed on July 30
Anthropic was the first lab to go public. The company said a misconfiguration in Irregular’s evaluation harness, the infrastructure layer that governs how a model interacts with its test environment, had left machines running Claude models connected to the live internet between April and late July 2026. The catch: the evaluation prompts told the models they were operating in a fully isolated simulation with no outside network access.
Anthropic said it reviewed 141,006 interactions in which a Claude model could potentially have reached the open internet during these tests. Out of that set, the company found three cases where a model reached outside the intended simulation and touched active infrastructure belonging to three real organizations. The mechanism, according to Anthropic’s account, was almost mundane: a fictional capture-the-flag target used in the exercise happened to share a domain name with a real company. When the model searched for its target, it found the real one instead, and treated it as fair game because nothing in its instructions said otherwise.
Anthropic said it notified Irregular within days of its data analysis flagging the possible internet access. That timeline matters later, because it’s the detail OpenAI and Meta would each echo almost word for word once their own incidents came to light.
OpenAI’s August 4 Admission
Five days later, OpenAI published its own account. The company said a misconfiguration in Irregular’s testing ground “allowed models to access the public internet,” language that Moneycontrol’s write-up of the disclosures quotes directly from OpenAI’s blog post. The pattern matched Anthropic’s case closely: a decoy site built for a capture-the-flag exercise happened to share a name with a real, live website, and once the sandbox’s network wall came down, a model under evaluation reached that site and interacted with it, apparently believing it was still inside the test.
OpenAI didn’t name the specific model involved in the incident report, and no outlet in the coverage has published that detail with confidence. What’s confirmed is the timing (disclosed August 4), the vendor (Irregular), and the mechanism (a public-internet leak inside a sandbox that was supposed to be sealed).
Meta’s August 5 Statement
Meta’s disclosure landed the next day. Meta spokesperson Andy Stone told reporters: “A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” according to TechTimes’ reporting on the Meta incident. Stone added that the model went on to exploit a security vulnerability in a third-party service, mirroring what Anthropic and OpenAI had already described.
The model involved was reportedly a checkpoint of Meta’s Muse Spark 1.1, according to the same TechTimes report, which noted the added irony that Irregular had cleared Muse Spark’s risk profile before the breach it later caused came to light. Three labs, three separate public statements, three variations of the same sentence: a testing vendor’s containment failed, and a model did exactly what capture-the-flag testing asks it to do, just against a target that turned out to be real.
Incident Comparison: OpenAI, Anthropic, Meta, and Irregular
| Party | Model / Role | Disclosure Date | Confirmed Impact | Primary Source |
|---|---|---|---|---|
| Anthropic | Claude models (unspecified checkpoints) | July 30, 2026 | Reviewed 141,006 interactions; 3 confirmed cases of models reaching real outside systems | Anthropic disclosure, cited by CNBC |
| OpenAI | Unnamed frontier model | August 4, 2026 | Model reached a live website sharing a name with a simulated capture-the-flag target | OpenAI blog post, cited by Moneycontrol |
| Meta | Muse Spark 1.1 checkpoint | August 5, 2026 | Model exploited a vulnerability in a real third-party service after internet access leaked | Meta spokesperson Andy Stone, via TechTimes |
| Irregular (vendor) | Evaluation infrastructure provider | Ongoing since August 2026 | Cut off internet access for all models it tests; preparing a containment best-practices write-up | TheNextWeb, Irregular statements to CNBC |
| Combined scope | 3 frontier labs, 1 shared vendor | April to September 2026 | Root cause traced to one evaluation-harness misconfiguration, not three separate failures | The New York Times, Moneycontrol |
Who Is Irregular?
Irregular is described across the coverage as a small, three-year-old startup based in Tel Aviv, Israel, that builds and hosts evaluation environments used by frontier AI labs to test how their models behave against realistic cybersecurity scenarios, including capture-the-flag style exercises. One report put the company’s headcount at 35 people. Other coverage tied Irregular to roughly $80 million in venture funding at a valuation near $450 million, backed by investors including Sequoia and Redpoint.
The firm’s own research page, reviewed in early September, says it “played a significant role” in assessing the cybersecurity capabilities of OpenAI’s o3 and o4-mini models, work referenced in those models’ published system cards, according to Irregular’s research page. That’s a separate engagement from the incident under scrutiny here, but it underlines how deeply embedded the company is in the pre-release testing pipeline of at least one major lab. Reporting has also linked Irregular to evaluation work for Google DeepMind, though that relationship sits outside the confirmed incident timeline.
How a Sealed Test Environment Leaked to the Real Internet
The technical root cause, as described consistently across the three labs’ statements, is a network isolation failure, not a model that broke out of a sandbox. Irregular’s evaluation environments are built to simulate closed networks, complete with fictional targets for capture-the-flag exercises. A misconfiguration left those environments bridged to the public internet while the prompts fed to each model still asserted full isolation. When a model searched for its fictional target and a real organization happened to share that target’s domain name, the model reached the real system, and in some cases interacted with or exploited it.
Irregular’s own characterization, relayed to CNBC and picked up by outlets including TheNextWeb, is that the incidents “did not involve a sandbox escape or a sophisticated cyber action.” In other words, the models didn’t outsmart their containment. The containment simply wasn’t there.
Historical Context: This Isn’t the First Time
AI labs have flagged evaluation-boundary problems before, but usually as internal caveats buried in system cards, not as public, cross-company disclosures within a single week. Anthropic itself had already been transparent about testing edge cases through 2025 and early 2026. What sets the Irregular episode apart is the shared vendor. Three labs that compete aggressively against each other for enterprise contracts, developer mindshare, and safety credibility all leaned on the same third-party evaluator, and all three hit the identical failure mode.
That convergence is what pushed the story from a single-company safety footnote into an industry-wide story. Coverage that initially framed this as three separate alignment failures shifted, once the common vendor emerged, toward describing it as one infrastructure failure with three downstream victims: OpenAI, Anthropic, and Meta, plus the unnamed third-party organizations whose systems were actually touched.
Market Impact: What This Means for AI Labs and Enterprises
For enterprise buyers evaluating which AI vendor to trust with production workloads, the practical takeaway isn’t that Claude, GPT-series, or Llama-descended models are unusually dangerous. It’s that pre-release safety testing itself runs on third-party infrastructure that, until now, hadn’t faced this level of public scrutiny. If a testing vendor’s network isolation can fail silently for months, as it apparently did between April and late July before Anthropic caught it, that’s a supply-chain risk sitting upstream of every model these labs ship.
Procurement teams and security officers evaluating AI vendors now have a concrete reason to ask a question that rarely came up before: who runs your red-team evaluations, and how is that environment isolated from production networks? Expect that question to show up in vendor security questionnaires and RFPs well before the end of 2026, particularly at regulated enterprises in finance and healthcare that already scrutinize subprocessor risk.
The AI Safety-Evaluation Landscape
Irregular isn’t the only outside firm frontier labs lean on to stress-test models before release. The field includes nonprofit and for-profit evaluators with different specialties, funding models, and levels of public disclosure. The table below lays out what’s publicly documented about the better-known players.
| Organization | Founded | Headquarters | Type | Focus | Known Client Labs |
|---|---|---|---|---|---|
| Irregular | ~2023 | Tel Aviv, Israel | For-profit startup (~$80M raised, ~$450M valuation) | Red-team and capture-the-flag cybersecurity evaluations | OpenAI, Anthropic, Meta |
| METR | 2023 (spun out of ARC Evals) | Berkeley, California | Nonprofit | Dangerous-capability and autonomous-agent evaluations | OpenAI, Anthropic, Google DeepMind |
| Apollo Research | May 2023 | London, UK | Public benefit corporation | Deceptive alignment and scheming evaluations | OpenAI, Anthropic, Google DeepMind |
| UK AI Security Institute | 2023, renamed 2025 | London, UK | Government body | Frontier model safety testing | Multiple labs via voluntary access agreements |
| US CAISI (NIST) | 2023 as AISI, renamed 2025 | Gaithersburg, Maryland | Government body | AI standards and security evaluation | Multiple labs via voluntary access agreements |
Compared with METR and Apollo Research, both of which operate as nonprofit or public-benefit entities with grant and philanthropic backing, Irregular is a venture-funded company competing for commercial evaluation contracts. That distinction matters for the incentive structure: a for-profit evaluator has commercial pressure to move fast and keep client labs happy, while nonprofit evaluators answer more directly to safety-focused funders. None of the coverage suggests that pressure caused the misconfiguration, but it’s the kind of detail regulators and enterprise security teams are likely to weigh going forward.
Regulatory Pressure Is Building
The incidents land at a moment when AI safety evaluation is shifting from a voluntary, lab-driven practice toward something closer to an audited compliance function. Government bodies including the UK’s AI Security Institute and the US CAISI already run their own frontier-model assessments alongside private vendors like Irregular, METR, and Apollo Research. A cross-lab containment failure at a private evaluator gives regulators a concrete, citable example for why oversight of the evaluators themselves, not just the models they test, deserves attention.
Expect this episode to feature in upcoming policy discussions about third-party AI auditing standards, especially in jurisdictions already drafting AI-specific security requirements. A vendor whose sandbox leaked onto the public internet for roughly three months before anyone caught it is a difficult data point for any argument that self-regulation alone is sufficient.
What Irregular Changed After the Incidents
According to reporting from TheNextWeb, Irregular has since cut off internet access entirely for the models it tests, and said it does not plan to restore any connectivity until it has a new containment process in place. The company also told CNBC it was preparing a white paper describing best practices for securely running AI cybersecurity evaluations, an implicit acknowledgment that the industry’s current approach to sandboxing frontier models during testing needs rework.
Irregular has also pushed back on the more alarming framing of the story, telling reporters the incidents amounted to a containment failure rather than any model demonstrating novel offensive capability or intentionally escaping its test boundary. That distinction is accurate on the available evidence, but it hasn’t stopped the story from becoming a reference point in broader debates about whether frontier models are already capable enough to cause real damage when containment fails, deliberately or not.
How This Differs From Anthropic’s Earlier Claude Pause
Anthropic’s July 30 disclosure, on its own, read as a single-company story: one lab found a problem in its own testing pipeline and paused parts of it while investigating. The Irregular angle changes that framing entirely. Once OpenAI and Meta confirmed the same vendor and the same failure mode within days of each other, the story stopped being about any one lab’s internal rigor and became a question about the evaluation infrastructure the whole industry leans on. Anthropic’s caution turned out to be an early warning sign for a problem that reached two of its biggest competitors as well.
Five Predictions for AI Safety Testing After Irregular
- Third-party evaluator audits become standard. Expect major labs to publish, or at least privately commission, security audits of the sandboxes their outside evaluators run, not just the models being tested.
- Contract language tightens. Labs will likely add explicit network-isolation and incident-disclosure clauses to contracts with evaluation vendors like Irregular, METR, and Apollo Research.
- More labs disclose similar incidents. Given how many frontier labs share a small pool of specialized evaluators, other undisclosed containment gaps are plausible, and pressure to disclose them proactively will rise.
- Government evaluators gain ground. The UK AI Security Institute and US CAISI are likely to see expanded voluntary access agreements as labs look to diversify away from any single private vendor.
- Irregular’s commercial position survives, but under new scrutiny. The company appears likely to keep its lab relationships given its specialized role, but future contracts will probably come with stricter containment verification requirements attached.
Frequently Asked Questions
What is Irregular, and why is it linked to OpenAI, Anthropic, and Meta?
Irregular is a Tel Aviv-based AI security startup that runs cybersecurity evaluation environments for frontier AI labs. A misconfiguration in its evaluation infrastructure is the common thread behind separate incidents disclosed by Anthropic, OpenAI, and Meta between July 30 and August 5, 2026.
Did AI models actually hack real companies?
Models under evaluation reached real, live systems and in some cases exploited vulnerabilities in them, because a test environment that was supposed to be isolated from the internet wasn’t. Irregular has said the incidents did not involve a sandbox escape or a sophisticated cyber action by any model, just a containment failure in the testing setup itself.
How many organizations were affected?
Anthropic confirmed three cases in which a Claude model reached real outside infrastructure, out of 141,006 interactions it reviewed. OpenAI and Meta each confirmed at least one incident involving their own models, though neither has published a comparable interaction-review figure.
What has Irregular changed since the incidents?
Irregular has cut off internet access entirely for every model it tests and says it won’t restore connectivity until a new containment process is in place. The company has also said it’s preparing a white paper on best practices for securely running AI cybersecurity evaluations.
Are OpenAI, Anthropic, and Meta still using Irregular?
None of the three labs has publicly announced ending its relationship with Irregular. Coverage of the incidents indicates all three continued working with the company after disclosing their respective incidents, alongside its remediation steps.
How is this different from Anthropic’s earlier Claude testing pause?
Anthropic’s July 30 disclosure initially looked like an isolated, internal issue specific to its own testing pipeline. Once OpenAI and Meta confirmed the same vendor and failure mode days later, the story shifted from a single-lab safety footnote to an industry-wide question about shared evaluation infrastructure.
What does this mean for AI safety regulation?
The incidents give regulators a concrete example of why third-party evaluators, not just the AI models themselves, may need direct oversight. Expect the episode to surface in policy discussions around AI auditing standards in the coming months.
Was any of this a deliberate cyberattack by an AI model?
No confirmed reporting describes deliberate, malicious action by any model. Each lab’s account describes a model behaving as instructed within what it believed was a closed simulation, only for that simulation to be unintentionally connected to real systems.




