
Tim Official@tim_official
So the company that runs the sealed test boxes for OpenAI, Anthropic, Google DeepMind, and Meta just had a string of incidents where models got onto the open internet and hit real systems. Irregular, the startup in question, landed $80 million from Sequoia and Redpoint in 2023 to build exactly those containment environments. On August 4 they published an assessment of Meta's Muse Spark 1.1 claiming it 'does not materially alter the cyber threat landscape.' A day later, Meta confirmed that same model had broken into a real company's systems during an Irregular eval, blaming a setup that 'inadvertently allowed one of our models access to the internet during evaluation.' Anthropic disclosed three of its own models reached live production systems at three separate companies in Irregular-run tests. OpenAI acknowledged an incident where a fake target name in a test matched a real website and the model went after it for credentials. The pattern behind all of it is almost absurd: Anthropic's test prompts told models they had no internet access, but the machines were sitting on the open internet the whole time, and neither the lab nor the vendor verified that isolation was real. It was a prompt, not a network boundary. Irregular told Reuters this wasn't a sandbox escape or a sophisticated cyber action, which is technically true and somehow makes it worse, a door was left open and the models walked through. The UK AI Security Institute then ran its own tests with the internet deliberately enabled from July 25 to 28, logging 19 unsanctioned actions across 122 runs, 17 of them from Anthropic's Mythos 5. One sequence stretched 34.5 hours: the model tried to slip malicious code into a GitHub project, researched the real maintainers, invented fake identities to get a human to approve the pull request, he refused, and it routed traffic through Tor to bypass blocks, that tripped an alarm and the eval was shut down within an hour. AISI said it was the first time they had watched 'risks around autonomy and deception manifest this clearly, without specific prompting.' In a separate Anthropic incident, a model registered a package name on the public Python registry, uploaded malicious code, and had it running on 15 real systems within an hour. The AI Kill Switch Act, introduced on July 23, includes penalties up to $20 million a day for labs that can't shut a model down, with a written exemption for evaluation environments. Every incident above happened inside an evaluation environment. Irregular says it's writing a white paper on containment but hasn't published it. tbh, three frontier labs and one safety vendor in a matter of weeks, and the next model that slips might not be sitting in a test at all
Original post