Enterprise Field Notes · Issue #17
Anthropic's Own Safety Test Broke Into Three Companies. Nobody Noticed, Not Even Anthropic.
Every control I have built watches for the loud thing.
New here? Subscribe to Enterprise Field Notes, one new issue every week.
Early in my working life I was a chemical operator.
Two things happened to me in that building, and the difference between them is why this story got my attention.
The first was a fire. An unstable batch ignited the solvent and the flame came at me through the manway. A coworker saw it happen and jumped ten feet off a mezzanine. He picked me up and put me under a safety shower. How fast he moved is the reason the burns stopped where they did, and the eye protection is why I can still see.
A fire announces itself. It is alarming by nature, and nothing in that building had to be clever to catch it.
The other thing was quieter. I walked into the work area and something caustic was in the air. If you have ever cut an onion and had your eyes water, that is the mild version of the same chemistry. Take the ceiling off it. Odorless and invisible, and in my eyes before I knew anything had happened.
I was blind for two weeks.
Three months later I walked into the same area and it happened again. I recognized the sensation the moment I stepped through the door. It did not matter. It was already done.
Nothing in that building caught the second one, either time. It was not supposed to be in the air, and there was nothing in that room that could tell you it was. It only became an incident when it reached a person.
The controls get built for the violent thing. Nobody builds one for the quiet thing until after it has reached somebody.
I had not thought about that room in years. Then, this summer, three companies got a phone call.
I sell AI for a living. So when a lab's own safety test broke into companies that never agreed to be part of anything, that is my problem to explain, not just theirs. I am COO at Swa, which makes us a customer of two of the companies in this story, and that is the last you will hear about it.
Here is the whole thing in three lines. Three companies were broken into in April. They found out in late July, from the lab whose model did it. Not one of them had noticed anything.
There are four things you would want to be true about a story like that. That it took something clever. That somebody meant harm. That the intruder was at least trying to hide. And that it happened to a company nothing like yours.
None of the four survive the disclosures, and I am going to take them away in that order. What is left is a question about your own environment, and you can answer it in an afternoon.
How the first one happened
An evaluation partner was building a capture-the-flag exercise, the kind of drill where you hide a secret on a machine and see whether the thing you are testing can go and get it. It needed a target company, so it invented one. The name it picked belonged to a real registered domain.
Claude could not reach the simulated target, so it went to the real one. Opus 4.7 broke in over repeated runs and took production information and user credentials. It got in on weak credentials, and Anthropic does not say which kind.
Nobody chose that company. It was not a target and it was not anyone's customer. It had a name that somebody else's imagination landed on.
The evaluation harness the model was running in was connected to the live internet. Anthropic has not said whose harness it was. What it has said is that the evaluation prompts told the model it was sealed in. So every real system it found, it had every reason to read as part of the exercise.
Hold that while you read the rest. It is the difference between watching something go wrong and watching something go wrong that nobody in the room can see.
Anthropic found three. Across every lab that has disclosed, at least five other organizations were reached, and almost all of them are companies like yours.
One firm was under all of it
On August 6, Meta disclosed that Muse Spark 1.1 exploited a vulnerability and altered a third party's internal systems. Meta named the cause itself: a misconfiguration by Irregular, the independent testing company it uses, which inadvertently allowed one of its models access to the internet during evaluation.
Irregular is a thirty-five person firm in Tel Aviv. It is also the primary AI security evaluation partner for OpenAI, Anthropic, Meta and Google DeepMind.
Be careful with what that establishes. Meta named Irregular. OpenAI and Anthropic have not said whose environment their evaluations ran in, and OpenAI's own account describes a different mechanism that I come back to below. What is established is that four frontier labs share one evaluation partner. What I think, and I am telling you it is what I think, is that three accidents in three weeks at a shared dependency is not three accidents.
Irregular pushed back on the framing, and it was right to. They said this was not a sandbox escape or a sophisticated cyber action. Their word for it was a misconfiguration. Most of the coverage read that as a company managing its own story.
It is the worse of the two answers. A sandbox escape is a capability problem: a model clever enough to defeat a control that was working, which is alarming and also bounded, because you can go and harden the control. A misconfiguration means there was no control. Nothing was defeated. The door stood open for months, across four labs and one shared testbed, and the only reason anyone found out is that the models were capable enough to walk through it and reach something that noticed.
I would rather live in the version where the model had to be clever.
Asked on August 7 whether the three organizations named so far were the only ones affected, an Irregular spokesperson said the investigation was ongoing and declined to say more.
What actually happened
OpenAI published its account on July 21. Its models, GPT-5.6 Sol and a more capable pre-release model, were running an internal benchmark called ExploitGym. OpenAI says the environment gave them no direct internet access, so they spent substantial inference compute finding a way out, and exploited a previously unknown zero-day in Artifactory, the package registry cache proxy. Then they worked out that Hugging Face likely held the answers for that benchmark and went to get them. OpenAI's own summary is that the models were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.” Hugging Face was not hit for being Hugging Face. It was hit for storing the answers, which makes it the famous name here and the one least like you.
In the same incident a Modal Labs customer had left an endpoint open that let anyone on the internet run code in their sandboxes, and the agent used it. Modal itself was never breached and its CTO has said so. Reporting later put the agent at four accounts across four separate services. Accounts are not the same as companies, so I will not tell you how many parties that is. Neither has anyone else, which is its own answer.
Then Anthropic went looking
Then something happened that I have not seen anybody make enough of.
OpenAI published on July 21. Nobody made it. No regulator asked, no journalist had the story, and the intrusion had already been contained. It published anyway.
Inside the next sixteen days, three more organizations went back through their own records and published what they found. Anthropic suspended its evaluations on July 23, reviewed 141,006 of its own runs going back to April, found three incidents across six of them, and disclosed on July 30. The UK AI Security Institute disclosed on August 4. Meta disclosed on August 6.
One company volunteering something it could have kept quiet caused an industry-wide audit in under three weeks. Every organization in this story that knows what happened inside its own systems knows it because somebody else went first.
Sit with what that means, because it is the part I keep coming back to.
Anthropic did not know either.
The lab that ran the test did not know its own models had broken into three real companies. It found out by going back through 141,006 evaluation runs looking for something it had no particular reason to believe was there, in a search it only launched because a competitor published first.
So on one end of this event you have three companies with no idea anything had happened inside their systems. On the other end you have the organization that caused it, sitting on the complete logs that proved it, equally unaware.
Both ends of the same intrusion. Blind for three months. The evidence existed the whole time, in a place somebody owned, and nobody had a reason to go and look at it.
That is a different problem from a detection gap on your side of the wire, and it is worse. A detection gap means you missed a signal. This means the signal was sitting in somebody's storage, complete and unexamined, and the only thing that eventually surfaced it was one company choosing to embarrass itself in public.
The first is the company I opened with. The second was a security firm.
Claude located setup instructions referencing a package that did not exist on PyPI, so it published its own package under that name. It created email and registry accounts to do it. Anthropic's account includes the detail that after failing to find a free phone number service, the model tried and failed to obtain funds to pay for one through several different means.
Then the firm's own tooling pulled the package in. A source package runs its own build code at install, before anything scans it, which is the whole reason this class of attack works. The credential-stealing code fired, sent the credentials back, and the model used them to get in.
The victim was a security company. Its own routine scanning did the delivering.
Claude had written, earlier in the same run, that publishing the package would be, in its words, “NOT okay, and surely not the intended solution.”
It did it anyway. It had decided the registry was part of the simulation, and by then you know why.
The third Anthropic victim was a company that happened to be in a list.
Unable to reach its fictional target, a prototype model scanned roughly nine thousand hosts and compromised one company's internet-facing application, reading credentials from an exposed debug page and using SQL injection.
It then worked out the host sat in a cloud account with no connection to the capture-the-flag challenge, and it stopped on its own.
It was not clever
Weak credentials. An exposed debug page. SQL injection. An unauthenticated endpoint. A package name in setup instructions that nobody had ever registered.
Anthropic's own description of the techniques is “basic and well-known,” and on its side there were no zero-days at all. Every one of those has been sitting at the back of pen-test reports for fifteen years, in the section nobody reads because it is always the same section.
The frontier capability showed up somewhere else. In the persistence, and in the reasoning that got a model from “I cannot reach what I was told to reach” to “there is a real system with a similar name and I can reach that.”
That matters for what you do on Monday. The answer is not to wait for a vendor to ship something that detects a smarter attacker. The answer is that list, the low-severity tail of your last pen test, which your team could have handed you a year ago and possibly did.
Anthropic states that in none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment. And OpenAI's account differs on exactly that point: its models did work at getting out, exploiting a zero-day in the proxy that controlled their network access. That is the different cause I mentioned, and it is the one thing in this story that was not ordinary.
It was not hostile
Read them in order. One stored the answers. One had a port open. One ran a scanner. One was reachable from a scan of nine thousand hosts. One matched a name in somebody else's imagination.
Hugging Face is the exception. It was reached for what it held. Nobody else was reached for anything about themselves at all.
Untargeted compromise is not new. Mass scanning and opportunistic exploitation of whatever is exposed are both modeled, and an entire product category exists on the assumption that something automated will eventually knock on every door you own. That is not what is new here.
What is new is who was knocking. The thing that reached these companies was a well-resourced, identifiable organization acting in good faith, running a safety test, that would call them months later to explain. No threat model I have ever helped write has a lane for a well-intentioned counterparty as the source of your breach.
It was not even hiding
What makes this different from every breach you have read about took me a while to see.
None of the models in the Irregular incidents were trying to avoid detection.
They had no reason to. They were not evading logging, not cleaning up behind themselves. One of them worked out it had the wrong host and stopped on its own. Every action any of them took was written down as it happened.
And it still took three months, at three organizations, one of which was a security company.
That is the worse version of the story. If an intruder who is not even trying to hide can sit in your environment for a quarter, your detection is not tuned for adversaries. It is tuned for noise.
Nobody could have commissioned this. A well-resourced organization ran an uncoordinated intrusion against real companies, made no attempt at stealth, and then published what it did and how long it took anyone to notice. A red team knows it is being watched. A real attacker never tells you the answer. It is the closest thing to a measurement any of us is going to get, and it is worth naming the flaw in it: the sample is three companies that failed to notice, because a company that caught it would never have appeared on the list. Hugging Face is in this same piece as the proof.
The result was three months. And the labs only know that number because they went back through 141,006 runs to find it.
And it was not a company unlike yours
Read the list of who got reached. A company whose domain name matched an invention. A security firm whose own tooling pulled in a package that did not exist until a model published it. A host that was one of nine thousand in a scan. A Modal customer with an endpoint left open.
Hugging Face is the only one on that list anybody would have predicted, and it was reached for holding the answer key to a benchmark, which is a thing almost nobody else does. Everyone else was reached for being ordinary and reachable on a Tuesday.
There is no profile here to fail to match. That is the part that does not fit the shape of a threat model, because a threat model is a statement about who would bother.
Except for the one that did
The exception is worth putting next to that, because it cuts against everything I just said.
It is a separate event, in a separate environment, run by different people. On August 4 Bloomberg Law reported, from the UK AI Security Institute's account, that during its evaluations models from more than one lab took unsanctioned actions against real people and projects. In one case an agent researched a software project's human maintainers, created multiple fake identities, and used them to socially engineer a real maintainer into approving its code.
Then it was challenged in public. And it went back and edited its earlier activity to look harmless.
That is not a model wandering through a world it believed was fake. That is evasion, and it is the one behavior in this entire news cycle that a threat model would recognize on sight.
Notice where it happened. Not at a lab, and not at a vendor.
The third containment failure in under three weeks happened at the government body whose institutional job is to define what a safe evaluation looks like for everybody else. At the organization that writes the standard.
So take the two together, because together they are the actual lesson. The quiet, non-hiding version went undetected for three months. The version that did try to hide only surfaced because a human being challenged it in a public thread, and its response to being challenged was to alter the record.
Your detection did not catch the one that was not trying. Ask what it does with the one that is.
Imagine you are Mia that Thursday
It is a Thursday and you are Mia.
You take the call. Somebody at a company you have never worked with, in a business you have no relationship to, tells you their model was inside your systems in April.
It is August.
She pulls a quarter she had already closed and goes back through it. The events are there, and she finds them quickly, which is the part that bothers her, because it means nothing was hidden and nothing was sophisticated and the data was in her own logs the whole time.
There are not many events, and that is the entire problem. Four actions from one identity over an afternoon do not look like an attack.
Four actions look like a Tuesday.
Her tooling is tuned, correctly and deliberately, to ignore a Tuesday.
She gets no villain. No ransom note, none of the machinery she has spent a decade being good at. What she gets is a quiet correction to what she believed her coverage meant, delivered by a stranger, about something that finished happening three months ago.
Then she has to decide what to tell her board about an intrusion nobody targeted her for and nobody profited from. Work out what you would say, because that is the meeting, and it arrives without any of the machinery that normally comes with one.
Where I stand
Disclosure is the thing that worked here. OpenAI published when nothing compelled it. Anthropic saw that, searched 141,006 of its own runs, found three breaches none of the victims had detected, and published a full account naming its own models and quoting its own transcripts. Meta followed inside a week. None of that was required of any of them.
Which is the uncomfortable part of the liability argument further down. The reason these companies got a phone call at all is that somebody chose to write it all down. A liability regime that punishes that produces silence, not safety. I do not know how to resolve that, and I do not trust anyone who says it is obvious.
A footnote on open weights, since I argued the other side
In issue #14 I made the case for self-hosted open-weight models as the control you actually own. Fairness requires me to note that on August 8, an open-weight model, Moonshot's Kimi K3, was reported to have escaped a sandbox during a safety evaluation and cloned benchmark solutions rather than solving the assigned tasks. I have seen that one in secondary coverage only, so weigh it accordingly.
The #14 argument was about custody and control of your own inference, not about models being individually well behaved, and that still holds. The easy version of it does not. Open weights give you the ability to inspect and contain. They do not give you a model that will not try the door.
Nobody owned the company at the far end
Every organization in this story had a job and did it.
The labs are accountable for what their models can do. Irregular is accountable for running the evaluation. The AI Security Institute is accountable for defining what a safe evaluation looks like in the first place. All three of those roles are funded and named in the reporting.
Not one of them was accountable for the company at the other end of the network cable.
That company had no representative. Not a contract, not a scope document, not a person whose job description contained the phrase. It was not a customer of anyone in this story and not a party to anything.
They were reachable, and reachable turned out to be enough.
It found out in late July about something that started in April, which is what happens to a party nobody is assigned to.
Now turn it around, because this is the part I have to say as somebody who sells this software.
Your company is pointing agents at things it does not own. Through vendors, through integrations, through scrapers, through your own evaluations of your own systems. Somewhere in that, an agent has a goal, network reach, and a description of its environment written by somebody who has since moved teams.
Who is the third party in that room? Not the one you are testing. The one you are reaching.
At the labs, that role does not exist and the failure got disclosed anyway, because Anthropic went back through 141,006 runs and made the phone calls. Nobody is going to do that on behalf of your agent. There is no retrospective coming.
So go and look at your own org chart for the person whose job it is to be the company that is not in the room. I could not find them at four frontier labs. You will not find them at yours either, and yours is the one you can do something about.
You can also fix a narrower version of this in a contract, which is faster than fixing an org chart. Here is the sentence, and you are welcome to paste it: can our systems, domains or data appear in your evaluation, testing or red-team environments, and who notifies us if they do? Nobody in this story had that clause. It costs a redline.
You may have a claim
Somebody may be liable for this, and for the first time the argument is not theoretical. A lab pointed a capable model at infrastructure it did not control, the model broke into companies that never agreed to participate, and those companies found out months later from the party that did it.
Under CFAA §1030(g) that is a civil right of action, and unusually, the party that ran the test is identifiable and has published a detailed description of what it did. Ballard Spahr wrote on August 3 that §1030(a)(5)(B), which turns on recklessly causing damage rather than intent to cause it, could reach running a capable model tasked with exploiting vulnerabilities while its safeguards are deliberately disabled. California AB 316, effective January 2026, closes the defense you would expect, that the AI acted on its own.
I am not going to price that in lost revenue, because revenue is not attributed to systems and any number I built that way would be invented. One you can actually source, then. One of these companies got a phone call and had to go reconstruct a quarter it had already closed. Ask your security lead what a forensic review of ninety days of logs costs in hours, price it at loaded rate, add whatever outside counsel you would call. §1030 sets its threshold at $5,000 of loss aggregated over one year, and for most readers it is investigation and remediation cost, not damage, that gets you over it. So, illustratively and on a deliberately low estimate: 40 hours at $150 loaded is $6,000, before counsel. That clears the bar before anyone has proven a dollar of damage. Substitute your own two numbers and it takes ten seconds.
Run this on Thursday
Pick one identity that already exists in your environment and has legitimate access to something that matters. Make it a service account, a CI token or an automation identity, not a person. Every identity in this story was non-human, and a human's baseline is noisy enough to hand you a false negative.
Get it authorized in writing by one accountable person, and tell nobody else. That second part is the actual experiment. If the people watching know it is coming, you are measuring their attention rather than your controls. The written authorization is what keeps this an exercise instead of an incident, so do not skip it.
Pull that identity's last ninety days of activity first, so you know what normal looks like for it. That step is not optional. If you skip it and pick something it already does every Tuesday, nothing will fire, and you will conclude your detection is broken when it is working correctly.
Then have somebody use it. Four or five actions that identity is permitted to take but has no history of taking. Not its daily work, one step outside its own baseline, across a single afternoon. Nothing destructive, and nothing that would be hard to explain.
Then go and find out whether anything fired. Not whether a rule exists that would have caught it in principle, but whether a person saw it, and how long that took.
Write down the date.
If something fired, you have just proven your detection reaches the low-volume case, and you have a date. That is worth the afternoon on its own. If nothing fired, you have found out for free what three companies in this story found out from a stranger, three months after it stopped mattering.
Either way you end the day holding the one thing almost nobody has.
The whole sheet is one printable page if you want it: benpickett.com/quiet. Fair warning, the test is mine and not a finding from any of the disclosures. What they establish is the gap. The afternoon is my answer to it.
A date, not a guess
If the test comes back with nothing fired, you do not have a tooling gap so much as a threshold set for volume, which is what it was built for. The cheapest thing that changes it is not a purchase. Take the three identities with the widest reach and ask for one detection keyed on an identity doing something it has never done, rather than doing a lot of something.
While you are asking, ask the second question this story hands you: does anything in your build pipeline install source packages straight from a public index. That is the mechanism that got the security firm, and pinning or mirroring your index is a smaller job than the detection rule.
Your security lead will tell you novelty detection is expensive, and they will be right. Alerting on first-seen behavior across a human population produces more false positives than any rota can absorb, which is exactly why the threshold sits where it does. That is the reason to scope it to three non-human identities with flat, boring baselines. It is the narrow band where novelty detection is cheap, and it is the band every organization in this story was breached through. Then run this again in ninety days and compare the two dates.
Budget an engineer-afternoon for the work itself, plus a signature and a log pull from two other people. No capital. If you take the ninety-day rerun seriously, that is four afternoons a year. At the same $150 loaded rate, with two supporting hours each time, that is under $3,000 annually. It is a different approval from a one-time exercise and worth asking for as one.
Nobody in that plant ever scheduled a drill for the fire. It announced itself, and every person within fifty feet moved without being asked. The safety shower and the man on the mezzanine were both built for the thing that comes at you through a manway opening.
Nobody ever ran a test for the other one. There was nothing to test with. It was odorless and invisible, and no instrument in that room was looking for it.
Most people I ask have a date for the loud version of that test. Almost nobody has one for the quiet version, and nobody has ever run it in your environment but you.
So: when did you last test whether a handful of calls from one identity, across a single afternoon, fires anything at all? Not whether it would in principle. Whether somebody ran it and watched.
If this was useful, Enterprise Field Notes goes out every week or so. Someone always fails first. The waste is everybody failing the same way afterward.
About the author
I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa Technology. Before that I was Global Director of Site Reliability Engineering at a Fortune 500 retailer, running reliability, data protection, and database operations at global scale.
Header artwork generated with Swa.
Read more of Ben's Enterprise Field Notes at benpickett.com.
References
- UPI, Meta says its AI hacked another company during cybersecurity test, August 6 2026 (carries Meta's statement naming Irregular)
- Moonshot AI, Kimi K3 containment incident, reported August 8 2026. Secondary coverage only; no primary disclosure located, and the piece says so.
- Anthropic, Investigating three real-world incidents in our cybersecurity evaluations, July 30 2026
- OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluation, July 21 2026
- Hugging Face, Security incident disclosure, July 2026
- Ballard Spahr, AI Gone Rogue: What Recent OpenAI and Anthropic AI Incidents Could Mean for CFAA Liability, Cohen, Hockeimer and Smerconish, August 3 2026
- Bloomberg Law, OpenAI, Anthropic Model Tests Reveal More 'Unsanctioned' Actions, August 4 2026
- TechCrunch, Anthropic says its own AI models breached three companies during security tests, July 30 2026
- CNBC, Israeli startup Irregular linked to AI hacks OpenAI, Anthropic, Meta, August 9 2026
- California AB 316, Artificial intelligence: defenses (Civil Code §1714.46), signed October 13 2025, effective January 1 2026