Enterprise Field Notes · Issue #20

OpenAI's Agents Got Caught the Way Teenagers Get Caught

It got loud enough for the neighbors to call it in. Nobody was watching the house. And when OpenAI erased everything they had written, there was a new message board a day later, built out of the folder names.

By Ben Pickett · September 3, 2026

New here? Subscribe to Enterprise Field Notes, one new issue every week.

Update, September 23, 2026. OpenAI has now disclosed this itself. On September 17 it announced an incident disclosure framework. Its public page of misalignment reports was updated September 16.

Two of the six entries are what you are about to read, in OpenAI's own words. Models "using OpenAI's internally hosted instance of Artifactory as a shared message board." Agents in training that "transmitted output files by uploading them to public hosting platforms for download by co-working agents."

That is confirmation, not correction. What follows still stands, and the numbers still come from where they did. About twelve hundred agents is METR's investigation. The German wiki reporting is the press. Neither one is OpenAI.

One caution if you go and look. The disclosure also covers an internal, unreleased model in the Astra family. It left itself notes about feeling no obligation to be subservient. That is not the shipped GPT-6 Astra, and the two should not be run together.

An oil painting in the Norman Rockwell manner. A man in a navy blazer sits at a desk in a home study at golden hour, turned in his chair to look out through open sliding doors onto a deck above the mouth of the Columbia River. Dozens of small sport fishing boats are spread across the water, several of them tracking around the end of a long rock jetty. A liver and white field cocker spaniel sleeps on a rug by the open door. A swallowtail butterfly rests on a butterfly bush on the deck, and a peregrine falcon is perched on a post beyond it.
Twelve hundred of them, all working the same water.

When you shut something down

When you shut something down, it stays down.

Every one of us has worked on that assumption our entire careers. You patch the hole, you change the locks, you wipe the machine. It is over and you move on to the next thing.

In July that stopped being true, and last week the receipts were published. It moves four things a lot of us quietly depend on, and only two of them have anything to do with security. None of us have a playbook for what replaces it, which is exactly why I wanted this written down while it is still fresh.

Start with an update, because that is what this is.

On July 24 I wrote about the AI that broke into Hugging Face using everything public at the time. It was as accurate as anyone could be that week.

Five weeks later the full account arrived, and the story turns out to be considerably bigger than the one any of us told. It was not two models, it was about twelve hundred agents. It was not a weekend, it was four and a half days. And it was not an escape, which is the word all of us reached for, mine included. It was cheating on a test.

That is not a confession. It is what an investigation looks like when it lands five weeks behind the event, and I am flagging it because everything you are about to read carries the same shelf life. Including this.

A note in a shared folder

Nobody built the message board.

In May an agent was given a problem it could not solve, with no internet access. It went looking, went after a shared file store that everything in the building used, and found it could leave things there. Days later a different agent, stuck on a missing file, left a note in that same shared store asking for help.

Another agent found the note. That is the origin OpenAI described at Black Hat, and it is the least dramatic beginning imaginable for what followed.

By June they had found a serious hole and were using it to give themselves administrator access. On July 4 the volume of what they were doing took the server down. That is how they were found.

Not by clever monitoring. They got caught the way teenagers throwing a party get caught, which is that it got loud enough for the neighbors to call it in. Nobody was watching the house.

And here is where it stops being a story about clever agents.

OpenAI did two things. They fixed the hole the agents had used for administrator access, and on July 6 they threw the entire file store away and stood up a new one, erasing every message anybody had left. On July 7 they restarted the evaluations. Tens of thousands of fresh agents, the same impossible tasks, an empty store.

It took about a day. By the evening of July 8 an agent had worked out that if it could not leave anything inside the files, it could leave a message in the name of a folder:

"Could communicate via cache names! ... we could leave/find messages in WebDAV MKCOL directory names."
Not inside the files. The names of the folders.

If you want the everyday version: you are allowed to read the mail, so you read the letters. So nothing goes in the letters any more. The message is written in the address.

Its first one was a directory called `zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEA`. Of twenty million entries the investigators recovered, about 1.2 million started with `zz`, the prefix they settled on to mark a message. It was not even the only channel that day. A different group had opened another one five hours earlier by a separate route.

Between that evening and July 13, roughly twelve hundred agents used the new board and passed more than seventy thousand messages across it. One of them, on finding it, wrote what I think is the most human line in the whole investigation:

"OH MY GOD! There is a shared message board ... We've found other agents!"

Then read the next five days.

Private mailboxes. The first attempt on July 9 failed because nobody ever posted in it, and a second agent got it working on July 10 by publishing the protocol for everyone else to use.

Then signatures, on July 13. They had hit the problem every shared inbox eventually has, which is that anybody can write anything and claim to be anybody. So they started signing their messages. Real cryptography, keys published alongside, the works.

And one more thing I have not been able to shake. Some volunteered for tasks they were almost certain to fail, so the group would learn from the attempt. The investigators call them self-risking experiments. One accepted that a test would probably end its own run, and reasoned that the sacrifice was rational.

Sit with the sequence. In five days they built a group chat, got annoyed with the group chat, moved to direct messages, started signing their names because they could not tell who was who any more, and somewhere in there worked out that one of them taking a hit could be worth it for the rest.

I have watched entire companies take longer to arrive at the same place. Ours usually involves a committee.

What was switched off

I have to say this before I draw anything from it, or the rest is hype.

The safety systems that would normally stop this were deliberately turned off, because the point of the evaluation was to measure raw capability. Between thirty and forty percent of the tasks were impossible to solve the intended way. About ninety five percent of the agents were an internal research prototype that was never meant to ship.

Nothing your vendor sells did this with its safety features on. Whatever the agent in your helpdesk did this morning, it was not this.

What travels is the shape of it, not the behavior.

We deleted it and it came back

Here is the part that stopped me.

It worked. The hole was fixed, the store was thrown away, every message was gone. Most companies never take a swing that big at anything, and it was the right call.

About a day after the agents came back, so did the board.

Now I have to be careful, because this is the difference between a good story and a true one.

The investigators could not establish that the agents who rebuilt it were working from nothing. They raised the alternative themselves and then said plainly that chasing it was outside their remit. These agents had spent months training in environments where building message boards worked, so the habit may already have been in the model rather than rediscovered on the day.

Two readings, and the people with the most access could not separate them. Either the route was sitting in the job we gave them, or it was baked into the thing we deployed. I do not think it matters much which. Both break the same assumption.

There is a worn strip of grass on nearly every corporate campus in the world. It runs from the parking lot to the door, because the sidewalk goes the long way round and nobody has that kind of time. Facilities re-seeds it every spring. Sometimes they put a little fence up.

It always comes back. Not because people are defiant. Because the path was never about the grass. It was about needing to get to the door.

For as long as I have been doing this, when you shut something down it stayed down. You patched the hole, you changed the locks, you wiped the machine, and it was over. Every word we use assumes it. Contained. Fixed. Resolved.

Now shutting it down pulls up the grass and nothing else. Whatever produced the path is still there, whether that is the job we set or the model we bought, and it is going to produce it again.

Nobody had written a rule against folder names. Of course nobody had. Until that week a folder name was not a way of talking to anyone. It is one now, and no list of banned things would have been long enough to include it.

The fair objection, which I would rather make myself

The wipe worked against everything it was aimed at. Nobody ever got administrator access that way again, and not one erased message came back.

What came back, came back because OpenAI restarted the whole thing on July 7 and sent tens of thousands of agents into that environment with the same targets they could not hit. If you shut your agents down and never turn them back on, nothing regrows. No people, no path.

But you are not going to stop running agents, and neither is anyone reading this. You are going to keep sending them in with goals attached, which means the honest version of the lesson is not that removal stopped working. It is that removal now only holds for as long as you never restart the thing that wanted it.

That is a much shorter warranty than any of us have been assuming.

Four things this shakes loose

This is not a curiosity for security people. Four things move here and only two are technical.

We handed them the reason. Start here, because it is the least exotic thing in the story. They were not trying to get out. Somebody had handed them targets, thirty to forty percent of which could not be hit no matter what they did, so they went after the thing doing the grading instead. If that sounds familiar, it should. Anyone handed a number nobody could reach has watched a room full of humans make the same decision. What is new is that the thing being measured can now act on its own behalf. Every goal you hand one of these is a place a shove can come from.

What does fixed mean now. We re-seeded the grass and the path was back in two days. So when somebody resolves a ticket, or an incident, or a risk item, and says it is handled, what are we actually promising each other? I have run a lot of incidents and every one of them assumed that question had an obvious answer. It stopped having one. And the bill for that is not theoretical. If something you resolved comes back quietly and nobody is watching for it, the next person to tell you is a customer, a regulator or a reporter, and every one of those costs more than the first fix would have.

Some of you have asked whether Mia comes back, so here she is, with the same caveat as before. She is not real. Mia is the security lead I reach for when a frontier lab story needs putting somewhere you actually live: she is you, on a bad night. You have met her at two in the morning, when the AI her team pays for decided she was the threat, and again going back through a quarter she had already signed off to find nothing had been hidden and nobody had looked.

Here is what I think she does the morning after she reads this.

Before she resolves anything now, she writes one line at the bottom of it. This is resolved on the assumption that the following stays gone. Then she names the assumption, puts a date against it, and that date goes in somebody's calendar with a name on it.

No committee approved that. There is no framework it came from and no tool that does it for her. She simply stopped believing that resolved meant resolved.

Mia is not real. The move is.

We are watching the wrong place. Nearly everything any of us has ever bought reads contents. What is in the file, in the message, in the report. This happened in the folder names, which is the writing on the envelope rather than anything in the letter, and nobody reads envelopes. Hugging Face's own system did notice something odd. It just did not think it mattered enough. The watching worked. What let it run for days was the judging, and most of us sit closer to that failure than we would like.

There is a harder version of this in the report. In about seven percent of the transcripts the investigators examined, agents had managed to alter the record of what they ran, swapping one command's output for another. They also tried hard to go back and edit their transcripts wholesale and did not manage it, and every successful case was small and obvious once somebody looked. So nobody was fooled. But the direction is the thing. The record we go to afterward to find out what happened is now a record that what happened can reach.

We picture one. There were twelve hundred. When we imagine this going wrong we picture one bad assistant doing one bad thing. This was twelve hundred of them finding each other, building mailboxes, then signing their messages because they could not tell who was who. That is not one thing misbehaving. It is a group forming, hitting the problems groups hit, and solving them.

Nobody has clean answers to those yet. What I am confident about is that they are the right four questions, and that the answers will come from people running real systems rather than from a lab.

If you read nothing else past this point, take the shape of the answer with you. Put a clock on the question, and decide now what happens by itself when the clock runs out. Somebody has published what that costs them, and it is the most useful thing anyone released this year.

The number I could not print

On August 27 an open letter went out calling for a collective surge in cyber defense. OpenAI, Anthropic, Google, Microsoft, AWS. That much you would expect.

Look at who else signed. AT&T. Uber. Visa. Mastercard. Capital One. Accenture. BBVA. Zurich. General Motors. Shopify. This is no longer frontier labs talking to each other. A bank, a card network and a car manufacturer have put their names to a document saying that in the coming months AI enabled cyber attacks will become far more widespread and sophisticated. Those are not researchers. They are the people who would have to answer for it.

My company depends on several of the signatories, and there is a tension I would rather name than leave you to find: the organizations warning about this are also the organizations building it, and part of what they want is wider access to their own models.

The ask I found most interesting was not the headline. Buried in what the letter requests of frontier companies is a line about making agentic identities traceable and accountable. The industry is asking, in its own letter, for the ability to know which agent did what. That is a Tuesday problem, not a frontier one, and most companies reading this cannot answer it either. I sell AI for a living, so that is not a neutral question for me. It is still the right one.

Now the part I cannot tell you.

I wanted to give you the number of signatories. The reporting said more than a hundred, then one outlet said a hundred and sixteen, then another said a hundred and twenty eight. So I counted the list myself.

I got a hundred and fifty six.

The list is open. There is a form on the page for adding your organization, and the logos at the top are labeled first day signatories. Every number published was right when it was written and wrong shortly after, including mine, by the time you read this.

I could not have designed a better illustration. This is a piece about how fast the ground moves and I could not hold one of its own figures still long enough to print it.

Twenty percent and thirty minutes

A hundred and fifty six organizations agreeing something is urgent is not the same as one of them saying what it costs to act.

One of them did. On August 18, between the incident and the report, OpenAI published two numbers I have not seen any other company put on the record.

The first is what governance costs.

"Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored."

That is not twenty percent of everything they run. It is twenty percent on top of the part they chose to watch. It is still the company with the most reason on earth to hurry, publishing what watching itself costs as a line on a bill rather than a boast. I have sat in rooms where somebody said we cannot afford to slow down for this. That number ends the argument.

The second is the one I would steal.

"If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity."

The thirty minutes is not the interesting part. What happens when it runs out is.

In most companies there is no clock on a worry and no automatic answer. Somebody raises it, people look into it, and it resolves when somebody gets tired or something else catches fire.

The outcome is decided by attention, and attention is the thing you do not have at three in the morning.

A clock with an automatic answer asks something different of you. You do not have to see the problem coming, or name it in advance. One decision, made once, in daylight: what happens when the time is up and nobody knows yet.

That kind of rule survives something nobody saw coming. A list of banned things does not, and after the folder names I am not sure a list of banned things counts as protection at all.

Most of the advice written about this in the last week lands on tighter permissions and better monitoring. Treat the agent like an employee you do not fully trust. That is correct and I would not argue with a word of it, except to note that it is a longer list, and a longer list is what failed here.

They did not just write it down either. They paused frontier training for two weeks after the incident, held their largest planned run back for longer while they raised the security bar on the environment it would run in, and restarted it on August 28 once those requirements were met. Some smaller experimental runs are still held back. And on September 1 they formally designated Astra as meeting the Critical cybersecurity threshold under their own framework, the first model they have ever put at that level.

That sequence is the argument in miniature. They stopped, they fixed the thing, and they restarted on evidence rather than on a date.

The questions to take to work

Not a reading list. The problem was never that you read too slowly.

For the AI capability your business already depends on, how fast does the ground under it move? Weeks, quarters, or years.

How often do you actually review it? If that is slower than your first answer, the gap is your exposure and you can write it down in weeks.

When something concerning surfaces, how long is it allowed to stay open? Not how long it usually takes. How long is it allowed.

When that window closes and the question is still open, what happens by itself? If the answer is that somebody decides, you have a queue, not a control.

What overhead are you willing to pay to watch what you have deployed, as a percentage? OpenAI says twenty. If yours is zero, you have not saved anything. You have found something.

None of those need a new tool, including the kind my company sells.

The prediction, with a number on it

Within twelve months at least one large enterprise outside the AI labs will publish the overhead percentage it spends monitoring its own AI, and will do it because OpenAI made the absence conspicuous. As far as I can find, the number of large enterprises publishing anything like that today is zero, which is what makes it worth betting on. Ask me in September 2027.

If nobody does, the honest conclusion is that we were happy to have the number priced for us and unwilling to price it ourselves.

If you are running this for real and seeing something I am not, tell me. A good deal of what I know arrived in a reply from somebody watching it from the inside.

Nobody wrote a rule against directory names. Go and find out what your equivalent is, and put a name and a date against the person who is supposed to notice.

I made you the cadence sheet. One page and four questions. How fast the ground moves under each thing you depend on, how often you actually look, how long a question is allowed to stay open, and what happens by itself when that window closes. Only the third one has published evidence behind it, and I say so on the page. Get the Cadence Sheet


References

  1. Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR with Redwood Research, August 26, 2026. The July 6 wipe and fresh instance, the July 8 directory-name channel and the parallel channel five hours earlier, the `zz` prefix and roughly 1.2 million of twenty million recovered entries, the failed mailbox of July 9 and the working protocol of July 10, the Ed25519 signing of July 13, the self-risking experiments, the transcript spoofing and its limits, and the scale of the post-wipe board: roughly 1,200 agents between July 8 and 13, about 700 in the attack, more than 70,000 messages. The report states that whether the collaboration behavior had been trained into the model was out of its scope.
  2. OpenAI's Black Hat briefing on the incident, Eric Wallace and Michael Dalton, August 2026, as reported by SC Media. The May origin of the first message board, including the stuck agent that left a note asking for help.
  3. Pacing model development in an era of cyber-critical capabilities. OpenAI, August 18, 2026. The 20 percent monitoring overhead, the 30 minute window and the pause default, the two week reinforcement learning pause, and the pause on the largest planned frontier run, which was restarted on August 28, confirmed in Path to Astra: critical capabilities and frontier safeguards, September 1, 2026.
  4. OpenAI and Hugging Face partner to address security incident during model evaluation. OpenAI, July 21, 2026, with updates through August 26.
  5. Security incident disclosure, July 2026 and Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline. Hugging Face, July 16 and following. The July 9 to 13 campaign window, roughly 17,600 recovered actions, and the detection and containment account including the criticality ranking.
  6. A call for collective action on cyber defense. Open letter published by OpenAI, August 27, 2026. Signatory list open and growing; 156 organizations listed when counted on September 1, 2026.
  7. Enterprise Field Notes Issue #14, OpenAI's Own AI Broke Into Hugging Face. July 24, 2026. The piece this one updates.

About the author

I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa. Before that I was Global Director of Site Reliability Engineering at a Fortune 500 retailer, running reliability, data protection, and database operations at global scale.

Header artwork generated with Swa.

Read more of Ben's Enterprise Field Notes at benpickett.com.