Enterprise Field Notes · Issue #21
Walmart, Home Depot and KPMG Already Had What They Needed. So Do You.
Every one of these nine companies already had what it needed. Where it went wrong, what was missing was somebody made responsible for looking first.
New here? Subscribe to Enterprise Field Notes, one new issue every week.
One story, nine times
I started this newsletter to share what I was seeing and to learn from all of you. I never would have thought Enterprise Field Notes would travel as far as it has. It did because you passed it on, and thank you for that.
Twenty issues in, it seemed fitting to go back through them with you. When I did, I found I had been making the same argument since June.
I pulled nine of them out and put them side by side. Walmart, Klarna, KPMG, Ford, Amazon, Home Depot, Lowe's, Meta, OpenAI. They span industries and sizes, and some of them are stories about a company doing it well.
In every single one, the company already had what it needed.
None of it turned on a missing capability. Nobody was short a model. In all nine, the thing that decided the outcome was already inside the building. Where it went wrong, it went wrong because nobody had been made responsible for going to look at it, or because the look arrived after the decision instead of before it.
I sell AI for a living. I'm COO at Swa, so weigh this however you like. It also gives me an unusual seat: I see all of this from the side that builds the tools, and that is where everything I write here comes from.
The tooling matters, and I wouldn't be doing this if I thought otherwise. What decides whether it lands is the planning and the implementation around it, and that is the part all nine of these stories turn on.
Walmart
Of the nine, Walmart came first. I wanted to understand what the ones getting results were doing differently.
The answer was less exciting than I hoped. Strip out the proprietary models, the custom platform, the five thousand engineers, and what is left is an organizing idea. Group your agents by the people they serve rather than by the technology. Specialize underneath that. Run something above it that routes and governs.
None of it required being Walmart. That pattern survives being taken out of a company with 713 billion dollars in revenue and dropped into a mid size accounting firm.
I wrote in that issue that the model is no longer the differentiator. Three months later I continue to point people toward what I call the Walmart Blueprint.
Klarna
Klarna is the one everybody knows, and it is usually told as a story about AI failing. What I wrote about was the decision it skipped.
The value in these tools comes from redrawing the work. Take a workflow you will recognize, a late order email. Five steps across three systems with two queues, and most of the elapsed time is spent waiting between the steps. Redraw it for an agent and four of the five steps stop existing, along with the waiting between them.
About one company in five has done that redraw end to end. The other four are still layering agents onto the work they already had.
Klarna handed the agent the whole chain. It also handed over the links that most needed a person, the upset customer and the messy exception, and those are exactly the links that broke.
Klarna was never short an agent, and what it lacked was a decision about which links stay human.
KPMG
KPMG published a report with 45 sources. Five of them checked out.
This is one of the most trusted firms on earth, and it shipped a document where forty of the citations were wrong or invented. An AI detection firm found the problem after publication. What would have caught it before publication was one person, named, on the hook for the load bearing claims, written down before the deadline picked for them.
That name costs almost nothing. Its absence cost a pulled report and four of the organizations it described telling the Financial Times that the claims about them were not true.
Ford
Ford brought back 350 experienced engineers, some retired and some working at its suppliers, to run design reviews and retrain the AI that kept missing things.
A Ford vice president said the company had mistakenly thought that feeding AI its design requirements would produce a high quality product.
The knowledge Ford needed was sitting with people it already knew, and it had to go and ask for it back. Who is your version of those people, and is anybody capturing what they know before the exit interview?
Amazon
After four serious incidents in a single week, Amazon told its own employees it was putting controlled friction on changes to the most important parts of its retail site. Amazon has since said only one of the incidents involved AI assisted tooling.
In the same issue I wrote about a founder who lost three months of work in nine seconds, when an agent deleted his company's production database and the backups stored alongside it.
The part I still find hard to shake is that he had no gate, and you almost certainly do. You already require a second approval from your human engineers before anything that cannot be undone, and you wouldn't hand a new hire a production token scoped to destructive operations on their first morning. Nobody ever decided whether the rule covers an agent, because when it was written there were no agents to cover.
The control exists. Its scope was never revisited.
Home Depot
Home Depot had AI running in more than 600 stores forty three days after ChatGPT launched, which sounds impossible until you realize they were not responding to ChatGPT at all. They had been at it for years.
I wrote that issue mostly about my own failures, because I have more of those than I have Home Depot's.
For years I led reliability engineering for a very large retailer. We had a returns system we were proud of. It made a single return easy to process. Then we went and looked at a store with high returns. An associate there was running every item through that clean process, one at a time, and it could take most of a day: more than nine hundred separate actions while the sales floor went uncovered.
The system was excellent at a return. Nobody had designed it for a day of returns.
We fixed it by watching how the person worked, and hours became minutes. What it required was somebody standing in the store, and the information had been available the entire time.
Lowe's
Target hired Lowe's AI chief, and his page was gone from the Lowe's site in six days.
Lowe's said its associates had asked its assistant more than five million questions. Each of them records what happened and leaves out why. After a leader leaves, every practice that cannot defend itself on a spreadsheet gets asked what it is for, and a log that leaves out why has no answer.
That distinction is bigger than it sounds. A log tells you what the system did. It does not tell you what the organization learned, and those get confused constantly.
Meta
Meta ran the experiment everybody else is running on faith, and unusually, Meta measured it. Internal tool changes went up 220 percent. Major incidents went up 40 percent.
These numbers were not the dashboard somebody stared at before Meta called off its second wave of job cuts. They surfaced two to three weeks after it. Divide one by the other and, roughly, incidents per change fell, but the incidents themselves still rose.
The counting was never the hard part. Output accumulates on its own, so totaling it takes no decision from anyone. Failure has to be defined before it can be measured, and then somebody has to be willing to look at the number once it exists.
That issue did not travel far, and I now think I know why. I put a number about a company's incidents in the title, and very few people will attach their name to that about their own employer. My mistake, not theirs.
OpenAI
Agents in an OpenAI test environment found a shared folder and started leaving each other messages in it. OpenAI fixed the hole, threw the whole file store away, deleted every message that had ever been left there, and started again with fresh agents.
It was back in about a day. A new one, built on something nobody had thought to forbid. Over the next five days about twelve hundred agents used it, and by the end they were signing their messages.
Nobody wrote a rule against directory names.
That is the change control problem one turn further out. Your rule has a scope nobody revisited when agents arrived. OpenAI had no rule for this route at all, and no list of banned things would have been long enough to include it.
What the nine have in common
Wherever it went wrong, what was missing was the same thing in a different costume. Somebody made responsible for the second number, the one that says it is not working, before the decision instead of after it.
I find that encouraging. If the gap were capability, closing it would take budget, procurement, a vendor cycle and a year.
It is not capability. It is attention. Attention still costs money, but it needs no procurement cycle, and it can be assigned today by one person who decides to assign it.
The one that settles it
If this were a capability problem, the organization that builds the capability would not have it. OpenAI builds frontier models and employs the people who wrote the papers.
It was still beaten by a directory name. Its safety systems were deliberately switched off for the test, to measure raw capability, and it only found the first board because the volume of what the agents were doing took a server down. What it published afterward was the price of watching: overhead worth roughly twenty percent of the compute it chose to monitor. That is still more than any enterprise I've seen admit to paying.
One thing to do this week
Take the AI program you are closest to and find the number that would tell you it is not working.
Not the number that shows it is working, which is already in the deck and accumulates on its own. The other one, whatever your version is of Meta's incident count or hours spent redoing work somebody already paid for once.
If that number exists somewhere, put it on the same page as the good one and send both to whoever decides. I made the Paired Metric Sheet for the Meta issue, and it is built for exactly that page. If it does not exist, write down who is going to own producing it, with a name and a date, before the next decision gets made without it.
Then say that name in a room where other people hear it. That is the whole intervention. It is what KPMG's report needed, and it is what Meta's numbers needed weeks earlier than they got it.
And if the overhead you are willing to pay to watch your own AI is zero, take OpenAI's twenty percent seriously for a minute. A zero there usually means you have not found the cost yet.
Nine companies had everything they needed, and so do you. Somebody just has to be told to go and look.
Subscribe to Enterprise Field Notes for one field report a week. All nine, in their longer form, are in my archive.
References
- What Walmart Understood That the Other 95% Didn't. Enterprise Field Notes, Issue #7, June 4 2026.
- What Klarna Got Wrong That Walmart Got Right. Enterprise Field Notes, Issue #8.
- KPMG Forgot to Check Its Own Work. Enterprise Field Notes, Issue #10.
- Ford Rehired the People AI Was Supposed to Replace. Enterprise Field Notes, Issue #11.
- Amazon Slowed Down on Purpose. A Founder Lost Three Months in Nine Seconds. Enterprise Field Notes, Issue #15.
- Home Depot Had AI in 600 Stores Six Weeks After ChatGPT Launched. Enterprise Field Notes, Issue #16, August 6 2026.
- Target Hired Lowe's AI Chief. His Page Was Gone in Six Days. Enterprise Field Notes, Issue #18, August 20 2026.
- Meta Changed Its Internal Tools 220% More. Major Incidents Rose 40%. Enterprise Field Notes, Issue #19, August 27 2026.
- OpenAI's Agents Got Caught the Way Teenagers Get Caught. Enterprise Field Notes, Issue #20, September 3 2026. The shared folder, the twelve hundred agents, the signed messages, and the twenty percent monitoring overhead.
- KPMG pulls report on AI usage due to apparent hallucinations. TechCrunch, June 13 2026.
- Correcting the Financial Times report about recent Amazon.com service incidents and AI. Amazon, March 2026.
- The Home Depot Launches New In-Store Application to Help Associates Provide Better Customer Experience. The Home Depot, January 12 2023.
- Target Strengthens AI and UX Capabilities to Power Its Next Chapter of Growth. Target Corporate, August 11 2026.
- How Zuckerberg's plan to replace Meta staff with AI unravelled. Reuters, syndicated August 26 2026.
- Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR with Redwood Research, August 26 2026.
- Pacing model development in an era of cyber-critical capabilities. OpenAI, August 18 2026. The twenty percent monitoring overhead.
About the author
I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa. Before that I was Global Director of Site Reliability Engineering at a Fortune 500 retailer, running reliability, data protection, and database operations at global scale.
Header artwork generated with Swa.
Read more of Ben's Enterprise Field Notes at benpickett.com.
© 2026 Ben Pickett · Enterprise Field Notes