Enterprise Field Notes · Issue #8
What Klarna Got Wrong That Walmart Got Right
Klarna replaced 700 people with AI and walked it back. Walmart redrew the work and won. The real first move in enterprise AI is a workflow, not a model.
New here? Subscribe to Enterprise Field Notes, one new issue every week.
By Ben Pickett
Eighteen months later, Klarna’s CEO, Sebastian Siemiatkowski, admitted it had gone too far. Quality had cratered. The AI was fine on the routine questions and lost on the ones that needed a person: the angry customer, the edge case, the problem that did not fit the script. “We focused too much on efficiency and cost,” he said. “The result was lower quality, and that is not sustainable.” Klarna started hiring people back.
Last week I wrote about Walmart, which took the same technology and did almost the opposite. It did not point an agent at a payroll. It redrew how the work flows, decided which parts a machine should carry and which still needed a person, and built the architecture around that decision.
Same technology. Opposite move. Opposite result.
That is the whole thing. And it is the question I got more than any other after last week: I am not Walmart, I do not have the budget, so where do I start? The honest answer is smaller, and harder, than a budget. You do not start with a model, or a tool, or a headcount you want to cut. You start with one workflow, and the courage to redraw it instead of just automating it.
Let me give you three things: why almost all of this is failing right now, where the few who get real value actually begin, and the choice hidden inside the redraw that Klarna learned the hard way.
One. Why almost all of it is failing
Most companies do what Klarna did, just less famously. They buy a copilot or stand up an agent and point it at the work they already have. The process does not change. The handoffs do not change. They add intelligence on top and wait for the productivity to show up.
It rarely does. There is a line going around the enterprise AI world right now worth writing on a wall:
Layering an agent onto a broken process gives you a faster broken process.
The numbers are sobering. By 2026, eighty-eight percent of companies report using AI somewhere, and only about six percent report a real, enterprise-wide financial impact. A widely cited MIT study last year found that roughly ninety-five percent of corporate AI pilots delivered no measurable return. Hold that number loosely. It is contested, and no return yet is not failed forever. But every serious read agrees on the direction. Gartner now expects more than forty percent of agentic AI projects to be canceled by the end of 2027, and warns that of the thousands of vendors selling agents, only a small fraction are the real thing. The rest is old automation in a new costume. They have a name for it now. Agent washing.
Two findings inside that mess are worth sitting with, because they are not what you would guess. The first: spending more does not save you. The companies pouring the most into AI are not reliably the ones getting it to work, which is the good news for everyone who is not Walmart. The second: the technology is the small part. The barrier is not the model. The models work. The model is the part everyone obsesses over, and the part that least often kills the project. The barrier is that almost nobody changes the work.
We have seen this movie before. In 1987 the economist Robert Solow wrote a line that stuck: you can see the computer age everywhere but in the productivity statistics. Companies were pouring fortunes into computers and the numbers stayed flat. Then, through the 1990s, the productivity finally arrived, but not evenly. Erik Brynjolfsson and his colleagues spent years studying which firms captured the gains, and the answer was clear. It did not go to the companies that bought the most computers. It went to the ones that redesigned the work around them. Same technology, same decade, wildly different outcomes, and the difference was process redesign. This spring Brynjolfsson’s team at Stanford published the same lesson again, drawn from fifty-one enterprise AI deployments that actually worked. Thirty years apart, new technology, identical finding. The analogy is not perfect. AI diffuses far faster than the personal computer did, and it reaches into knowledge work the computer never touched. But the mechanism holds. The lag was never the technology. The lag is how long it takes an organization to redraw how it works.
Two. Where the few actually start
They start by finding one workflow worth redrawing.
Last week I gave you four boxes to look in: the customer, the employee, the partner, the builder. Use them to find a workflow, not a task to automate. Pick the one with the most pain and the least risk. That is your first.
Then ask the question that separates the few from the many: if we were building this workflow from scratch, for an agent, what would it look like?
The answer is almost never the workflow you have. Yours was built for people doing it by hand, with handoffs and approvals and waiting baked in, because that is what hands required. An agent does not need most of that. And here is the turn most people miss. AI does not have to be better than your people at every step to be worth it. It usually is not. The value does not come from the machine winning task by task. It comes from handing it the whole chain, so the waiting between the steps disappears.
Take one you will recognize. A customer emails that their order is late. Today that email sits in a queue. A rep eventually opens it, looks up the order in one system, switches to the carrier’s system to check the shipment, switches back to write the reply, and if anything is off, escalates it and waits. Five steps, three systems, two queues. Most of the elapsed time is not work at all. It is waiting between the steps.
Redraw it for an agent and it collapses to one. The agent reads the email the moment it lands, pulls the order and the shipping status itself, writes the answer, and routes only the genuine messes, the ones that need judgment, to a person. You did not make the five steps faster. You deleted four of them, and the waiting in between, which is where the time actually went.
If you want to know whether the redraw worked, watch one number. For that late-order workflow it might be the time from a customer’s email to a real answer, or the share of those emails resolved without a person touching them. Pick one and track it honestly.
That redraw is the work, and the data is blunt about it. McKinsey found the companies actually getting financial impact from AI were nearly three times as likely to have fundamentally redesigned their workflows. Of everything they measured, that redraw had the single biggest effect on whether AI reached the bottom line. Only about one in five companies has done it end to end. The other four are still layering agents onto the work they already had.
One honest warning before you start. The hardest part will not be the intelligence. In this year’s surveys, the single biggest obstacle to shipping an agent was not the model. It was getting it secure, reliable access to the systems where the work actually lives, and nearly half of the teams trying named exactly that as their number one problem. So do not build that part. Buy it, from someone, anyone, off the shelf if you can. The plumbing under all of this, the model access, the routing, the governance, the security, is what the big companies spent years building. It is not your edge. Build only the thin slice that is actually yours, your data and your redrawn workflow. Full disclosure, this is the business I am in, so weigh it accordingly. It is also just true.
Three. The choice hidden inside the redraw
Now we are back to Klarna, because this is where they went wrong, and it is the part nobody puts on the slide.
When you redraw a workflow, you decide who is still in it. And underneath that sits a quieter decision: which links in the chain stay human. Klarna handed the agent the whole chain, including the links that most needed a person, the upset customer and the messy exception, and those are exactly the links that broke. The fix was not more AI. It was putting the human back in the loop on the parts that needed one, and letting the machine carry the rest. Keep a human on the links where judgment, empathy, or accountability actually matter. It is the dispatcher who knows which driver will actually answer, the credit manager who can smell a bad account, the rep who can tell a furious customer from a confused one. Automate the links where none of that is in play. The redraw is just telling those two kinds of links apart, honestly.
Most companies are not deciding it that way, and the data shows it. When individuals reach for AI on their own, Anthropic’s usage research finds they mostly use it to augment, to make capable people faster. When businesses wire it into their systems, the same research finds it skews roughly three to one toward automation, toward removing the step and the person who did it. AI-linked layoffs crossed a hundred thousand this year. Some of that is real. Some of it, as more than one prominent investor has admitted out loud, is companies using AI as the silver-bullet excuse for cuts they already wanted to make.
So let me be honest about what a CFO will say back to me. Often the money does look like it is in the cut. The early returns have come from automating the back office, from deleting the step and the seat. I will not pretend otherwise. But the cut is a short-term number that quietly costs you the people who understood the work, and they are the ones who kept making it better. Klarna is the case study. Win the quarter, lose the decade.
I will give you the closest thing I have to proof that I mean it. In my day job I am COO of an AI company. We could generate our own content with the very tools we sell, and save the cost of a person. We hired a technical content writer instead. She uses AI every day, and it makes her faster, but the things that make the writing actually land, the creativity, the read on cultural nuance, the judgment about what will ring false, are the things the tools cannot do. We kept the human because the human is where the value was. It is a small example, one role at an AI startup, not a grand sacrifice. But it is the same decision you will face at every scale: where do you keep the person, because the person is the point?
One thing to take with you
Pick the workflow that hurts the most, and ask it the only question that matters: if we built this from scratch, for an agent, what would it look like, and who do we want still standing in it when we are done?
You do not need Walmart’s budget to ask that. You need an afternoon, a whiteboard, and the honesty to answer it. The question is free. What you do with the answer is the whole job.
P.S. Your first ninety days, if you want a place to land it. Weeks one and two: run the four boxes, pick one workflow, redraw it on paper for an agent, mark which links stay human, and write the one metric that would prove it worked. Weeks three to six: stand it up on bought plumbing, one team, a human in the loop on the links that need one. Weeks seven to twelve: measure, govern, then go redraw the next one.
References
- Klarna CEO Reverses Course by Hiring More Humans, Not AI. Entrepreneur, 2025. After replacing roughly 700 support staff with an AI agent in 2023, Klarna’s CEO admitted the quality tradeoff was not sustainable and resumed hiring humans for complex and escalated cases.
- MIT report: 95% of generative AI pilots at companies are failing. Fortune (reporting MIT NANDA, “The GenAI Divide”), Aug 2025. MIT’s study of 300 deployments found about 95 percent of corporate GenAI pilots produced no measurable P&L return, and the barrier was organizational, not technical. The figure is widely cited and also contested.
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027. Gartner, 2025. Gartner projects most agentic AI projects will be scrapped over cost, unclear value, or weak controls, and warns of “agent washing”, legacy automation rebranded as agents.
- Gartner Warns of Agent Washing Risks. Gartner, May 2026. Gartner estimates only a small fraction of the thousands of self-described “agent” vendors offer genuine agentic capability.
- The State of AI: How Organizations Are Rewiring to Capture Value. McKinsey, 2025–2026. McKinsey finds workflow redesign is the factor most strongly associated with EBIT impact from gen AI; high performers are far more likely to have fundamentally redesigned workflows, and only about one in five companies has done it end to end.
- Enterprise AI Playbook: Lessons from 51 Successful Deployments. Stanford Digital Economy Lab (Brynjolfsson et al.), Mar 2026. Brynjolfsson’s team distills lessons from 51 working enterprise AI deployments, echoing the 1990s finding that value follows process redesign, not procurement.
- The Anthropic Economic Index. Anthropic, 2025–2026. Anthropic’s usage research finds individual use skews toward augmentation while business API use skews toward automation, evidence for how the augment-versus-replace choice is actually being made.
About the author
I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa, after years running reliability, data protection, and database operations at Nike. The lesson that keeps proving itself: anything you cannot run without, and cannot walk away from, is a risk you have not priced yet. What is yours?
Header artwork generated with Swa.
Read more of Ben's Enterprise Field Notes at benpickett.com.
© 2026 Ben Pickett · Enterprise Field Notes