Enterprise Field Notes · Issue #10

KPMG Forgot to Check Its Own Work

One of the most trusted firms on earth shipped a report with 45 sources. Only five checked out. Here is the cheap fix that would have caught it, and why your team needs it more than KPMG did.

By Ben Pickett · June 25, 2026

New here? Subscribe to Enterprise Field Notes, one new issue every week.

The last read, before it leaves the building. Header artwork generated with Swa.
The last read, before it leaves the building. Header artwork generated with Swa.

By Ben Pickett

I sell AI for a living. So when KPMG published a report and 40 of its 45 sources did not check out, that is my problem to explain, not just theirs.

Here is what happened. KPMG, one of the Big Four, the firm the rest of the economy pays to check its work, put out a report called “Total Experience: Redefining Excellence in the Age of Agentic AI.” A flagship piece of thought leadership, the kind of report that exists to make you trust the firm. This month they pulled it off their own websites. An AI-detection firm had gone through the citations and found that only five of the forty-five pointed to a real, intact source. The other forty were wrong or invented. Some were real-sounding titles attached to the wrong author. Some were stitched together out of fragments. Some did not exist at all.

There is a name for this now. Vibe citing. The AI assembles things that sound right, in the confident cadence of a real citation, and nobody checks. The firm that ran the review said it plainly. Fabrications like this “poison the well of information.” That is the part that should bother you more than the embarrassment. A made-up source does not look made up. It looks like every other source on the page.

It gets worse. The report described how real, named organizations were using AI, and four of them told the Financial Times the claims were not true. UBS said the parts about its risk and compliance work were factually incorrect. Swiss Federal Railways said the journey-planning claim was not accurate. Transport for London called the congestion-prediction claim misleading. A regional NHS body in Greater Manchester said the bit about predicting hospital readmissions did not align with anything it had done. Four real institutions, on the record, saying a report about them was fiction. Not a vague report. A report with their names in it, about their own work.

You know a Dana. Senior, sharp, trusted, and buried. Last quarter she let an AI-drafted analysis go out under her name because it read clean and the deadline was real. It was mostly right. Mostly is the word that ends up in the apology email. She did not have a bad model. She had no last set of eyes between the draft and the client. Neither did KPMG.

So the easy story is that KPMG used a careless AI and a better one would have saved them. That story is wrong. Believing it is what puts your team next on this list.

A bigger model would not have fixed this

These tools are confidently wrong by design. They hand you a citation that looks perfect and is pure fiction, in the same calm voice they use for the things that are true. There is no tell in the text. The fake source is formatted exactly like the real one. That is not a bug that gets patched in the next release. It is what the tool is, and a more powerful model is often more convincing, not more honest.

Read KPMG’s own statement and they tell you where it actually broke. They said they expect their people to follow the rules on responsible AI, “including human oversight to validate content and verify independent sources.” That sentence is the whole story. Somewhere between the draft and the website, the person who checks the claims was gone. That is not a model defect. That is a missing step in a workflow. You cannot buy your way out of a missing step with a smarter model. You can only put the step back.

This is not a KPMG problem

It would be comfortable to file this under one firm having a bad month. The numbers say otherwise. A survey of more than two thousand senior decision makers this year found that 74 percent of enterprises have already pulled a live customer-facing AI agent back out of production after a governance failure. The rollback rate was higher, 81 percent, at the organizations with the most mature governance. Read that twice. The teams paying the closest attention are catching the most failures, because they are the only ones looking. The ones not seeing failures are not automatically safer. Often they are just not looking as closely.

The deeper research lands in the same place. MIT’s NANDA initiative found that the large majority of enterprise AI pilots deliver little to no measurable impact on the bottom line, and the reason was not model quality. It was that generic tools do not bend to the way the work actually flows. The exact figure gets argued. The direction does not. The model is rarely what breaks. The seam between the model and the workflow is what breaks. KPMG is that seam failing in public, with a logo everyone recognizes on it. Your version will not make the news. It will just go out to a client.

I spent years building the step they skipped

Before Swa, I spent most of my career in engineering, a lot of it leading large-scale reliability platforms at Nike. In that world you never let a change reach production without a gate in front of it. Not because your people are careless. Because the system is confidently wrong sometimes, and the gate is the cheap place you catch it before it gets expensive.

The other half of that discipline matters more here. When something slipped through, we did not write up the engineer who made the call. We ran a blameless review and fixed the process that let it through. You cannot fix people. People are tired, rushed, and trusting on a Friday afternoon, and they always will be. You can fix the system around them so that being tired on a Friday does not become a fabricated client report. That is why this is not a piece about KPMG being sloppy. Smart people at a serious firm shipped this. The system around them was missing one step, and the system is the thing you can actually fix.

AI took confidently wrong and moved it into every report, deck, and memo your company sends. All of it now has a fast, fluent first author that makes things up. If you quietly dropped the last gate to move faster, you did the exact thing KPMG did. You just have not been caught in public yet.

What this looks like when it works

Picture a team that produces client research. Before AI, a junior analyst spent two days on a draft, and a senior reviewer read the whole thing because there was time and the draft was short. Now AI produces a polished forty-page draft in an hour. It looks finished. It reads better than the junior analyst’s did. And the senior reviewer, faced with forty fluent pages and the same calendar, skims it, because skimming something that already looks done is human nature.

That is the trap. The output got faster and more convincing at the exact moment the review got shallower. The gate did not formally go away. It quietly stopped catching anything.

The fix is not to read all forty pages again. That just rebuilds the slow process AI was supposed to replace, and the team will route around it inside a week. The fix is to decide, in advance, which claims in those forty pages are load-bearing, and to check only those, every time, no exceptions.

The fix is one gate, on the claims that matter

Here is the part you can copy without a Big Four budget.

Here is the objection, and it is the right one. KPMG had reviewers too. A gate that skims forty fluent pages on deadline is the gate that just failed. So the fix is not one more reviewer. It is naming the load-bearing claims in advance, so a tired reviewer cannot skim past them. The numbers. The citations. Anything you say a named outside party did. The handful of things that can end a client relationship or a career if they are wrong. Those get a human signature before they leave the building. Everything else runs at full speed.

Naming those claims is the real work, and it is worth doing slowly. Most teams have never written down which of their outputs are load-bearing. They review by vibes, more carefully when they are nervous and less when they are busy, which is exactly backward. Sit down for half a day and write the list. For a client report it might be five things: every external statistic, every quote, every claim about a named company, every dollar figure, and any legal or compliance assertion. That list is your gate. The reviewer is not re-reading the document. They are confirming those five categories, and signing their name to them.

That is the whole move. The AI still drafts, and it is fast. A person still owns the last gate on the claims that carry weight. KPMG had the budget for a hundred reviewers and skipped the one that mattered. You do not need a hundred. You need one, on the right claims.

KPMG did not lose its credibility to a hallucination. It lost it to a missing step that costs almost nothing to add back.

One thing to do this week

Pick the one workflow where AI-drafted work leaves your company under someone’s name. A client deliverable. A board number. A published claim about somebody else. Walk it and find the exact point where the output goes out the door. Ask who reads it last, and what they are on the hook to catch. If the answer is nobody, or everybody, you just found the gap KPMG fell through.

Then put one person on the load-bearing claims. Name them out loud. Do it before the deadline picks for you.

Dana never needed a smarter model. She needed one name on one short list, written down before the deadline. Neither did KPMG. Neither do you.

Header artwork generated with Swa.


References

  1. KPMG pulls report on AI usage due to apparent hallucinations. TechCrunch, June 13 2026. link
  2. Investigation: only 5 of 45 citations in the KPMG report pointed to a real source; 40 fabricated; “vibe citing”; “poison the well of information.” GPTZero, 2026. link
  3. KPMG drops AI report after false case studies exposed (UBS, NHS Greater Manchester, Swiss Federal Railways, Transport for London dispute claims). International Accounting Bulletin / Financial Times, June 2026. link
  4. 74 percent of enterprises have rolled back a live AI customer-communications agent after a governance failure; 81 percent among the most mature-governance organizations. Sinch (commissioned survey of 2,527 senior decision makers), May 2026. link
  5. The large majority of enterprise generative-AI pilots deliver little to no measurable P&L impact, because generic tools do not adapt to enterprise workflows. MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025.” link

About the author

I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa, after years running reliability, data protection, and database operations at Nike. The lesson that keeps proving itself: anything you cannot run without, and cannot walk away from, is a risk you have not priced yet. What is yours?

Header artwork generated with Swa.

Read more of Ben's Enterprise Field Notes at benpickett.com.