Enterprise Field Notes · Issue #19

Meta Changed Its Internal Tools 220% More. Major Incidents Rose 40%.

Meta measured the experiment everyone else is running on faith. The numbers surfaced three weeks after the decision they are assumed to explain.

By Ben Pickett · August 27, 2026

New here? Subscribe to Enterprise Field Notes, one new issue every week.

An oil painting in the Norman Rockwell manner. A man in a navy blazer sits at a table beside a large window looking out over an Oregon coastal bay at sunset, two sheets of paper in front of him. Beyond the glass an arched bridge crosses the bay, a raft of logs waits beside a timber ship, and a peregrine falcon is perched on the branch of a wind-bent shore pine, poised to leave. A swallowtail butterfly passes outside the window.
Two sheets on the table. Only one of them ever gets read.

The only question on-call ever asked

For most of my career I was accountable for how much of what we shipped woke somebody up.

I have also been a falconer since I was fifteen. That is how I came to be president of the Oregon Falconers Association, and how I ended up holding a document nobody would look at.

We almost lost the peregrine falcon.

What brought it back was decades of extraordinary effort by a great many people, falconers among them, who bred the first birds for release and climbed to nest sites and kept at it for longer than most careers run. The recovery goal was written down in 1982, by a team that had no idea whether it would ever be met. By the time anyone got around to counting properly, the population in Oregon had not simply reached that goal. It had gone well past it, and past the historic recorded numbers the goal had been set against in the first place.

The document was the proof of that. A retired biologist from the US Fish and Wildlife Service had spent years producing it, a book length petition to take the peregrine off Oregon's endangered species list altogether, and he had it peer reviewed by biologists at universities, outside the agency that would be ruling on it.

So it was not really a request. It was an acknowledgement of a milestone that a lot of people had spent their working lives earning, and he had the counts to prove it.

None of that mattered on its own, because the petition had to be submitted at the commission meeting itself, in front of them, and we were told plainly that we would not be given the floor. Whether that was procedure or discretion I could not tell you now. What I know is that it was the gate, because I would not have spent those months collecting letters if there had been a simpler way through.

So I went and got letters. Hunting groups, shooting associations, conservation groups, organizations that serve different Oregonians and answer to different priorities. By the time we were finished we were carrying letters representing tens of thousands of Oregonians, and the stack had gotten large enough that declining to look at it would have been its own decision, made by somebody, on the record.

They granted the audience. He presented. They accepted the petition, and in April 2007 the commission removed the peregrine from Oregon's endangered species list.

I did not write that petition and I want to be careful about the credit. Years of rigorous work belong to the biologist who did them. What I did was make the number impossible to leave unread.

It had existed the whole time. Nobody was obliged to look at it.

That is the shape I keep running into, and it is not how most organizations behave. Every engineering org I have worked in could tell you its output. Commits, pull requests, story points, releases. Very few could tell you, without going and looking, what all that output cost them in incidents, or how many hours the team spent last month cleaning up work it had already been paid for once.

Output is easy to count because it accumulates. Failure is hard to count because somebody has to define it first, and then somebody has to be willing to look at the number after they define it.

And this week it turned up inside a company that did write its failures down. Reuters published the best evidence I have seen of what output actually costs at enormous scale, inside an organization that was measuring both while it happened.

What Meta was actually doing

At Mark Zuckerberg's annual leadership retreat in Hawaii in January, Meta's executives conceived a plan called Project OT, for Organization Transformation. The goal was an AI native company: work restructured into small pods, AI agents absorbing a large share of what employees had been doing, and smaller groups of humans supervising the agents.

Internal planning documents modeled reducing some teams by as much as 60 percent. The restructuring was to run in two waves, one in May and a second in November.

In April, Meta mandated tracking software on US employees' machines to capture their keystrokes and mouse clicks, in order to teach its AI agents to replicate how humans interact with computers. Reuters reported that at the time.

People were asked to install software that recorded how they worked so that a system could learn to do their work.

Meta's response to Reuters is worth quoting rather than characterizing: "As part of our company restructuring earlier this year, we asked some teams to conduct a scenario planning exercise looking at the potential impact of redeployments, open role closures and cuts," and the company "didn't move forward with every scenario from the exercise, and it was never assumed we would."

On the evening of May 19, hours before the first wave, Zuckerberg called off the November one. The first wave went ahead the next day and cut about 10 percent of the company.

The four numbers

Here is what Meta's own internal reporting showed about the period in question.

Code changes made to the internal software platforms and infrastructure employees use on the job were up 220 percent year over year.

Changes that led to new or upgraded features reaching Meta users were up 36 percent.

Major technical and security incidents, the kind Reuters describes as service disruptions and possible data leaks, rose 40 percent.

Time spent dealing with those problems grew 70 percent.

It is the cleanest picture anybody has published of what happens when you put a lot of generative capacity into an engineering organization and do not change anything else about how that organization works.

Start with the first two. Code changes grew roughly six times faster than delivered features. Both of those are counts of changes. What separates them is scope: the first counts changes to the platforms employees work on, the second counts changes that put something in front of a user, and nothing published says one feeds the other. There is also a reading where much of that 220 percent is infrastructure work, refactors and test coverage that pays off over a longer horizon than one year of feature delivery. I do not think that accounts for a six-fold spread, and I cannot rule it out from outside. So the comparison is fair in kind and still imperfect, because a great deal of legitimate work on internal platforms never was going to reach a user. What I can say is that the two moved very differently, and only one of them is the reason the work exists.

One caveat belongs here before the third and fourth numbers, and it cuts against my argument. An organization that spends a year instrumenting itself finds more of what it is looking for, so some part of a 40 percent rise in recorded incidents can be a detection improvement rather than a reliability decline, and nothing published tells us how much.

Nor is the rise necessarily about AI at all. A 10 percent workforce reduction, a reorganization into new team structures and a large jump in change volume all landed in roughly the same window, and any of those raises incidents on its own. Reuters reported a correlation. I am telling you it is worth your attention, not that the mechanism is settled.

Even so, the third and fourth numbers tell you where a good deal of the work went. Incidents up 40 percent, and the time spent on them up 70 percent, which means incidents did not just get more frequent, they got more expensive to clear. That is the signature of failures arriving faster than the understanding needed to fix them.

The division I owe you before I go further

If you spent any time on an SRE team you already ran this in your head while reading, and if I skip it you should stop trusting me.

Changes went up 220 percent, which is 3.2 times. Major incidents went up 40 percent, which is 1.4 times. Divide the second by the first and incidents per change fell by roughly 56 percent.

The 220 percent is scoped to internal platforms and infrastructure. The incident figure reads as company-wide. So the division is indicative rather than exact, and anybody presenting it as a clean rate, including me, is rounding off a scope mismatch.

On the rate that reliability engineering actually uses, Meta got substantially safer during this period. Each individual change was far less likely to break something than a change had been the year before.

I am not going to pretend that away, because it is real and it is the single strongest thing anyone can say in defense of what Meta was doing.

Here is why the absolute numbers still matter more, and it is the whole argument.

Nobody pays a rate. You pay the absolute. The on-call engineer woken at three in the morning is not consoled that the per-change failure rate improved. And that elevated volume of failure now lands on a headcount that was cut by 10 percent in May, absorbed by people who have other work to do. Rates are how you evaluate a process. Volumes are how you staff one, and Meta increased the volume of failure by 40 percent in the same year it reduced the number of people available to absorb it.

This is an old trap and it is not specific to AI. Improve the defect rate, raise the throughput faster than the rate improves, and you have a better process producing more damage.

Google's own DORA research has been circling this for two years. Its 2024 report estimated that every 25 percent rise in AI adoption came with roughly a 7.2 percent drop in delivery stability. The 2025 edition found throughput had turned positive while the negative relationship with stability persisted. Meta's numbers are a single company's version of a pattern that shows up across thousands.

Output accumulates whether or not anything is working. That is exactly why it is the number everybody reports.

Why a reliability person reads this differently

Code changes is an odometer. It counts up, it never counts down, and it rises whether the system is healthy or on fire. Every organization publishes something in this family, because it is easy to instrument and it always looks like progress.

Incidents and remediation time are the other kind of number. They can go down. That is what makes them worth having, and it is also why they are so often absent from the slide.

What is genuinely unusual about Meta here is that they had both, and that somebody wrote them down where they could be read.

Reuters traces all four figures to posts by Andrew Bosworth in early June, which is two to three weeks after the November wave was called off on May 19. So these numbers were not the dashboard somebody stared at before making the decision. They surfaced afterward.

It forecloses the story everyone wants this to be.

What the reporting does and does not establish

Reuters reported the conditions around the decision: the productivity gains had not materialized, investors were unhappy about AI spending, and employees were openly unhappy. Other outlets have written that those pressures forced the reversal. What I have not seen is a document or a quote that says Zuckerberg canceled the second wave because of the incident numbers.

So I am not going to tell you the four numbers killed the November layoffs. The numbers exist, the cancellation happened the evening before the first wave, and the public record does not connect them for me.

That distinction matters more than it might seem. If you take from this that data changed a CEO's mind, you will go looking for a dashboard that wins arguments, and the timing above rules that reading out anyway. The version I would actually stand behind is smaller. Meta is a company that generated these numbers at all, published them internally where colleagues could read them, and did not suppress them when they went the wrong way. That is a property of the organization rather than of any single decision, and it is the property most companies running this experiment do not have.

Also worth saying, since everything above comes from one investigation: this is a single Reuters report, and the April tracking story was Reuters as well. What hardens it is that Meta answered on the record and confirmed the program existed rather than denying it. That is more than most stories of this kind get. It is still one newsroom.

Meta is not the only one who published half of this

If Meta were the only data point, I would be more careful with it. It is not.

On its second quarter earnings call on July 29, Cognizant's CEO Ravi Kumar told analysts the company has "over 8,000 AI engagements," and, separately on the same call, that "over 40% of our software development is now AI-assisted." In the same call he gave the outcome number: "our revenue and adjusted operating income per associate increased 4.6% and 7.1%, respectively."

Forty percent of software development assisted by AI, and a 4.6 percent move in revenue per person. A 4.6 percent move across 356,700 people is a large amount of money, and Cognizant is not presenting it as a failure. I will not call it a productivity gain either, since revenue per associate rises on its own when the denominator shrinks, and the two numbers are not commensurable enough to divide. What matters is this: the input metric was said loudly and the outcome metric was said quietly, on the same call, by a CEO who had no obligation to give the second one at all.

Two companies, two industries, the same shape. A large, easily counted input, and an outcome that moved by much less than the input did. Meta is the only one of the two that also published what it broke.

The part that has nothing to do with AI

Employee sentiment in Meta's internal Pulse survey fell 19 points, from 74 percent favorable to 55 percent, after the tracking software was announced.

I would not skip past that on the way to the engineering numbers, because it is a measurement too, and it is the one most likely to be repeated elsewhere.

Consider what was being asked. Install this so we can learn how you work. The stated purpose was to train agents to replicate that work, during a year in which the company had modeled cutting some teams by more than half. Whatever the intent, the arrangement asks people to contribute to something they have reason to believe is aimed at them.

There is a version of instrumenting work that people accept, and most engineers have lived inside one. Error tracking, performance monitoring, incident review. What makes those tolerable is that the subject of the measurement is the system, and the beneficiary of the improvement is the person doing the job. Reverse either of those and you get a different thing wearing the same clothes.

Nineteen points is what Meta's own half-year Pulse survey recorded. I say recorded rather than caused deliberately, and Reuters places the drop against the whole restructuring period rather than the tracking mandate alone. By the time that survey ran, the May layoffs had happened and the scenario planning was known inside the company. Any one of those moves a sentiment score.

The questions to take to work

I am not going to hand you Meta's lesson, because you have a different company and I cannot see it. But their numbers make five questions answerable inside your own organization this week, and none of them need a new tool, including the kind my company sells.

How much more is your engineering organization producing this year than last, and where is that number reported?

Do you have the matching number for what broke, and does anyone present them on the same page?

If change volume went up and delivered outcomes did not, who would notice, and how long would that take?

What did your last quarter cost in remediation time, and is that a number anybody is accountable for?

And if you are measuring how your people work in order to automate it, has anybody asked them what they think that is for?

That last one is not a soft question, and it is the one I would ask first. What Meta has done is show you the size of the thing you are not currently measuring.

The prediction, with a number on it

Here is what I expect. Within twelve months, at least one large company will publicly report an incident or defect rate alongside an AI productivity claim, specifically because this reporting made the absence conspicuous. If that happens, the 220 and the 40 will have done more for this field than any vendor benchmark published this year.

If nobody does, then the honest conclusion is that output numbers are reported because they flatter, and failure numbers stay internal because they do not, and we should stop pretending the reporting gap is a technical problem.

Either way, you do not have to wait for anyone else to publish first. Both numbers already exist inside your company, and only one of them has ever made it onto a slide.

Which brings me back to the falcon, and to the thing I actually learned in Oregon.

Somebody wrote that number down in 1982 and then died waiting, or retired, or moved on. The counting happened. The goal was passed. And the petition still sat finished for years, because what was missing was never the science. What changed was that declining to look stopped being free. Once the letters were on the table, somebody had to put their name to the choice not to read it.

Numbers do not act on their own. Somebody has to make ignoring them cost something.

Meta's four numbers arrived three weeks after the decision they are supposed to explain. Not because nobody could count.

I made you the sheet. One page, one row per function: the output number your company already reports, and the number that belongs next to it. Only the engineering row has published evidence behind it, and I say so on the page. The rest is the same rule applied where nobody has run the experiment yet. Get the Paired Metric Sheet


References

  1. How Zuckerberg's plan to replace Meta staff with AI unravelled. Katie Paul, Reuters, syndicated August 26, 2026. Project OT, the January Hawaii retreat, the up to 60 percent team-cut scenarios, the two waves, the four internal metrics, the Pulse survey drop, and Meta's on-record response.
  2. Meta reportedly abandoned an AI-focused restructuring plan that would have laid off thousands. Will Shanklin, Engadget, August 26, 2026. Independent restatement of the metrics, and the note that Reuters did not establish Zuckerberg's specific reasons for canceling the second wave.
  3. Meta will start tracking employees' screens and keystrokes to train AI. Fortune, April 21, 2026, reporting on and crediting Reuters. Contemporaneous coverage of the tracking mandate, not an independent confirmation of it.
  4. Announcing the 2025 DORA Report. Google Cloud, 2025, and the 2024 Accelerate State of DevOps Report. Source of the AI adoption to delivery-stability relationship across thousands of organizations.
  5. Cognizant (CTSH) Q2 2026 Earnings Call Transcript. July 29, 2026. Ravi Kumar on 8,000 AI engagements, over 40 percent of software development AI-assisted, and revenue and adjusted operating income per associate up 4.6 and 7.1 percent.
  6. Findings on Removal of Both Subspecies from the State List of Threatened and Endangered Species. Oregon Department of Fish and Wildlife, April 13, 2007. The Commission's removal of the American and Arctic peregrine falcons from Oregon's Threatened and Endangered Species list.
  7. Employee revolt and failing agents forced Meta to scrap its AI layoff plan. The Decoder, August 2026. An example of the causal framing this issue deliberately does not adopt.

About the author

I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa. Before that I was Global Director of Site Reliability Engineering at a Fortune 500 retailer, running reliability, data protection, and database operations at global scale.

Header artwork generated with Swa.

Read more of Ben's Enterprise Field Notes at benpickett.com.