Enterprise Field Notes · Issue #23
Oracle Sped Up Coding With AI. The Wait Moved to Release. Has Yours?
Oracle's co-CEO told staff quicker code doesn't mean everything is quicker. Where the wait went, and how to find yours.
New here? Subscribe to Enterprise Field Notes, one new issue every week.
A town hall in September
At an internal town hall in the week of September 14, Oracle co-CEO Clay Magouyrk started with where the company had been a year earlier. "When I think about where we were a year ago," Magouyrk said, "I don't think we figured out how do we make AI really that useful for ourselves." Business Insider reported what was said.
As Business Insider put it, Oracle spent billions over the past two years helping other companies run AI, and its own companywide rollout came much later. It had made some progress with AI in customer support, but it hadn't found broad use with developers, finance or sales.
That changed in April and May, CIO Jae Evans told employees, when Oracle rolled out ChatGPT Enterprise and OpenAI's coding tool, Codex. According to Business Insider, the company paired the tools with its own standards, security controls and policies, and adoption reached 80 percent within three months.
Then Evans described what developers were doing with it. "We see developers able to generate code using this tool in like a week's time," Evans said. "Normally, that would have taken a team of developers two to three quarters to go develop."
Magouyrk put the other half next to it. "When you make the actual act of writing the code quicker," Magouyrk said, "it doesn't mean that suddenly everything is 1000 times faster." Business Insider's summary was that for now, AI hasn't gotten Oracle's products to customers any faster. Oracle is still working out how to redesign testing, validation, deployment and release management to keep up.
What I'll give Oracle credit for is the method. They put standards and controls around the tools instead of just handing them out, and the co-CEO told Oracle's own people that quicker code hadn't made everything faster. Evans was just as plain about cost. Adoption was so easy, Evans said, "we might have gotten a little bit of sticker shock," and Oracle now shows employees which models they use and what they cost. One model costs about two and a half times what the others do, so routine work can go to a cheaper one.
Before I go on, I should say I sell AI for a living, so weigh what follows however you like.
Then the review queue got longer
A change goes from written, to reviewed, to tested, to approved, to deployed, to running in front of a customer. For years, writing was slow enough that it set the pace for everything behind it. Speed it up and the work doesn't vanish. It lands on the next step, and it lands in bigger piles, because more code shows up at once.
DORA (DevOps Research and Assessment), the Google Cloud research program behind the software delivery metrics many engineering teams already track, found the same thing in its 2024 data. "Because AI allows developers to generate code much faster," its researchers wrote, "it often leads to larger batch sizes, which are slower to review and more prone to creating system instability."
So the review queue gets longer, and the test environment gets booked further out.
None of those steps got slower. They just became the slowest thing left.
There's a little arithmetic under this that I find useful. Little's Law, proved in 1961, says the average number of items in a queue equals the rate they arrive multiplied by how long each one waits, on average. Put in pipeline terms, the time a change spends between commit and customer equals the number of changes in flight divided by how many you ship a week. If AI doubles what arrives and your test step still ships the same number a week, there's no steady state left. The pile just grows.
It also grows faster than people expect. Donald Reinertsen's The Principles of Product Development Flow puts it in one line: "Capacity utilization increases queues exponentially." For a single reviewer or one shared environment, the standard approximation puts the wait factor at 4 when the step is booked 80 percent of the time, and at 19 when it's booked 95 percent. A pool of people or environments softens those numbers, but the curve keeps its shape. The same step that felt fine last quarter can feel stuck this one, with nothing about the step itself having changed.
A few test environments
Years ago, at a very large retailer, my constraint was pre-production.
We had a few test environments, and each one was only a partial copy of production. Refreshing one was slow, manual work. Someone had to restore and mask the data and then make sure all the dependent data was there. Even then, it was a subset, and its shape wasn't an accurate picture of production. Indexes and code that performed well in pre-production didn't perform the same way once they met production, and from where engineering sat, everything had passed. Teams that were ready to test also waited for a slot in one of those few environments.
I got the first attempt at fixing this wrong. I started with a single database. It was one of many, so our ability to stand up a full application stack on it was very limited, and it helped only a small number of teams. I hadn't spent enough time with the teams who owned those systems to understand how big the problem was or to earn their help.
The second time, I started with those teams. I kept regular touch points with them, and I respected their expertise and took it into account at every step. They're the reason the rest worked.
A weekend before Tuesday
The other wait was on a decision. We had a few days before a go-forward call on which technical direction to take. I was sure that virtual copies of masked production data were the fastest and most resilient option, and I knew a slide wouldn't win that argument, so we built it over a weekend.
On Tuesday the people making the call could see it working.
Forty, because the queue said so
After that, we built full pre-production environments, each running a complete application stack on virtual copies of masked production data. Production was tens of terabytes then, a lot of enterprise data at the time. Because nothing had to be manually copied for each one, every team could refresh its environment on demand, and some built refreshes into their automated path to production. A few teams still needed change control for their own reasons, so they refreshed in a more measured, structured way. The option was there for every team, and it was used often. We also built synthetic products and scenarios that matched the shape of production data into the refresh, so they were created automatically, and they turned out to be a very valuable addition.
We didn't pick how many to build. Application and quality teams kept asking for time to test, and we kept adding environments until there were enough slots for the testing and development they needed. We'd discouraged performance testing on those environments, but the performance teams wanted the data, so we enabled them too. We went from a few partial copies to forty full ones, and from helping a handful of teams to helping all of them.
Faster engineers would only have made that queue longer. We added capacity at the step where the work was waiting, which is the testing and validation step Oracle is redesigning now, and the queue told us how much. Reinertsen has a principle for that too: "Don't control capacity utilization, control queue size."
That was years ago, when there weren't many good options for any of this. Much of it is far easier to get now, in the cloud and elsewhere. The complexity is still real, though. With that many teams and applications to account for, masking data while keeping a shape that's still meaningful for testing, and getting it into enough full environments to keep every team moving, is hard work even with today's tools.
A refresh button away
One lesson came after the fact. As new people rotated in, mostly contractors, we found each new group had to be taught the faster way. It was so different from what they knew that they didn't go looking for it, and some went back to writing scripts to copy data that was one refresh button away.
So we wrote onboarding documents and kept regular check-ins with the teams. We also watched the environments themselves.
If one wasn't being refreshed, we asked why.
Oracle's 80 percent is an adoption number. We learned we needed a second one, whether the new way was actually being used, and for us that meant watching the refreshes.
A year of findings in two weeks
Oracle also got early access to Anthropic's Mythos Preview, which Business Insider describes as a model that scans code for security vulnerabilities. In its first two weeks it turned up more potential security issues than Oracle had found in an entire year.
Evans estimated that roughly 60 to 70 percent of them were false positives, and Oracle built new processes to verify vulnerabilities before engineers began fixing them.
It's the same shape as the coding story. The model raised the count of potential issues, and checking them became the slow step, so Oracle put a verification step in front of its engineers. Without it, most of their fixing time would have gone to issues that weren't real.
Meta moved the tests
Oracle is far from alone in what it found. In January, Meta CFO Susan Li reported a 30 percent increase in output per engineer since the start of 2025, most of it from agentic coding. Two weeks later, Meta's engineering blog described what that pace did to testing. "Agentic development dramatically increases the pace of code change," it said, "straining test development burden and scaling the cost of false positives and test maintenance to breaking point."
Meta's answer was to move the testing to where the new code arrives. It described generating tests automatically for each change, the moment a pull request is submitted, and those tests "only require human review when a bug is actually caught." We made the same move with environments years earlier, at a different step. Put the capacity where the new work lands, and let the teams use it themselves.
The volume isn't slowing down. Alphabet's CFO, Anat Ashkenazi, said on the company's February earnings call that about half of Alphabet's code is now written by coding agents and then reviewed by its own engineers.
The broader data points the same way. Faros AI, which sells engineering analytics, published telemetry this April from about 22,000 developers across more than 4,000 teams on its platform, comparing each organization's periods of lowest and highest AI adoption. Median time in review was 441.5 percent higher in the high-adoption periods, and pull requests merged with no review at all, human or agent, were up 31.3 percent.
DORA's more recent work adds the honest complication. Its researchers wrote in March that "higher AI adoption is associated with an increase in both software delivery throughput and software delivery instability."
The questions to take to work
DORA defines change lead time as the time it takes a change to go from committed to version control to deployed in production. That clock starts after the code is written, so it covers the part of the job Oracle says AI hasn't sped up yet. I'd run it one step further than DORA does, to the moment a customer can actually use the change.
For one change your team shipped last month, how much of the time from commit to customer was someone working on it, and how much was it waiting for review, a test environment, data, an approval or a release window?
If your developers produced twice as much code next quarter, which of those waits would double first?
If your company is cutting headcount this year, is any of it at the step where the work already waits?
When a test passes in your pre-production environment, how much does that environment's data look like production's?
If you added capacity at your slowest step tomorrow, who would tell you how much was enough, and who would notice if people kept working the old way?
If you run release or test, where does your pipeline wait? Repost this with the step you'd name, and add the working and waiting split for one real change if you have it.
Subscribe to Enterprise Field Notes for one field report a week. Everything I've written is in my archive.
References
- Oracle exec tells workers that its own AI rollout didn't go so smoothly. Business Insider, Ashley Stewart, September 18 2026 (exclusive, subscriber content). Magouyrk: "I don't think we figured out how do we make AI really that useful for ourselves" and "it doesn't mean that suddenly everything is 1000 times faster." Evans: "We see developers able to generate code using this tool in like a week's time. Normally, that would have taken a team of developers two to three quarters to go develop" and "we might have gotten a little bit of sticker shock." Also the April and May rollout, the 80 percent adoption within three months, the standards and controls, the model cost comparison, Mythos Preview "which scans code for security vulnerabilities," and the 60 to 70 percent false positive estimate.
- Impact of Generative AI in Software Development. DORA, 2025, from its 2024 survey data. "Because AI allows developers to generate code much faster, it often leads to larger batch sizes, which are slower to review and more prone to creating system instability."
- Little's Law as Viewed on Its 50th Anniversary. John D. C. Little, Operations Research 59(3), 2011, pp. 536 to 549. Original proof: "A Proof for the Queuing Formula: L = λW," Operations Research 9(3), 1961.
- The Principles of Product Development Flow. Donald G. Reinertsen, Celeritas, 2009. Principle Q3, "Capacity utilization increases queues exponentially" (p. 59), and Q13, "Don't control capacity utilization, control queue size" (p. 75).
- The single server queue in heavy traffic. J. F. C. Kingman, Mathematical Proceedings of the Cambridge Philosophical Society 57(4), 1961. Kingman approximates mean wait as three terms multiplied together; the utilization term, utilization divided by one minus utilization, is the one the body quotes at 80 and 95 percent.
- Alphabet Q4 2025 earnings call transcript. February 4 2026. Anat Ashkenazi on code written by coding agents and reviewed by Alphabet's engineers.
- Meta Q4 2025 earnings call transcript. January 28 2026. Susan Li: "Since the beginning of 2025, we've seen a 30% increase in output per engineer, but the majority of that growth coming from the adoption of agentic coding."
- The death of traditional testing: agentic development and the JiT testing revival. Meta Engineering, Mark Harman, February 11 2026. Paper: arXiv 2601.22832.
- AI Engineering Report 2026: The Acceleration Whiplash, takeaways. Faros AI, April 12 2026. About 22,000 developers and more than 4,000 teams, comparing each organization's lowest and highest AI adoption periods.
- Balancing AI tensions. DORA, Jessica Baolin and Nathen Harvey, March 10 2026.
- DORA's software delivery performance metrics. DORA. Change lead time: "The amount of time it takes for a change to go from committed to version control to deployed in production."
About the author
I'm Ben. I write Enterprise Field Notes, and by day I'm COO at Swa. Before that I was Global Director of Site Reliability Engineering at a Fortune 500 retailer, running reliability, data protection, and database operations at global scale.
Header artwork generated with Swa.
Read more of Ben's Enterprise Field Notes at benpickett.com.
© 2026 Ben Pickett · Enterprise Field Notes