Skip to main content
Test Double company logo
Services
Pragmatic Services Overview
Holistic software investment consulting
Acccelerate Software Delivery
Balance efficiency and quality
Improve Product Impact
Drive results that matter
Upgrade Rails Seamlessly
Update Ruby and Rails versions
Scale DevOps
Dev experience and infrastructure
Technical Recruitment
Build tech & product teams
Case Studies
Solutions
Legacy Modernization
Renovate legacy software systems
Pragmatic AI
Solve business problems without hype
Technical & Product Assessments
Uncover root causes & improvements
About
About
What's a test double?
Approach
Meeting you where you are
Founder's Story
The origin of our mission
Culture
Culture & Careers
Double Agents decoded
Great Causes
Great code for great causes
EDI
Equity, diversity & inclusion
Insights
All Insights
Hot takes and tips for all things software
Leadership
Bold opinions and insights for tech leaders
Developer
Essential coding tutorials and tools
Product Manager
Practical advice for real-world challenges
Say Hello
Test Double logo
Menu
Services
BackGrid of dots icon
Services Overview
Holistic software investment consulting
Software Delivery
Accelerate quality software development
Product Impact
Drive results that matter
Cycle icon
DevOps
Scale infrastructure smoothly
Upgrade Rails
Update Rails versions seamlessly
Technical Recruitment
Build tech & product teams
Case Studies
Solutions
Solutions
Legacy Modernization
Renovate legacy software systems
Pragmatic AI
Solve business problems without hype
Technical & Product Assessments
Uncover root causes & improvements
About
About
About
What's a test double?
Approach
Meeting you where you are
Founder's Story
The origin of our mission
Culture
Culture
Culture & Careers
Double Agents decoded
Great Causes
Great code for great causes
EDI
Equity, diversity & inclusion
Insights
Insights
All Insights
Hot takes and tips for all things software
Leadership
Bold opinions and insights for tech leaders
Developer
Essential coding tutorials and tools
Product Manager
Practical advice for real-world challenges
Say hello
Developers
Developers
Developers
AI

AI tooling may have solved coding, but not solutions delivery

AI made writing code faster, and software delivery didn't speed up to match. The Theory of Constraints predicted exactly this outcome forty years ago, and it also tells you what to do about it: measure your own end-to-end flow before you name the constraint or buy the tooling.
River Lynn Bailey
|
September 29, 2026
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

There's an interesting pattern that the software engineering industry has fallen into, with the sudden explosion of AI coding and tooling. A team brings in AI coding tools. Individual developers start producing more code than anyone can remember, and pull requests land faster than they ever have. Then a few months go by, and the release calendar either looks the same as it did before the tools showed up, or more likely, it's starting to show signs of slowing down. More code is being written. But the software isn't reaching users any sooner.

The interesting part, to me, isn't the pattern itself. Rather, it's that this exact outcome was predicted more than forty years ago, by a physicist writing about a factory.

Goldratt saw this coming, forty years before AI

In 1984, Eliyahu Goldratt published The Goal, a business novel about a factory manager trying to save his failing plant. The book teaches the Theory of Constraints (ToC): at any given time, every system has one constraint, or a small number of them, that sets the throughput of the whole system. Speed up any other stage and you have not made the system faster. You've made a bigger pile of work-in-progress in front of the constraint. Adding a lane to the on-ramp doesn't get anyone downtown sooner when the constraint is the bridge.

Goldratt never said a word about AI. He didn't need to. The theory's prediction for what happens when you dramatically accelerate one stage of a pipeline is unambiguous: if writing code was ever the constraint, speeding it up moves the constraint somewhere else, and after that, more coding speed stops paying. An hour saved at a stage that isn't the constraint does nothing for the system as a whole.

These ideas reached software through several paths. David J. Anderson applied ToC to software delivery directly in Agile Management for Software Engineering in 2003. The pull system he built at Microsoft in 2004, modeled on ToC's Drum-Buffer-Rope scheduling, grew into the Kanban Method. Mary and Tom Poppendieck arrived at the same warning about local optimization through the Toyota lean lineage, without citing Goldratt at all. And The Phoenix Project carried the factory-floor framing to a generation of DevOps readers. The warning has been sitting on our bookshelves the whole time.

"Coding was never the bottleneck" is an old claim, newly relevant

Rob Bowley traces the claim back through Fred Brooks in 1975, Steve McConnell in 1993, and Joel Spolsky in 2000. Developers have been telling each other that typing was never the slow part for as long as I've been in this industry.

In 2026 the claim resurfaced as a cluster. O'Reilly Radar, LeadDev, and an interview with DX's Brian Houck all argued a version of it within months of Bowley's post. Four essays across fifty years, all saying the same thing, is a discourse moment and not just four independent measurements. What is becoming more apparent  is that the delivery is data finally catching up to the claim.

The speedup at the keyboard may be real, but it's contested

The studies on individual coding speed don't agree with each other. The original GitHub Copilot experiment measured a 55.8% speedup on a narrow, synthetic task. Three field experiments at Microsoft, Accenture, and a Fortune 100 company found a pooled 26.08% increase in completed tasks across thousands of developers.

Then METR's randomized controlled trial found the opposite. Sixteen experienced open-source developers, working in early 2025 on real tasks in mature repositories they knew well, took 19% longer with AI. They had forecast a 24% speedup going in, and they believed they'd been 20% faster coming out. The contradiction between these studies is open and unresolved.

That perception gap - feeling 20% faster while measuring 19% slower - is exactly why you measure instead of trusting your gut. And it applies to more than coding speed.

The delivery-level data moves the wrong direction, across independent sources

Whatever is happening at the keyboard, the organization-level numbers agree on direction:

  • DORA's 2024 report found AI adoption raising individual productivity while delivery throughput and stability both got worse. The 2025 report saw throughput recover while instability still rose with adoption, and framed AI as "an amplifier, magnifying an organization's existing strengths and weaknesses."
  • Faros AI's telemetry across more than 10,000 developers found review time rising 91% and pull request size growing 154% on high-AI-adoption teams. Because Faros sells engineering analytics, though, it's fair to be skeptical of the precision. It is, however, reasonable to assume they are pointing at real changes in PR activity.
  • CircleCI's analysis of 28 million CI workflows found feature-branch throughput up 15% while main-branch throughput fell 7%, and main-branch build success hit a five-year low of 70.8%. Their conclusion names the destination plainly: "Writing code is no longer the constraint. Review, validation, integration, recovery" are where the delays accumulate.
  • AWS's enterprise strategy team describes AI-accelerated code production overwhelming merge, test, and deploy processes that were sized for a slower era - and prescribes eliminating constraints one at a time, without ever naming Goldratt.

‍

Code volume goes up. Review load goes up. Merge success goes down. And system stability goes down. That's a constraint moving downstream, right on schedule.

Four candidates for where the constraint landed

The sources agree the constraint moved. What they don't agree on is where it ended up. Generally, there are four candidates that recur in the studies and results. However, they're not mutually exclusive - nor are they an exhaustive list. To understand where the constraint actually landed, measurement of the real software development and delivery system must be done. These candidates are just that - candidates. They may be a common pattern in the studies and reviews, and they should be examined as possible locations, but they are not an exhaustive list. They represent a starting point for consideration, not an actual result from analysis.

Review and verification

Simon Willison puts it plainly: "the bottleneck is no longer how fast you write code, it is how fast a trusted human can be confident in a review." The Faros and CircleCI numbers above point in the same direction. This is the most visible queue in the whole pipeline, which makes it the loudest candidate, and loud is not the same as binding.

Product decisions and specification

Andrew Ng argues that teams can now build faster than they can decide what to build, and that product judgment is the scarce input. OpenCode's Dax Raad says the same thing from the trenches: "the daily bottleneck is still deciding what you should do, not doing it." Everything in this cluster is a convergent opinion from senior voices. Nobody has measured decision speed as a bottleneck yet.

Deployment and integration batching

A contrarian piece at The New Stack argues that the review bottleneck is a misdiagnosis, and that the real constraint is deployment batching after review. Its author works for Octopus Deploy, a deployment-automation vendor, so consider the source of the argument. But CircleCI's main-branch data and AWS's pipeline analysis are consistent with a constraint at or past integration, so it can't be waved away either.

Comprehension debt and coordination

Anthropic's own controlled study found junior engineers learning an unfamiliar library with AI assistance scored 50% on a comprehension quiz versus 67% without it, while finishing only about two minutes faster. One practitioner essay names the resulting condition "comprehension debt": code entering the system faster than human understanding of it. LeadDev's reporting frames the same gap as organizational toil - planning, handoffs, and decision ownership eating the gains. This candidate is the quietest of the four, because it doesn't show up on a dashboard until much later.

The trap: diagnosing by the loudest complaint

Before you pick a favorite from that list, look at how the evidence echoes. AWS's headline statistic - that 77% of organizations still deploy once a day or less - is DORA's number, re-cited. Faros's review figures resurface in write-ups at InfoQ and LogRocket until one vendor dataset starts to look like three independent ones. When the same few numbers bounce around the industry, the echo isn't corroboration.

Notice who's selling what, too. The review-bottleneck diagnosis arrives bundled with review tooling. The deployment-batching diagnosis comes from a deployment-automation vendor. The analytics vendor's data says you need analytics. Even the comprehension study comes from a company that sells the coding tools, though that finding cuts against its own interest. The one diagnosis with no clean product attached - product decisions - is also the one with no measurement behind it. None of this makes the diagnoses wrong. It makes them a bad substitute for looking at your own pipeline.

The Goldratt answer: measure first

Goldratt didn't leave us guessing about what to do when the constraint moves. The Five Focusing Steps are the whole playbook, and Yuval Yeret has walked them through this exact scenario. Applied to your delivery pipeline, they look like this:

  1. Identify the constraint - with end-to-end flow measurement across every stage, from "we should build this" to "a user is using it", not by listening for where the complaints are loudest
  2. Exploit the constraint - get more out of the constrained stage with what you already have, like making changes smaller and cheaper to review if review turns out to be your constraint
  3. Subordinate everything else - cap the work entering the constrained stage instead of piling more up in front of it, even if that means the fast stages sit idle sometimes
  4. Elevate the constraint - invest in real added capacity only after the cheaper steps are exhausted
  5. "Lather, rinse, repeat. Always repeat." - Homer Jay Simpson

‍

The constraint is the stage that limits end-to-end throughput, and that's not necessarily the busiest stage, the loudest one, or the one your favorite vendor has a fix for. Review might be your constraint. So might deployment batching. So might a product decision queue that nobody has ever measured because it lives in meetings instead of dashboards.

The entire point is that a system must be measured at the input and output of each step, to understand where the work backs up. This is the actual constraint - not a theory or guess; measured input and output, with data disclosing the system bottleneck.

Instrument your flow before you buy the tooling

If your team adopted AI coding tools and delivery didn't speed up, you aren't doing it wrong, and the fix is not more coding speed. The fix is knowing where your constraint went. Measure the time from idea to running-in-production, stage by stage: discovery, specification, coding, review, integration, deployment, validation. Find the stage that governs the whole. Exploit it by subordinating all other steps to its throughput, improve that step's throughput, then start over when the slowest step has moved. That's less exciting than buying a tool, and it's the only version of this that Goldratt would recognize as taking the theory seriously.

The developers in METR's study felt 20% faster while measuring 19% slower. Feelings about speed are cheap right now, and measurements are not. So when your team names its bottleneck, will it be naming the one it measured? Or the one that complains the loudest?

‍

River Lynn Bailey is a Senior Software Consultant at Test Double, and has experience in AI coding tools and delivery.

AI Insights and More

Sign up for the Test Double newsletter to receive more tips for pragmatic AI approaches.

Subscribe

Related Insights

🔗
AI amplifies everything (including what's broken). Start there.

Explore our insights

See all insights
Developers
Developers
Developers
Defend your downtime to improve agentic output

With agentic tools writing code, it feels natural to fill the waiting time by reaching for the next task. Before you do, though, have you considered what you might be losing by trying to parallelize more work? River Bailey digs into the question to understand why we should protect that downtime instead of filling it with more activity.

by
River Lynn Bailey
Developers
Developers
Developers
What Feels Easy Might Be Your Greatest Strength

We often mistake effort for value. By combining feedback, work history, and AI, James Walker discovered that the work that felt easiest was often what others valued most.

by
James Walker
Developers
Developers
Developers
Agentic Tooling as Accessibility

Most organizations expect agentic tooling to greatly increase the speed at which we write code. But there's another benefit in this new world that may not get as much as attention as it deserves: accessibility. River Bailey looks at her own experiences, and that of other developers, to better understand accessibility in the age of AI.

by
River Lynn Bailey
Letter art spelling out NEAT

Join the conversation

Technology is a means to an end: answers to very human questions. That’s why we created a community for developers and product managers.

Explore the community
Test Double Executive Leadership Team

Learn about our team

Like what we have to say about building great software and great teams?

Get to know us
Test Double company logo
Improving the way the world builds software.
What we do
Services OverviewSoftware DeliveryProduct StrategyLegacy ModernizationPragmatic AIDevOpsUpgrade RailsTechnical RecruitmentAssessments
Who WE ARE
About UsCulture & CareersGreat CausesEDIOur TeamContact UsNews & AwardsN.E.A.T.
Resources
Case StudiesAll InsightsLeadership InsightsDeveloper InsightsProduct InsightsPairing & Office Hours
NEWSLETTER
Sign up hear about our latest innovations.
Your email has been added!
Oops! Something went wrong while submitting the form.
Standard Ruby badge
614.349.4279hello@testdouble.com
Privacy PolicyTerms & Conditions
© 2020 Test Double. All Rights Reserved.

AI Insights and More

Sign up for the Test Double newsletter to receive more tips for pragmatic AI approaches.

Subscribe