Software companies do not ship lines of code. They ship reviewed, tested, integrated and maintainable changes — and that distinction is quietly dismantling the most popular productivity story in technology. A sweeping new study of real engineering telemetry has found that after firms adopt AI coding agents, the volume of code produced rises substantially while the amount of finished software barely moves. The gains, it turns out, are real — they are just being consumed by a part of the process nobody has yet managed to automate: deciding whether the code should be trusted.
The research is notable for its scale and its source. Rather than relying on surveys or self-reported feelings, the study draws on engineering telemetry from hundreds of firms covering roughly 300 million individual work events — commits, pull requests, reviews, deployments. That kind of evidence is rare in the debate over AI productivity, which has so far been dominated by vendor claims on one side and developer scepticism on the other. This time, the numbers come from the machinery of software development itself.
The study: what the data actually shows
The headline figures are striking. After companies introduced AI coding agents — tools that write, edit and refactor code, often autonomously from a prompt — lines of code rose by about 30%, commits climbed roughly 20% and pull requests jumped around 23%. By every upstream measure, the machines were doing exactly what their demos promised: generating code at unprecedented speed.
Yet the researchers found little evidence that completed software output rose alongside the volume. Releases, features and finished projects did not follow the same upward curve. This is the productivity paradox at its purest: an intervention that dramatically accelerates the earliest phase of production can leave the final output untouched if every downstream step remains human-paced.
The findings fit into a growing body of research telling a consistent story. An earlier study by Wharton and MIT researchers tracking more than 100,000 developers on GitHub from 2022 to 2026 found the same pattern across three successive generations of AI tools: autocomplete systems lifted coding activity by 40%, synchronous agents that edit alongside developers pushed the cumulative gain to 140%, and fully autonomous async agents drove it to 180% — yet software projects rose only about 50% and releases just 30%. As that study's authors put it, the binding constraint in software appears to be shifting from writing code to reviewing, integrating and ultimately distributing it. The new telemetry research confirms the shift in fine detail.
The review bottleneck, measured
The most revealing numbers concern what happens after AI-generated code is submitted. The average time between a pull request being opened and merged into the codebase balloons by about 49% once AI coding agents enter the picture. The share of pull requests sent back with changes requested nearly doubles. The number of comments reviewers leave per pull request rises roughly 35%. In response, firms have been forced to reallocate talent: the share of workers spending their time on code reviews has increased by about 14%.
Think about what that means in human terms. The senior engineers who used to spend their mornings designing systems and their afternoons writing the tricky parts now spend those hours auditing machine output — line by line, checking for subtle bugs, architectural mismatches, security flaws and the kind of quiet incorrectness that compiles cleanly but behaves wrongly. It is cognitively demanding work, it is expensive, and it scales badly: the more code the agents produce, the more review capacity the organisation must find.
Why the gains pile up at the wrong end
The paradox makes sense once you look at where AI helps and where it does not. Writing code was already the cheapest phase of software development — fast, automatable, and in many cases already faster than the thinking that precedes it. Generating a function that sorts a list or wires up an API endpoint was never the bottleneck; deciding what the function should do, proving it does it correctly, integrating it with a system of a million moving parts, and being willing to sign your name on it at 3am when it breaks — that was always the hard part. AI coding agents have now made the cheap part cheaper while leaving the expensive part untouched.
There is also a subtler dynamic at work. Code written by a human carries the author's understanding: the reviewer can rely on shared context, established patterns and the knowledge that the writer knows the system. AI-generated code arrives without that context. It may be stylistically flawless and structurally alien at the same time — using unfamiliar patterns, reinventing internal conventions, or solving the wrong problem very elegantly. That unfamiliarity forces reviewers to slow down precisely when the volume is telling them to speed up.
Can AI review its own output?
The obvious question is whether AI can simply take over the downstream work too. So far, the data suggests not. The study found that although some 80% of measured firms used AI-assisted code review by March 2026, AI agents were responsible for only about 23% of review comments and a mere 11% of pull requests — humans were still doing the vast majority of the trusting work. And there is a philosophical limit here: delegating review to the same class of systems that wrote the code concentrates risk rather than distributing it. If the generator hallucinates a subtle vulnerability, will the reviewer — trained on the same data, prone to the same blind spots — reliably catch it? The industry has not yet demonstrated that it will.
Notably, the study offers one piece of reassurance amid the gloom: the researchers said they could not attribute significant employment changes to AI adoption, after comparing workforce data across the firms and cross-referencing with professional network data. The machines are not replacing the engineers — but they are changing what engineers do, shifting more of them into the role of auditors of synthetic work.
What teams should do differently
For engineering leaders, the findings suggest the current generation of AI tooling is best treated not as a headcount replacement but as a rebalancing instrument — and one that needs deliberate management:
- Measure downstream, not upstream. Teams that track lines of code or commit counts will see only the AI's success story. The metrics that matter now are review latency, change-request rates, defect escape rates and time-to-release.
- Invest in review capacity deliberately. If agents multiply submission volume by a third, review capacity must grow in step — or the pipeline simply moves its queue from writing to reviewing, as the data shows it has.
- Constrain what agents produce. Smaller, well-scoped agent tasks — a single function, a defined refactor — generate output a human can actually verify. Unbounded autonomous work generates output nobody fully understands, and that is precisely the output that clogs review.
- Keep humans writing the hard parts. Architecture, security-sensitive code and novel problem-solving are where human authorship pays for itself — not because AI cannot produce them, but because a human author creates reviewable, trustworthy context.
The study does not argue that AI coding tools are useless; used well, they clearly accelerate genuine work. What it demolishes is the lazy assumption that faster typing equals faster shipping. In software, as in so much else, the bottleneck simply moved. The teams that thrive will be the ones that noticed where it went — and started investing there.
🤖 This article was rewritten by Feed and Figures' editorial AI from a report originally published by Ars Technica. Facts and quotes are preserved from the original; the rewrite focuses on clarity and structure. For the unedited original, see the source link below.