The short version
- AI coding agents increased lines of code by thirty percent and pull requests by twenty-three percent across studied firms.
- The surge in generated code led to a forty-nine percent increase in review time, creating a bottleneck that absorbed efficiency gains.
- Despite widespread adoption, there is no statistically significant evidence that these tools have increased overall software output or reduced headcount.
The rapid integration of artificial intelligence into software development workflows has yielded a paradoxical result: while machines write code faster than ever before, the actual delivery of functional software features has not accelerated. A recent study conducted by researchers at Harvard University indicates that the efficiency gains achieved during the initial coding phase are effectively neutralized by downstream constraints in the production process. Specifically, the human review mechanism acts as a significant bottleneck, absorbing the time saved by automated generation and preventing firms from realizing broader productivity improvements.
The analysis draws on extensive aggregated analytics data from Jellyfish, a platform that tracks granular engineering outputs. The dataset encompasses three hundred million individual work events, including commits and pull requests, alongside issue management records. This information covers more than seven hundred thousand employees working across over seven hundred software development firms. The period under review spans from 2021 through March of 2026, capturing the trajectory of AI adoption as it moved from experimental use to widespread implementation within the industry.
Researchers Fiona Chen and James Stratton utilized a combination of directly measured AI usage metrics and GitHub activity analyses to pinpoint when individual companies introduced AI coding assistants or autonomous agents into their operations. By employing difference-in-differences regression models, they compared key performance variables before and after these technological shifts. This method allowed them to isolate the impact of AI tools from other concurrent changes in the software development landscape, providing a clearer picture of how automation affects team dynamics and output.
The raw data regarding code generation is striking. The introduction of AI coding agents resulted in a thirty percent increase in total lines of code produced. Additionally, firms saw a twenty percent rise in the number of commits and a twenty-three percent increase in pull requests on average. These figures suggest that AI tools are highly effective at accelerating the mechanical aspects of programming. However, this surge in volume did not translate into a proportional increase in completed software features or resolved issues.
Instead of streamlining development, the influx of AI-generated code placed substantial pressure on the review process. The average time required to review and merge pull requests ballooned by forty-nine percent after firms adopted these agents. Granular data further illustrates this strain: the share of pull requests requiring changes nearly doubled, and the number of comments left by reviewers increased by thirty-five percent. This indicates that human engineers are spending significantly more time scrutinizing and correcting machine-generated output than they previously did with human-authored code.
In response to this growing workload, organizations adjusted their staffing allocations. The study found a fourteen percent increase in the proportion of workers dedicated to performing code reviews following the introduction of AI agents. Despite this reallocation of labor, the overall resolution rate for major software features tracked by project management tools did not change in a statistically significant way. There was also no evidence of a compositional shift in the size or complexity of these issues, suggesting that the nature of the work remained consistent even as the volume of intermediate steps increased.
The potential for AI to alleviate this review burden remains limited at present. Although eighty percent of the firms in the study had implemented some form of AI-assisted code review by March 2026, these tools accounted for only twenty-three point three percent of all review comments and ten point eight percent of all pull requests. Humans continue to bear the vast majority of the responsibility for ensuring code quality and functionality. Consequently, the researchers found no significant changes in total active worker counts when cross-referencing Jellyfish data with LinkedIn employment records.
The findings highlight a transitional phase in software engineering where technological capabilities outpace organizational adaptation. While AI agents are now present in ninety-five percent of the firms studied, many teams are still navigating the complexities of integrating these tools effectively. The trade-off between coding speed and review time may improve as engineers gain experience with when and how to deploy autonomous agents. For now, however, the data suggests that increased coding velocity is counteracted by proportional increases in human oversight effort.
This dynamic presents a double-edged sword for the industry. The promise of AI lies in its ability to automate routine tasks and accelerate development cycles. Yet, without corresponding advancements in automated quality assurance or review processes, the benefits are diluted. As firms continue to refine their workflows, the challenge will be to balance the volume of generated code with the capacity for effective human verification. Until that equilibrium is reached, the overall output of software features may remain stagnant despite the apparent efficiency of individual coding tasks.
The study underscores the importance of looking beyond raw metrics like lines of code or commit frequency when evaluating technological impact. True productivity in software development depends on the seamless integration of generation, review, and deployment. Current AI tools excel at the first stage but struggle to deliver end-to-end improvements due to human-centric bottlenecks. Future research will likely focus on how evolving AI capabilities can address these downstream constraints, potentially unlocking the full potential of automated coding assistance.
Sources behind this briefing
Go to the original reporting
- Ars Technica↗AI coding agents generate more code, but not more software