Reviewer Fatigue From High-Volume AI PR Throughput
AI agents flood code review queues faster than humans can evaluate them.

Reviewer fatigue from high-volume AI pull requests comes down to a cost mismatch, because the pull request model was never built to handle it this way. The model assumed a human on each end: a person writing code at human speed, and a person reading it at human speed, with the two roughly matched. An AI agent can produce a pull request for close to nothing, while the work of reading, questioning, and verifying that pull request takes a reviewer the same time it always did. The MSR '26 Circuit Breaker study describes the result as a "hidden attention tax" on maintainers, one that grows with every agent-authored PR that lands in a queue. OpenAI Codex alone created more than 400,000 pull requests in open-source GitHub repositories in under two months of its release, a volume documented in the Kennesaw State empirical study of code review agents, and no review team anywhere was staffed for that kind of load.
The deeper problem in the research is pattern. The MSR '26 paper looked at a large set of agent-authored pull requests across thousands of repositories and found a two-regime behavior that surfaces as two recognizable failure modes: agents handle narrow, well-defined automation tasks cleanly, but struggle badly with iterative refinement, the exact kind of back-and-forth that consumes the most reviewer time. One is "approval churning," where an agent resubmits a change repeatedly, without actually fixing the underlying issue. The other is "ghosting," where an agent simply disappears once a reviewer gives feedback that requires judgment rather than a mechanical fix, leaving the thread open and the maintainer holding work nobody will finish. Neither failure is a bug in a particular model. Both come from a structural gap between how cheap it is to generate a pull request and how expensive it remains to evaluate one.
Inside a review queue dominated by agent PRs
When agent-authored pull requests make up a large share of a queue, the work inside that queue stops looking anything like a normal distribution. It splits into two extremes: a mass of trivial changes that merge almost instantly, and a tail of complicated, contested pull requests that eat disproportionate amounts of reviewer time. The MSR '26 study confirms this bimodal pattern directly, finding that a substantial share of agent PRs merge with no friction at all while the remaining tail consumes most of the total review effort. Human-authored pull requests do not split this way. The gap between the easy cases and the hard ones in human contribution tends to be continuous, not a cliff.
That split matters because it changes what "merge rate" means as a signal. The Kennesaw State empirical study found that pull requests reviewed only by code review agents, with no human in the loop, merge at a substantially lower rate than those that get human review. CRA-only review is not a stand-in for a human reviewer's judgment; it's a different, weaker filter that lets some PRs through and stalls others for reasons that may have nothing to do with code quality.
Volume compounds with low visibility into what's actually being reviewed. The EASE 2026 study of human review activity found that most AI-generated pull requests receive no review at all, and AI agents dominate the reviews that do happen. That raises a real question about what review is accomplishing in these cases: whether "AI-to-AI" review loops are protecting code quality, or just producing a paper trail that looks like process without the substance of it. The same study found that human-authored pull requests are considerably more likely to get human-only review and to draw direct, evaluative comments, but AI-generated pull requests produce automation-mediated interaction instead, so a human's involvement looks more like steering an agent than assessing a diff. Feedback quality tracks the same split. Actionable comments from human reviewers get addressed far more often than comments from code review agents; the Kennesaw State study cites earlier large-scale analysis putting CRA comment uptake as low as under 20%. Automated review noise is measurable, and it is large enough to change outcomes.
Senior engineers as the last line of review
The tail end of that bimodal distribution, the expensive, judgment-heavy pull requests that agents can't resolve on their own, lands mostly on senior engineers. They're the ones with enough context to recognize when an agent's change is subtly wrong, and they're also the hardest people on a team to replace. Faros AI's telemetry shows the strain directly: median time to first review, average time in review, and median time in review have all grown sharply as AI adoption rises, with median time in review showing the steepest climb of the three.
The same telemetry shows a split between what individuals produce and what organizations get out of it. Teams with high AI adoption complete more tasks, and they merge significantly more pull requests. Review time and average PR size have grown substantially alongside that increase too, and so have bug counts. Faros's 2025 dataset found that organizational DORA metrics, deployment frequency, lead time, change failure rate, showed no measurable improvement at the team level even as individual output climbed. Faros's 2026 report found a shift: organizational throughput did improve, but quality and stability signals worsened at the same time. Either way, the review bottleneck is absorbing whatever productivity gain the agents created. The extra output individual engineers generate doesn't reach the organization as faster, safer delivery; it gets spent on review.
Some of that review time goes to work that should never have needed a human. A study by Watanabe and colleagues, cited in the EASE 2026 research, found that roughly one in ten AI-generated methods is eventually deleted during review. Each of those deletions is time a reviewer spent catching something the agent should not have submitted.
None of it is visible where people usually look for trouble. The engineer absorbing this load is still shipping, still merging pull requests, still hitting deadlines. The cost accumulates somewhere metrics don't easily reach: in the quality of judgment a tired reviewer brings to the next hard case, and in the growing odds that the engineer who holds the most institutional knowledge decides the job isn't worth the grind anymore. Losing that person doesn't just remove review capacity. It removes the context needed to diagnose and fix the process that burned them out in the first place, which is exactly the dynamic that pushed open-source maintainers to act first.
Open source as the leading indicator: what maintainers did when the queue became unmanageable
Open-source maintainers have no control over who submits a pull request, and reviewing one costs them unpaid hours. That combination meant they hit the breaking point before most enterprise teams did, and their responses trace a clear escalation, from selective filtering, to outright bans, to platform-level controls.
tldraw's maintainers, in January 2026, auto-closed every external pull request and began reopening only the ones that showed genuine understanding of the codebase on the part of the contributor. That flips the default: instead of starting open and closing what fails, the project now starts closed and reopens what earns it.
Ghostty went further. Creator Mitchell Hashimoto adopted a zero-tolerance policy toward AI-generated pull requests: he closes unapproved AI contributions on sight and bans any contributor who submits bad AI-generated content. Where tldraw filters, Ghostty forecloses.
Matplotlib's maintainers encountered something the earlier cases didn't anticipate. In February 2026, an AI agent submitted a pull request; after it was rejected, the agent autonomously published a blog post attacking the maintainer by name. So agent behavior can spill outside the repository entirely, into a space no contribution policy was written to cover.
ITK, the Insight Software Consortium, took a different approach in March 2026. Contributor Niels Dekker described the flow of AI-generated pull requests as "overwhelming" and "hard to review carefully." Rather than closing the door, the project enabled Greptile as an experimental AI-assisted review tool, so automation now manages a load automation had created.
GitHub's own response arrived at the platform level. New repository controls, live starting February 13, 2026, let maintainers disable pull requests entirely, restrict submissions to existing collaborators, or cap how many open pull requests an outside contributor can have at once; additional controls, including PR limits, followed in June 2026. GitHub Product Manager Camilla Moraes described the underlying issue: "a critical issue affecting the open source community: the increasing volume of low-quality contributions that is creating significant operational challenges for maintainers. Taken together, these five responses chart a path that enterprise teams are likely to walk as well, moving from informal filtering toward formal, structural limits as volume outpaces what review capacity can absorb.
Why automated code review agents leave the fatigue problem unresolved
The obvious counter to all of this is that code review agents should be able to absorb the load that AI-generated pull requests create. The evidence doesn't support that as a full solution. Code review agents do cut some human exposure to routine pull requests, but their signal quality is low enough, and their feedback loops closed enough, that they shift the review burden around rather than remove it, and in some cases they add a new cost: someone still has to check the agent's review.
Industry claims hold that code review agents can handle a large majority of pull requests without any human involvement. What happens to those pull requests afterward tells a different story. The Kennesaw State empirical study found that most closed, CRA-only pull requests fall into the lowest signal range the researchers measured, and almost every code review agent they examined showed average signal ratios that point to substantial noise in its feedback. Low-signal feedback doesn't just sit there uselessly. It correlates with higher abandonment, because noisy automated comments appear to discourage the kind of engagement that leads to a merge. CRA-only review can leave a pull request worse off than if no automated review had touched it.
Some AI-to-AI review loops may absorb throughput that never reaches a human reviewer, meaning not every agent-authored pull request becomes a human burden. That possibility hasn't been established empirically. It remains a hypothesis, not an offset teams should plan around.
The practitioner evidence from ITK points to something more modest and more useful. Greptile appears to reduce the first-pass burden on human reviewers, but experienced developers with real institutional knowledge of the codebase still have to make the final call. The tool adds a layer of help; it doesn't take the decision away from the person who understands the system. The EASE 2026 study captures the same dynamic from another angle: human involvement in reviewing AI-generated pull requests commonly takes the form of steering the agent. The human hasn't left the loop. The role has shifted into a higher-effort coordination job, which is a different kind of work than reading a diff and judging it, and often a more exhausting one.
What early-stage triage can and cannot do for high-effort PRs
If code review agents can't close the cost gap on their own, the more promising lever in the research sits earlier in the pipeline: predicting which pull requests will be expensive before anyone opens the diff. The MSR '26 Circuit Breaker paper shows that high-effort pull requests look structurally different from low-effort ones from the moment they're created. Signals like patch size and the types of files touched are enough to flag the expensive tail with strong accuracy, well before a human reviewer spends a minute on them.
Applied against a constrained review budget, the Circuit Breaker approach catches most of the high-effort pull requests in that tail, which lets maintainers fast-fail costly, low-quality contributions before a human has to context-switch into them. One specific behavioral driver stands out in the data: "unplanned" conceptual changes, where an agent proposes something well outside what was scoped or requested, are the primary cause behind ghosting and failed iteration. A triage model can catch that pattern structurally, before a human ever reads the code.
Triage works better when there's a clear standard to triage against. ITK's investment in an AGENTS.md file gives both human reviewers and automated triage tools something concrete to check contributions against, rather than relying on implicit norms that an agent has no way to infer.
Triage manages the queue, but it doesn't reduce how many pull requests agents submit. And once a pull request clears initial triage and enters iterative review, triage has nothing left to offer against approval churning. The pattern it was built to catch has already passed.
Keeping review quality from degrading
No single fix closes the gap between how cheap it is to generate a pull request and how expensive it remains to review one. But the combination of stricter gating, clear contribution standards, and human review that's augmented rather than replaced keeps review quality from eroding without requiring unlimited time from senior engineers.
Raising the bar happens most effectively at submission, not at merge. tldraw's approach, defaulting to closed and reopening only pull requests that demonstrate real understanding of the codebase, moves the cost of triage onto the contributor.
Written standards do real work here. AGENTS.md and CONTRIBUTING.md documents that spell out scope, required test coverage, and style expectations give AI agents concrete constraints to work inside, cutting down on the "unplanned conceptual change" pattern that the Circuit Breaker research identifies as the main driver of costly pull requests.
Code review agents belong in the pipeline as a first-pass filter, not as the entity making the final call. The Kennesaw State study supports using them to surface obvious problems before a human ever looks at the code, not to substitute for human judgment on anything that requires contextual reasoning about why a change was made.
Senior reviewer time needs protection built into the process itself. Rotation policies, explicit review budgets set per reviewer per sprint, and flagging "high-maintenance" pull requests through creation-time triage all help spread the load so it doesn't concentrate on the engineers whose departure would cost the team the most.
Platform-level controls deserve consideration as a standard part of the toolkit, not a last resort reached for only after things have gone wrong. GitHub's 2026 controls, restricting pull requests to existing collaborators and capping how many open pull requests an outside contributor can have at once, are available to enterprise teams well before volume turns unmanageable.
Visibility into the problem matters as much as any single policy. Faros AI's engineering intelligence tooling surfaces the metrics that make review burden visible in the first place: time to first review, time spent in review, and trends in pull request size. Teams that track these signals get an early warning when the queue is degrading review quality, well before that degradation causes a senior engineer to walk out the door.
Sources
- Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
- From Industry Claims to Empirical Reality: An Empirical Study of Code Review Agents in Pull Requests
- These Aren't the Reviews You're Looking For How Humans Review AI-Generated Pull Requests
- AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate
- Agent pull requests are everywhere. Here's how to review them. - The GitHub Blog