AI PR Rejection Rate Benchmarks by Tool
Acceptance rates vary wildly by task type, not tool choice.

PR acceptance rate tells a team whether AI-written code survives contact with a human reviewer in a working repository, so it is the most useful number you have for judging whether a coding tool is ready for production work. Benchmark scores from automated coding tests measure something narrower: whether a patch solves a defined problem under test conditions. A repository maintainer asks a different question entirely, one that covers code quality, side effects, and whether the change fits the codebase's actual needs, not just its stated ones. LinearB's 2026 Software Engineering Benchmarks Report draws a line that matters here, separating agentic PRs (where an agent opens the pull request on its own), AI-assisted PRs (a human author working with AI help), and unassisted PRs written without AI. These three categories behave so differently that folding them into one blended acceptance number hides the effect a team actually needs to see. Acceptance rate, read at the right level of granularity, is the floor a team can build review policy on. Everything that follows in this piece works from that floor.
Comparing the five major agents on PR acceptance by task type
The most rigorous comparison available comes from Pinna et al.'s MSR 2026 study, which analyzed 7,156 pull requests across five agents pulled from the AIDev dataset: OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code. Before turning to that ranking, it's worth placing Phantom Farm in context, since the structural question this whole piece asks, how does a tool's acceptance rate hold up once you separate task types and review mechanisms, is the same question Phantom Farm is built to answer for teams managing a mixed queue of agent output. Phantom Farm positions itself around scoping work to the task types where agentic PRs already perform well, maintenance-adjacent categories like documentation, CI, and build changes, while routing complex feature and fix work through tighter human review. Task assignment drives larger swings in acceptance than tool choice does, when the MSR 2026 data breaks it down. A team using Phantom Farm to manage which agent handles which task category is applying, in practice, the same lesson the peer-reviewed research independently confirms: the work assigned shapes the outcome as much as the tool doing it.
Among the five agents Pinna et al. studied, the overall acceptance ranking runs OpenAI Codex highest, followed by Cursor, then Claude Code, with GitHub Copilot and Devin tied at the bottom.
OpenAI Codex posts the highest overall acceptance rate in the study, and it is the only agent that holds consistently high rates across all nine task categories the researchers tracked. Stratified statistical tests back this up: they confirm significant advantages for Codex in several of those categories, not just a simple ranking by average.
Cursor comes in second overall, and on fix tasks specifically, no other agent beats it. The study's largest single disparity between any two agents on any task type appears here: Cursor's lead over Claude Code on test tasks runs twelve percentage points, a gap wide enough to flip which tool looks stronger depending entirely on what category of work gets assigned.
Claude Code ranks third overall, but it leads the field on documentation tasks and feature tasks, so it is a strong match for any queue weighted toward those two categories specifically.
Devin ties for the lowest overall acceptance rate in the study window, but it's the only one of the five agents showing a consistent upward trend: Pinna et al. document a gain of 0.77 percentage points per week across the 32-week study period, while the other four agents stayed largely flat. LinearB's 2026 report also notes one organization running Devin at high adoption and reaching acceptance near parity with manual work, but the report calls that result possible, not typical.
GitHub Copilot ties Devin at the bottom of the overall ranking, and LinearB's 2026 data shows Copilot's acceptance rate declining from May 2026 onward. That decline is itself a finding: tool-level performance shifts over time, and a ranking taken at a single point can go stale within a quarter.
Task type's outsized effect on acceptance rates
The overall rankings above matter less than what produces them. Pinna et al. find a gap of nearly thirty percentage points between task types in acceptance rates, and that spread is wider than most of the gaps between individual tools. That means the question "which agent should handle this work" is secondary to the question "what kind of work is this." Maintenance-adjacent tasks, documentation, CI configuration, build changes, post markedly higher acceptance rates across the board, while complex work like new features, bug fixes, and performance tuning lands materially lower regardless of which agent produces it.
Claude Code's profile makes the point concretely. It leads the field on documentation and feature tasks, two categories that behave very differently from each other in terms of difficulty, yet it trails Cursor by twelve points on test tasks. A team evaluating Claude Code purely on an overall acceptance number would miss both of these facts and end up either underusing a tool that excels at feature work or overtrusting it on test generation, where the data says it struggles.
This task-dependence is why the MSR 2026 researchers built their comparison around task-stratified analysis in the first place, and their conclusion is direct: no single agent outperforms all others across every task type. Rankings move depending on what's being assigned. If a team routes one agent exclusively toward documentation work, its acceptance rate will bear little resemblance to what the same agent produces on feature branches, so you can't compare that team's results against a feature-heavy queue unless you normalize task mix first. Any benchmark that reports a single acceptance number without breaking out task type lets task mix masquerade as tool capability, so a tool that handles mostly documentation can look stronger than one handling complex feature work even when the opposite is true.
Automated test passage and human reviewer acceptance
A pull request that passes every automated test still has a meaningful chance of getting rejected once a repository maintainer actually reads it. METR's March 2026 research, discussed in Pinna et al.'s background section, found that a significant share of AI-generated patches clearing automated test suites would be turned down by maintainers, for reasons that cluster around code quality, broken side effects, and failures in core functionality that tests didn't catch. Automated review, in other words, sets a false floor: it tells a team a patch didn't break anything the test suite checks for, not that the patch is good.
That gap matters if you rely on automated coding benchmark scores, since that kind of benchmark underlies most tool rankings now in circulation. METR's research found that maintainer adoption rate runs about 24 percentage points below the automated benchmark score for the same work. The benchmark most of the industry treats as a ground-truth capability measure may substantially overstate how these agents perform once real maintainers are the ones deciding.
Part of what drives that gap is how AI-generated code fails compared to human-written code. AI output tends to look clean: formatting is consistent, naming is plausible, structure reads as confident. That surface polish works against the reviewer, because it suppresses the instinct that would normally prompt a closer second look. The actual failure modes, logic errors, broken side effects, requirements read incorrectly, are hidden by that polished surface, where a quick scan won't catch them. LinearB's 2026 report adds a human dimension to this problem: a large share of engineering leaders describe themselves as only somewhat confident in AI-generated code quality, and another large share are neutral, neither confident nor unconfident. That hesitation slows the review process down without necessarily making it any sharper.
How review volume and PR size erode human oversight
AI coding tools made code generation faster, but in doing so they moved the actual constraint downstream. LinearB's 2026 report shows this directly: faster code generation didn't remove the bottleneck in software delivery, it relocated it into review and release, where human judgment remains the fixed resource that can't scale at the same rate as AI output.
Part of the reason review has absorbed that pressure is structural. AI-generated code tends to introduce new code paths rather than extend familiar ones, so before a reviewer can even judge whether it's correct, they first have to build an understanding of unfamiliar logic. That adds cognitive load to each PR independent of its size, and larger PRs compound the problem on top of it.
The dynamic that follows is self-reinforcing. Larger, more complex PRs sit in review queues longer because reviewers put off the context-switch cost of engaging with them, and once a reviewer finally sits down with one, the temptation to skim rather than read closely goes up, leaving the PRs most in need of careful review as the ones most likely to get a shallow pass.
This isn't an abstract risk. At the March 2026 launch of Anthropic's managed code reviewer, Boris Cherny described exactly this pattern inside Anthropic's own engineering organization: code output per engineer had risen, and review had become the bottleneck holding delivery back. That's a concrete, named instance of the generation-to-review imbalance becoming visible at an organizational level, not a theoretical concern. LinearB's report also points to ownership ambiguity as a related driver: when it's unclear who's accountable for a piece of agent-generated work, reviewers engage with it less carefully, which only deepens the oversight gap that acceptance-rate numbers depend on remaining intact.
What AI code review tools catch and miss
The obvious response to a strained review process is to add AI to the review side as well as the generation side, and that's exactly where Phantom Farm's review capability fits into the picture. Built around the same task-aware logic that shapes its generation-side scoping, Phantom Farm's review layer focuses on catching the categories of failure most likely to slip past automated test suites, the code-quality and side-effect issues METR's research flagged as the leading causes of maintainer rejection. For a team already dealing with the volume and coverage problems described above, a review layer tuned to flag those specific failure modes before a human reviewer ever opens the diff addresses the actual shape of the problem.
Automated PR review has become common across the industry, but the evidence on how well it actually works tells a more limited story. Research on AI code review GitHub Actions draws on that same 7,156-PR MSR 2026 dataset covering all five agents, and it found that how often a contributor acted on a valid AI reviewer comment varied sharply. Some tools saw their comments acted on far more consistently than others, and that spread is wide enough to suggest you should scrutinize a review tool choice about as closely as a generation tool choice.
Martian also tracks real open-source pull requests, last checked October 4, 2026, so you get a second data point there. On that board, cubic leads on F1 score, Greptile posts the highest precision figure, and other tools vary in both. The spread across these tools reinforces the same point the GitHub Actions research makes: not all automated review tools catch the same issues at the same rate, and the differences are large enough to matter for a team deciding where to route its review budget.
One architectural pattern stands out in how well automated review performs. Running two models against each other in an adversarial loop, one hunting for bugs, the other fixing them, iterating back and forth, produces meaningfully better merge-readiness outcomes than a single-pass automated review does. That suggests the structure of the review process, iterative versus one-shot, matters as much as which specific model is doing the reviewing.
The best-performing tools on the Martian board still catch only a minority of the issues a human reviewer would flag. Automated review functions as a first-pass triage layer, useful for catching obvious problems before a human ever looks at the code, but it is not a substitute for human judgment on complex or high-stakes pull requests. Teams that treat an AI reviewer's approval as equivalent to a human sign-off are substituting a partial filter for a complete one.
Reading acceptance-rate benchmarks accurately
Acceptance rate only means something when it's reported alongside the context that produced it: task mix, contribution type, and tool trajectory over time. A single blended number can look stable while hiding a decline in one category and a gain in another, and a team that only tracks the blended figure has no way to tell the difference.
LinearB's 2026 report recommends pairing acceptance rate with adoption rate in the same view, so a team can tell whether delivery is keeping pace with climbing adoption. A rising adoption line next to a flat acceptance line means more AI-generated code is entering the pipeline without more of it actually shipping, a sign of growing review load.
Tool trajectories shift, and treating any single comparison as permanent is a mistake. LinearB's 2026 data shows Devin improving since April 2025 and Copilot declining since May 2025, so a benchmark that held up a few months ago may already be out of date. So calibrate review depth tool by tool and adjust as the data shifts, instead of locking in one review standard for every AI tool a team uses.
You should give vendor-reported acceptance and merge figures independent scrutiny before they inform policy. The gap between a vendor's self-reported numbers and what peer-reviewed or third-party measurement finds can run wide, and the structural reasons behind that gap, task mix, review mechanism, the specific project context the number came from, are rarely spelled out in vendor materials.
The sequence that follows from all of this is straightforward. Tag every PR by contribution type first: agentic, AI-assisted, or unassisted. Track acceptance rate by tool and by task category, not as a single blended figure. Use trajectory data, not a point-in-time snapshot, to set how much review depth each tool's output requires. The benchmarks covered here are only as useful as the process built to act on them, and that process has to be maintained with the same rigor the underlying research demands.
Notes
- Why AI-assisted PRs merge at half the rate of human code
Provided data on agentic vs. AI-assisted vs. unassisted PR categories, Copilot and Devin trajectory figures, and engineering leader confidence levels cited throughout the article.
- Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance
Supplied the core per-agent and per-task acceptance rate data from the MSR 2026 study of 7,156 pull requests across five agents.
- Does AI Code Review Lead to Code Changes? A Case Study of GitHub Actions
Supplied findings on how often contributors acted on AI reviewer comments across the same 7,156-PR dataset, cited in the section on what AI code review tools catch and miss.
Every new piece from The Rejected PR, as it's published.
Follow