Skip to content

AI code review

Updated 26 September 2026 Reviewed by Teemu Malinen

What is AI code review?

An automated first pass over a pull request, where a model flags bugs, security issues, style problems and missing tests, then comments inline before a human reviews. It does not replace the human reviewer. It clears the routine findings so people can spend their attention on design and intent. The most widely used tools now handle tens of millions of pull requests this way.

Why AI code review matters now

AI code review matters because code now arrives faster than people can read it. Coding assistants raise the number of changes each developer produces, and coding agents can open pull requests on their own, so an automated reviewer may be the only thing that reads every line before merge. GitHub reported in March 2026 that its Copilot code review had completed 60 million reviews since launching in April 2025. By then it accounted for more than one in five code reviews on GitHub and ran automatically on every pull request in over 12,000 organisations. In September 2026 GitHub went a step further: administrators can now authorise Copilot to submit an approval that counts towards a repository’s required-approvals rule. That changes the question teams face. It is no longer whether a machine comments on the code, but which changes a machine may sign off and which still need a person.

What does an AI code reviewer check?

An AI code reviewer reads the diff of a pull request, pulls in surrounding code from the repository for context and leaves inline comments where it sees a problem. Typical findings are logic errors, missing null checks, unhandled edge cases, missing tests, maintainability problems, inconsistent naming and known insecure patterns such as a hard-coded key. Many tools also propose a fix that the author can apply with one click. GitHub says its reviewer surfaced actionable feedback in 71% of reviews, averaging about 5.1 comments per review, and left no comment in the other 29%. Google calibrated its internal system for resolving reviewer comments to 50% precision and reported in 2023 that authors applied 40–50% of the suggested edits they previewed. Both figures come from the companies that built the tools and have not been independently audited.

How is AI code review different from human code review?

Code review works on two levels, and AI code review belongs mostly to the first. The machine level checks what rules and patterns can settle: does the code build, do the tests pass, does it follow agreed conventions, does it contain known vulnerabilities or obvious bugs. The human level checks whether the change does what it was meant to do, whether it fits the design and whether the remaining risk is acceptable for this system. NIST’s Secure Software Development Framework (SP 800-218) draws the same line. It separates code review, where a person looks directly at the code, from code analysis, where tools find issues either fully automatically or alongside a person, and leaves each organisation to decide which to use where. OWASP’s Secure Code Review Cheat Sheet places manual review where human judgement adds most, such as business logic and authorisation models. Google’s engineering practices open their reviewer checklist with design: do the interactions of the pieces in a change make sense? Answering that takes someone who knows what the software is for.

Why is code review becoming the bottleneck?

AI speeds up writing code far more than it speeds up accepting it, so the queue forms at review. Faros AI’s AI Engineering Report 2026 (April 2026) analysed two years of telemetry from 22,000 developers in more than 4,000 teams, comparing each organisation’s periods of lowest and highest AI adoption. The acceptance rate of AI-generated code rose from 20% to 60%, while pull requests merged per developer rose by only 16.2%. Median time in review rose by 441.5%, pull requests merged without any review, human or agentic, by 31.3%, and the ratio of incidents to pull requests by 242.7%. Faros sells software for managing and measuring engineering work, so the findings suit its product. The report is not peer reviewed, and the comparison shows association rather than cause. Google and DORA’s 2025 report, based on nearly 5,000 survey respondents, points the same way: AI adoption correlated positively with delivery throughput but negatively with delivery stability. The report’s own summary is that AI amplifies what a team already has.

Automated review scales with the volume of code; human review does not. That gap is the strongest argument for AI code review and also its main risk. The tool can clear the routine findings so reviewers keep their attention for intent, or it can become the reason nobody reads the change at all.

Does AI code review work in practice?

Independent evidence from real teams is still thin, and what exists points to modest quality gains rather than faster delivery. Umut Cihan and colleagues studied an LLM-based review tool in an industrial setting and presented the results at ICSE 2025, in the Software Engineering in Practice track. Around 238 practitioners across ten projects had access to the tool. In the three projects studied, with 4,335 pull requests of which 1,568 received automated reviews, developers resolved 73.8% of the automated comments. Most practitioners reported a minor improvement in code quality, better bug detection and more awareness of good practice. Average pull request closure time, however, rose from 5 hours 52 minutes to 8 hours 20 minutes, and developers named faulty reviews, unnecessary corrections and irrelevant comments as the main drawbacks. A resolved comment is not the same as a prevented defect, and every extra comment is work for the author before merge. Expect routine problems to be caught more often, not a shorter review cycle by default.

What are the limits of AI code review?

The main limits of AI code review are missing context, unreliable reasoning and noise. A model sees the code, not the conversation with the customer, the regulatory constraint or the reason a strange-looking workaround exists, so it cannot judge whether a change is the right one. Its reasoning about security is not yet dependable. In a study presented at the IEEE Symposium on Security and Privacy 2024, Saad Ullah and colleagues tested eight large language models on 228 code scenarios and found non-deterministic answers, incorrect and unfaithful reasoning and poor performance on real-world cases. GPT-4 and PaLM2 gave wrong answers in up to 26% of cases when function or variable names were changed or library functions added. Noise is the third limit. A reviewer that flags too many low-value issues trains people to scroll past its comments, and the genuine findings go with them. A review tool nobody reads is worse than none, because it looks like a safety net while catching nothing.

How should a team introduce AI code review?

Introduce AI code review as the first pass on every pull request and keep a person accountable for the merge decision. Start by deciding in writing which changes the machine may approve on its own, if any, and which always need a human; changes to authentication, payments, personal data and infrastructure belong in the second group. Then tune suppression hard. Switch off style comments that a linter or formatter already handles, because once developers start dismissing comments by reflex, the one that mattered goes with them. Route findings into the tools the team already uses: NIST’s SSDF asks organisations to record and triage discovered issues in the development team’s workflow or issue tracker, not leave them in a comment thread. Finally, measure the review process itself: time to first review, time in review, the share of automated comments acted on and the share of pull requests merged without human review. Plan human review capacity before scaling code generation, since faster writing lands on the reviewers first.

What should you look for in an AI code review tool?

Judge AI code review tools on your own repositories over several weeks rather than on a demo. Five criteria decide most of the value:

  • Repository context. Does the tool read beyond the diff, so it can spot a broken caller elsewhere in the codebase?
  • Signal-to-noise. How many comments does it leave per pull request, and what share does the team act on?
  • Configurability. Can you switch off categories, add house rules and set severity thresholds?
  • Approval controls. Can you stop the tool approving, or limit approval to low-risk paths?
  • Data handling. Where is the code sent, is it retained, and is it used for training?

Tools range from features built into the code host, such as GitHub Copilot code review, to dedicated review services and self-hosted models. Running two candidates side by side on the same pull requests for a fortnight gives a fairer comparison than any benchmark a vendor publishes.

AI code review and neighbouring terms

  • AI code generation produces the code; AI code review checks it before merge.
  • AI-generated code security covers the vulnerabilities AI-written code tends to introduce. AI code review is one control against them, alongside static analysis, dependency scanning and human review.
  • AI-augmented testing writes and maintains tests; review asks whether those tests check the right thing.
  • AI-assisted software development is the umbrella term for AI across the whole workflow, and review is one stage of it.
  • Vibe coding, as Simon Willison defines it, means building software with a language model without reviewing the code it writes. Review is the dividing line: reviewed, tested and understood code is ordinary software development.

Frequently asked questions

Can AI code review replace human reviewers?

Not for the decision that matters. AI code review catches routine bugs, missing tests and known insecure patterns, but it cannot judge whether a change is what the business needed or whether the remaining risk is acceptable. The workable division is that the machine clears routine findings and a person owns the merge decision.

How accurate is AI code review?

It is useful but imperfect. In an industrial study of 4,335 pull requests presented at ICSE 2025, developers resolved 73.8% of automated review comments, yet also reported faulty and irrelevant ones. Accuracy varies by language, codebase and configuration, so track the share of comments your own team acts on rather than relying on vendor figures.

Can AI code review find security vulnerabilities?

It can flag common insecure patterns, but it should not be the only security control. Research presented at IEEE S&P 2024 found that large language models gave non-deterministic and sometimes wrong answers about vulnerabilities, and that renaming functions or variables was enough to change results. Pair AI review with static analysis, dependency scanning and human review of security-critical changes.

Does AI code review make merging faster?

Adding the tool alone does not. In the ICSE 2025 industrial study, average pull request closure time rose from just under six hours to over eight after the tool was introduced, because authors had more comments to work through. Time savings come from cutting noise and letting reviewers skip routine checks.

Should AI be allowed to approve pull requests?

Only within limits the team sets in advance. Since September 2026 GitHub has let administrators authorise Copilot approvals that count towards required-approvals rules, which turns the question into a policy decision rather than a technical one. A sensible starting point is to allow machine approval, if at all, only for low-risk changes and to keep human approval for anything touching authentication, payments, personal data or infrastructure.

Sources

Otto Sunnari, myynti ja kumppanuudet, Sofokus / Otto Sunnari, Sales and partnerships at Sofokus

Ready to start leveraging AI?

Call, email, or book a time straight from my calendar.

Otto Sunnari

Sales and partnerships