Skip to content
    All posts
    AI & Tools

    AI Code Review: What It Catches, and What It Will Never Catch

    After a few hundred AI-assisted reviews, a clear pattern: it's excellent at local correctness and blind to whether you built the right thing.

    3 min readby

    We've had AI review in our pipeline for about a year — running on every PR before a human looks. It's been clearly net positive and it has a sharply defined ceiling. Both halves are worth being precise about.

    What it reliably catches

    Unhandled edge cases in code you can see

    Empty arrays, null returns, division by a value that could be zero, a Promise.all with no rejection handling, an index that can be -1. It's essentially exhaustive about this in a way tired humans are not, and it doesn't get bored on line 400 of a diff.

    Async and race conditions

    This is where it has genuinely saved us. Missing await, a state update after unmount, a useEffect that fires twice, a fetch without cancellation. Bugs that are hard to see, hard to reproduce, and expensive in production.

    Inconsistency with the surrounding code

    "Every other endpoint in this file returns { data, error }; this one throws." A human reviewer notices this only if they happen to have the file's conventions loaded. It always has them loaded, because it just read the file.

    The security basics

    SQL built by string concatenation, an unescaped user value in a template, a secret in a config file, a missing authorisation check on an endpoint that has one everywhere else. Not a replacement for real security review, but it catches the embarrassing tier reliably.

    What it never catches

    Whether the feature is right

    It will review a perfectly-implemented feature that solves the wrong problem and find nothing wrong. It has no access to the customer conversation, the support ticket that prompted it, or the fact that we tried this two years ago and it confused everyone.

    Whether the abstraction will hold

    "This is over-engineered for what we need" and "this will need to change the moment we add the second payment provider" are judgements requiring knowledge of a roadmap. It sees the code; the roadmap is in someone's head.

    Organisational context

    "The payments team owns this module — loop them in." "We deliberately duplicated this last quarter; don't re-unify it." "This is deprecated and being deleted next sprint." All invisible.

    Whether the tests test anything

    It'll confirm tests exist and cover the branches. It's much weaker on whether the assertions are meaningful — a test that mocks the thing under test and asserts the mock was called will often pass review.

    The failure mode to watch for

    Not false positives — you learn to dismiss those fast. It's the reviewer who sees a clean AI pass and skims. The AI found nothing, so it's probably fine.

    But the AI checked the layer where problems are cheap. The expensive problems — wrong feature, wrong abstraction, wrong team, wrong moment — are exactly what the human was there for. If the tool causes humans to review less carefully, it's a net negative even while every individual comment it makes is correct.

    How we set it up

    1. 01AI review runs automatically on PR open, posting inline comments.
    2. 02The author addresses or dismisses them before requesting human review.
    3. 03Human reviewers are explicitly told to review for design, fit, and consequence — not correctness.
    4. 04We tune the prompt to skip style entirely. That's the linter's job and duplicate comments train people to ignore all of them.

    Step 3 is the whole thing. Framing it as a division of labour rather than as a first pass keeps the human attention on the layer where it's irreplaceable.

    The most useful reframe I've heard: it's not a reviewer, it's a very thorough linter that understands semantics. Excellent at that. Not a substitute for someone who knows why you're building this.

    AICode ReviewQualityProcess

    Keep reading

    PAPulapa Arun Kumar

    Full-stack developer building performant, clean, and user-friendly web & mobile applications. Also written as Arun Kumar Pulapa — same person, surname first.

    Built with

    ReactTypeScriptTailwind CSSVite

    © 2026 Pulapa Arun Kumar. All rights reserved.