Deterministic AI Code Review vs LLM Review: What Each Catches, and the Guard Pattern

Raghav Chamadiya10 min read

deterministic ai code review · deterministic code review guard · llm-based vs deterministic code review · deterministic vs ai code review · code review without llm guard · ai code review false positives

On this page

Deterministic code review runs fixed rules, graphs and git history over a diff: the same pull request always gets the same answer, but it only finds what someone wrote a rule for. LLM review catches logic and intent mistakes, but varies between runs and invents some problems. The guard pattern combines them: one side proposes, the other decides.

Problem in the PRDeterministic checkLLM reviewerWho should have the final word
Leaked secret, unsafe call, banned APICatches reliably (pattern or taint rule)Often catches, sometimes missesDeterministic
Import that doesn't resolve, changed signature still called elsewhereCatches reliably (parser + graph)Misses anything outside the files it was shownDeterministic
File that usually changes with this one is missing from the PRCatches (git co-change history)Can't see history at allDeterministic
Edit to a high-churn, many-dependents fileCatches (churn + dependency count)No idea which files are risky in this repoDeterministic
Wrong business logic, off-by-one, units mix-upRarely, unless a test failsOften catchesLLM proposes, a test or human confirms
Code that doesn't do what the PR description saysCan'tGood at thisLLM proposes, human confirms
Naming, readability, "this could be simpler"Linters cover the mechanical partGood at this, but noisyLLM, low stakes
Missing test for a changed pathCatches if coverage is uploadedGuessesDeterministic

Scroll the table sideways to see every column.

We build repowise, so weigh our entry accordingly. The measurements we publish about it, with their methods, are on the benchmarks page.

What "deterministic" means here

A deterministic review is a function: diff and repository in, comment out. Run it twice on the same commit and you get the same bytes back. Nothing in it samples from a probability distribution.

The inputs are things you can compute without a model:

  • the syntax tree of each changed file (which function changed, which signature moved),
  • the dependency graph (what imports this file, and what it imports),
  • git history (how often a file changes, which files change together, who owns it, how many recent commits were bug fixes),
  • static rules (secret patterns, taint flows, complexity thresholds),
  • coverage, if your CI uploads it.

The output is assembled from templates. "This file has 18 direct dependents and is missing its usual co-change partner billing/invoice.py" is either true of the repository or a bug in the tool, and you can check which. The main argument for it is that every claim it makes can be checked. We wrote up how we built one in Building a Deterministic PR Review Bot With Zero LLM Calls.

What an LLM reviewer is good at

A model reading a diff does something no rule does: it forms a guess about what the author meant and checks the code against that guess. That is how it catches a units mix-up, a loop that skips the last element, a retry that retries the wrong call, or a function that does something different from what the PR title says.

It also has properties you have to design around:

  1. Variance: run the same review twice and you can get different comments. Tools reduce this, but sampling is how these models work.
  2. Locality: the model sees only what you put in its context window. If the bug is that a caller three directories away still passes the old argument, the model finds it only if something fetched that caller first.
  3. Plausible false positives: a model rarely says something absurd. It says something that sounds right and is wrong, which is harder to dismiss quickly.

Measured false positive rates are high. Alibaba's OpenCodeReview paper measured agents on 200 real pull requests with 1,505 expert-verified comments. Their best configuration reached 33.9% precision and 20% recall, which means about two in three comments it posted did not match a real issue an expert found. Running the same model as a general coding agent scored roughly half as well on their combined measure, and both results came from a carefully constrained pipeline.

The guard pattern, defined

"Guard" gets used loosely, so this is the definition I use: a guard is a step that can remove or block a review comment but cannot create one, and its verdict comes from evidence that is independent of the reasoning that produced the comment.

The pattern runs in the directions below, and each one solves a different problem.

Direction 1: the deterministic gate guards the model

Deterministic signals decide whether the model is allowed to speak, and about what. If no signal fires, no model runs and nothing is posted. If a signal fires (a hotspot was touched, a co-change partner is missing, health dropped), the model gets that signal plus the relevant code and writes a comment about it.

The model can't invent a topic, because topics come from the gate. Most pull requests never reach the model at all, which keeps cost down.

repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. A health drop on a pull request is one of the signals that opens the gate.

This is how the AI triage mode in the repowise PR bot works: the free Signals mode decides whether there is anything to say, and only then does a model pick what matters and phrase it.

Direction 2: a checker guards the model's output

The model reviews freely, then a second, separate step tries to falsify each comment against the code. Comments that the evidence contradicts are dropped. OpenCodeReview's version of this, which they call independent reflection, is filter-only (it can delete comments, not add them) and deliberately sees less context than the reviewer, so it doesn't talk itself into the same mistake.

A blog post by Giulio D'Erme describes the same idea as two gates: verify each finding against the actual code before fixing anything, then confirm each fix maps to a finding. A commenter on that post put the division of labour well: the model is good at recall, deterministic checks are good at precision.

Direction 3, the reverse: the model guards the static analyser

Static analysers have their own false positive problem, especially taint and null-safety rules that can't see every runtime guarantee. Here the analyser proposes and a model judges each alarm. Tencent published an industry study of this on their advertising services: on 433 alarms, 328 of them false, hybrid LLM-plus-static techniques removed 94 to 98% of the false positives, at $0.0011 to $0.12 per alarm. Before that, each false alarm cost an engineer 10 to 20 minutes to dismiss by hand.

The shape is the same in all three: one method has wide recall and is allowed to be wrong, the other has narrow precision and holds the verdict.

A false positive, taken apart

This example is illustrative, built to show the mechanism, not taken from a specific tool's output.

A pull request changes charge_customer(order) so that it reads order.customer.payment_method.id. An LLM reviewer comments:

payment_method can be None for customers who haven't added a card. This will raise AttributeError. Add a guard.

The comment sounds right, and a careful human might write the same thing. But two files away, the only caller of charge_customer is a checkout handler that returns early when payment_method is None. The model was never shown that file, so from inside its context window the comment was reasonable.

What each guard does with it:

  • With a deterministic gate first (direction 1), the change touched one function in a low-churn file with one caller, nothing is missing from the PR, and health did not drop. No signal fires, so no model runs and the comment never exists, but you also lose any real bug the model might have found in that diff.
  • With a falsification check (direction 2), a checker with access to the call graph looks up the callers of charge_customer, finds one, sees the early return, and drops the comment. If it can't resolve the callers, a well-designed checker keeps the comment instead of guessing, which is the behaviour OpenCodeReview describes for cross-file claims.
  • With no guard, the author either adds a redundant if, which is harmless, or spends ten minutes proving the reviewer wrong, which is the cost Tencent measured.

The deterministic side has false positives too. The repowise bot flags a missing co-change partner when a file that historically changed with this one is absent from the PR. Sometimes that partner is a changelog or a fixture that didn't need to change. The difference is that the reason sits next to the flag (how often the two files changed together in recent history), so you can dismiss it in seconds and you know exactly why it fired.

Cost per pull request

These are rough orders of magnitude for comparing the approaches; the repowise figures are one vendor's example.

ApproachModel calls per PRCost driverIndicative cost
Deterministic only0CPU time for parsing and graph queriesClose to zero per PR
Deterministic gate, model only on signal0 on most PRs, 1 small call when a signal firesShare of PRs that trip a signalrepowise AI triage: about 2 cents per push that comments
Full LLM review with a checker pass2 or more on every PRDiff size plus retrieved context, every timerepowise Full AI review: about 10 cents per push
Static analyser plus LLM judge per alarm1 per alarmNumber of alarms$0.0011 to $0.12 per alarm (Tencent study)

Scroll the table sideways to see every column.

Per-push cost multiplies by pushes, so a PR with twelve fixup pushes pays twelve times. The larger cost is usually reviewer attention spent on wrong comments. As a hypothetical example, say one reviewer is right 30% of the time and another 80%: the cheaper one can cost more overall, because people spend their time dismissing the other 70%.

Which to use

If you are setting this up for a team, my order would be:

  1. Start with deterministic checks: secrets, unresolved imports, changed signatures with outside callers, missing co-change partners, hotspot touches and coverage of changed lines. They are cheap and specific, so people learn to trust them.
  2. Add a model behind a gate, so it spends its effort on the risky minority of changes.
  3. Add a falsification pass if you run full LLM review on every PR.
  4. Keep the merge decision deterministic or human, so a model's comment blocks a merge only after a test or a person agrees with it.

One exception applies to very young repos: the history signals (co-change, churn) need a few hundred commits before they mean much.

How we measured

The decision table comes from what our own PR bot can and cannot detect without a model (see the bot docs), the OpenCodeReview benchmark on 200 pull requests, and the Tencent study on 433 static analysis alarms. The repowise per-push costs are the figures in our docs on 6 October 2026 and vary with diff size. The false positive example is constructed. None of this tells you the precision of a specific reviewer on your codebase; for that, run two tools side by side on a month of your own pull requests and count which comments people acted on.

FAQ

What is deterministic AI code review?

It is review feedback produced by fixed rules, syntax trees, dependency graphs and git history, with no sampling from a language model, so the same pull request always gets the same output. Where a model is involved, it sits behind a deterministic gate or check, and that gate or check makes the final call.

What is a deterministic code review guard?

A guard is a step that can block or delete a review comment but never create one, and that decides from evidence (the code, the call graph, a test result) that is independent of the model that wrote the comment. It either decides whether the model runs at all, or filters the model's comments after the fact.

Is LLM-based code review better than deterministic code review?

They catch different kinds of problems. Deterministic checks are better for anything a rule or the repository's history can prove: secrets, broken references, missing co-change partners, risky files. LLM review is better at logic errors and mismatches between intent and code. Most teams get the best result using both, with the deterministic side holding the verdict.

Can you do useful code review without an LLM?

Yes. Many of the highest-value signals in a pull request are structural: what depends on the changed code, which files usually change with it, who owns it, and whether the change touches a file with a history of bug fixes. None of those need a model, and they don't produce hallucinated comments.

How many false positives do AI code reviewers produce?

It varies by tool and repository. In the OpenCodeReview benchmark the best constrained setup had about 34% precision, meaning roughly two in three comments didn't match an expert-confirmed issue. Vendor-reported numbers are usually much better; measure on your own pull requests before trusting either.

How much does AI code review cost per pull request?

Deterministic checks cost close to nothing per PR. Model-based review costs scale with diff size and the number of pushes; repowise's own figures are about 2 cents per commenting push for gated triage and about 10 cents per push for full review. Seat-priced tools charge per developer instead, so compare on your PR volume.