On this page
- What "deterministic" means here
- What an LLM reviewer is good at
- The guard pattern, defined
- Direction 1: the deterministic gate guards the model
- Direction 2: a checker guards the model's output
- Direction 3, the reverse: the model guards the static analyser
- A false positive, taken apart
- Cost per pull request
- Which to use
- How we measured
- FAQ
Deterministic code review runs fixed rules, graphs and git history over a diff: the same pull request always gets the same answer, but it only finds what someone wrote a rule for. LLM review catches logic and intent mistakes, but varies between runs and invents some problems. The guard pattern combines them: one side proposes, the other decides.
| Problem in the PR | Deterministic check | LLM reviewer | Who should have the final word |
|---|---|---|---|
| Leaked secret, unsafe call, banned API | Catches reliably (pattern or taint rule) | Often catches, sometimes misses | Deterministic |
| Import that doesn't resolve, changed signature still called elsewhere | Catches reliably (parser + graph) | Misses anything outside the files it was shown | Deterministic |
| File that usually changes with this one is missing from the PR | Catches (git co-change history) | Can't see history at all | Deterministic |
| Edit to a high-churn, many-dependents file | Catches (churn + dependency count) | No idea which files are risky in this repo | Deterministic |
| Wrong business logic, off-by-one, units mix-up | Rarely, unless a test fails | Often catches | LLM proposes, a test or human confirms |
| Code that doesn't do what the PR description says | Can't | Good at this | LLM proposes, human confirms |
| Naming, readability, "this could be simpler" | Linters cover the mechanical part | Good at this, but noisy | LLM, low stakes |
| Missing test for a changed path | Catches if coverage is uploaded | Guesses | Deterministic |
Scroll the table sideways to see every column.
We build repowise, so weigh our entry accordingly. The measurements we publish about it, with their methods, are on the benchmarks page.
What "deterministic" means here
A deterministic review is a function: diff and repository in, comment out. Run it twice on the same commit and you get the same bytes back. Nothing in it samples from a probability distribution.
The inputs are things you can compute without a model:
- the syntax tree of each changed file (which function changed, which signature moved),
- the dependency graph (what imports this file, and what it imports),
- git history (how often a file changes, which files change together, who owns it, how many recent commits were bug fixes),
- static rules (secret patterns, taint flows, complexity thresholds),
- coverage, if your CI uploads it.
The output is assembled from templates. "This file has 18 direct dependents and is missing its usual co-change partner billing/invoice.py" is either true of the repository or a bug in the tool, and you can check which. The main argument for it is that every claim it makes can be checked. We wrote up how we built one in Building a Deterministic PR Review Bot With Zero LLM Calls.
What an LLM reviewer is good at
A model reading a diff does something no rule does: it forms a guess about what the author meant and checks the code against that guess. That is how it catches a units mix-up, a loop that skips the last element, a retry that retries the wrong call, or a function that does something different from what the PR title says.
It also has properties you have to design around:
- Variance: run the same review twice and you can get different comments. Tools reduce this, but sampling is how these models work.
- Locality: the model sees only what you put in its context window. If the bug is that a caller three directories away still passes the old argument, the model finds it only if something fetched that caller first.
- Plausible false positives: a model rarely says something absurd. It says something that sounds right and is wrong, which is harder to dismiss quickly.
Measured false positive rates are high. Alibaba's OpenCodeReview paper measured agents on 200 real pull requests with 1,505 expert-verified comments. Their best configuration reached 33.9% precision and 20% recall, which means about two in three comments it posted did not match a real issue an expert found. Running the same model as a general coding agent scored roughly half as well on their combined measure, and both results came from a carefully constrained pipeline.
The guard pattern, defined
"Guard" gets used loosely, so this is the definition I use: a guard is a step that can remove or block a review comment but cannot create one, and its verdict comes from evidence that is independent of the reasoning that produced the comment.
The pattern runs in the directions below, and each one solves a different problem.
Direction 1: the deterministic gate guards the model
Deterministic signals decide whether the model is allowed to speak, and about what. If no signal fires, no model runs and nothing is posted. If a signal fires (a hotspot was touched, a co-change partner is missing, health dropped), the model gets that signal plus the relevant code and writes a comment about it.
The model can't invent a topic, because topics come from the gate. Most pull requests never reach the model at all, which keeps cost down.
repowise's health score (1 to 10) rates how likely a file is to cause bugs and how hard it is to change, using static checks plus git history; the repo score is the average across files, weighted by lines of code. A health drop on a pull request is one of the signals that opens the gate.
This is how the AI triage mode in the repowise PR bot works: the free Signals mode decides whether there is anything to say, and only then does a model pick what matters and phrase it.
Direction 2: a checker guards the model's output
The model reviews freely, then a second, separate step tries to falsify each comment against the code. Comments that the evidence contradicts are dropped. OpenCodeReview's version of this, which they call independent reflection, is filter-only (it can delete comments, not add them) and deliberately sees less context than the reviewer, so it doesn't talk itself into the same mistake.
A blog post by Giulio D'Erme describes the same idea as two gates: verify each finding against the actual code before fixing anything, then confirm each fix maps to a finding. A commenter on that post put the division of labour well: the model is good at recall, deterministic checks are good at precision.
Direction 3, the reverse: the model guards the static analyser
Static analysers have their own false positive problem, especially taint and null-safety rules that can't see every runtime guarantee. Here the analyser proposes and a model judges each alarm. Tencent published an industry study of this on their advertising services: on 433 alarms, 328 of them false, hybrid LLM-plus-static techniques removed 94 to 98% of the false positives, at $0.0011 to $0.12 per alarm. Before that, each false alarm cost an engineer 10 to 20 minutes to dismiss by hand.
The shape is the same in all three: one method has wide recall and is allowed to be wrong, the other has narrow precision and holds the verdict.
A false positive, taken apart
This example is illustrative, built to show the mechanism, not taken from a specific tool's output.
A pull request changes charge_customer(order) so that it reads order.customer.payment_method.id. An LLM reviewer comments:
payment_methodcan beNonefor customers who haven't added a card. This will raiseAttributeError. Add a guard.
The comment sounds right, and a careful human might write the same thing. But two files away, the only caller of charge_customer is a checkout handler that returns early when payment_method is None. The model was never shown that file, so from inside its context window the comment was reasonable.
What each guard does with it:
- With a deterministic gate first (direction 1), the change touched one function in a low-churn file with one caller, nothing is missing from the PR, and health did not drop. No signal fires, so no model runs and the comment never exists, but you also lose any real bug the model might have found in that diff.
- With a falsification check (direction 2), a checker with access to the call graph looks up the callers of
charge_customer, finds one, sees the early return, and drops the comment. If it can't resolve the callers, a well-designed checker keeps the comment instead of guessing, which is the behaviour OpenCodeReview describes for cross-file claims. - With no guard, the author either adds a redundant
if, which is harmless, or spends ten minutes proving the reviewer wrong, which is the cost Tencent measured.
The deterministic side has false positives too. The repowise bot flags a missing co-change partner when a file that historically changed with this one is absent from the PR. Sometimes that partner is a changelog or a fixture that didn't need to change. The difference is that the reason sits next to the flag (how often the two files changed together in recent history), so you can dismiss it in seconds and you know exactly why it fired.
Cost per pull request
These are rough orders of magnitude for comparing the approaches; the repowise figures are one vendor's example.
| Approach | Model calls per PR | Cost driver | Indicative cost |
|---|---|---|---|
| Deterministic only | 0 | CPU time for parsing and graph queries | Close to zero per PR |
| Deterministic gate, model only on signal | 0 on most PRs, 1 small call when a signal fires | Share of PRs that trip a signal | repowise AI triage: about 2 cents per push that comments |
| Full LLM review with a checker pass | 2 or more on every PR | Diff size plus retrieved context, every time | repowise Full AI review: about 10 cents per push |
| Static analyser plus LLM judge per alarm | 1 per alarm | Number of alarms | $0.0011 to $0.12 per alarm (Tencent study) |
Scroll the table sideways to see every column.
Per-push cost multiplies by pushes, so a PR with twelve fixup pushes pays twelve times. The larger cost is usually reviewer attention spent on wrong comments. As a hypothetical example, say one reviewer is right 30% of the time and another 80%: the cheaper one can cost more overall, because people spend their time dismissing the other 70%.
Which to use
If you are setting this up for a team, my order would be:
- Start with deterministic checks: secrets, unresolved imports, changed signatures with outside callers, missing co-change partners, hotspot touches and coverage of changed lines. They are cheap and specific, so people learn to trust them.
- Add a model behind a gate, so it spends its effort on the risky minority of changes.
- Add a falsification pass if you run full LLM review on every PR.
- Keep the merge decision deterministic or human, so a model's comment blocks a merge only after a test or a person agrees with it.
One exception applies to very young repos: the history signals (co-change, churn) need a few hundred commits before they mean much.
How we measured
The decision table comes from what our own PR bot can and cannot detect without a model (see the bot docs), the OpenCodeReview benchmark on 200 pull requests, and the Tencent study on 433 static analysis alarms. The repowise per-push costs are the figures in our docs on 6 October 2026 and vary with diff size. The false positive example is constructed. None of this tells you the precision of a specific reviewer on your codebase; for that, run two tools side by side on a month of your own pull requests and count which comments people acted on.
FAQ
What is deterministic AI code review?
It is review feedback produced by fixed rules, syntax trees, dependency graphs and git history, with no sampling from a language model, so the same pull request always gets the same output. Where a model is involved, it sits behind a deterministic gate or check, and that gate or check makes the final call.
What is a deterministic code review guard?
A guard is a step that can block or delete a review comment but never create one, and that decides from evidence (the code, the call graph, a test result) that is independent of the model that wrote the comment. It either decides whether the model runs at all, or filters the model's comments after the fact.
Is LLM-based code review better than deterministic code review?
They catch different kinds of problems. Deterministic checks are better for anything a rule or the repository's history can prove: secrets, broken references, missing co-change partners, risky files. LLM review is better at logic errors and mismatches between intent and code. Most teams get the best result using both, with the deterministic side holding the verdict.
Can you do useful code review without an LLM?
Yes. Many of the highest-value signals in a pull request are structural: what depends on the changed code, which files usually change with it, who owns it, and whether the change touches a file with a history of bug fixes. None of those need a model, and they don't produce hallucinated comments.
How many false positives do AI code reviewers produce?
It varies by tool and repository. In the OpenCodeReview benchmark the best constrained setup had about 34% precision, meaning roughly two in three comments didn't match an expert-confirmed issue. Vendor-reported numbers are usually much better; measure on your own pull requests before trusting either.
How much does AI code review cost per pull request?
Deterministic checks cost close to nothing per PR. Model-based review costs scale with diff size and the number of pushes; repowise's own figures are about 2 cents per commenting push for gated triage and about 10 cents per push for full review. Seat-priced tools charge per developer instead, so compare on your PR volume.