← Leodegario Magsino Jr.
Agentic engineering · Open source

Three-Pass Review

A code review process for teams where AI writes much of the code. Two AI passes on every pull request, and a third only where a mistake would really hurt. AI reads every line. People own the risk.

3
Passes, effort follows risk
2
People sign off a deep PR
0
Approvals by an AI
How a pull request moves through review: it triggers pass 1 (Copilot line comments), pass 2 (the light review, using Addy Osmani's skill or Claude Code's pr-review-toolkit) and a tier check. The light review can escalate. The tier check sends the PR to one approval, or to pass 3, the deep review: an AI reviewer brief, then the code owner and a second reviewer. A review-gate check enforces both before merge.
The flow. Every PR gets passes 1 and 2; the tier check sends risky ones to pass 3, and review-gate enforces both before merge.
Why

Why I built it

Once agents write most of a pull request, review becomes the slow part. Teams answer that in one of two ways: rubber-stamp the change, or read every line as if a person had written it. Neither holds up. Running AI review on a production codebase taught me three things, and each one became part of the process.

Reading every line doesn’t scale. Judgement does.

AI can read every line of every push, cheaply and without getting tired. What it can’t do is own the decision. So the process gives the reading to AI and keeps the decision with people.

Not every change deserves the same review

A copy fix and a payments migration shouldn’t get the same review. Uniform review means either rubber-stamping the dangerous change or spending an hour on the trivial one. The effort has to follow the risk.

Reviewers need orientation, not more comments

Even a good AI review hands the human a pile of comments. On a big, risky change, what a reviewer needs first is the shape of it: what it does, what it touches, what could go wrong, and where to look. That became pass 3’s reviewer brief.

Three-Pass Review is that process, packaged as a guide and a kit you install into any GitHub repository.

The idea

Effort follows risk

Every pull request gets two AI reviews of different kinds: Copilot’s fast, line-level pass, and a structured review of the whole change. A tier check, with rules read from the base branch, decides which PRs also get the third pass. Routine PRs need one quick approval. Risky ones get an AI reviewer brief, then two careful reviewers.

The AI never approves anything. It can comment, and it can raise a PR to deep review, but nothing automated can lower a tier, and if the author removes an escalation the tier check puts it back. A required review-gate check turns the tiers from labels into merge rules: a light PR merges once its light review is clean on the latest commit and one person approves; a deep PR needs two approvals of the latest commit.

PassWhoRuns onWhat you get
1 · Auto review GitHub Copilot code review Every push Line comments with one-click fixes
2 · Light review An AI reviewer you choosecode-review-and-quality · Addy Osmanipr-review-toolkit · Claude Code Every ready PR One verified summary, inline comments for real problems, and an escalation when the PR needs more
3 · Deep review An AI reviewer brief, then the code owner and a second reviewerthreepass Deep-tier PRs only The brief does the reading; people judge and sign off
Pass 1

Auto review: catch the small stuff early

Copilot is already in the pull request view, it’s quick, and its suggestions commit with one click. It isn’t the reviewer of record: it leaves comment-only reviews, so it never counts toward approvals, and path-specific instructions point it harder at the sensitive parts of the codebase. Its job is to clear the cheap mistakes before anyone spends attention on them.

Pass 1, auto review: on an example diff, Copilot comments that reviews.find returns null for an unknown ID, so review.author throws instead of returning a 404, and offers a suggested change committed with one click. It runs on every push, takes seconds to minutes, and can't approve, block or change the tier.
Pass 1 on the running example, PR #482. Illustrative, not a real review.
Pass 2

Light review: a second, different AI

Two copies of the same reviewer mostly repeat each other, so pass 2 is a different model with a different method. Teams choose it in the policy: Addy Osmani’s code-review-and-quality skill, vendored at a pinned commit, or Claude Code’s pr-review-toolkit, whose specialist agents report candidates that the reviewer verifies before posting.

Either way it follows one protocol: intent first, tests before implementation, then correctness and security, blast radius (who still calls what changed) and whether the change can be rolled back. Every finding needs evidence — a path and line, the input that triggers it, or the caller that breaks — and anything it can’t prove goes under “Could not verify” instead of becoming a comment.

It returns its verdict as data. A plain script posts the summary and, on any Critical or Required finding, adds the escalate label. The model can’t approve, push, merge or change labels, and its instructions come from the base branch, so a pull request can’t rewrite its own reviewer.

Pass 2, light review: the team picks Addy Osmani's skill or Claude Code's pr-review-toolkit. The light reviewer checks intent, tests, correctness and security, blast radius and rollback, and posts one verified summary: dismiss() was removed but actions.ts line 88 still calls it. Clean means one approval; a Critical or Required finding escalates the PR to pass 3.
The same PR: pass 2 finds a caller left behind and escalates it to deep review.
Pass 3

Deep review: the AI reads, people judge

Deep reviews are where human time goes, so pass 3 starts by doing the reading. When the tier check routes a pull request deep, it starts a workflow that runs threepass and posts a reviewer brief, updated on every new commit. It’s built to be read in five minutes, before anyone opens the code.

Pass 3, deep review: for a deep-tier PR, the AI computes a change map from the diff, writes a brief, and runs three independent checks, which become one reviewer brief pointing first at actions.ts line 88. Then the code owner and a second reviewer start from the brief, verify the tests, check contracts, confirm rollback and sign off; review-gate needs two approvals before merge.
The top row is automated; the bottom row is people. The brief never approves or blocks.
In the briefWhat it gives the reviewer
The changeWhat the PR does and why, in plain words, and whether that matches its description
Change mapAreas touched, files and lines per area, and warnings for migrations, dependency, CI and infra changes and deletions. Computed from the diff, with no model, so it’s exact
ImpactWhat changes for users, callers and operators: behaviour, API contracts, data, config, dependencies
Risks and rollbackWhat could go wrong in production, and whether a plain revert undoes it
TestsWhich changed behaviour the tests cover, and the gaps
Where to look firstUp to five places, most important first: the top findings, then the brief’s own picks
Questions for the authorWhat the reviewer needs answered before approving
Sign-off draftThe deep-review note, pre-filled with rollback, test coverage and the questions, with placeholders for what only a person can confirm

Then the code owner and a second reviewer take over: verify the tests would catch a regression, check every changed contract, confirm the change can be rolled back, and sign off. The draft says what the AI saw. The approval has to say what the person verified.

The engine

Three independent checks, measured

The findings inside the brief come from threepass, a Ruby CLI in the same repository. It splits review into three checks — correctness, security and data safety, and architecture — that run in parallel on one frozen input and never see each other’s output. A reviewer shown earlier findings tends to anchor on them; three independent reads turn agreement into a signal. When two checks flag the same lines on their own, their combined confidence goes up.

A deterministic reconciler validates every finding against the diff, merges duplicates, scores agreement, drops low-confidence findings and caps the list, so the same input always gives the same comment. The brief is a fourth independent call, and the change map is computed without a model at all.

threepass: a pull request goes to a context builder that makes one frozen input. A cost ceiling estimates every call first, and shrinks context, skips a check or refuses rather than overspend. Three isolated checks, correctness, security and data safety, and architecture, review it in parallel. A deterministic reconciler validates, merges duplicates, scores agreement, thresholds, ranks and caps the findings into one comment with its cost.
Inside pass 3: one frozen input, a cost ceiling before any call, three isolated checks, one reconciled comment.

A cost ceiling that holds

Every call is priced before it’s made, using the API’s own token counter and assuming the maximum output. Over budget, the review shrinks its context, then drops a check, then the brief, and finally refuses rather than overspend. The real cost, from the API’s usage report, is printed in the comment.

Measured, not claimed

Independent checks are a claim, so the repository ships an eval harness that tests it against chained checks, a single combined prompt, and that prompt sampled three times as the cost-matched control. It reports recall and precision per check with 95% Wilson intervals, counts a defect as caught only when most runs catch it, hides precision until people have labelled the findings, and can’t publish numbers from fake or stale runs. The first real run hasn’t happened yet, so there are no numbers to quote, and the results table can only be filled in by the harness.

Guardrails

Guarantees, and what holds them

A review process for AI-written code is only as good as what stops it from being gamed, by a pull request or by the AI itself. Each guarantee here is enforced by code or by GitHub, not by a sentence in a prompt.

GuaranteeHow it’s enforced
A PR can’t weaken its own reviewThe routing policy, the choice of light reviewer, the reviewer instructions and the AI budget are all read from the base branch. Any change to the review setup is routed deep.
AI never approvesReviewers return verdicts as data and plain scripts act on them. Escalation only goes up, and only a maintainer can remove it — never the PR author.
A failed review is never a clean oneNo result means a “didn’t finish” comment, a failed run, and a merge gate that stays pending.
Untrusted input stays dataThe deep-review job checks the PR out as data and never runs it. threepass gives the model no tools, wraps input in markers it can’t forge, never follows symlinks or reads .git, and escapes everything it posts.
Spend is boundedA hard ceiling per review, estimated before any call. Retries are off for billed calls, so one call can’t be billed twice.
Checks stay independentA test fails if any check’s request contains another check’s output or prompt.
Tested without the networkThe Ruby suite refuses every socket and runs on recorded responses. The routing rules have their own Node tests, and every workflow template passes actionlint.
Rollout

One setup, every project

The installer copies the kit into a repository, vendors the review skill at a pinned commit, sets the pass 2 reviewer, and pins pass 3 to an exact commit of threepass, so every project reviews with a known version. When the kit improves, --upgrade refreshes the kit’s own files and moves the pin, and never touches what the team edits: the routing policy, CODEOWNERS, the PR template, the Copilot instructions. The upgrade touches the review setup, so its own pull request gets a deep review.

The guide recommends a shadow mode first. Tiers, AI reviews and briefs all appear while people keep reviewing the way they do today, and the merge gate goes live only once the light review’s precision and the tier routing hold up against what reviewers actually found.

Credits

Built on tools that already exist

Pass 1 is GitHub Copilot code review. Pass 2 runs on claude-code-action with either Addy Osmani’s MIT-licensed code-review-and-quality skill or Claude Code’s pr-review-toolkit, used as they are.

Three-Pass Review adds the process around them: the tier check and merge gate, the routing policy, the pass 3 brief and the threepass engine behind it, the installer, and the guide that ties it together.

Scope

What it deliberately doesn’t do

It never approves or merges; people do. threepass posts one summary comment, not inline comments on each line. Pull requests from forks get no AI deep review, because fork runs get no secrets; they go deep, and people review them with the checklist. It’s GitHub-only, and it isn’t a hosted service: everything runs in your own Actions, with your own API key.

And it’s early. Every part is built and tested end to end against recorded responses, but it hasn’t yet been run on a live repository or evaluated against the real API.

Try it

Install the process

You need admin access to a GitHub repository, the GitHub CLI, and an Anthropic API key. Setup takes about 30 minutes per repository.

# get the kit
git clone https://github.com/lmagsino/three-pass-review.git

# install it in a project, picking the pass 2 reviewer
cd your-repo && git switch -c three-pass-review
bash ../three-pass-review/guide/scripts/install.sh . --light-reviewer addy   # or pr-review-toolkit

# labels, the API key, then the branch rules (see the setup guide)
bash ../three-pass-review/guide/scripts/create-labels.sh your-org/your-repo
gh secret set ANTHROPIC_API_KEY -R your-org/your-repo

# later, when the kit improves
bash ../three-pass-review/guide/scripts/install.sh . --upgrade