Why I built it
Once agents write most of a pull request, review becomes the slow part. Teams answer that in one of two ways: rubber-stamp the change, or read every line as if a person had written it. Neither holds up. Running AI review on a production codebase taught me three things, and each one became part of the process.
Reading every line doesn’t scale. Judgement does.
AI can read every line of every push, cheaply and without getting tired. What it can’t do is own the decision. So the process gives the reading to AI and keeps the decision with people.
Not every change deserves the same review
A copy fix and a payments migration shouldn’t get the same review. Uniform review means either rubber-stamping the dangerous change or spending an hour on the trivial one. The effort has to follow the risk.
Reviewers need orientation, not more comments
Even a good AI review hands the human a pile of comments. On a big, risky change, what a reviewer needs first is the shape of it: what it does, what it touches, what could go wrong, and where to look. That became pass 3’s reviewer brief.
Three-Pass Review is that process, packaged as a guide and a kit you install into any GitHub repository.
Effort follows risk
Every pull request gets two AI reviews of different kinds: Copilot’s fast, line-level pass, and a structured review of the whole change. A tier check, with rules read from the base branch, decides which PRs also get the third pass. Routine PRs need one quick approval. Risky ones get an AI reviewer brief, then two careful reviewers.
The AI never approves anything. It can comment, and it can raise a PR to deep review, but nothing automated can lower a tier, and if the author removes an escalation the tier check puts it back. A required review-gate check turns the tiers from labels into merge rules: a light PR merges once its light review is clean on the latest commit and one person approves; a deep PR needs two approvals of the latest commit.
| Pass | Who | Runs on | What you get |
|---|---|---|---|
| 1 · Auto review | GitHub Copilot code review | Every push | Line comments with one-click fixes |
| 2 · Light review | An AI reviewer you choosecode-review-and-quality · Addy Osmanipr-review-toolkit · Claude Code | Every ready PR | One verified summary, inline comments for real problems, and an escalation when the PR needs more |
| 3 · Deep review | An AI reviewer brief, then the code owner and a second reviewerthreepass | Deep-tier PRs only | The brief does the reading; people judge and sign off |
Auto review: catch the small stuff early
Copilot is already in the pull request view, it’s quick, and its suggestions commit with one click. It isn’t the reviewer of record: it leaves comment-only reviews, so it never counts toward approvals, and path-specific instructions point it harder at the sensitive parts of the codebase. Its job is to clear the cheap mistakes before anyone spends attention on them.
Light review: a second, different AI
Two copies of the same reviewer mostly repeat each other, so pass 2 is a different model with a different method. Teams choose it in the policy: Addy Osmani’s code-review-and-quality skill, vendored at a pinned commit, or Claude Code’s pr-review-toolkit, whose specialist agents report candidates that the reviewer verifies before posting.
Either way it follows one protocol: intent first, tests before implementation, then correctness and security, blast radius (who still calls what changed) and whether the change can be rolled back. Every finding needs evidence — a path and line, the input that triggers it, or the caller that breaks — and anything it can’t prove goes under “Could not verify” instead of becoming a comment.
It returns its verdict as data. A plain script posts the summary and, on any Critical or Required finding, adds the escalate label. The model can’t approve, push, merge or change labels, and its instructions come from the base branch, so a pull request can’t rewrite its own reviewer.
Deep review: the AI reads, people judge
Deep reviews are where human time goes, so pass 3 starts by doing the reading. When the tier check routes a pull request deep, it starts a workflow that runs threepass and posts a reviewer brief, updated on every new commit. It’s built to be read in five minutes, before anyone opens the code.
| In the brief | What it gives the reviewer |
|---|---|
| The change | What the PR does and why, in plain words, and whether that matches its description |
| Change map | Areas touched, files and lines per area, and warnings for migrations, dependency, CI and infra changes and deletions. Computed from the diff, with no model, so it’s exact |
| Impact | What changes for users, callers and operators: behaviour, API contracts, data, config, dependencies |
| Risks and rollback | What could go wrong in production, and whether a plain revert undoes it |
| Tests | Which changed behaviour the tests cover, and the gaps |
| Where to look first | Up to five places, most important first: the top findings, then the brief’s own picks |
| Questions for the author | What the reviewer needs answered before approving |
| Sign-off draft | The deep-review note, pre-filled with rollback, test coverage and the questions, with placeholders for what only a person can confirm |
Then the code owner and a second reviewer take over: verify the tests would catch a regression, check every changed contract, confirm the change can be rolled back, and sign off. The draft says what the AI saw. The approval has to say what the person verified.
Three independent checks, measured
The findings inside the brief come from threepass, a Ruby CLI in the same repository. It splits review into three checks — correctness, security and data safety, and architecture — that run in parallel on one frozen input and never see each other’s output. A reviewer shown earlier findings tends to anchor on them; three independent reads turn agreement into a signal. When two checks flag the same lines on their own, their combined confidence goes up.
A deterministic reconciler validates every finding against the diff, merges duplicates, scores agreement, drops low-confidence findings and caps the list, so the same input always gives the same comment. The brief is a fourth independent call, and the change map is computed without a model at all.
A cost ceiling that holds
Every call is priced before it’s made, using the API’s own token counter and assuming the maximum output. Over budget, the review shrinks its context, then drops a check, then the brief, and finally refuses rather than overspend. The real cost, from the API’s usage report, is printed in the comment.
Measured, not claimed
Independent checks are a claim, so the repository ships an eval harness that tests it against chained checks, a single combined prompt, and that prompt sampled three times as the cost-matched control. It reports recall and precision per check with 95% Wilson intervals, counts a defect as caught only when most runs catch it, hides precision until people have labelled the findings, and can’t publish numbers from fake or stale runs. The first real run hasn’t happened yet, so there are no numbers to quote, and the results table can only be filled in by the harness.
Guarantees, and what holds them
A review process for AI-written code is only as good as what stops it from being gamed, by a pull request or by the AI itself. Each guarantee here is enforced by code or by GitHub, not by a sentence in a prompt.
| Guarantee | How it’s enforced |
|---|---|
| A PR can’t weaken its own review | The routing policy, the choice of light reviewer, the reviewer instructions and the AI budget are all read from the base branch. Any change to the review setup is routed deep. |
| AI never approves | Reviewers return verdicts as data and plain scripts act on them. Escalation only goes up, and only a maintainer can remove it — never the PR author. |
| A failed review is never a clean one | No result means a “didn’t finish” comment, a failed run, and a merge gate that stays pending. |
| Untrusted input stays data | The deep-review job checks the PR out as data and never runs it. threepass gives the model no tools, wraps input in markers it can’t forge, never follows symlinks or reads .git, and escapes everything it posts. |
| Spend is bounded | A hard ceiling per review, estimated before any call. Retries are off for billed calls, so one call can’t be billed twice. |
| Checks stay independent | A test fails if any check’s request contains another check’s output or prompt. |
| Tested without the network | The Ruby suite refuses every socket and runs on recorded responses. The routing rules have their own Node tests, and every workflow template passes actionlint. |
One setup, every project
The installer copies the kit into a repository, vendors the review skill at a pinned commit, sets the pass 2 reviewer, and pins pass 3 to an exact commit of threepass, so every project reviews with a known version. When the kit improves, --upgrade refreshes the kit’s own files and moves the pin, and never touches what the team edits: the routing policy, CODEOWNERS, the PR template, the Copilot instructions. The upgrade touches the review setup, so its own pull request gets a deep review.
The guide recommends a shadow mode first. Tiers, AI reviews and briefs all appear while people keep reviewing the way they do today, and the merge gate goes live only once the light review’s precision and the tier routing hold up against what reviewers actually found.
Built on tools that already exist
Pass 1 is GitHub Copilot code review. Pass 2 runs on claude-code-action with either Addy Osmani’s MIT-licensed code-review-and-quality skill or Claude Code’s pr-review-toolkit, used as they are.
Three-Pass Review adds the process around them: the tier check and merge gate, the routing policy, the pass 3 brief and the threepass engine behind it, the installer, and the guide that ties it together.
What it deliberately doesn’t do
It never approves or merges; people do. threepass posts one summary comment, not inline comments on each line. Pull requests from forks get no AI deep review, because fork runs get no secrets; they go deep, and people review them with the checklist. It’s GitHub-only, and it isn’t a hosted service: everything runs in your own Actions, with your own API key.
And it’s early. Every part is built and tested end to end against recorded responses, but it hasn’t yet been run on a live repository or evaluated against the real API.
Install the process
You need admin access to a GitHub repository, the GitHub CLI, and an Anthropic API key. Setup takes about 30 minutes per repository.
# get the kit git clone https://github.com/lmagsino/three-pass-review.git # install it in a project, picking the pass 2 reviewer cd your-repo && git switch -c three-pass-review bash ../three-pass-review/guide/scripts/install.sh . --light-reviewer addy # or pr-review-toolkit # labels, the API key, then the branch rules (see the setup guide) bash ../three-pass-review/guide/scripts/create-labels.sh your-org/your-repo gh secret set ANTHROPIC_API_KEY -R your-org/your-repo # later, when the kit improves bash ../three-pass-review/guide/scripts/install.sh . --upgrade