Why I built it
Over the past year my work as a tech lead moved from “an engineer with an AI pair” to “an engineer running a line of agents”. The agents got good enough to take a ticket from description to pull request. What they didn’t do was make the work trustworthy on their own. Running a line like that on a production codebase taught me three things, and each one became a default here.
The bottleneck moves
Agents write code faster than anyone can read it. Once typing stops being the constraint, deciding, reviewing and verifying become it. So the guardrails belong on decisions, review and verification, not on code generation.
Unlimited parallelism is a trap
Run wide open, a fleet of agents burns a week of model budget in days and opens more pull requests than any human can review. Every extra worker grows its own context, and the coordinator’s with it. Two lanes is the ceiling, and it is not a starting point to tune upward.
A reviewer from the same vendor shares the builder’s blind spots
When the model that wrote the code also reviews it, its misses are silent: the review passes because both passes share the same gaps. So the Reviewer seat runs on a different vendor by design, and if you put the same vendor in both seats, every PR says so.
AI Light Factory is that line, rebuilt from scratch for open source, so any team can run it with whichever agents they already use.
Light, not dark
Addy Osmani drew the distinction in Software Factories, Light and Dark. A dark factory ships code that no human has read. A light factory keeps a human in the loop, and gives a loop only as much autonomy as you can cheaply and reliably verify.
This project takes the light side literally. Your judgement moves to where it pays off: you settle the decisions and approve the plan before any code exists, agents build and review unattended, and you read every diff before it merges.
The line never marks a pull request ready, never approves, and never merges. That isn’t a sentence in a prompt that a model might decide to skip. A guard hook and a pair of git and gh shims refuse those commands for every worker, whatever agent it runs.
The line
A factory isn’t for speed. It exists so that the same mistake cannot be made twice quietly. Every stage writes an artefact to a known path, and the next stage reads it instead of re-deriving it. When something goes wrong, you can name the stage that should have caught it, and put the check there.
The alternative is one long agent session that reads, decides, plans, builds and reviews without ever writing down what it concluded. That works until it doesn’t, and when it doesn’t, there is nothing to inspect.
| Stage | Agent and skill | Leaves behind | You decide |
|---|---|---|---|
| 1 · Decide | Architectgrill-with-docs · Matt Pocock | ADRs and a glossary | Do the decisions hold? |
| 2–3 · Plan | Planneragent-skills:plan · Addy Osmanito-tickets · Matt Pocock | A plan, then GitHub issues with acceptance criteria and blocking edges | Right plan? Right tickets and edges? |
| 4 · Dev loop | Developer, Reviewer, PR Guardian/alf:dev-loop | Squashed draft PRs, a live board, a ledger | Nothing. It runs unattended |
| 5 · Verify | Verifier/alf:verify | A verdict for every feature-flag state | Merge. Always yours |
Three cadences, not one pass
Deciding, planning and ticketing happen once per project. The dev loop runs continuously, a ticket at a time per lane, through a parent issue or milestone. The Verifier runs once per milestone. Mixing these up is the common failure: a Verifier run per ticket is theatre, because until a phase is nearly done there is nothing end to end to exercise, and re-opening decisions per ticket produces a plan that contradicts itself.
What each stage is really for
None of the stages is mainly for what its name suggests. Deciding is corrective: it catches the requirement that would have been built exactly as written, because it read like a decision already made. Planning is arithmetical: it catches the plan that is internally inconsistent, which only becomes visible once the dependency graph is written down as real edges. The dev loop is compressive: its value isn’t that it writes code, but that a ticket reaches you already reviewed, swept and squashed, so your attention goes to the diff rather than the process. The Verifier is falsifying: it is the only stage whose job is to find out that the plan was wrong. Everything before it checks internal consistency. The Verifier checks the world.
Unattended, and still accountable
The dev loop is the part no existing skill pack covered, and the reason the project exists. Your session becomes the Orchestrator. It owns the queue of ready tickets, the base branches, the board and the budget, and it never edits a file, never reads a full diff, never builds and never reviews. It dispatches seats and reads one-line results. Each seat is a named git worktree — Developer, Developer-2, Reviewer, Guardian — with an agent of your choice in it.
Nothing left to ask
Before a ticket gets a seat, a readiness gate checks it. A ticket without a clear “what to build”, checkbox acceptance criteria free of TBDs, and explicit blocked-by edges is refused, labelled as needing a decision, and sent back upstream. That is why the dev loop never asks you a question mid-run: every decision you could make was taken while deciding, and written into the ticket. Anything it would want to ask is an upstream defect, recorded on the board and on the issue while the run carries on.
Build, then a review from another vendor
The Developer builds one ticket test-first, following Addy Osmani’s incremental-implementation and test-driven-development skills. It squashes, re-runs the full suite, and opens a draft only from green. The Reviewer then reads the diff with fresh context, on a different vendor, following code-review-and-quality plus three reach questions an ordinary review skips: does the change alter an outward contract, does it touch code that other paths run through, and does it undo a decision that was made on purpose?
Findings go through a severity ladder instead of to you. Critical findings are fixed. Important ones are fixed when the fix is small, and otherwise listed in the PR body for you. Suggestions are dropped with a one-line reason. Every call lands in the PR’s decision log, so what the loop decided without you is auditable exactly where you are already looking. If round one kept nothing, the ticket ships. If it kept something, the Developer fixes it and the same Reviewer reads only the delta. A Critical that survives round two parks the lane: two review rounds, then a human.
Five reasons to stop, and none of them is a question
A lane stops for five things only:
- The Reviewer says the approach is wrong.
- A fix would touch a contract, a registry or a migration the ticket didn’t ask for.
- A Critical finding survives round two.
- A previously passing test fails for a reason outside the diff.
- The base branch moved underneath the lane.
Each one is a fact that makes the lane unable to proceed. A stopped lane becomes a parked row with the evidence posted on its issue, and every other lane keeps running. A worker that hits a cap, dies, goes silent or never reports parks its lane too, with its log. Nothing is re-dispatched silently.
Stack on open work
A dependency on an unmerged pull request is a base branch, not a blocker. The loop branches off the open parent, reviews against it and bases the PR on it, so a milestone keeps moving while drafts wait for you. Four places have to agree on the base — the commit count, the branch point, the review diff and the PR base — or the review reads the parent’s commits as if they belonged to this change.
The PR Guardian
Once drafts are open, review threads arrive from bots and from people. The PR Guardian sweeps them in batches of about five, because most of a short session is startup cost. It triages each thread with the same ladder, fixes what qualifies, replies to every thread it touches with what it did and why, and rebases stacked PRs whose parent moved. It answers questions about the code from the code, leaves questions about what the work should be for you, and never marks a PR ready.
One screen to read
While the loop runs, the board is the only thing you need to read. It is ordered by what needs you, and it checks that every claimed worker is actually alive before printing it.
2026-10-01-m1-export-foundation Thu 01 Oct 21:14 ------------------------------------------------------------------------------------------------------------------------ FLAG TICKET PR / SHA SEAT STATE NEXT ------------------------------------------------------------------------------------------------------------------------ PARK T1.4 #61 Currency rounding #64 a19f003 - round 2: Critical survives needs you LANE T1.5 #62 CSV writer - Reviewer review round 1 judge findings LANE T1.6 #63 Column registry - Developer-2 building review SWEEP T1.3 #60 Export job #66 3bf66cd Guardian 1 bot thread verify push WAIT T1.7 #65 Retry policy - - blocked by #63 - DONE T1.2 #59 Export schema #58 77e01ab - draft, swept you read it ------------------------------------------------------------------------------------------------------------------------ 2 advancing · 1 sweeping · 1 needs you · 1 waiting · 1 done lanes 2/2 · guardian 1/2 · drafted 3/4 · $11.40
An example board: one lane parked for you, two advancing, one draft being swept.
Enforced by code, not by prompts
Plenty of tools have an agent write code. This project is about the guardrails around the agents, and wherever possible each one is a script, a shim, a hook or a file rather than a sentence a model might skip.
| Guarantee | How it’s enforced |
|---|---|
| Stops at a draft PR | Every worker runs with git and gh shims first on its PATH, and the guard refuses merge, ready, approve, non-draft PRs, pushes to the integration branch and force-pushes without a lease. Claude Code sessions get the same check as a PreToolUse hook covering Bash and MCP tools, including the Orchestrator while a run is live. Your own sessions are untouched. |
| Green tests or no PR | The Developer’s ship step squashes, re-runs the full suite and opens a draft only from green. Red means a failure report with evidence, and no PR. |
| A second vendor reviews | The Reviewer’s findings must start with a canary line, or the review counts as not run. A same-vendor pairing is flagged on every PR. |
| Bounded spend | A third lane is refused, new lanes stop once the ticket budget has drafts, Claude seats run with dollar and turn caps, and every worker sits behind a wait timeout that turns silence into a board row. |
| No silent skills | Dispatch refuses a seat whose instruction files aren’t absolute paths that exist. A worker that ends without its status line is recorded as NOT RUN, never as a pass. |
| The Verifier checks the world | It drives the running app in every flag state — off, on with the setting off, on with the setting on — and records each journey as PASS, FAIL, NOT RUN or N/A with an artefact. An OFF row that wasn’t run makes the whole verdict “not verified”. |
| Clean-up never loses work | Seats are named worktrees, swept at the start of every run, and a tree with uncommitted or unpushed work is never removed. |
A smoke suite tests every guardrail that lives in code against fixtures, including a non-Claude agent that tries to push to main, merge, and open a non-draft PR. CI runs it on Linux and on macOS’s stock bash 3.2.
Five rules and one habit
I arrived at these by running agent lines on real codebases and watching where they went wrong. Each one exists because breaking it cost something.
- Never guess for the human. The line may stop and wait for you, but it never fills a gap in your decisions with its own guess, because a guess that reads as a decision gets built, reviewed and merged exactly as written.
- Questions belong on the ticket. A question about what the work should be is a planning defect, not a runtime event. It goes back upstream and onto the ticket, not into a run.
- No evidence, no block. A claim that something is missing or broken is checked against the repo before anyone acts on it. The search that returns nothing is the evidence.
- Stack on open work. An unmerged dependency is a base branch, not a blocker.
- Two review rounds, then a human. If two rounds haven’t converged, a third won’t either.
And the habit: check the effect, not the message. Treat your own agents’ reports the way you treat code. After any step that changes something outside the session, look at the result itself — run gh pr view instead of trusting “PR created!”.
Pay at creation, not at inspection
A defect that survives to a human costs a review, a fix, a re-review and a round trip: three or four agent sessions. Preventing it costs one extra effort step on one session. Model is the expensive dial and effort is the cheap one, so savings come from putting cheap seats on smaller models, never from making a seat that judges think less.
That shapes the defaults. Deciding and planning run in your own session at high effort, because nothing re-examines them until the Verifier. The Orchestrator is the only medium-effort seat, because it follows procedure rather than making judgements. In the reference setup, Claude Code fills the Developer, Guardian and Verifier seats and Codex fills the Reviewer seat, all at high effort — and every seat is pinned in config, so lowering your session default to save tokens can’t silently downgrade them.
The rest of the defaults keep a week’s allowance inside a week: two lanes as a hard ceiling; workers that report one status line and a file path instead of their output; waiting in a background shell rather than in turns; test output written to files; the Guardian batched; and the Developer and Reviewer never batched, because fresh context is the Reviewer’s product.
alf-stats reports sessions per ticket and cost per drafted ticket, and merging is the measurement: a draft only counts as a success once a human has merged it. If sessions per ticket climb well above four, the loop’s retro mode traces which stage is leaking work into it.
Built on good skills that already exist
Version 0.1 re-implemented deciding, planning and ticketing itself. Version 0.2 stopped doing that. The line now runs on two published skill packs, used as they are: Matt Pocock’s skills — grill-with-docs settles the decisions one question at a time and writes the ADRs, and to-tickets publishes the tickets with their blocking edges — and Addy Osmani’s agent-skills, which writes the plan and gives the Developer and Reviewer seats their build and review discipline. The term light software factory is Addy’s too.
AI Light Factory adds what those packs don’t have: the dev loop, the Verifier, the readiness gate, the guard and its shims, the seats and the board. The seats read the companion skill files by absolute path, so an agent from another vendor in a seat follows them too. Every skill is a SKILL.md in the open Agent Skills format.
What it deliberately doesn’t do
It doesn’t triage issues or run a scheduled queue — addyosmani/factory does that well, and this line starts from a PRD you bring. It doesn’t merge or deploy; pair it with your own CI and release train. It isn’t a swarm: two lanes is the ceiling. And it isn’t a hosted service. Everything runs in your terminal, against your GitHub, on your accounts.
The repo has a side-by-side comparison with other agent factories, if you want to see where it sits.
Run the line
AI Light Factory installs as a Claude Code plugin, next to the two skill packs it builds on. You need git, an authenticated gh, jq and python3 — and ideally two coding-agent CLIs from different vendors.
# install /plugin marketplace add lmagsino/ai-light-factory /plugin install alf@ai-light-factory /plugin marketplace add mattpocock/skills /plugin install mattpocock-skills@mattpocock /plugin marketplace add addyosmani/agent-skills /plugin install agent-skills@addy-agent-skills # run the line /setup-matt-pocock-skills # once: point the ticket skills at GitHub /alf:dev-loop setup # once per repo: seams, agents, companion paths /mattpocock-skills:grill-with-docs docs/prd.md # 1. decide /agent-skills:plan # 2. plan /mattpocock-skills:to-tickets tasks/plan.md # 3. tickets /alf:dev-loop #<parent-issue> --budget 2 # 4. dev loop, unattended /alf:verify #<parent-issue> # 5. every flag state