← Leodegario Magsino Jr.
Agentic engineering · Open source

AI Light Factory

A light software factory for coding agents. PRD in, reviewed draft PRs out. You make the decisions and approve the plan. Agents build and review unattended, and stop at a draft PR. You merge.

2
Lanes, hard ceiling
2
Review rounds, then you
0
Merges by an agent
The line: you write a PRD; the Architect (grill-with-docs) settles the decisions; the Planner writes the plan (agent-skills plan) and publishes tickets with blocking edges (to-tickets); the dev loop, where one agent builds, an agent from another vendor reviews, and a PR Guardian answers review threads, opens draft PRs; the Verifier (/alf:verify) checks every flag state; you merge.
The line. Dashed boxes are agents, solid boxes are you, and every ◆ is a gate you approve.
Why

Why I built it

Over the past year my work as a tech lead moved from “an engineer with an AI pair” to “an engineer running a line of agents”. The agents got good enough to take a ticket from description to pull request. What they didn’t do was make the work trustworthy on their own. Running a line like that on a production codebase taught me three things, and each one became a default here.

The bottleneck moves

Agents write code faster than anyone can read it. Once typing stops being the constraint, deciding, reviewing and verifying become it. So the guardrails belong on decisions, review and verification, not on code generation.

Unlimited parallelism is a trap

Run wide open, a fleet of agents burns a week of model budget in days and opens more pull requests than any human can review. Every extra worker grows its own context, and the coordinator’s with it. Two lanes is the ceiling, and it is not a starting point to tune upward.

A reviewer from the same vendor shares the builder’s blind spots

When the model that wrote the code also reviews it, its misses are silent: the review passes because both passes share the same gaps. So the Reviewer seat runs on a different vendor by design, and if you put the same vendor in both seats, every PR says so.

AI Light Factory is that line, rebuilt from scratch for open source, so any team can run it with whichever agents they already use.

The idea

Light, not dark

Addy Osmani drew the distinction in Software Factories, Light and Dark. A dark factory ships code that no human has read. A light factory keeps a human in the loop, and gives a loop only as much autonomy as you can cheaply and reliably verify.

This project takes the light side literally. Your judgement moves to where it pays off: you settle the decisions and approve the plan before any code exists, agents build and review unattended, and you read every diff before it merges.

The line never marks a pull request ready, never approves, and never merges. That isn’t a sentence in a prompt that a model might decide to skip. A guard hook and a pair of git and gh shims refuse those commands for every worker, whatever agent it runs.

How it works

The line

A factory isn’t for speed. It exists so that the same mistake cannot be made twice quietly. Every stage writes an artefact to a known path, and the next stage reads it instead of re-deriving it. When something goes wrong, you can name the stage that should have caught it, and put the check there.

The alternative is one long agent session that reads, decides, plans, builds and reviews without ever writing down what it concluded. That works until it doesn’t, and when it doesn’t, there is nothing to inspect.

StageAgent and skillLeaves behindYou decide
1 · Decide Architectgrill-with-docs · Matt Pocock ADRs and a glossary Do the decisions hold?
2–3 · Plan Planneragent-skills:plan · Addy Osmanito-tickets · Matt Pocock A plan, then GitHub issues with acceptance criteria and blocking edges Right plan? Right tickets and edges?
4 · Dev loop Developer, Reviewer, PR Guardian/alf:dev-loop Squashed draft PRs, a live board, a ledger Nothing. It runs unattended
5 · Verify Verifier/alf:verify A verdict for every feature-flag state Merge. Always yours

Three cadences, not one pass

Deciding, planning and ticketing happen once per project. The dev loop runs continuously, a ticket at a time per lane, through a parent issue or milestone. The Verifier runs once per milestone. Mixing these up is the common failure: a Verifier run per ticket is theatre, because until a phase is nearly done there is nothing end to end to exercise, and re-opening decisions per ticket produces a plan that contradicts itself.

What each stage is really for

None of the stages is mainly for what its name suggests. Deciding is corrective: it catches the requirement that would have been built exactly as written, because it read like a decision already made. Planning is arithmetical: it catches the plan that is internally inconsistent, which only becomes visible once the dependency graph is written down as real edges. The dev loop is compressive: its value isn’t that it writes code, but that a ticket reaches you already reviewed, swept and squashed, so your attention goes to the diff rather than the process. The Verifier is falsifying: it is the only stage whose job is to find out that the plan was wrong. Everything before it checks internal consistency. The Verifier checks the world.

The dev loop

Unattended, and still accountable

The dev loop is the part no existing skill pack covered, and the reason the project exists. Your session becomes the Orchestrator. It owns the queue of ready tickets, the base branches, the board and the budget, and it never edits a file, never reads a full diff, never builds and never reviews. It dispatches seats and reads one-line results. Each seat is a named git worktree — Developer, Developer-2, Reviewer, Guardian — with an agent of your choice in it.

Nothing left to ask

Before a ticket gets a seat, a readiness gate checks it. A ticket without a clear “what to build”, checkbox acceptance criteria free of TBDs, and explicit blocked-by edges is refused, labelled as needing a decision, and sent back upstream. That is why the dev loop never asks you a question mid-run: every decision you could make was taken while deciding, and written into the ticket. Anything it would want to ask is an upstream defect, recorded on the board and on the issue while the run carries on.

Build, then a review from another vendor

The Developer builds one ticket test-first, following Addy Osmani’s incremental-implementation and test-driven-development skills. It squashes, re-runs the full suite, and opens a draft only from green. The Reviewer then reads the diff with fresh context, on a different vendor, following code-review-and-quality plus three reach questions an ordinary review skips: does the change alter an outward contract, does it touch code that other paths run through, and does it undo a decision that was made on purpose?

Findings go through a severity ladder instead of to you. Critical findings are fixed. Important ones are fixed when the fix is small, and otherwise listed in the PR body for you. Suggestions are dropped with a one-line reason. Every call lands in the PR’s decision log, so what the loop decided without you is auditable exactly where you are already looking. If round one kept nothing, the ticket ships. If it kept something, the Developer fixes it and the same Reviewer reads only the delta. A Critical that survives round two parks the lane: two review rounds, then a human.

Five reasons to stop, and none of them is a question

A lane stops for five things only:

  1. The Reviewer says the approach is wrong.
  2. A fix would touch a contract, a registry or a migration the ticket didn’t ask for.
  3. A Critical finding survives round two.
  4. A previously passing test fails for a reason outside the diff.
  5. The base branch moved underneath the lane.

Each one is a fact that makes the lane unable to proceed. A stopped lane becomes a parked row with the evidence posted on its issue, and every other lane keeps running. A worker that hits a cap, dies, goes silent or never reports parks its lane too, with its log. Nothing is re-dispatched silently.

Stack on open work

A dependency on an unmerged pull request is a base branch, not a blocker. The loop branches off the open parent, reviews against it and bases the PR on it, so a milestone keeps moving while drafts wait for you. Four places have to agree on the base — the commit count, the branch point, the review diff and the PR base — or the review reads the parent’s commits as if they belonged to this change.

The PR Guardian

Once drafts are open, review threads arrive from bots and from people. The PR Guardian sweeps them in batches of about five, because most of a short session is startup cost. It triages each thread with the same ladder, fixes what qualifies, replies to every thread it touches with what it did and why, and rebases stacked PRs whose parent moved. It answers questions about the code from the code, leaves questions about what the work should be for you, and never marks a PR ready.

One screen to read

While the loop runs, the board is the only thing you need to read. It is ordered by what needs you, and it checks that every claimed worker is actually alive before printing it.

2026-10-01-m1-export-foundation  Thu 01 Oct 21:14
------------------------------------------------------------------------------------------------------------------------
FLAG   TICKET                     PR / SHA      SEAT         STATE                       NEXT
------------------------------------------------------------------------------------------------------------------------
PARK   T1.4 #61 Currency rounding #64 a19f003   -            round 2: Critical survives  needs you
LANE   T1.5 #62 CSV writer        -             Reviewer     review round 1              judge findings
LANE   T1.6 #63 Column registry   -             Developer-2  building                    review
SWEEP  T1.3 #60 Export job        #66 3bf66cd   Guardian     1 bot thread                verify push
WAIT   T1.7 #65 Retry policy      -             -            blocked by #63              -
DONE   T1.2 #59 Export schema     #58 77e01ab   -            draft, swept                you read it
------------------------------------------------------------------------------------------------------------------------
2 advancing · 1 sweeping · 1 needs you · 1 waiting · 1 done     lanes 2/2 · guardian 1/2 · drafted 3/4 · $11.40

An example board: one lane parked for you, two advancing, one draft being swept.

Guardrails

Enforced by code, not by prompts

Plenty of tools have an agent write code. This project is about the guardrails around the agents, and wherever possible each one is a script, a shim, a hook or a file rather than a sentence a model might skip.

GuaranteeHow it’s enforced
Stops at a draft PREvery worker runs with git and gh shims first on its PATH, and the guard refuses merge, ready, approve, non-draft PRs, pushes to the integration branch and force-pushes without a lease. Claude Code sessions get the same check as a PreToolUse hook covering Bash and MCP tools, including the Orchestrator while a run is live. Your own sessions are untouched.
Green tests or no PRThe Developer’s ship step squashes, re-runs the full suite and opens a draft only from green. Red means a failure report with evidence, and no PR.
A second vendor reviewsThe Reviewer’s findings must start with a canary line, or the review counts as not run. A same-vendor pairing is flagged on every PR.
Bounded spendA third lane is refused, new lanes stop once the ticket budget has drafts, Claude seats run with dollar and turn caps, and every worker sits behind a wait timeout that turns silence into a board row.
No silent skillsDispatch refuses a seat whose instruction files aren’t absolute paths that exist. A worker that ends without its status line is recorded as NOT RUN, never as a pass.
The Verifier checks the worldIt drives the running app in every flag state — off, on with the setting off, on with the setting on — and records each journey as PASS, FAIL, NOT RUN or N/A with an artefact. An OFF row that wasn’t run makes the whole verdict “not verified”.
Clean-up never loses workSeats are named worktrees, swept at the start of every run, and a tree with uncommitted or unpushed work is never removed.

A smoke suite tests every guardrail that lives in code against fixtures, including a non-Claude agent that tries to push to main, merge, and open a non-draft PR. CI runs it on Linux and on macOS’s stock bash 3.2.

Rules

Five rules and one habit

I arrived at these by running agent lines on real codebases and watching where they went wrong. Each one exists because breaking it cost something.

  1. Never guess for the human. The line may stop and wait for you, but it never fills a gap in your decisions with its own guess, because a guess that reads as a decision gets built, reviewed and merged exactly as written.
  2. Questions belong on the ticket. A question about what the work should be is a planning defect, not a runtime event. It goes back upstream and onto the ticket, not into a run.
  3. No evidence, no block. A claim that something is missing or broken is checked against the repo before anyone acts on it. The search that returns nothing is the evidence.
  4. Stack on open work. An unmerged dependency is a base branch, not a blocker.
  5. Two review rounds, then a human. If two rounds haven’t converged, a third won’t either.

And the habit: check the effect, not the message. Treat your own agents’ reports the way you treat code. After any step that changes something outside the session, look at the result itself — run gh pr view instead of trusting “PR created!”.

Cost

Pay at creation, not at inspection

A defect that survives to a human costs a review, a fix, a re-review and a round trip: three or four agent sessions. Preventing it costs one extra effort step on one session. Model is the expensive dial and effort is the cheap one, so savings come from putting cheap seats on smaller models, never from making a seat that judges think less.

That shapes the defaults. Deciding and planning run in your own session at high effort, because nothing re-examines them until the Verifier. The Orchestrator is the only medium-effort seat, because it follows procedure rather than making judgements. In the reference setup, Claude Code fills the Developer, Guardian and Verifier seats and Codex fills the Reviewer seat, all at high effort — and every seat is pinned in config, so lowering your session default to save tokens can’t silently downgrade them.

The rest of the defaults keep a week’s allowance inside a week: two lanes as a hard ceiling; workers that report one status line and a file path instead of their output; waiting in a background shell rather than in turns; test output written to files; the Guardian batched; and the Developer and Reviewer never batched, because fresh context is the Reviewer’s product.

alf-stats reports sessions per ticket and cost per drafted ticket, and merging is the measurement: a draft only counts as a success once a human has merged it. If sessions per ticket climb well above four, the loop’s retro mode traces which stage is leaking work into it.

Credits

Built on good skills that already exist

Version 0.1 re-implemented deciding, planning and ticketing itself. Version 0.2 stopped doing that. The line now runs on two published skill packs, used as they are: Matt Pocock’s skills — grill-with-docs settles the decisions one question at a time and writes the ADRs, and to-tickets publishes the tickets with their blocking edges — and Addy Osmani’s agent-skills, which writes the plan and gives the Developer and Reviewer seats their build and review discipline. The term light software factory is Addy’s too.

AI Light Factory adds what those packs don’t have: the dev loop, the Verifier, the readiness gate, the guard and its shims, the seats and the board. The seats read the companion skill files by absolute path, so an agent from another vendor in a seat follows them too. Every skill is a SKILL.md in the open Agent Skills format.

Scope

What it deliberately doesn’t do

It doesn’t triage issues or run a scheduled queue — addyosmani/factory does that well, and this line starts from a PRD you bring. It doesn’t merge or deploy; pair it with your own CI and release train. It isn’t a swarm: two lanes is the ceiling. And it isn’t a hosted service. Everything runs in your terminal, against your GitHub, on your accounts.

The repo has a side-by-side comparison with other agent factories, if you want to see where it sits.

Try it

Run the line

AI Light Factory installs as a Claude Code plugin, next to the two skill packs it builds on. You need git, an authenticated gh, jq and python3 — and ideally two coding-agent CLIs from different vendors.

# install
/plugin marketplace add lmagsino/ai-light-factory
/plugin install alf@ai-light-factory
/plugin marketplace add mattpocock/skills
/plugin install mattpocock-skills@mattpocock
/plugin marketplace add addyosmani/agent-skills
/plugin install agent-skills@addy-agent-skills

# run the line
/setup-matt-pocock-skills                         # once: point the ticket skills at GitHub
/alf:dev-loop setup                               # once per repo: seams, agents, companion paths
/mattpocock-skills:grill-with-docs docs/prd.md    # 1. decide
/agent-skills:plan                                # 2. plan
/mattpocock-skills:to-tickets tasks/plan.md       # 3. tickets
/alf:dev-loop #<parent-issue> --budget 2          # 4. dev loop, unattended
/alf:verify #<parent-issue>                       # 5. every flag state