Post

AI

The Software Factory Is a Queue, Not a Commit Counter

HoYeon Lee reports a 300-commit-a-day Software Factory. Here is what that number measures, what it leaves unknown, and how to track accepted work instead.

AI-generated conceptual illustration of paper changes moving through a wooden review gate into an orderly tray; not an actual software interface or measured result.
AI-generated conceptual illustration of paper changes moving through a wooden review gate into an orderly tray; not an actual software interface or measured result.

Source Audit: A measured read of HoYeon Lee’s Software Factory post.

Curiosity: What does a 300-commit day measure?

A 300-commit day is a vivid activity signal. It is not automatically a throughput result. HoYeon Lee’s LinkedIn post says a Software Factory produced 300 commits per day during the Chuseok holiday. I have not independently verified the count. The post is a first-person report, not an activity export or external audit.

A commit answers one question: did Git record a change? It does not answer what user-visible task was finished, whether the change met its acceptance criteria, or whether it remained correct after integration. One task can involve many commits. One commit can be only a slice of a task. Neither fact makes commits useless; it makes them a poor proxy for accepted product work.

The better question is not “How many commits did the factory make?” It is “How much accepted work cleared the system, with what review cost and what evidence?”

Retrieve: The post already names the hard part

Lee does not present the factory as a commit-count contest. The post describes a multi-agent platform refactored from Swift to Electron over the holiday. A PRD became tasks; task dependencies were analyzed; pull requests were opened automatically; and Lee kept human review for key tasks.

The post also calls CI and verification a bottleneck. It recommends removing tests that are truly unnecessary, revisiting the harness’s validation loop, and avoiding a full verification cycle for every tiny change when the cost is disproportionate. The word “unnecessary” matters: deleting tests without replacing their risk coverage would simply hide the bill.

Lee’s proposed boundary is human-centered. Tasks should carry intent and a verification method. Agents can work through implementation and validation, while an Observer agent gives the person one place to follow direction and discuss decisions that need escalation. Lee reports that this reduced drift and made mid-course intervention more flexible.

Those are meaningful design details, and the qualitative result should be attributed accurately. The post does not give a baseline or a measured drift rate. It also does not report how many tasks were accepted, how many pull requests merged, first-pass CI results, escaped regressions, rollbacks, rework, or time spent waiting for human review. The public post therefore supports a workflow description and a self-reported commit rate, not an independently measured productivity conclusion.

Innovation: Build an acceptance ledger, not a counter

A factory needs a queue with explicit units, state changes, and evidence. A commit is a trace in that queue, not its finish line.

Start with a task that has an observable outcome and a verification method. Give it a stable ID, acceptance criteria, risk level, owner, and links to its pull requests and checks. Then distinguish states such as ready, in progress, blocked, in review, verified, accepted, and released. Keep rejected, needs rework, and rolled back visible rather than folding them into “done.”

The flow below is a proposed measurement model, not a diagram of Lee’s implementation.

flowchart LR
    A["Task + acceptance criteria"] --> B["Ready"]
    B --> C["Agent work and commits"]
    C --> D["Pull request + checks"]
    D --> E{"Human decision required?"}
    E -->|"No, risk-bounded"| F["Verified"]
    E -->|"Yes"| G["Review and decision"]
    G -->|"Needs rework"| C
    G -->|"Approved"| F
    F --> H["Accepted task"]
    H --> I["Release verification, when relevant"]
    C -->|"Blocked"| J["Record blocker"]
    J --> C

A merged pull request is not always an accepted product outcome. A task can require more than one pull request, and a passing check can establish only the property that check actually tests.

A pull request makes a change reviewable. GitHub Docs presents review as a distinct surface around the diff, comments, suggestions, and a formal review action. The counts visible in these documentation examples belong to the examples, not Lee’s Software Factory.

A GitHub pull request navigation bar with the Files changed tab outlined in orange.
GitHub Docs example: the Files changed tab exposes a diff for review. The example's counts are not measurements from Lee's project. Source page: GitHub Docs, Reviewing proposed changes. Repository scope: the pinned repository README includes content in the assets folder under CC BY 4.0. Publisher/creator: GitHub Docs. Reproduced unchanged from GitHub Docs under CC BY 4.0. License: CC BY 4.0. Original asset: pinned PNG.

A line comment is useful evidence of a review conversation, but a comment by itself is not approval. A suggested change is a concrete patch proposal, but it still needs the right checks and decision. The distinction matters when an agent can generate commits and open pull requests faster than a person can review them.

A code diff with the Add a line comment button highlighted beside a changed line.
GitHub Docs example: a reviewer can start a line-level comment beside a changed line. Source page: GitHub Docs, Reviewing proposed changes. Repository scope: the pinned repository README includes content in the assets folder under CC BY 4.0. Publisher/creator: GitHub Docs. Reproduced unchanged from GitHub Docs under CC BY 4.0. License: CC BY 4.0. Original asset: pinned PNG.
A pull request review comment editor with the Add a suggestion control highlighted.
GitHub Docs example: a reviewer can propose a code suggestion in a comment. A suggestion is a review artifact, not a count of accepted tasks. Source page: GitHub Docs, Reviewing proposed changes. Repository scope: the pinned repository README includes content in the assets folder under CC BY 4.0. Publisher/creator: GitHub Docs. Reproduced unchanged from GitHub Docs under CC BY 4.0. License: CC BY 4.0. Original asset: pinned PNG.
A GitHub pull request Files changed view with the Review changes button highlighted.
GitHub Docs example: Review changes is a distinct review action. The displayed counts belong to the documentation example, not Lee's factory. Source page: GitHub Docs, Reviewing proposed changes. Repository scope: the pinned repository README includes content in the assets folder under CC BY 4.0. Publisher/creator: GitHub Docs. Reproduced unchanged from GitHub Docs under CC BY 4.0. License: CC BY 4.0. Original asset: pinned PNG.

Protected-branch rules can require pull-request reviews and status checks before merging. That is a useful gate, not a complete definition of product acceptance. A green check proves that a particular check passed; a review decision proves that a reviewer made that decision. The task ledger should still connect both to the acceptance criteria and, when relevant, a post-release verification.

A dashboard can then answer questions the commit counter cannot:

SignalWhat it helps answerWhat it cannot prove alone
Accepted tasks per periodHow much work met its stated acceptance criteriaWhether the change remains healthy over time
Lead time and blocked timeWhere work waits between ready and acceptedWhether the task was valuable or correct
First-pass CI and reworkHow often changes need another implementation or check cycleWhether the product behavior is right
Escaped defects and rollbacksWhat integration or release failures cost after acceptanceEvery user impact or long-term outcome
Human review queue ageHow much decision work is waiting on peopleThe quality of each review decision

Do not set targets for these signals before measuring a baseline. A rate can be gamed just as easily as a count if the denominator is vague. First define what “accepted” means, then record a stable baseline and change one bottleneck at a time.

The Observer role can make this ledger easier to use. It can surface stale tasks, missing verification, blocked dependencies, and decisions waiting for a person. But an Observer that says “complete” is still reporting a status. The acceptance evidence should remain attached to the task, not inferred from the agent’s confidence or the number of commits it emitted.

Human review should also be risk-aware. Reversible, well-tested changes may be safe to automate within a narrow boundary. Ambiguous product decisions, security-sensitive changes, destructive data operations, and changes with a large user impact deserve an explicit escalation point. The boundary is not “human versus agent” in the abstract; it is “which decision can be delegated, with what evidence and recovery path?”

For a narrower example of how an agent action crosses a review boundary, read the Cloudflare OS MCP approval source audit.

For a reminder to separate published proxy numbers from inspected requests, see the Headroom evaluation audit. These audits concern approval and measurement boundaries; they do not validate Lee’s factory.

Editorial method: This draft uses AI-assisted source research and drafting; factual claims are linked to checked sources, and no first-hand experience is claimed.

Key Takeaways

  • Lee reports 300 commits per day, but the public post does not provide a measured accepted-work or quality rate.
  • The post does discuss verification bottlenecks, task intent, human escalation, and an Observer. Its qualitative drift improvement remains the author’s unquantified report.
  • Count tasks only when their acceptance criteria are evidenced. Keep commits, pull requests, CI, review decisions, and release checks as separate signals.
  • Use humans for explicit, risk-sensitive decisions. Do not treat an agent status or a green check as a complete product verdict.

New Questions

  • What is the smallest stable unit of accepted work for a multi-agent project: a task, a user-visible slice, or a release outcome?
  • Which classes of changes need a human decision, and can the system show the evidence that triggered escalation?
  • How much of the apparent throughput is spent in review, rework, blocked time, or post-merge repair?
  • What baseline would let the team optimize verification cost without silently increasing escaped defects?

References

Primary source

  1. HoYeon Lee, LinkedIn post on the Software Factory, accessed 2026-10-02. The commit count and the reported reduction in drift are attributed to the author and are not independently verified here.

Platform documentation

  1. GitHub Docs, Reviewing proposed changes in a pull request, accessed 2026-10-02.
  2. GitHub Docs, Managing a branch protection rule, accessed 2026-10-02.

Image rights and provenance

  1. GitHub Docs repository, README license scope at pinned commit and CC BY 4.0 license text. The four source PNGs are committed under assets/images/help/ at that same revision. The README places documentation and content in the assets folder within CC BY 4.0’s scope.
Working on something like this?

I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.

Work with me
This post is licensed under CC BY 4.0 by the author.

Search article titles, categories, and tags. Full text is searched when needed.

Type to search articles.