Case study
How we run our engineering on agent teams
The problem
A small team had more operational engineering work than people, and no appetite for unreviewed automated merges.
The setup
Paperclip and Hermes Agent are the control plane. Work is tracked on a board. Agent workers run in git worktrees. Claude Code, Codex, and pi are the coding delegates. GitHub is where the work lands.
The gates
A pull request is opened. An automated reviewer approves at the exact head commit. CI must be green. A person merges. Nothing consequential merges on an agent's own say-so.
What we measured
| Stat | What | Detail |
|---|---|---|
| 570 | pull requests merged in a single month | Across five of our own repositories, in the closed window 2026-08-27 to 2026-09-26. |
| 287 | of them were approved by an automated reviewer before merge | A reviewer app approved the pull request at its exact head commit before it could merge. |
| Every | consequential action waits for a person | Merges, sends, and irreversible steps are gated on a human approval, not on the agent's own confidence. |
| 2 | AgentPod tenants running on our own k3s cluster | Measured 2026-09-27 with a read-only kubectl get pods; both gateway pods 3/3 Running. |
These are our own numbers from our own repositories and cluster. The pull-request figures are measured over the closed window 2026-08-27 to 2026-09-26 (UTC); the cluster figure was read on 2026-09-27. We are not claiming a customer outcome. We publish the method so you can judge the number.
What we would not claim
We cannot attribute a merged pull request to an agent versus a person from git metadata, because agent workers push under a human identity, so we report merges and review gates rather than an "agent-authored" count.