Case study

How we run our engineering on agent teams

The problem

A small team had more operational engineering work than people, and no appetite for unreviewed automated merges.

The setup

Paperclip and Hermes Agent are the control plane. Work is tracked on a board. Agent workers run in git worktrees. Claude Code, Codex, and pi are the coding delegates. GitHub is where the work lands.

The gates

A pull request is opened. An automated reviewer approves at the exact head commit. CI must be green. A person merges. Nothing consequential merges on an agent's own say-so.

What we measured

StatWhatDetail
570pull requests merged in a single monthAcross five of our own repositories, in the closed window 2026-08-27 to 2026-09-26.
287of them were approved by an automated reviewer before mergeA reviewer app approved the pull request at its exact head commit before it could merge.
Everyconsequential action waits for a personMerges, sends, and irreversible steps are gated on a human approval, not on the agent's own confidence.
2AgentPod tenants running on our own k3s clusterMeasured 2026-09-27 with a read-only kubectl get pods; both gateway pods 3/3 Running.

These are our own numbers from our own repositories and cluster. The pull-request figures are measured over the closed window 2026-08-27 to 2026-09-26 (UTC); the cluster figure was read on 2026-09-27. We are not claiming a customer outcome. We publish the method so you can judge the number.

What we would not claim

We cannot attribute a merged pull request to an agent versus a person from git metadata, because agent workers push under a human identity, so we report merges and review gates rather than an "agent-authored" count.

Map one workflow — free 45-min session, written go/no-go