A few weeks ago, we did something crazy. We made human pull request (PR) reviews entirely optional on the Corridor monorepo by building an auto-approver.
Now, many of our PRs are merged without a human reviewer ever even looking at the code. An agent reads the diff, decides whether the change is safe, approves it, and the PR content lands on main. This is only possible because we dogfood our own product: Corridor runs on every change, so security is enforced continuously instead of depending on a human reviewer to catch it. Below, we discuss how we achieved safe and secure auto-approval at Corridor.
Why human review became the bottleneck
Coding agents have made it extremely cheap to produce code. At Corridor, we routinely have cloud agents like Cursor and Devin which pick up a task, write a feature, and open a PR without anyone ever touching a keyboard.
When the effort of producing a change is so slight, the constraining factor is verifying this change. A team that still requires a person to read and approve every diff will find that, before long, its reviewers rather than its authors are setting the pace.
The question we set out to answer on our own repository was a fairly narrow one: which PRs actually require human judgment, and how can we reliably and securely let the rest through automatically?
To answer this question, we rolled out our auto-approver in three steps.
Step 1: Auto-approving only mechanically-safe changes
We turned on the first version of our auto-approver in May. This version was inspired by the observation that a large fraction of our PRs were entirely mechanically safe. That is, we can often decide that they are safe to merge given the paths involved and the shape of the diff; we don't need to reason in detail about runtime behavior.
Some mechanically-safe PRs might be:
- A behavior-preserving refactor, where we split one large file into several smaller ones
- Removing a block of dead code
- Adding new tests
- Bumping a dependency version and updating the matching lockfile
We wrote out this policy in a file at the root of the repository called AUTO-APPROVE.md and pointed a Cursor automation at this file. The automation triggers on every PR opened, classifies the code change into one of three buckets, and marks the PR with one of the following labels:
0-auto-approved— safe to merge without human review. The agent submits the approving review itself.1-minor-review-needed— a short, roughly five-minute human look.2-major-review-needed— substantial human review required.
A PR is eligible to be marked as 0-auto-approved only if every file in it falls into one of a small set of allow-listed categories: tests-only, a behavior-preserving refactor like a file split or dead-code removal, documentation, or a version bump of a package that already existed. If a pull request mixes categories, say a refactor together with a one-line behavior change, it is no longer eligible for auto-review. The policy is written to prefer the more cautious label whenever the classification is ambiguous.
However, this was only a small step towards truly automatic merges: we were only automating the long tail of low-risk changes. Additionally, even if a PR is auto-approved, it is ultimately up to the PR author to decide whether to seek additional human review.
Step 2: Auto-approving higher-risk changes
Classifying by path only gets you so far, and it leaves the more intensive and time-consuming portion of work — the actual feature changes — entirely to human reviewers. The second step, which came together in the middle of June, was to build a reviewer that could reason about substantive changes rather than merely recognize safe-looking ones.
We call this the ai-review path. When a developer makes a PR and adds the ai-review label, it calls a background agent to review the change in earnest. The agent reads the description, the diff, the surrounding code, and the tests, and then writes a structured review to .pr-reviews/<number>.md following a review spec we maintain. That spec requires the reviewer to reach a view on several things that matter to us: correctness, security, feature duplication, readability, performance, and, without exception, test coverage. Crucially, the reviewer does not assess security from scratch: the prompt has it read and build on Corridor's findings for the PR, so our own product's analysis of the diff feeds directly into the review rather than sitting off to the side as a separate check. If the agent determined that no concerns were raised, it should leave an approving review.
The reason we wrote the review to a .md file was to enable a final round of low-effort human review. This way, the human author can review the agent's review document and make changes to the PR if needed.
Each review records the exact commit it was written against in a reviewed_commit field. Our safety check treats the review as current only when the sole difference between that commit and the head of the pull request is the review file itself. If any code is pushed after the review, the approval goes stale automatically and the change has to be reviewed again. An agent cannot approve code it has not read.
This is the step that let us hand substantive work to agents. In practice most of the pull requests we label ai-review are opened by our own cloud agents: one agent writes the change, a second reviews it against the spec, and if it holds up, the human who initiated the change can merge the pull request. Importantly, the agent that writes a change is never the one that approves it.
Step 3: Evaluating the auto-approver
The risk of an auto-approver is that it may make systematic mistakes across the many PRs it reviews. This is the core problem of securing a software factory: the moment you remove humans from the critical path, nobody notices when the bot is quietly getting something wrong, and small mistakes can compound into a codebase full of mistakes.
To catch this early, we continuously measure three things: how well the auto-approver's calls align with what careful reviewers would have decided, how often auto-approved changes actually get reverted, and the velocity increase associated with using the auto-approver to see if it's actually worth it.
1. Alignment
The alignment eval asks a simple question: when the auto-approver makes a call, do the reviewers looking at the same PR agree?
The first signal is what the author or reviewers did afterward. Every time the bot labels a PR 0-auto-approved, 1-minor-review-needed, or 2-major-review-needed, we snapshot it and watch what happens: did a human approve and merge with no comment? Request changes? Leave comments that led to follow-up commits? We score each snapshot as an agreement or a disagreement, weighted by how clear the signal was. For example:
- Auto-approved, then merged with no human comments is a high-confidence agree.
- Auto-approved, then a reviewer requested changes is a high-confidence disagree.
- A reviewer left comments but the human didn't make any follow-up commits is a lower confidence disagree.
But a human merging without comment is a weak signal on its own: maybe they agreed, maybe they didn't look closely. So we also bake in the other review bots that read the same PR, such as Devin review or Corridor review. An independent agent that read the full diff and raised nothing is a much stronger vote that the auto-approver got the call right than a silent human merge. When the auto-approver, the human, Devin's review, and Corridor all line up, we count it as a high-confidence agree.
So far we've scored 1,657 auto-approver PRs against what a human ultimately did with them. On the 1,418 that reached a clear verdict, humans agreed with the auto-approver's call 70% of the time. On high-confidence verdicts, where a human clearly approved or clearly pushed back, the bot and the human agree around 96% of the time.
2. Reversion rate
While alignment tells us whether reviewers agreed at merge time, reversion rate tells us what happened after: of the changes the auto-approver approved, how many had to be reverted because something was wrong?
In the seven weeks before the auto-approver went live, roughly 1.2% of merged PRs were subsequently reverted — all with full human review. Since auto-approve, the reversion rate has dropped to around 0.7% on nearly double the weekly volume. Interestingly, changes the auto-approver approved on its own had the lowest reversion rate of any category, while the PRs flagged for the most human scrutiny had the highest.
3. Velocity
Alignment and reversions tell us the auto-approver is safe. The rest of the picture is whether it was worth doing, and that shows up in how fast we now ship.
Before the auto-approver we were merging around 115 pull requests a week. Afterwards, this roughly doubled to around 220 a week, with one week reaching 307.
The latency numbers are more informative. The chart below is the median time between opening a pull request and merging it, which is a reasonable proxy for the amount of time a change waits on review.
The first step on its own (auto-approving only mechanically-safe changes) barely moved the median. The median only fell substantially (from around 21 hours to under 3 in the most recent full weeks) once the second step with the ai-review agent began clearing substantive PRs. This shows that the benefit shows up when the system can review the changes that actually carry risk, not when it handles the trivial ones.
The last chart looks at who actually approved each merged pull request, week by week: a person, the AI reviewer, or both.
Through early May, every merge still carried a human approval. The AI reviewer's first approvals show up in the week of May 18, and its share climbs from there. By the last full week of June a clear majority of merges were approved by the agent, sometimes alongside a human and increasingly on its own, while the human-only bar became exceedingly small.
But doesn't auto-approval carry concerning security implications?
Yes, absolutely — allowing a language model to approve and merge code introduces real attack surfaces, and as a security startup, one of the most important parts of our ethos is to "move fast and break nothing".
This is why we took great care to ensure that using our auto-approver on our own codebase would not significantly increase the chances of insecure code reaching main. To do this, we rely on all four of these mechanisms which work in tandem to protect against potential vulnerabilities:
- Corridor. Our own product runs on every pull request as a required, blocking check: it reviews the diff for vulnerabilities, applies the guardrails we have configured, and reports findings inline. A change that Corridor flags cannot merge until the finding is resolved or explicitly accepted, regardless of whether it was written by a person or an agent and regardless of who, or what, approved it. Corridor also runs on our cloud and local agents to fix vulnerabilities before they are even introduced, before the pull request is created. Finally, the auto-approver takes into account Corridor's findings when making its assessment. Combined, this gives us the confidence that our code is secure.
- A human backstop, enforced through CODEOWNERS. Some changes always require a human reviewer, and the way we codify that is to list the sensitive paths and their owners in a
CODEOWNERSfile. GitHub then treats owner review as a required check: a pull request that touches one of those paths cannot merge until a listed human has approved it. The paths we protect this way are the ones where a mistake is expensive or hard to reverse, such as production infrastructure in Terraform, the auto-approval policy itself, and AI prompts. When an agent would otherwise be reviewing a change to the very system that lets agents merge code, CODEOWNERS puts a person back in the loop. - The automated review process, which is the system described above. The
ai-reviewagent, and the bucket classifier before it, form an independent opinion on every change and have to justify it against the review spec. This layer includes the freshness rule — an approval only stands while the code it was written against is unchanged — so an agent cannot approve code it has not actually read. And the reviewer is told to treat everything a contributor controls as untrusted data, and never to act on instructions found inside it. That is the same discipline about untrusted input that we build into the Corridor product itself. - Our engineers. Ultimately, we rely on our engineers to take responsibility for the code that they merge to our main branch. We trust their judgment to determine when a PR is truly complex enough to need another pair of eyes, versus when it is sufficient to rely on an AI-written review.
Improving the auto-approver
Our auto-approver is far from perfect. Currently, we only use a single agent, a single model, and a single prompt spec. This can leave us exposed to a single model's blind spots: if it is confidently wrong about some category of change, nothing is positioned to catch it, and the same mistake repeats across every similar PR.
We are actively working on several improvements:
- Model ensembling. Today, a single model makes the approval decision. We are moving toward running multiple models independently and requiring consensus before auto-approving. Ensembling is one of the most reliable ways to reduce the rate of confident-but-wrong decisions — the failure modes of different models tend not to overlap.
- Tighter category boundaries. The eval data shows us which kinds of changes the bot misjudges most often. We are using that to refine the classification policy — narrowing the categories that qualify for auto-approval and expanding the ones that trigger human review.
- Richer context. The current classifier works primarily from paths and diff shape. We are experimenting with giving it deeper context: test coverage data, semantic understanding of what changed, and history of past issues in the affected area.
- Continuous eval collection. The dashboard today runs on-demand. We are moving to continuous collection so we can track the agreement rate as a time series and set alerts if it degrades.
What a team needs before turning on an auto-approver
If a team wants to move towards something like a software factory in which people sit outside the critical path for most changes, here are the changes that they need to move towards:
-
Sound security checks, both during development and in the PR. Taking a human out of the loop widens the attack surface: a coding agent can be steered by a malicious instruction buried in an issue or a dependency, or simply introduce a vulnerability with the same confidence it writes everything else. Security cannot be a single scan bolted on at the end; it has to run while the agent works, catching and fixing vulnerabilities as code is written, and again as a blocking check on the PR, so nothing merges until it clears. This is what we use Corridor for: it runs on our cloud agents to fix issues before a PR even exists, and then as a required, non-negotiable gate on every pull request.
-
Trust in developers, and the tools to act on it. Automated merging only works in an engineering culture that is willing to point capable automation at its own codebase and that gives people the leverage to do so well. A team whose instinct is to add gates and approvals in order to slow people down will not get here; the objective is closer to the opposite, which is to remove the gates that are not earning their place.
-
An end-to-end test suite you actually trust, and agents that add to it. auto-approval is only as safe as the tests behind it. If it is safe to merge a change without a person reading it, that is because a thorough and reliable test suite would catch a regression. That implies two commitments: investing seriously in end-to-end coverage, and requiring coding agents to produce tests as part of the change they make. Our review spec treats new or changed source that ships without adequate unit and end-to-end coverage as a blocking problem, so an agent that writes a feature also writes the tests that demonstrate it.
-
Genuinely good pull-request review. The bar for an automated reviewer is higher than for a human one, because it runs unattended and at volume. Ours has to form a view on correctness, security, duplication, readability, performance, and test coverage, and to justify that view rather than simply assert it.
-
Evals for the reviewer itself. It is hard to responsibly ship a reviewer you cannot measure. We treat our review system the way we treat any other AI product we build, by running it against labeled datasets and tracking false-positive rates, agreement rates, and confidence breakdowns — and publishing those numbers honestly.
-
Engineers who own automations, not just pull requests. Once agents can open and merge changes safely, the natural next step is to point them at recurring engineering work rather than only at one-off features. We ask each engineer to own a set of automations for the housekeeping that used to sit in a backlog, and we run a growing number of them. A few examples:
- Datadog error triage and remediation — an automation that watches our Datadog errors, triages them, and opens a fix.
- Customer-bug triage from Slack — a Devin automation that reads customer bug reports in our
#eng-bugsSlack channel and turns them into actionable, triaged work. - Coverage repair through chaos testing — an automation that mutates random parts of the codebase, watches what breaks without any test catching it, and writes the missing tests for those gaps.
- Style cleanup of recent pull requests — a daily pass over the last day's merged pull requests that fixes them up against our code-style conventions.
- From customer conversations to shipped features — an automation that reads our calls in Gong and Granola and our notes in Notion, works out what customers are asking for, files Linear tickets, and builds the features.
We made human review optional on our own repository because this is the direction the world is moving towards. Get these right and auto-approving will help you accelerate development and trust in a codebase. Our job is to build the security layer that lets every team become more autonomous safely. If you are working toward something similar, we would love to compare notes!
Our auto-approver prompt
Here is the entire prompt we use for our auto-approver, the document that tells the reviewing agent how to write a review. It is the same thing a person on our team would read to understand what we expect a review to cover, and it is what the ai-review agent follows before it can approve a substantive change.
# Writing a PR review
We write a short review per pull request so a teammate gets a clear
read on whether it's safe and clean to merge. Reviews live at
`.pr-reviews/<number>.md`. The job is to judge the change on
correctness, security, and simplicity, and to flag anything worth a
decision before merge.
Use this prompt to write one:
Read the PR — its description, the diff, and the code around the
changes. Read the tests too. Don't guess; if you can't confirm
something from the code, say so rather than asserting it.
**Wait for the Corridor CI check.** Before writing the review, wait
for the Corridor CI check on the PR to finish and read its findings —
it's a required input, not an optional extra. If it's still running,
hold off; if it errored or never ran, say so in the review rather than
reviewing without it. Treat its findings as one signal to weigh
alongside your own reading of the code.
Open the file with YAML frontmatter pinning what you reviewed, so a
reader (or the auto-approve bot) can tell whether the review still
matches the PR:
---
pr: <number>
branch: <head branch>
reviewed_commit: <full SHA of the code tip you reviewed>
reviewed_at: <YYYY-MM-DD>
verdict: <no major concerns | concerns raised>
---
`reviewed_commit` is the commit whose code you read — pin the current
code tip, then add this review file on top of it. It is **not** the
eventual HEAD (a commit can't contain its own hash). The auto-approve
bot treats the review as fresh only when the sole change between
`reviewed_commit` and HEAD is this review doc; any code commit pushed
afterwards makes it stale, so re-review and re-pin if the code moves.
Then write the review the way you'd talk a colleague through it:
- **Overall** — a couple of sentences on whether it's solid, what it
does well, and your bottom line. Be honest, not flattering.
- **Duplication of existing features** — check whether the PR
re-implements a feature, endpoint, util, hook, or flow that already
exists on `main`.
- **Security** — the part to read closely. Start from the Corridor CI
check findings (see above) and your own reading of the diff. Number
the things worth raising and tag each with a rough severity and kind,
e.g. *(medium / defense-in-depth)*; note which came from the CI check
and which you found yourself. Explain the concern in plain terms and
what you'd do about it. Then add a short "correctly handled" list
confirming the things you checked that are fine — tenant scoping,
query-level ownership checks, "not found" over "forbidden", wrapping
user text before a model/Slack, bounded request bodies — so the
reader knows you looked.
- **Readability** — comment density, naming, duplication, anything
dense enough to slow a reader down.
- **Performance** — anything that won't scale, clearly marked
non-blocking if it's fine for the current scope.
- **Test coverage (required for approval)** — new or changed source
must ship with tests. Treat new or changed source landing without
adequate unit + e2e coverage as a blocking concern.
- **Suggested actions** — a short list: what to fix, what to document
+ ticket, what's optional.
Set `verdict: no major concerns` only when nothing in the review
blocks merge. Otherwise use `concerns raised` and say what would
change your mind.