AI Code Review and Human Production Approval

AI Code Review and Human Production Approval

GitHub Copilot can now do something that would have sounded slightly absurd not long ago: approve a pull request in a way that can satisfy a repository’s required-approval rule. The feature entered public preview in September 2026. It is off by default, but an organization can turn it on and allow Copilot’s approval to count much like a teammate’s.

That is a meaningful shift. AI is no longer only writing the first draft, suggesting a test, or leaving review comments. It is moving into the part of the workflow where teams traditionally asked a human to say, “Yes, this change is ready.”

I still think that is fine. What I would not do is let that same automation own the final production decision. The distinction sounds small, but it changes how I think about the whole AI coding pipeline. Code review and production approval are not the same job. One is mostly about judgment on the code. The other is about accepting blast radius.

AI is getting very good at the first one. The second one is where I still want a name attached.

Access without medium partner: AI Can Approve Pull Requests ??

The usual “human in the loop” advice is too vague

A lot of AI safety advice for software teams boils down to “keep a human in the loop.” I do not think that is useful enough anymore. Where exactly?

If an agent writes 600 lines, another agent reviews them, a test suite checks them, static analysis runs, staging deploys successfully, and then a tired engineer scrolls to the bottom of the pull request and clicks Approve without really reading it, did the human make the system safer? Probably not.

That is approval theater.

The better question is which decision still needs a person, and what is that person supposed to judge?

GitHub’s own product design is already pushing teams toward this question. Copilot code review can run automatically, it can make approval assessments, and admins can now allow it to submit approvals that count toward merge requirements. GitHub is also measuring how often Copilot-reviewed pull requests get merged and how long they take to merge.

The old line between “AI writes code” and “humans review it” is already disappearing. So keeping a human on every review just because that is where humans used to sit is not much of a strategy.

AI review can be better than human review in some boring but important ways

Human code review has a fatigue problem. The first diff of the morning gets attention. The sixth gets less. The twenty-third one on a Friday afternoon sometimes gets a quick scan because the tests are green and the author is trusted.

An automated reviewer does not get annoyed by the fourth retry path in the same file. It can check the same security rule every time. It can compare the change against repository instructions, inspect surrounding files, look for missing error handling, and do it again after the author pushes another commit.

GitHub now offers different review-effort levels for Copilot, including a deeper mode for complex logic, security-sensitive code, and cross-service changes. OpenAI’s Codex review workflow similarly tells users to inspect pull-request changes, checks, conflicts, and specific behaviors, and it can prepare fixes from review findings.

This is exactly the kind of work agents are good at. Consistency matters, and a reviewer that never gets bored has value. I would rather have an AI check every database write for error handling than hope a human remembers to do it on diff number 47.

But five green checks can share the same blind spot

Here is the uncomfortable part. Automation can give you several independent-looking approvals that are not truly independent.

Imagine an agent writes a request handler. It updates an account record and returns success. A second agent reviews the diff. Unit tests pass. Types pass. Staging looks fine. A smoke test sends one request and gets the correct response.

Five green checks.

Then production gets a retry at exactly the wrong moment. The service acknowledges success before the write is durably persisted. The client believes the operation succeeded, the row never lands, and the retry path leaves the system in a state nobody tested.

The code can look clean. The tests can be honest. The review can be honest too. The problem is that every check inherited roughly the same model of reality: one request, normal timing, no awkward partial failure.

That kind of failure is easy to miss because the bug is not necessarily an ugly line of code. It is a bad assumption spread across the whole pipeline.

The source article that prompted this piece describes almost exactly that shape of incident: a write path acknowledged a request before persistence, automated review and tests all passed, and the failure appeared only when a real retry hit production. The author’s point was not that the agents were useless. The agents had caught other problems. The failure came from confusing review quality with ownership of the release decision.

That distinction is the part worth keeping.

Review asks “is this correct?” Production asks “what happens if it isn’t?”

Those are different questions.

A code reviewer looks for correctness, maintainability, security issues, missing cases, unclear names, broken interfaces, and bad assumptions. An AI system can do a lot of that work and will probably keep improving.

A production approver should be looking at something else: what does this change touch, can it lose data, does it change authentication or permissions, does it move money, does it alter persistent writes, and what happens if rollback is incomplete?

The person at the production boundary should not spend 20 minutes re-reading variable names that an agent already checked. They should be judging the risk envelope. That is a much better use of a human.

The merge button is not always the real boundary

This is where I would change one thing from the common “never automate merge” argument.

A Git merge is not magically irreversible. You can revert a commit, reset a branch, or open another pull request. Plenty of teams merge continuously into main without immediately exposing the change to customers.

The real boundary is the first moment the change can create an external side effect that a revert cannot fully erase.

For one team, that is merge-to-main because every merge automatically deploys to production. For another, it is a separate production deployment approval. For a mobile app, it might be a release submission. For infrastructure, it could be the apply step.

That is where I want the human signature. Not because humans are magically better at reading code, but because that is where the consequence becomes real.

Rollback is not the same thing as undo

Software teams sometimes talk about production changes as if rollback makes everything reversible. It does not.

If a deployment sends 40,000 wrong emails, rolling back the code does not unsend them. If a migration drops data, rollback may not restore the missing rows. If an authorization bug exposes private records for ten minutes, reverting the commit does not make those ten minutes disappear.

The same is true for billing, external APIs, destructive writes, and one-way migrations. You can fix the software and still be left with real-world cleanup.

This is why I think the useful dividing line is reversibility of consequences, not reversibility of Git history. An AI agent can be excellent at deciding whether a diff looks safe. A human still needs to own the decision to cross a boundary where “we can revert” stops being the full answer.

GitHub’s current Copilot design quietly admits this

There is an interesting tension in GitHub’s own documentation.

On one side, Copilot code review can now submit an approval that counts toward required approvals when administrators enable the feature. That is a clear move toward letting AI perform more of the formal review work.

On the other side, GitHub’s documentation for Copilot-generated pull requests still tells users to review the changes thoroughly before merging. In repositories that require pull-request approvals, GitHub also has extra rules around how approvals for agent-created changes are counted.

That makes sense. GitHub is separating agent judgment from merge authority instead of pretending they are automatically the same thing.

I think teams should do the same, even if their exact policy differs.

OpenAI and Anthropic are solving approval fatigue too

This is not only a GitHub problem.

OpenAI introduced Auto-review for Codex earlier this year because constant human approval prompts create friction. Instead of interrupting the user at every sandbox boundary, a separate reviewer model can approve or deny many actions. OpenAI says this reduced human approval stops by roughly 200× in its internal measurements.

Anthropic has been working on the same basic problem in Claude Code. Its auto mode uses classifiers to automate some permission decisions because users approve the overwhelming majority of prompts, and constant “yes, yes, yes” clicking trains people to stop paying attention.

That is the approval-fatigue problem in one sentence: a human gate that fires too often becomes a decorative gate.

So I do not want humans approving every file read, command, test, generated diff, and harmless branch update. Automate the boring boundaries. Save human attention for the expensive one.

The best production gate should be short

This is another place teams get it wrong.

If the final human approval requires reading a 3,000-line diff, checking six dashboards, manually re-running tests and opening three tickets, people will eventually rubber-stamp it.

The production gate should be compact. The agents should prepare the evidence.

For every change, I would want an automatically generated release packet with the things a production owner actually needs: blast radius, changed services, data writes, migrations, auth changes, retry behavior, rollback plan, observability, and known test gaps.

Then the person approving production can answer a much narrower question:

Am I comfortable sending this risk to real users now?

That is a useful human decision. “Did I personally re-read all 3,000 lines?” usually is not.

Risk-based approval is better than forcing a human onto every change

I also would not make every production change equally manual.

A typo on a documentation page does not need the same approval process as a payment-write change. Neither does a static asset with an instant rollback path.

The pipeline should classify risk.

A change touching authentication, permissions, money movement, persistent writes, schema migrations, secrets, infrastructure policy, or destructive operations should get a human production sign-off. A low-risk change with no persistent side effects, clean canary metrics and a reliable automatic rollback path can probably ship with much less friction.

This is where AI can help again. Let an agent assign a risk score and explain why. Let a second agent challenge that score. If both classify the change as low risk and the policy agrees, automate more of the path.

But for the high-risk class, I still want a person to accept the release.

That is a more practical policy than “humans approve everything forever.”

The human should review blast radius, not syntax

This changes what senior engineers need to get good at.

The valuable human skill is moving upward.

Less time asking whether a loop should be extracted into a helper. More time asking whether the change can corrupt state across a retry.

Less time checking formatting. More time checking whether a database migration can be rolled back after partial completion.

Less time spotting obvious missing tests. More time deciding whether the current monitoring is enough to detect the failure before customers call.

AI review makes code-level judgment cheaper. That does not make engineers less useful. It makes system ownership more important.

Accountability is not a philosophical argument. It is an operational control.

I do not mean that an AI cannot feel guilty. That discussion goes nowhere.

The practical point is much simpler: companies need a clear answer to who had authority to release the change.

Incident response needs an owner. Audit systems need an owner. Regulated environments often need separation of duties. Customers need someone who can explain why the change shipped and what will change after the incident.

“The agent approved it” may describe what happened inside the pipeline. It is not a complete governance model.

A named production owner creates a place where responsibility lands, and that changes behavior before the click. People ask different questions when they know they are the person who will have to explain the outcome later.

That is useful.

Agents should make the human’s decision harder, not easier

There is one failure mode I would watch carefully.

Agents are good at producing polished confidence.

A release summary that says “all checks passed, no issues found, low risk” can make the final approval feel routine. That is exactly what you do not want.

The final agent in the pipeline should be a skeptic. Its job should be to argue against release.

What assumptions are shared across the tests? What was not exercised in staging? Which side effects cannot be reversed? What happens on duplicate delivery, timeout, partial write, stale cache, or retry?

If the release still looks safe after that, the human approval becomes much more meaningful.

I would rather have an AI tell me why I should not ship than another AI congratulate the first one for clean code.

The best AI coding pipeline is not fully autonomous

I know “fully autonomous software engineering” makes a better demo.

Give an agent a ticket. It writes the code. Another agent reviews it. Tests run. The change deploys. Nobody touches anything.

That is a beautiful loop until the system reaches a part of reality the tests did not model.

I think the better design is almost fully autonomous. Let agents generate, review, run tests, stage changes, inspect telemetry, prepare rollback commands, and challenge each other.

Then stop at the production consequence boundary when the change can materially affect users, data, money, permissions, or infrastructure.

That one stop does not ruin the automation. It gives the automation somewhere to end responsibly.

The strange future is that the merge approver may read less code

This is the part that will make some engineers uncomfortable.

The human who approves production may eventually read less of the diff, not more.

If agents become better at code review than tired humans, redoing their work is wasteful. The human should instead read the risk summary, inspect the high-risk lines, check the failure plan, and understand what will happen if the system is wrong.

That is closer to an aviation checklist than a traditional code review.

I think that is where mature agentic development is heading. The person is not there because the machine cannot reason. The person is there because production is where reasoning becomes consequence.

Copilot approving a PR is not the scary part

The September Copilot update will probably make some teams uncomfortable because the AI can now submit an approval that counts toward merge requirements.

I am less worried about that than I would have been a year ago.

Automated review is getting good enough that teams should use it aggressively.

What worries me is a pipeline where the same automated confidence flows all the way from generation to production with no independent owner at the end.

That is how five honest green checks become one bad night.

So yes, let the agent write the code. Let another one attack it. Let Copilot approve the pull request if your team is comfortable with that. Let CI do everything it can.

But when the change is about to cross the boundary where a customer, a database, a payment, or an access-control rule can feel it, I still want one person to say:

Ship it.

Not because humans are better reviewers.

Because somebody has to own what happens next.

Sources I checked

GitHub: Copilot code review can now approve pull requests

GitHub Docs: Using Copilot code review

GitHub Docs: Review output from Copilot

OpenAI: Auto-review of agent actions

Anthropic: How we built Claude Code auto mode

Post a Comment

Previous Post Next Post