Reviewing agent-written pull requests without becoming the bottleneck
When an agent writes most of the diff, the review queue becomes the slowest part of the team. Here's how to review for the things only a human can catch, push the rest into the harness, and keep your judgment sharp while you do it.
Six months ago a pull request was a conversation between two engineers who both understood what the code was for. Now the author is often an agent, the diff is three times larger than it used to be, and there are four of them waiting for you by lunch. The instinct is to read every line harder. That doesn't scale, and it isn't the point of review anyway. Review exists to catch what the author couldn't see — and an agent can't see a lot. Here's how to review for the things only a human can catch, hand the rest to the harness, and keep the queue moving without rubber-stamping.
Why reading harder doesn't work
Human-written PRs came with a hidden property: the author had already spent their judgment. They'd thought about whether the change belonged, whether it fit the system, whether it would surprise the on-call engineer. Review was a second opinion on a decision someone had already made carefully. An agent-written PR arrives without that first opinion. It's syntactically clean, it passes the tests it wrote, the description is fluent — and nobody has yet asked whether it should exist in this shape. If you review it the old way, line by line, you'll catch typos and miss the fact that the whole approach is wrong.
The volume makes it worse. Reviewers who try to hold the old standard across three times the diff do one of two things: they burn out and the queue grows, or they start skimming and approving, which is the worst of both worlds — the appearance of review without the protection. The way out is to change what you're reviewing for.
Decide what a human is for
Sit down and write the list of things a careful reviewer on your team catches that no linter, type checker, or test would. It's shorter than people expect, and it's almost entirely about intent and context rather than correctness. Does the change solve the problem that was actually asked for, or a nearby easier one? Does it fit how this codebase does things, or does it import a second way of doing the same thing? What happens at the edges — empty input, a retry, a slow dependency, a user who does the thing nobody designed for? Will the person paged at 2am understand it? Those questions are your review. Everything else is a job for the harness.
- Problem fit. The PR solves the ticket as written, not a simpler problem the agent quietly substituted. Read the ticket before the diff.
- Shape. New code follows the conventions already in the repo — same error handling, same logging, same module boundaries. One new pattern per PR is a smell.
- Edges. Empty, huge, duplicated, slow, and concurrent. Pick the two most likely and trace them by hand.
- Blast radius. What else calls this? What else reads this table? The agent rarely looks; you must.
- Operability. Can it be rolled back, flagged off, and debugged from logs by someone who didn't write it?
Let the machine review the machine's work
Everything you removed from the human list goes into automation — and the bar for "automated" is that it runs before a human ever opens the PR. Formatting, lint, types, coverage thresholds, dependency and secret scanning, size limits on the diff. If your team has been meaning to add any of these for years, agent volume is the forcing function. Then add the layer most teams skip: a second agent as the first reviewer. Give it the same checklist you wrote above and ask it to leave comments, not approvals. It will catch a surprising share of the edge cases and convention drift, and it will do it at three in the morning so the human review starts from a cleaner diff.
One rule keeps this honest: a machine can block a merge but never approve one. The approval is the human's signature that the intent is right, and that is exactly the judgment you've been protecting.
Read the description first, and demand a good one
An agent-written PR description tells you how the agent understood the task, which is the single most useful thing in the whole submission. If the description says "added caching to the user endpoint" and the ticket said "fix the slow dashboard," you've already found the most important problem without opening a file. Make the description a required artifact with a required shape: what was asked, what was changed, what was deliberately not changed, and how to verify it. If the agent — or the engineer driving it — can't fill that in, the PR isn't ready for a human, and sending it back costs you thirty seconds instead of thirty minutes.
Size the PR to the review, not the task
Agents are happy to produce a 1,400-line change because nothing about the task told them not to. Reviewers are not happy to receive one, and worse, a large diff hides the real risk behind a wall of plausible code. Put a size limit in the harness and make the engineer driving the agent split the work: the mechanical refactor in one PR, the behavior change in another, the migration in a third. The mechanical ones get a fast, mostly automated review. The behavior change gets your full attention on a diff small enough to actually hold in your head. Same total code, a fraction of the review cost, and far better odds of catching the thing that matters.
Keep your judgment from atrophying
There's a quiet risk in all of this. If you only ever review, and the agent only ever writes, your sense of how the system behaves starts to fade — and that sense is the whole reason your review is worth anything. Protect it on purpose. Once a week, take an agent-written PR and rewrite one function by hand before reading the agent's version. Pick one edge case per review and run it, don't just reason about it. Rotate who reviews what, so no one becomes the only person who understands a corner. Review stays valuable exactly as long as the reviewer still knows what good looks like from the inside.
What this looks like when it works
A team that does this well has a review queue that moves in hours, not days, and approvals that mean something. Agents produce the volume; the harness absorbs the mechanical checks; humans spend their attention on the small number of questions that decide whether a change is right. Nobody is reading harder. Everyone is reading for the right things.
That shift — from checking code to judging intent — is the review skill that matters now, and it's worth more than any amount of speed at reading diffs. The agent will keep getting faster. Your job is to stay the person who knows whether it built the right thing.
Want a coach in your corner?
Book a 1:1 call — we'll map your next step and pressure-test your plan. Group courses coming soon.
Get one like it in your inbox each week
Practical guides and roadmaps for the AI shift — free, no hype, written by the same people.




