As AI agents accelerate code production, engineering teams increasingly face a massive pull request (PR) bottleneck. Left unchecked, the continuous acceleration of agent-generated code turns repositories into a “slop cannon” of fragile, low-quality changes. To counter this, engineering teams must introduce three deliberate quality “brakes”—automated checks, automated agent reviews, and focused human review—designed to systematically prevent software entropy while speeding up overall delivery.
The Fallibility of Automated Checks and Codebase Design
Deterministic checks such as linting, type-checking, and unit tests are cost-effective because they run on CPU cycles rather than expensive tokens. However, automated tests can easily lie. AI agents frequently generate tautological tests that simply re-assert implementation details or create over-mocked tests that cannot fail in real runtime environments. To prevent agents from writing structure-sensitive tests, codebases must adopt deep module architecture—hiding complex logic behind minimal, robust interfaces. Designing clean boundaries forces agents to write behavior-focused tests rather than tests coupled to internal implementations.
Separating Implementation from Code Review
Overloading an implementation agent with architectural rules, bug fixing, and coding standards leads to poor execution because its context window is already strained by exploration and debugging. Instead, teams should implement a two-step process: make it work, then make it good. The implementation agent writes the initial functionality, while an underloaded reviewer sub-agent evaluates the diff against a dedicated coding standards document. Rather than leaving verbose comments for human developers to resolve, the automated reviewer agent should directly commit the necessary style and standard fixes.
Streamlining Human Review with Risk Triage
Human review should serve as the final filter, focusing only on high-impact changes. PRs should be categorized based on merge danger:
- One-Way Doors: High-risk modifications (e.g., irreversible database migrations, external notifications, potential data loss) requiring meticulous manual review.
- Two-Way Doors: Low-risk, easily revertible changes with localized blast radius that can be approved with minimal friction.
To reduce cognitive load, PR descriptions should prioritize visual aids, pseudocode, and system diagrams rather than walls of text to explain the context behind the diff.
Reviewing the System, Not Just the Code
Reviewing agentic PRs requires shifting focus from correcting individual mistakes to fixing the system that created them. By implementing post-review retrospectives, teams can analyze review sessions to identify recurring patterns. These insights should continuously feed back into updated coding standards, improved automated checks, and refined agent steering instructions, ensuring developers never have to leave the same PR comment twice.
Mentoring question
How much of your current PR review process is spent pointing out repetitive style or design issues that could be converted into automated agent rules, and how do you differentiate review depth between one-way and two-way door decisions?
Source: https://youtube.com/watch?v=LlgiOCmFG_w&is=RJjKwdbkHROw4M5U