The practice of code review, a cornerstone of software development for the past two decades, is undergoing a profound transformation driven by the rapid advancement of artificial intelligence. While the concept of a pull request has been standard for approximately 20 years, and internal code reviews at major tech companies like Google began around 2006, the fundamental nature and efficacy of these processes are now being challenged. This evolution is not merely a refinement but a necessary re-evaluation of what code review was intended to achieve and how those objectives can be met in a landscape where machines are increasingly responsible for generating significant portions of code.
The traditional code review model, often characterized by a rigid approval gate positioned immediately before code integration, is struggling to keep pace with the accelerating velocity and volume of software development. Historically, code review served critical purposes: identifying defects before they entered the main codebase, facilitating knowledge transfer and mentorship for junior engineers, and ensuring distributed understanding of project context. However, over time, these noble intentions have often devolved into bureaucratic compliance measures, gatekeeping, and, in many organizations, a performative ritual.
As the industry grapples with the implications of AI-generated code, the dilemma is no longer a simple choice between human oversight and unchecked machine output. Instead, it necessitates a fundamental reimagining of the core functions of code review—defect prevention, knowledge sharing, and quality assurance—for a future where AI agents are primary code contributors.
The Stranglehold of the Merge Point
For years, the answer to "where does code review happen?" has been universally understood: right before merging code. This singular focus on the pre-integration stage has created a bottleneck, making it the sole point of scrutiny in the development pipeline. Practices like trunk-based development, test-driven development (TDD), and pair programming, championed by firms like Thoughtworks, have long sought to distribute review and validation throughout the development lifecycle. In these methodologies, the benefits of review are realized continuously, often within the context of a pair programming session, rather than weeks later when a substantial diff is presented. When development avoids long-lived branches and is supported by robust testing, a green build pipeline often serves as the primary indicator of code quality.
The idea that pre-integration should not be the exclusive domain of code review is not new. What has been conspicuously absent is a compelling catalyst for change. The emergence of powerful agentic AI and AI-powered Integrated Development Environments (IDEs) capable of producing code at an unprecedented scale has provided precisely that catalyst. When an AI agent can generate an entire feature’s worth of code in a matter of hours, the critical points for identifying potential issues must shift upstream. This means moving the focus from the final code artifact to the very moment a developer articulates their intent to the AI tool. By the time a diff is generated, the underlying decisions and potential misalignments, often representing thousands of lines of code and hours of computational effort, have already been made.
One prominent approach addressing this shift is "intent-driven development," exemplified by companies like Aviator. This paradigm emphasizes capturing developer intent at its inception. This can manifest as a concise description of the feature’s scope, explicit exclusions, or a detailed list of acceptance criteria. The crucial insight is that intent is most accurately and effectively captured directly from the prompts and decision-making processes that occur when a developer collaborates with an AI agent.
Re-evaluating the Object of Review
The shift towards AI-generated code fundamentally alters what constitutes a "reviewable artifact." The traditional focus on lines of code is becoming increasingly impractical. Instead, the process of articulating intent now produces a diverse array of documentation, including markdown files detailing feature specifications, logs of clarifying questions, and the intricate decision trees embedded within complex prompt conversations. Many teams currently do not version control these artifacts, leading to a fragmentation of the reviewable context. Consequently, the very definition of what is deemed "worth reviewing" is in flux.
This recalibration has led to the emergence of distinct approaches among development teams. Some organizations are adopting a minimalist strategy, reviewing only the initial specifications and placing a high degree of trust in the subsequent code generation. Others are attempting to maintain the comprehensive review of all artifacts—specifications, generated code, and everything in between—though they readily acknowledge this approach increasingly positions them as the primary bottleneck. A third group, facing the sheer volume of AI-generated code, is opting to review little to nothing directly, relying instead on extensive testing of the running system as their primary quality assurance mechanism.
The sheer volume of code being produced by AI agents is fundamentally reshaping how code reviews are conducted. A diff that spans more than five files already presents a significant challenge for human reviewers attempting to reconcile intended changes with actual implementation. When this volume is multiplied, the reviewer’s need extends beyond a simple diff and a ticket. They require visibility into the original intent, the reasoning process employed by the AI agent, and, of course, the generated code itself.
Prioritizing Intent Over Implementation Detail
A more effective utilization of senior engineers’ time involves shifting the focus from scrutinizing hundreds of lines of code to examining a concise set of intent statements and acceptance criteria. The reviewer’s primary question becomes: "Is this addressing the right problem with the appropriate constraints?" This approach not only optimizes reviewer effort but also preserves the vital knowledge-sharing aspect of code review.
For instance, a seasoned reviewer who is aware of a long-standing, well-tested date-handling library within the platform can codify this knowledge into the organization’s "AI slop register." This register acts as a repository of established best practices and common pitfalls. By analyzing thousands of past review comments, clustering recurring themes, and seeking human approval for invariant candidates, teams can proactively prevent future errors. Each codified invariant represents a potential code review comment that will never need to be manually written again, significantly scaling the effectiveness of quality assurance.
This process is also accelerating the practical realization of collective code ownership—the principle that no single individual should hold all project context. As AI agents require explicit and well-defined context to function effectively, this knowledge must be externalized from individual developers’ minds and integrated into the project’s accessible knowledge base.
A recent discussion featuring Vanitha Kumar from Thoughtworks on "What Happens to Code Review When Agents Write the Code" further underscores this paradigm shift. The conversation highlighted the increasing need for review processes that can adapt to the accelerated pace of AI-assisted development.
The Evolution of Automated Review Mechanisms
Current AI code review tools often function as add-ons to existing platforms like GitHub or GitLab, generating comments and suggestions. While this automates aspects of the review process, it can inadvertently perpetuate the "theater" of human-driven reviews—an agent reads comments, engages in automated debates, proposes changes, or defends its output. This automated loop, while efficient in its own way, does not fundamentally alter the downstream nature of the review. Crucially, this form of automated review no longer necessitates a human-facing user interface.
The core principle is that code review, whether performed by humans or AI, must occur much earlier in the development cycle. This early intervention can take various forms, from informal "rubber ducking" sessions to targeted teaching moments. By pushing review activities as far "left" as possible in the development pipeline, the amount of corrective work required downstream is minimized. An advisory or adversarial AI agent that actively monitors and guides the code generation process, identifying and rectifying anti-patterns as they emerge, offers far greater value than an agent that simply comments on completed code.
At Thoughtworks, a code review agent was developed, initially serving as a pedagogical tool for junior developers. It was trained on the team’s established archetypes and conventions documentation, with the directive to identify deviations. This agent subsequently evolved into a more comprehensive review agent.
Tools like Aviator Verify are pushing this evolution further by simulating server environments, injecting real-world traffic, and driving UI interactions to confirm that the generated code aligns with the stated intent, rather than merely appearing correct. The ultimate goal is to provide reviewers with concrete evidence, transforming the review from a line-by-line inspection to a critical assessment of whether the provided evidence and the initial intent are sound.
Gradual Transformation, Not Overnight Revolution
This profound shift in code review practices does not occur through top-down mandates or organizational memos. Cultivating a culture that embraces new methodologies requires demonstrating their efficacy through practical application and tangible results.
One notable instance involved a Thoughtworks client that maintained a strict policy requiring all consultant-produced code to undergo review. However, a pilot program employing spec-driven development resulted in an unprecedented volume of markdown specifications and unusually large code change sets. The existing review process quickly proved unsustainable, forcing the organization to confront the reality that traditional, exhaustive reviews would inevitably render them a bottleneck.
The core philosophy emerging from this evolution can be summarized as: "Verification goes to the machines. Judgment and knowledge stay with people." While machines excel at consistent, high-speed verification, human judgment and domain knowledge remain indispensable for deeper understanding, mentorship, and strategic decision-making.
The development of an "AI slop register" is an iterative process. In its initial stages, it may feel like duplicating effort—performing traditional code reviews while simultaneously working to codify invariants. However, once established, this register automates the prevention of recurring mistakes, freeing reviewers from having to address the same issues repeatedly.
The Enduring Purpose of Review in an AI-Driven World
The fundamental reasons for code review—defect detection, knowledge transfer, and maintaining visibility into development decisions—remain constant. What is changing is the locus of these activities. Verification tasks, which are highly amenable to automation due to their reliance on pattern recognition and rule adherence, are increasingly being delegated to machines. These tools are demonstrably faster and more consistent than human reviewers in executing such tasks.
Conversely, judgment and knowledge-based aspects of review, which are crucial for learning, mentorship, and strategic alignment, will continue to reside with human engineers. The evolution, therefore, is not a move away from review altogether, but a strategic reallocation of effort. The emphasis is shifting from meticulously "reading the code" to critically "reading the intent" behind the code. This subtle yet significant change promises to make software development more efficient, robust, and adaptable to the accelerating pace of technological advancement. The journey of code review is far from over; it is entering a dynamic new chapter shaped by the capabilities and challenges of artificial intelligence.
