We use cookies. Analytics and advertising, only if you accept.

ArkusNexus
October 2, 2026

The Review Bottleneck: the same ticket, run two ways

When coding agents produce pull requests faster than a team can review them, the problem is not simply that reviewers need to work faster. The deeper problem is that review is being asked to answer questions that should have been settled earlier:

When coding agents produce pull requests faster than a team can review them, the problem is not simply that reviewers need to work faster. The deeper problem is that review is being asked to answer questions that should have been settled earlier:

  • What was the agent allowed to change?
  • Which contracts had to remain stable?
  • What evidence would prove the feature worked?
  • Which configuration or security controls were in scope?

That is the review bottleneck: too many decisions arriving at the same place, after implementation is already complete.

A recent ArkusNexus webinar, The Review Bottleneck, examined a practical alternative: three explicit decisions distributed across the delivery process. The practice does not depend on a new platform. It uses a plan in the repository, CI, a review agent, pull requests, and a named person accountable for the final decision.

The Review Bottleneck: the same ticket, run two ways

The demonstration used the same ticket and the same model in two runs:

Let a listener share today’s Daily Mix with a friend through a link.

The difference was whether the work had to pass through explicit gates before merge.

Run 1: the ticket moves through implementation, green CI, human review, merge and deployment, but exposes private listening data

Run 1: no plan

The first run followed a familiar path:

  1. The agent implemented the ticket and opened a pull request.
  2. Type-checking, linting, tests, and the build all passed.
  3. A reviewer validated the change and made a decision.
  4. The change merged and deployed.

The result: the public share page exposed the owner’s private listening data.

This was a demo, but it reads as a real failure, and that is the point. CI was green because the defect was not a type error, lint error, or failing test. Nobody had written down which data was allowed to cross the new boundary: a public link that someone could paste anywhere.

The defect was a planning defect wearing a review costume.

CI checks what it was told to check. A new public boundary requires a statement of what may cross it before implementation begins. That statement is the plan.

Run 2: the ticket passes through a plan checkpoint, implementation, CI, review checkpoint, and final decision before merge

Run 2: with the gates

The second run added three decisions:

  1. The agent drafted a plan and interviewed the person accountable for the feature.
  2. The plan was approved before implementation.
  3. A review agent checked the implementation against the approved plan and assembled evidence for a final decision.

The flow was:

Ticket → /plan → plan decision → implementation → CI → review agent → implementation decision → merge and deploy

The outcome: the public page carried exactly the allowlist.

The difference was explicit boundaries around what automation was allowed to do.

The three decisions

These are three places where someone says yes or no.

Decision 1: What may the agent implement?

Before implementation, create a one-page plan that defines both the intended change and its limits.

The plan should cover:

  • Files that may change
  • Data stores that may change
  • Contracts that must remain stable, including systems inside and outside the repository that consume them
  • Configuration and permissions touched
  • Required security controls
  • Risk tier
  • Acceptance criteria, with the evidence that will prove each one

For a new boundary (such as a public link, export, webhook, or administrative response reused elsewhere) the plan must include an allowlist: what crosses and what does not.

A planning agent can read the repository, draft the mechanical sections, and ask only the questions that require a product or technical decision. The plan itself remains a file in the repository.

Approval is enforced through Git:

  • The plan is merged through its own pull request with status: approved.
  • That merge is the approval.
  • A plan-gate check on the implementation pull request verifies that the approved plan exists on main.

This is where Run 1 would have been stopped. The owner would have had to state what the public response could contain before the agent implemented the route.

Decision 2: Was the plan followed?

At review time, the review agent receives four inputs:

  • The approved plan
  • The diff and the full content of every touched file
  • The real CI output
  • Contract documentation and captured evidence

It then reports five checks, each as PASS, FAIL, or UNKNOWN, with a file and line for every finding:

  1. Files outside the plan
    Were any unplanned files changed?

  2. Scope and contracts
    Did the implementation touch anything marked out of scope or any contract that was supposed to remain stable? A stable contract exposed through a new route counts as touched.

  3. Evidence for acceptance criteria
    Is there evidence for every criterion? A test asserting what a response must not contain counts. The existence of a test file does not; its assertions must be read.

  4. New dependencies
    Are all new dependencies listed? CI or a lockfile can prove that a dependency exists. Suitability and security are separate questions.

  5. Configuration, permissions, and controls
    Were these changed, and did the plan approve those changes? If the plan requires a control that was not built, this check is a FAIL.

Every claim is labelled:

  • OBSERVED: seen in the diff, CI output, or captured response
  • DOCUMENTED: stated in the plan or supporting documentation
  • ASSUMED: inferred rather than verified

UNKNOWN is never green. If the available evidence is insufficient, the review should identify the exact evidence needed to resolve the uncertainty.

The review agent does not approve or reject the pull request. It assembles evidence so a named decision-maker can decide. The public prompt explicitly tells the agent not to evaluate code style, suggest refactors, or decide the merge. It is available verbatim in scripts/review-gate/PROMPT.md.

Decision 3: Is the implementation approved?

The final decision uses the review result as a routing signal:

Result Route Who decides
All five checks PASS Green: eligible for focused review. Review depth comes from the plan’s risk tier. A named owner with the plan, verdicts, and evidence
Check 3 or 4 FAIL, or any UNKNOWN Yellow: return for evidence and re-run. The author or agent supplies the evidence
Check 1, 2, or 5 FAIL Red: stop and return to the plan owner to amend or reject. An unapproved control change is treated as a control issue first. The plan owner

Precedence is Red over Yellow over Green.

Green is a status check, not an approval. A named person still reads the change and merges it. The gates reduce the amount of evidence that must be assembled manually; they do not remove accountability from the final decision.

If you want to see where these three decisions land in your own work: book a 30-minute backlog review. Thirty minutes on two or three items from your own backlog, no deck — we map where the plan gate, the review agent, and the final decision go, and who owns each one. Book a backlog review

This is a practice, not a platform

The three decisions map onto work most teams already do:

  • The plan is a file in the repository.
  • Approval is a merged pull request.
  • Review verdicts are a pull request comment.
  • Enforcement in the kit comes from GitHub Actions and branch protection.
  • The review agent can be the service or model the team already uses.

What changes is the discipline. The team writes down what may cross a new boundary before an agent builds it, then assigns a named owner to each decision.

That distinction matters for adoption. The gates do not eliminate review work; they make the work more specific and better timed.

How we run this in our pods

Our AI-Native Development Pods are built around this model. A Pod starts with two Product Engineers plus agents. Each engineer runs the whole cycle on his own slices: spec, plan, implementation, review and release. A named human accepts the agents' work at every gate. Your team starts out owning two of those gates: approving the spec and plan, and the production go/no-go. You hand them to the Pod only when you decide it knows your product well enough. Every change that touches auth, permissions, data exposure or new dependencies also needs a blocking security sign-off from the second engineer.

How to adopt the gates without slowing the team

Start with one repository and the plan gate.

Do not roll out all three decisions at once. The plan gate is the highest-leverage and cheapest step: one page and one merged pull request per change. It is also where the demonstration defect would have been caught.

The plan gate does add a step before implementation. It will not necessarily make the first commit arrive faster. Its value is preventing rework caused by unclear scope, unrecorded contracts, and missing acceptance evidence.

Add the review agent second. Add routing and branch protection third.

For an organization-wide rollout, the central decision is not which tool to buy. It is who owns each plan and who has authority to send it back. The workflow works when ownership is explicit.

The open-source kit

The webinar’s kit includes:

  • Two GitHub Actions workflows
  • The /plan command
  • The plan template
  • The review prompt

It is MIT licensed and available in the arkusnexus-labs/daily-mix-demo repository. Read README-GATES.md first; it explains the setup, flow, fail-closed behavior, and local execution.

The accompanying Gate Map worksheet is one page and works with GitHub and whatever review agent your team already uses.

ArkusNexus has been building software for 23+ years, with a senior engineering bench of 250+ engineers and 92% client satisfaction (industry average ~70%).

The full replay is available here.

FAQ

Does this replace code review?

No. Green is a status check, not an approval. A named person still reads and merges the change, with review depth determined by the plan’s risk tier.

Does it work with existing CI and review agents?

Yes. The workflow uses GitHub Actions, the repository’s existing check output, and the review agent the team already pays for. The kit provides the workflow and prompt, and works with GitHub today; the practice carries over to GitLab or other CI.

Is the plan gate worth the extra step?

For changes involving new boundaries, data exposure, contracts, permissions, or security controls, it is often the cheapest place to catch ambiguity. It may slow the first commit; it is designed to prevent later rework and unsafe assumptions.

What happens when the review agent returns UNKNOWN?

UNKNOWN is never green. The pull request follows the Yellow route until the author or agent supplies the missing evidence and the checks run again.

Is the kit actually open source?

Yes. The repository is public and MIT licensed. Start with README-GATES.md.

Who owns the plan?

A named person accountable for the feature owns the decision to approve, return, amend, or reject the plan. The planning agent can draft it, but it does not own the approval.

The practical takeaway is simple: decide what may change before the agent changes it, assemble evidence against that decision, and give a named person the final yes or no.

If you want the three decisions mapped onto your own backlog, book a 30-minute backlog review — thirty minutes on two or three items, no deck. If you would rather try it yourself first, start with the Gate Map worksheet and the open-source kit.

About the Author

Jorge Orenday

Jorge Orenday

Consultant Delivery Manager at ArkusNexus