RunnableLog in
A line of candidate objects moving beneath a magnifying instrument toward an inspected resultEvery incoming finding needs an outcomeConfirm or Reject

SEPTEMBER 28, 2026 · RUNNABLE TEAM

Why Multi-Agent Code Review Needs Explicit Adjudication

Engineering

Adding reviewers to an AI code review sounds like a straightforward way to improve it. Let several agents inspect the change independently, challenge one another, and ask a fresh adjudicator to decide what should be published. The extra perspectives can help, but they also create a less visible problem: someone has to account for every candidate finding from the moment it enters the debate until the final review is assembled.

Without that accounting, a finding can disappear because its title changed, appear twice because its line range moved, or return after an adjudicator rejected it. None of those failures require a bad code analysis. They can happen in the coordination layer around otherwise sensible reviewers.

Runnable now treats that coordination as a versioned protocol. When a finding enters evidence depth from the earlier review, it receives a stable identity and must end in exactly one explicit disposition: confirmed or rejected. Evidence depth can then test the strongest surviving claims against the merge-base and pull request head.

More reviewers create a coordination problem

HOW A THREE-REVIEWER CONCLAVE WORKS
ONE INPUTPinned PR snapshotIdentical base and head revisions for every pass
01 / Independent reviewNo reviewer sees the others' conclusions
PASS 1Reviewer 1Proposes evidence-backed findings
PASS 2Reviewer 2Proposes evidence-backed findings
PASS 3Reviewer 3Proposes evidence-backed findings
SHARED CANDIDATE POOLAll proposed findingsEvery challenger receives the complete set
02 / Adversarial challengeLook for counterexamples and equivalent behavior
CHALLENGE 1Challenger 1Tests every candidate
CHALLENGE 2Challenger 2Tests every candidate
CHALLENGE 3Challenger 3Tests every candidate
03 / FRESH ADJUDICATORWeighs the evidenceRejects weak or contradicted claims
ONE OUTPUTAdvisory GitHub reviewOnly surviving findings on the pinned head
03 / CONCLAVE Two to four reviewers are selectable; every pass shares the configured cost ceiling.

Independent passes are useful because one reviewer does not anchor the next. Each reviewer sees the same pinned base and head revisions, proposes findings from its own context, and contributes them to a shared candidate pool. Challengers can then look for counterexamples, unchanged behavior, or assumptions that the diff does not support.

The shared pool is where the system stops being a collection of model calls and becomes a protocol. Candidates may overlap. A challenger may restate a bug more precisely. An adjudicator may move the affected range after tracing the actual execution path. The final wording can be better than the original while still referring to the same underlying claim.

A reliable review therefore needs more than a final list of comments. It needs continuity between the claims that entered adjudication and the decisions that came out.

Text is not identity

It is tempting to match a final finding back to its source by title, file, and line range. Those fields look distinctive, but they are exactly the fields a good adjudicator may improve. A vague title can become specific. A range can move from the changed line to the condition that makes the defect reachable. Severity can change when the blast radius becomes clearer.

Textual matching forces an unsafe choice. Restore every unmatched candidate and a deliberately rejected finding can reappear. Drop every unmatched candidate and a confirmed but reworded finding can vanish. Use fuzzy matching and two related concerns can collapse into one—or one concern can be published twice.

Runnable assigns an opaque, stable identifier before adjudication begins. Reviewers can improve the public finding without changing which candidate they are deciding. Identity follows the claim; presentation remains editable.

Every incoming finding must end in exactly one state

EVERY INCOMING FINDING MUST BE ACCOUNTED FOR
STABLE ID / F-7C91Tenant scope can be skippedThe wording and line range may change; the claim's identity does not.
DISPOSITION 01ConfirmedPublish one linked finding, with clearer evidence if available.
DISPOSITION 02RejectedOmit it deliberately. No title or line matching can restore it.
PROTOCOL GUARDExactly one dispositionMissing, duplicate, or unknown IDs fail the review.
04 / ADJUDICATION Rewording can improve a finding without changing which incoming claim was decided.

The evidence adjudicator returns two things: the findings that survived, each linked to its incoming identifier when applicable, and the identifiers it explicitly rejected. The control plane accepts the result only when every finding passed in from the earlier review appears exactly once across those two sets.

A duplicate disposition, an unknown identifier, or an omitted candidate makes the review operationally invalid. The system does not guess which result the adjudicator intended, restore the missing input, or reinterpret a malformed response as a clean review.

This is a small protocol rule with an important effect: silence is no longer a decision. Every proposed finding is accounted for before publication.

  • ConfirmedThe candidate survives and may be reworded, reranged, or reprioritized without losing its identity.
  • RejectedThe candidate is deliberately absent from the final review and cannot be restored by fallback matching.
  • InvalidMissing, duplicate, or unknown dispositions fail the operation instead of producing a deceptively clean result.

Adjudication is not majority voting

A candidate does not become correct because several agents repeated it. Reviewers can share the same blind spot, infer the same nonexistent invariant, or anchor on the same visually prominent line. Counting votes would turn correlated confidence into apparent evidence.

The conclave instead separates roles. Independent reviewers propose claims. Adversarial challengers receive the complete candidate pool and try to disprove each claim. At evidence depth, a fresh adjudicator sees the earlier findings, specialist proposals, and challenges, then produces the explicit disposition ledger for every incoming finding.

The goal is not consensus. It is a traceable decision under a constrained procedure: what was claimed, what survived opposition, and what was rejected before a human ever sees the review.

Evidence turns a strong claim into a testable one

For eligible high-risk changes, Runnable raises the same PR review to evidence depth. Roslyn and the TypeScript compiler map changed symbols, public contracts, routes, authorization boundaries, schemas, migrations, owners, and affected consumers. Relevant language specialists inspect that map alongside shared security, concurrency, compatibility, performance, and failure-mode review.

The system can generate up to three focused regression tests and run them in secretless, offline containers against both the pinned merge-base and pull request head. A test that passes on the base and fails on the head is labeled a verified regression. A test that fails on both sides, passes on both sides, or cannot run is not promoted to verified evidence.

Verification does not erase uncertainty. Unsupported languages, incomplete maps, missing repository configuration, and inconclusive tests remain visible as incomplete coverage and require human review.

The protocol is versioned because trust is versioned

Changing a prompt is easy. Changing the contract that determines which findings may reach a pull request is a product change. Runnable pins the evidence protocol version into the review snapshot so an in-flight review cannot begin under one disposition contract and finish under another.

The explicit-adjudication change advances the evidence protocol to v1.1.0. A model qualified for the previous protocol does not automatically remain qualified. The evaluation suite now includes both sides of the decision: a real regression whose identity must survive rewording, and an equivalent refactor whose initial false positive must be explicitly rejected.

Until a model passes the exact protocol evaluation, evidence depth stays unavailable for that model. A stricter schema without a corresponding evaluation would make the interface look safer without establishing that the reviewer can follow it.

One action owns all three depths

Explicit adjudication does not introduce another workflow action. Standard, thorough, and evidence are depths of runnable/pr-review@v1. The minimum-depth input sets the floor, while a deterministic risk policy can recommend—or, when active, apply—a deeper path for the pinned change.

Low-risk changes can stay with one focused reviewer. Medium-risk changes can use a two-reviewer adversarial conclave. Eligible high-risk changes can use four reviewers, compiler-derived mapping, and differential verification. The paths share one consent, one cost ceiling, one advisory GitHub review, and one authenticated report.

name: PR reviewon:  pull_request:    types: [opened, synchronize, reopened, ready_for_review] permissions: {} jobs:  review:    runs-on: ubuntu-latest    steps:      - uses: runnable/pr-review@v1        with:          model: recommended          minimum-depth: standard          max-cost-usd: "5.00"

What explicit adjudication does not prove

A complete disposition ledger is not a proof that every confirmed finding is correct or that every real bug was found. Models can still misunderstand code, miss an affected consumer, or accept a weak challenge. Differential tests cover only hypotheses that can be expressed safely inside the bounded harness.

The protocol solves a narrower but necessary problem: it prevents the orchestration layer from losing track of what the reviewers decided. Combined with pinned revisions, fail-closed validation, visible coverage, and model evaluation, that makes the final review easier to interpret and harder to accidentally overstate.

More agents do not create trust by themselves. Trust begins when every claim keeps its identity, every decision is explicit, and uncertainty remains visible all the way to the pull request.

RUN THE EVIDENCE

Your workflows are already runnable.

Scan one before you move it. The report names what runs, what needs review, and what stays put.

Check a workflow