Runnable

Paid preview

Adaptive PR review

One action chooses standard, thorough, or evidence depth from a fixed risk policy, then publishes one advisory GitHub review and one authenticated report.

What the report proves

Runnable separates deterministic facts, reproduced regressions, and unverified concerns instead of compressing them into one score.

EvidenceMeaning
Change mapCompiler-derived symbols, contracts, routes, authorization, schemas, migrations, package boundaries, consumers, and CODEOWNERS guidance.
Verified regressionA focused generated test passed on the pinned merge-base and failed on the pinned PR head.
Pre-existingThe same focused failure occurred on both revisions and is suppressed as a PR finding.
DisprovedThe test passed on both revisions and the candidate is suppressed.
InconclusiveThe harness was unsupported, failed to build, timed out, or produced non-equivalent failures; the concern stays visibly unverified.

Advisory by default

Runnable publishes one head-SHA-pinned GitHub COMMENT review. It never approves, requests changes, assigns reviewers, modifies source, or merges a pull request.

Enable the paid preview

Adaptive review is available to allowlisted post-trial paid organizations after an owner accepts AI pricing, source processing, and secretless execution terms.

  1. 1

    Accept one billing and execution consent

    An organization owner enables PR review under Plan & budgets and sets one per-review and billing-period AI ceiling.
  2. 2

    Choose a minimum depth

    standard, thorough, or evidence. The risk policy can recommend or raise the depth, depending on rollout mode.
  3. 3

    Create the workflow

    Generate YAML or let Runnable open a setup pull request. Existing workflow files are never overwritten.
.runnable/workflows/pr-review.ymlYAML
name: PR review
on:
  pull_request:
    types: [opened, synchronize, reopened, ready_for_review]

permissions: {}
concurrency:
  group: pr-review-${{ github.event.pull_request.number }}
  cancel-in-progress: true

jobs:
  review:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - id: review
        uses: runnable/pr-review@v1
        with:
          model: recommended
          max-cost-usd: "5.00"
          minimum-depth: evidence

Project discovery is automatic. For monorepos or ambiguous layouts, add optional base-revision configuration:

.runnable/evidence.ymlYAML
version: 1
projects:
  - root: apps/api
    language: dotnet
    project: Api.sln
    test-project: tests/Api.Tests/Api.Tests.csproj
    test-runner: xunit
  - root: apps/web
    language: typescript
    project: package.json
    package-manager: npm
    test-runner: vitest

Configuration is data, not a shell

Paths must stay inside the repository. Build and test overrides are argv arrays. Shell strings, environment injection, network changes, unsupported runners, and unsafe paths are rejected.

Action contract

The public action requires no checkout and receives no workflow-visible GitHub token.

FieldContract
modelA curated, evidence-evaluated Vercel AI Gateway model ID or recommended.
minimum-depthstandard (default), thorough, or evidence. The result stays advisory at every depth.
reviewersDeprecated one-release alias: 1 → standard, 2–3 → thorough, 4 → evidence.
max-cost-usd$5.00 by default. Reserved up front for the worst-case path; includes marked-up confirmed model usage and the $0.25 delivery fee.
OutputsEvidence/report/review IDs and URLs, status, risk, finding/contract/test counts, and actual AI cost.

Mapping, specialist review, and differential verification

All evidence is pinned to the merge-base and PR head, never the synthetic merge commit.

  • Roslyn maps .NET public APIs, callers, ASP.NET routes and authorization, DI lifetimes, serialization/EF surfaces, migrations, and changed behavior.
  • The TypeScript compiler maps changed declarations and behavior, imports/exports, public types, runtime schemas, React/Next boundaries, routes, and consumers.
  • Relevant .NET and TypeScript specialists run alongside shared security, concurrency, failure, performance, and compatibility review; a fresh adjudicator rejects weak candidates.
  • The evidence subphase admits at most six paid model calls and three generated test recipes; the deepest full review path admits at most seventeen model calls. Fifty eligible files and 500 KiB of cumulative source/diff context are the initial hard limits.
  • Unsupported languages, missing evidence configuration, or inconclusive tests publish with incomplete coverage and require a human. Ambiguous infrastructure outcomes fail or enter reconciliation.

Secretless verification

Repository code runs only after inference has proposed a bounded recipe, and only in disposable offline containers.

  • System child jobs strip organization and repository secrets, variables, OIDC, workflow environment, reusable inputs, and workflow-visible repository tokens.
  • A trusted outer runner fetches exact revisions with a short-lived backend credential. Generated tests modify disposable copies only.
  • Dependency restore receives manifests and lockfiles—not source—uses allowlisted public registries with install hooks disabled, and records a resolution digest.
  • Build and test run with no network or Docker socket, as non-root, with a read-only container root, dropped capabilities, process/memory/CPU limits, and bounded logs.

Billing, reuse, and gates

AI and compute remain separately visible, while admission evaluates both against the organization hard cap under one billing lock.

  • Confirmed Gateway cost is charged with the accepted pricing margin. Failed or cancelled reports keep confirmed model charges but waive the delivery fee.
  • A delivered report—including a clean one—adds $0.25 exactly once. Evidence runner compute is metered normally and shown separately.
  • Identical pinned revisions, model, base configuration, CODEOWNERS, analyzer, harness, and reviewer versions reuse completed work without another charge.
  • Ambiguous inference, sandbox, publication, or Stripe outcomes hold their reservation for reconciliation instead of repeating paid work.
NextBilling and limitsUnderstand normalized compute, AI ceilings, reservations, invoice previews, alerts, and the combined hard cap.