A green build tells you that your checks passed. It does not tell you whether a user can sign in, follow a navigation link, or reach the screen your change was supposed to deliver. Those questions start with the running application.
Runnable Browser QA brings that check into the workflow. The runnable/qa@v1 action opens a deployed preview or staging URL in real Chrome and exercises cases you write in plain language. Each case describes what to do and what should be observable afterward.
Browser QA is in a gated preview. An organization owner must enable it and accept its separate AI pricing consent before a QA job can start. Start with a few important journeys on a staging environment, alongside your existing automated tests.
Deploy first. Check the experience next.
Make the QA job depend on your deployment job and pass its preview URL into the action. Keep the cases in your repository so the expected behavior can change in the same pull request as the application.
This example assumes a deploy job exports a public HTTPS URL. The QA job checks out the repository, loads the suite, and fails if a case fails or cannot reach a supported conclusion.
qa: needs: deploy runs-on: ubuntu-24.04 steps: - uses: actions/checkout@v4 - uses: runnable/qa@v1 with: url: ${{ needs.deploy.outputs.url }} cases-file: .runnable/qa.yml model: recommended max-cost-usd: "5.00" gate: failWrite the journey and the observable result
A suite contains up to 20 cases, each with steps and expectations. Be concrete: name the page, the interaction, and the visible result. A pricing link opening a page with the current plans is a stronger expectation than a general instruction to check that the site works.
Each case gets a fresh browser session. The agent can navigate, inspect the page, click, fill fields, press keys, wait, and observe text or URLs. It has no shell or browser JavaScript evaluation tool.
version: 1cases: - id: pricing-navigation name: Visitors can find pricing steps: - Open the home page. - Follow the Pricing link. expect: - The pricing page shows the current plans.A result needs evidence, including when it is uncertain
Logs show per-case progress, and the step summary lists each result. A tenant-scoped artifact contains a JSON report with a redacted excerpt of the last browser observation for each case, plus safe final screenshots. QA reports and artifacts expire after one day.
An ordinary failed case does not stop later cases. A timeout, exhausted budget, ambiguous model billing, or insufficient evidence produces an inconclusive result. The workflow can distinguish an observed failure from a check that could not finish.
- PassedThe browser check reports that the expected behavior was observed.
- FailedThe case reports that the application did not meet an expectation.
- InconclusiveThe check could not establish a supported verdict within its limits.
Choose how QA affects the workflow
The default gate: fail fails the step when any case is failed or inconclusive. Choose gate: advisory to collect results without failing the job for those case outcomes. Setup and artifact-upload errors still fail the action.
Browser QA complements deterministic tests and human review. It checks the journeys you describe on the deployment you provide; it does not establish that every route works or that a release has no defects.
Keep credentials, navigation, and spend bounded
The target must be a public HTTPS URL without embedded credentials. Additional login and asset hosts must be explicitly allowed. Browser navigation and subresources outside the allowed domains are blocked.
For authenticated journeys, pass dedicated, low-privilege test credentials as QA_* step environment variables and reference them with placeholders such as {{QA_EMAIL}}. Values are substituted only for field fills and redacted from observations, logs, and JSON reports. Screenshots are suppressed for suites using secrets. Use staging accounts and data: the browser may submit forms and make the changes your cases request.
Set max-cost-usd for each run. Owner QA limits, subscription-period AI limits, and the combined compute and AI hard cap also apply. Each case is limited to 40 browser actions and five minutes; the whole action is limited to 100 model calls and 20 minutes.
AI usage is billed at confirmed AI Gateway cost divided by 0.70, plus normal runner compute, with no fixed QA delivery fee. An interrupted or ambiguous model call is recorded for reconciliation and is not blindly retried.


