Running Playwright in CI

A Playwright suite delivers value when it runs automatically on every pull request and blocks regressions. CI environments add their own challenges, though: browsers and system libraries must be installed, the app must be started and reachable, large suites need sharding across machines, and failures must be debuggable without reproducing them locally, which is what traces, screenshots, videos, and HTML reports are for.

A well-tuned pipeline runs quickly, fails for real reasons, and gives developers everything they need to fix a failure in minutes.

TL;DR

Quick Example

A sharded GitHub Actions workflow with merged reports:

Core Concepts

Browsers and System Dependencies

Playwright downloads specific browser builds matching its version. In CI:

Starting the Application

Retries, Traces, and Artifacts

Sharding

--shard=x/y splits tests across y machines. Each shard writes a blob report, and a final job merges them into one HTML report with merge-reports. Combine it with a CI matrix (see matrix builds). Balance shards by keeping test files similarly sized, since sharding splits by test and file, and with fullyParallel: true it splits more evenly.

Workers

CI machines often have few CPUs, and too many workers cause resource contention and timeouts. Start with 1–2 workers per shard on standard runners and tune from there. Scale horizontally with shards rather than vertically with workers.

Flaky Test Management

Visual Regression Testing

await expect(page).toHaveScreenshot() compares against baselines. Rendering differs across OSes, fonts, and GPUs, so generate and compare snapshots in the same environment, typically the Docker image, both locally and in CI. Mask dynamic regions (timestamps, ads), disable animations, and set thresholds thoughtfully.

Speeding Up Suites

Best Practices

Make CI Match Local

Pin Playwright and browser versions, use the same Docker image for visual tests, and provide a script (npm run test:e2e) that developers run locally the same way CI does.

Keep E2E Focused on Critical Journeys

End-to-end tests are the most expensive tests. Cover key user journeys, and push detailed logic coverage down to unit and integration tests.

Gate Merges on E2E Results

Make the E2E job a required status check. A suite that can be ignored will be ignored, especially when it's flaky.

Protect Secrets in Artifacts

Traces and videos can contain credentials and personal data from test accounts. Mask inputs, use test-only accounts, and set short artifact retention.

Common Mistakes

Committing test.only

A stray test.only makes CI run a single test and report green. Set forbidOnly: !!process.env.CI.

Too Many Workers on Small Runners

Running 8 workers on a 2-vCPU runner makes every test slower and timeouts frequent, which looks like flakiness. Reduce workers and add shards.

Comparing Screenshots Across Operating Systems

Baselines generated on a Mac fail on Linux CI because of font rendering differences. Generate baselines in the CI environment (or the Docker image), and commit those.

FAQ

How do I install Playwright browsers in CI?

Run npx playwright install --with-deps (optionally naming specific browsers) after installing npm dependencies, or use the official Playwright Docker image, which includes browsers and system libraries. Pin the image version to match your @playwright/test version.

How do I debug a test that only fails in CI?

Enable traces (trace: 'on-first-retry' or 'retain-on-failure'), upload them as artifacts, and open them in Trace Viewer to inspect each action, DOM snapshot, network request, and console message. Screenshots and videos help too. Reproducing locally in the same Docker image often reveals environment differences.

How does sharding work in Playwright?

--shard=1/4 runs the first quarter of tests, and so on. Run each shard in a parallel CI job with the blob reporter, then merge the blob reports into one HTML report with npx playwright merge-reports.

Should I use retries in CI?

Yes, a small number (1–2) to absorb rare infrastructure hiccups, while tracking tests that pass only on retry as flaky and fixing them. Retries shouldn't hide persistent flakiness.

Related Topics

References