Overview

I avoided end-to-end tests for years because every setup I tried was fragile. Selenium needed a driver per browser. Cypress was pleasant until I needed multiple tabs. TestCafe worked but had a small ecosystem. Playwright is the first one where I've written tests that still pass six months later without me touching them.

This is the setup I've landed on, and the things that took me too long to figure out.

Why Playwright over the alternatives

PlaywrightCypressSelenium
LanguagesJS, TS, Python, Java, .NETJS, TSAlmost everything
Browser supportChromium, Firefox, WebKitChromium-based, Firefox, WebKit (experimental)All major
Multiple tabsYesNoYes
Auto-waitingYes, built inYes, built inManual
SpeedFastFastSlower
Parallel by defaultYesWith paid plan or parallel configWith Selenium Grid

The auto-waiting is the feature that makes the difference. Playwright waits for elements to be actionable before clicking — visible, stable, enabled, and not obscured by anything. Cypress does this too. Selenium doesn't, which is why Selenium tests are full of sleep() calls that make the suite slow and still flaky.

Setup

npm init playwright@latest

The init creates a project structure, installs browsers, and adds a GitHub Actions workflow if you want one. The whole thing takes a minute.

my-project/
├── tests/
│   └── example.spec.ts
├── playwright.config.ts
└── package.json

A first test

import { test, expect } from "@playwright/test";

test("user can log in", async ({ page }) => {
  await page.goto("http://localhost:3000/login");

  await page.getByLabel("Email").fill("user@example.com");
  await page.getByLabel("Password").fill("password123");
  await page.getByRole("button", { name: "Sign in" }).click();

  await expect(page).toHaveURL("/dashboard");
  await expect(page.getByRole("heading", { name: "Welcome back" })).toBeVisible();
});

Two things worth noticing. First, the locators are semantic — getByRole, getByLabel. They match how a user would identify the element, which means the test breaks when the UI actually changes, not when you rename a CSS class.

Second, there's no waiting. Playwright's assertions retry until they pass or time out. You don't write waitFor before expect.

Locators: the part that determines if tests survive

Playwright offers several locator strategies, and choosing the right one is what makes tests stable.

LocatorUse forStability
getByRoleButtons, links, headings, inputsVery stable
getByLabelForm fieldsVery stable
getByTextText contentStable if copy doesn't change
getByPlaceholderInputs with placeholdersMedium
getByTestIdAnything with data-testidVery stable, but requires adding attributes
locator("css")CSS selectorsFragile

My rule: getByRole and getByLabel for everything possible. Add data-testid when the semantic locator isn't unique or when the element has no accessible role. Never use CSS selectors with class names — they change with every styling refactor.

// Fragile — breaks on any CSS refactor
await page.locator(".btn-primary.submit").click();

// Stable — semantic
await page.getByRole("button", { name: "Submit" }).click();

// Stable — explicit test ID when needed
await page.getByTestId("submit-order").click();

Configuration that matters

import { defineConfig, devices } from "@playwright/test";

export default defineConfig({
  testDir: "./tests",

  fullyParallel: true,
  forbidOnly: !!process.env.CI,
  retries: process.env.CI ? 2 : 0,
  workers: process.env.CI ? 2 : undefined,

  reporter: process.env.CI ? "github" : "html",

  use: {
    baseURL: "http://localhost:3000",
    trace: "on-first-retry",
    screenshot: "only-on-failure",
    video: "retain-on-failure",
  },

  projects: [
    {
      name: "chromium",
      use: { ...devices["Desktop Chrome"] },
    },
    {
      name: "firefox",
      use: { ...devices["Desktop Firefox"] },
    },
    {
      name: "mobile",
      use: { ...devices["iPhone 14"] },
    },
  ],

  webServer: {
    command: "npm run dev",
    url: "http://localhost:3000",
    reuseExistingServer: !process.env.CI,
  },
});

Points worth calling out:

  • trace: "on-first-retry" is the single most useful setting. When a test fails and then passes on retry, Playwright saves a trace file with screenshots, DOM snapshots, network log, and console output. Open it with npx playwright show-trace and you can step through exactly what happened.
  • retries: 2 in CI only — locally, failures should fail immediately so you fix them. In CI, retries filter out genuinely flaky tests. If a test needs more than 2 retries, it's a real problem, not a flake.
  • webServer config starts your app automatically before tests. No more "did I remember to start the dev server?"

Test fixtures: setup that doesn't repeat

import { test as base } from "@playwright/test";

type Fixtures = {
  authenticatedPage: Page;
};

export const test = base.extend<Fixtures>({
  authenticatedPage: async ({ page, context }, use) => {
    await page.goto("/login");
    await page.getByLabel("Email").fill("user@example.com");
    await page.getByLabel("Password").fill("password123");
    await page.getByRole("button", { name: "Sign in" }).click();
    await page.waitForURL("/dashboard");

    await use(page);

    // cleanup runs after the test
  },
});
import { test, expect } from "./fixtures";

test("shows user profile", async ({ authenticatedPage }) => {
  await authenticatedPage.goto("/profile");
  await expect(authenticatedPage.getByRole("heading")).toContainText("Profile");
});

The fixture handles login once, and every test that needs it declares authenticatedPage instead of page. Better than a beforeEach hook because you opt in explicitly and the dependency is visible in the test signature.

Handling network

// Mock an API response
await page.route("**/api/users", (route) => {
  route.fulfill({
    status: 200,
    body: JSON.stringify([{ id: 1, name: "Test User" }]),
  });
});

// Fail a request to test error handling
await page.route("**/api/payment", (route) => route.abort("failed"));

// Wait for a real response
const responsePromise = page.waitForResponse("**/api/orders");
await page.getByRole("button", { name: "Submit" }).click();
const response = await responsePromise;
expect(response.status()).toBe(201);

Mocking at the network layer keeps tests fast and independent of backend state. But use it sparingly — the point of E2E tests is testing the whole system, and if you mock everything, you've written an integration test in the wrong place.

My rule: mock third-party APIs (Stripe, SendGrid, analytics), never mock your own backend except when testing specific error states.

Parallelism and isolation

Playwright runs test files in parallel by default, with one browser context per file. Tests within a file run sequentially unless you explicitly parallelize them.

test.describe.configure({ mode: "parallel" });

test.describe("checkout flow", () => {
  test("adds item to cart", async ({ page }) => { /* ... */ });
  test("applies coupon", async ({ page }) => { /* ... */ });
});

The catch: parallel tests can't share state. If test A creates a user and test B expects to find it, running them in parallel breaks. Either use separate users per test (with fixtures), or run them sequentially.

What made my tests stop being flaky

Three things, in order of impact:

  1. Stopped using waitForTimeout. Every await page.waitForTimeout(2000) I removed made a test faster and less likely to fail on a slow CI box. The replacements are expect(...).toBeVisible() and page.waitForURL(), which retry until success or timeout.
  2. Moved to semantic locators. Class-name selectors broke whenever the CSS changed. Role and label locators only break when functionality changes.
  3. Enabled tracing. When a test failed on CI, I could see the DOM state, network calls, and console output at the moment of failure. Before traces, I was guessing based on screenshots.

Running in CI

name: Playwright

on: [push, pull_request]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
      - run: npm ci
      - run: npx playwright install --with-deps
      - run: npx playwright test
      - uses: actions/upload-artifact@v4
        if: failure()
        with:
          name: playwright-report
          path: playwright-report/
          retention-days: 7

--with-deps installs the OS-level dependencies that browsers need. Without it, the tests fail on a clean CI runner with missing shared libraries, and the error is unhelpful.

The artifact upload on failure is what makes CI debugging tolerable. The HTML report includes traces, screenshots, and videos, and it's download-and-open.

When not to use Playwright

  • Testing business logic. Unit tests are faster and give better failure messages. E2E tests should cover user journeys, not edge cases in a date parser.
  • Visual regression on every commit. Playwright has screenshot comparison, but it's slow and produces false positives from font rendering differences across OSes. Use it sparingly or on a schedule.
  • Load testing. Playwright launches real browsers. Use k6 or Locust instead.

The most common mistake is over-testing. A suite of 500 E2E tests takes 20 minutes and fails for environmental reasons half the time. A suite of 30 covering the critical paths takes 90 seconds and passes reliably. The second one is what actually prevents regressions.