Overview
I avoided end-to-end tests for years because every setup I tried was fragile. Selenium needed a driver per browser. Cypress was pleasant until I needed multiple tabs. TestCafe worked but had a small ecosystem. Playwright is the first one where I've written tests that still pass six months later without me touching them.
This is the setup I've landed on, and the things that took me too long to figure out.
Why Playwright over the alternatives
| Playwright | Cypress | Selenium | |
|---|---|---|---|
| Languages | JS, TS, Python, Java, .NET | JS, TS | Almost everything |
| Browser support | Chromium, Firefox, WebKit | Chromium-based, Firefox, WebKit (experimental) | All major |
| Multiple tabs | Yes | No | Yes |
| Auto-waiting | Yes, built in | Yes, built in | Manual |
| Speed | Fast | Fast | Slower |
| Parallel by default | Yes | With paid plan or parallel config | With Selenium Grid |
The auto-waiting is the feature that makes the difference. Playwright waits for elements to be actionable before clicking — visible, stable, enabled, and not obscured by anything. Cypress does this too. Selenium doesn't, which is why Selenium tests are full of sleep() calls that make the suite slow and still flaky.
Setup
npm init playwright@latest
The init creates a project structure, installs browsers, and adds a GitHub Actions workflow if you want one. The whole thing takes a minute.
my-project/
├── tests/
│ └── example.spec.ts
├── playwright.config.ts
└── package.json
A first test
import { test, expect } from "@playwright/test";
test("user can log in", async ({ page }) => {
await page.goto("http://localhost:3000/login");
await page.getByLabel("Email").fill("user@example.com");
await page.getByLabel("Password").fill("password123");
await page.getByRole("button", { name: "Sign in" }).click();
await expect(page).toHaveURL("/dashboard");
await expect(page.getByRole("heading", { name: "Welcome back" })).toBeVisible();
});
Two things worth noticing. First, the locators are semantic — getByRole, getByLabel. They match how a user would identify the element, which means the test breaks when the UI actually changes, not when you rename a CSS class.
Second, there's no waiting. Playwright's assertions retry until they pass or time out. You don't write waitFor before expect.
Locators: the part that determines if tests survive
Playwright offers several locator strategies, and choosing the right one is what makes tests stable.
| Locator | Use for | Stability |
|---|---|---|
getByRole | Buttons, links, headings, inputs | Very stable |
getByLabel | Form fields | Very stable |
getByText | Text content | Stable if copy doesn't change |
getByPlaceholder | Inputs with placeholders | Medium |
getByTestId | Anything with data-testid | Very stable, but requires adding attributes |
locator("css") | CSS selectors | Fragile |
My rule: getByRole and getByLabel for everything possible. Add data-testid when the semantic locator isn't unique or when the element has no accessible role. Never use CSS selectors with class names — they change with every styling refactor.
// Fragile — breaks on any CSS refactor
await page.locator(".btn-primary.submit").click();
// Stable — semantic
await page.getByRole("button", { name: "Submit" }).click();
// Stable — explicit test ID when needed
await page.getByTestId("submit-order").click();
Configuration that matters
import { defineConfig, devices } from "@playwright/test";
export default defineConfig({
testDir: "./tests",
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 2 : undefined,
reporter: process.env.CI ? "github" : "html",
use: {
baseURL: "http://localhost:3000",
trace: "on-first-retry",
screenshot: "only-on-failure",
video: "retain-on-failure",
},
projects: [
{
name: "chromium",
use: { ...devices["Desktop Chrome"] },
},
{
name: "firefox",
use: { ...devices["Desktop Firefox"] },
},
{
name: "mobile",
use: { ...devices["iPhone 14"] },
},
],
webServer: {
command: "npm run dev",
url: "http://localhost:3000",
reuseExistingServer: !process.env.CI,
},
});
Points worth calling out:
trace: "on-first-retry"is the single most useful setting. When a test fails and then passes on retry, Playwright saves a trace file with screenshots, DOM snapshots, network log, and console output. Open it withnpx playwright show-traceand you can step through exactly what happened.retries: 2in CI only — locally, failures should fail immediately so you fix them. In CI, retries filter out genuinely flaky tests. If a test needs more than 2 retries, it's a real problem, not a flake.webServerconfig starts your app automatically before tests. No more "did I remember to start the dev server?"
Test fixtures: setup that doesn't repeat
import { test as base } from "@playwright/test";
type Fixtures = {
authenticatedPage: Page;
};
export const test = base.extend<Fixtures>({
authenticatedPage: async ({ page, context }, use) => {
await page.goto("/login");
await page.getByLabel("Email").fill("user@example.com");
await page.getByLabel("Password").fill("password123");
await page.getByRole("button", { name: "Sign in" }).click();
await page.waitForURL("/dashboard");
await use(page);
// cleanup runs after the test
},
});
import { test, expect } from "./fixtures";
test("shows user profile", async ({ authenticatedPage }) => {
await authenticatedPage.goto("/profile");
await expect(authenticatedPage.getByRole("heading")).toContainText("Profile");
});
The fixture handles login once, and every test that needs it declares authenticatedPage instead of page. Better than a beforeEach hook because you opt in explicitly and the dependency is visible in the test signature.
Handling network
// Mock an API response
await page.route("**/api/users", (route) => {
route.fulfill({
status: 200,
body: JSON.stringify([{ id: 1, name: "Test User" }]),
});
});
// Fail a request to test error handling
await page.route("**/api/payment", (route) => route.abort("failed"));
// Wait for a real response
const responsePromise = page.waitForResponse("**/api/orders");
await page.getByRole("button", { name: "Submit" }).click();
const response = await responsePromise;
expect(response.status()).toBe(201);
Mocking at the network layer keeps tests fast and independent of backend state. But use it sparingly — the point of E2E tests is testing the whole system, and if you mock everything, you've written an integration test in the wrong place.
My rule: mock third-party APIs (Stripe, SendGrid, analytics), never mock your own backend except when testing specific error states.
Parallelism and isolation
Playwright runs test files in parallel by default, with one browser context per file. Tests within a file run sequentially unless you explicitly parallelize them.
test.describe.configure({ mode: "parallel" });
test.describe("checkout flow", () => {
test("adds item to cart", async ({ page }) => { /* ... */ });
test("applies coupon", async ({ page }) => { /* ... */ });
});
The catch: parallel tests can't share state. If test A creates a user and test B expects to find it, running them in parallel breaks. Either use separate users per test (with fixtures), or run them sequentially.
What made my tests stop being flaky
Three things, in order of impact:
- Stopped using
waitForTimeout. Everyawait page.waitForTimeout(2000)I removed made a test faster and less likely to fail on a slow CI box. The replacements areexpect(...).toBeVisible()andpage.waitForURL(), which retry until success or timeout. - Moved to semantic locators. Class-name selectors broke whenever the CSS changed. Role and label locators only break when functionality changes.
- Enabled tracing. When a test failed on CI, I could see the DOM state, network calls, and console output at the moment of failure. Before traces, I was guessing based on screenshots.
Running in CI
name: Playwright
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 20
- run: npm ci
- run: npx playwright install --with-deps
- run: npx playwright test
- uses: actions/upload-artifact@v4
if: failure()
with:
name: playwright-report
path: playwright-report/
retention-days: 7
--with-deps installs the OS-level dependencies that browsers need. Without it, the tests fail on a clean CI runner with missing shared libraries, and the error is unhelpful.
The artifact upload on failure is what makes CI debugging tolerable. The HTML report includes traces, screenshots, and videos, and it's download-and-open.
When not to use Playwright
- Testing business logic. Unit tests are faster and give better failure messages. E2E tests should cover user journeys, not edge cases in a date parser.
- Visual regression on every commit. Playwright has screenshot comparison, but it's slow and produces false positives from font rendering differences across OSes. Use it sparingly or on a schedule.
- Load testing. Playwright launches real browsers. Use k6 or Locust instead.
The most common mistake is over-testing. A suite of 500 E2E tests takes 20 minutes and fails for environmental reasons half the time. A suite of 30 covering the critical paths takes 90 seconds and passes reliably. The second one is what actually prevents regressions.
