Before you startBefore You Start#
This book is for people who can write decent code, have probably written Selenium tests, and now have to test a large, branching, multi-locale web app with Playwright and TypeScript. It gets you from "I can make a test pass" to "I can design a suite that a team can live with for years".
It is short on purpose: fourteen chapters of about eight minutes each, plus the code, so about two hours in all. Every chapter teaches one idea properly, and the epilogue covers the rest of Playwright in one-line entries that link to the official documentation.
Kestrel and Tomas
Kestrel Insure sells car insurance online in four markets: the UK (en-GB), the US (en-US), Germany (de-DE) and Japan (ja-JP). Its quote journey is about forty screens long:
- Vehicle: registration, modifications, mileage.
- Drivers: names, dates of birth, licences.
- History: claims, convictions.
- Cover: cover level, add-ons.
- Payment: frequency, method, declarations.
Half the screens appear only if an earlier answer calls for them. Every market adds, removes or reshapes a few. The Kestrel app is fictional, but the shape is very real: most enterprise web apps are a long form with opinions.
Tomas has written Selenium tests for nine years and has just joined Kestrel's QA team. The old suite has 600 tests, a third of them skipped, and it takes 70 minutes on a good day. Tomas has been asked to replace it. Each chapter opens with something going wrong for Tomas, and closes with what Tomas learned from fixing it.
What you will build
By the end, Tomas's suite has five layers. Each chapter adds or sharpens one of them:
Here is what each layer does: - Specs say what is being tested. - Scenarios hold the variation as data: the answers a customer gives, and what the app should do with them. - Journeys know how to get through the app, whichever way it routes. - Page and component objects know how to find things. - Fixtures hand each test exactly what it asks for. - Projects multiply the suite across locales and browsers without a line of extra test code.
How to read it
Read Part I in order: it is the vocabulary for the rest. After that, jump to what hurts. Every chapter has the same sections, so you always know where to look:
| Section | What it gives you |
|---|---|
| The problem · The idea | The story and the Playwright way, with code |
| How it works | The mechanism underneath, usually as a diagram |
| Selenium habit → Playwright way | What to unlearn |
| The traps | Anti-patterns: the tell-tale sign, why it hurts, what to do instead, and the lint rule that catches it |
| Spot it | A snippet with one planted flaw. The answer is folded |
| At the whiteboard | Three questions for code review, and one point that is still argued |
| What Tomas learned · …but | Three takeaways, and the crack the next chapter opens |
All 43 traps are collected in the epilogue's Anti-pattern Catalog, next to the Concept Atlas, a baseline configuration, and a glossary for Selenium users.
What you need
Node 22 or later and Playwright 1.63. Every TypeScript snippet in this book is compiled against Playwright 1.63's types in strict mode before publication. Snippets start with a // file: comment when they belong to a particular file in the suite. The few that are deliberately wrong (the anti-patterns) are marked as such.
Chapter 1 · Part IThe Sleep That Passed on My Machine#
The problem
Tomas's first job is to port the old suite's most important test: get a quote and check the premium. The Selenium version had an explicit wait wrapped in a helper, and a Thread.sleep(2000) "for the animation". Tomas translates it line by line:
await page.goto('/quote/resume/Q-1042');
await page.click('#get-quote');
await page.waitForTimeout(2000); // the premium is calculated server-side
const premium = await page.locator('#premium').textContent();
expect(premium).toBe('£412.50');
It passes on Tomas's laptop ten times in a row. In CI it fails one run in six. The pricing service takes 2.3 seconds when the build agents are busy, and when it is fast the test still spends two seconds doing nothing. Multiply that by 600 tests and you get the old suite's 70 minutes.
The idea
Playwright waits for you, in two separate ways.
Actions auto-wait. Before click(), fill() or check() does anything, Playwright finds the element and waits until it is actionable: attached, visible, stable (not animating), able to receive events (not covered by something else) and enabled. Only then does it act. You never wait before an action.
Assertions retry. expect(locator).toHaveText(…) is a web-first assertion. It re-reads the element and re-checks the condition until it passes or the assertion timeout (five seconds by default) runs out. You never wait before an assertion either.
// file: tests/specs/summary.spec.ts
import { test, expect } from '@playwright/test';
test('shows the annual premium', async ({ page }) => {
await page.goto('/quote/resume/Q-1042');
await page.getByRole('button', { name: 'Get my quote' }).click();
await expect(page.getByTestId('premium')).toHaveText('£412.50');
});
That is the whole test. No sleeps, no waits, no helper. If pricing takes 300 ms, the test takes 300 ms. If it takes 4 seconds, the test takes 4 seconds. If it never answers, the assertion fails after its timeout and the error names the locator, the expected text and what was actually on the page.
Remember it as: Actions wait until the element is ready; assertions retry until the page agrees. Neither needs a clock.
There is a second difference hiding in async ({ page }). Each test receives a brand-new page in a brand-new browser context: its own cookies, storage and cache, created in milliseconds inside an already-running browser. No test can leak a logged-in session or a half-filled basket into the next one. Selenium suites often reuse one browser for speed and pay for it with order-dependent failures. Playwright gives you isolation and speed together.
How it works
Your test code runs in Node. It talks to the browsers over one long-lived connection. For Chromium that is the Chrome DevTools Protocol, and for Firefox and WebKit it is Playwright's own patched protocol. It is not the one-HTTP-request-per-command WebDriver protocol. Playwright can therefore ask the browser to do the waiting itself, close to the page, instead of polling across the network.
In theatre terms: the actor does not stand in the wings counting to two. The actor waits for the cue, and the cue is the element being ready. And every performance gets a freshly reset stage, so nothing left over from the last show trips anyone up.
Three timeouts matter:
| Timeout | Default | Covers |
|---|---|---|
| Test | 30 s | The whole test, including fixtures |
Assertion (expect) |
5 s | Each web-first assertion |
| Action and navigation | none | Bounded by the test timeout unless you set actionTimeout |
When something is genuinely slow, raise the timeout for that one assertion (toHaveText(x, { timeout: 15_000 })), not for everything.
What Tomas learned
- Every sleep and explicit wait from the old suite can go: actions wait for actionability, and web-first assertions retry.
awaitgoes outsideexpect(), never inside it.- Each test gets its own browser context, so isolation costs nothing.
…but
The next morning the test fails with a clear message: no element matches getByTestId('premium'). A front-end refactor has renamed the attribute, and the old suite's 600 CSS selectors, #get-quote, .btn-primary > span and //div[3]/button, are next. Waiting correctly is worth nothing if you are waiting for the wrong thing.
Chapter 2 · Part IThe Button That Moved#
The problem
Kestrel's front-end team ships a redesign. Nothing a customer would notice changes: the same questions, the same buttons, the same words. Underneath, every div gets a new utility class and the form is restructured for a new component library.
The old suite loses 212 tests overnight. Its selectors described the markup, .quote-form > div:nth-child(3) .btn-primary and //div[@class='addon-card'][2]//input, and the markup is exactly what changed. Tomas's new tests from chapter 1 used #get-quote and data-testid="premium", and one of those broke too.
The idea
Find elements the way a customer would describe them. Nobody says "the third div with class btn-primary". They say "the Continue button" or "the Postcode box". Playwright's locators let you say exactly that:
await page.getByRole('button', { name: 'Continue' }).click();
await page.getByLabel('Postcode').fill('BS1 4DJ');
await page.getByRole('radio', { name: 'Comprehensive' }).check();
await expect(page.getByRole('heading', { name: 'Your quote' })).toBeVisible();
getByRole matches the element's ARIA role (button, link, checkbox, radio, heading, dialog, row…) and its accessible name, which is what a screen reader announces: the button text, the label that belongs to a field, an aria-label. That is the most stable description of an element there is, because changing it changes what users experience.
The built-in locators, in the order to reach for them:
| Locator | Use it for |
|---|---|
getByRole(role, { name }) |
Almost everything interactive, plus headings, dialogs, tables and rows |
getByLabel(text) |
Form fields with a visible label |
getByPlaceholder(text) |
Fields whose only hint is a placeholder |
getByText(text) |
Non-interactive content: messages, paragraphs |
getByAltText / getByTitle |
Images; elements with a title attribute |
getByTestId(id) |
When no user-facing description exists or is stable |
Remember it as: Locate by what the user perceives, role and name first, and fall back to test ids only when there is nothing to perceive.
Real pages repeat themselves. Kestrel's add-ons screen shows a card per add-on, and each card has a "Choose" button. Narrow down by chaining and filtering:
const addOn = page.getByRole('listitem').filter({ hasText: 'Breakdown cover' });
await addOn.getByRole('button', { name: 'Choose' }).click();
// cards that contain a price but no "Unavailable" badge
const available = page.getByRole('listitem')
.filter({ has: page.getByTestId('price') })
.filter({ hasNot: page.getByText('Unavailable') });
await expect(available).toHaveCount(3);
// either the new or the old sign-in button, during a migration
const signIn = page.getByRole('button', { name: 'Sign in' })
.or(page.getByRole('link', { name: 'Sign in' }));
await signIn.click();
filter() takes hasText, hasNotText, has, hasNot and visible. .and() requires both locators to match the same element, and .or() accepts either. .describe('breakdown add-on card') gives a locator a readable name in traces and reports.
When there really is nothing to perceive, such as a premium figure in a <span> with no label, ask for a test id and use getByTestId('premium'). If your app already uses another attribute, set testIdAttribute in the config (chapter 3). A test id is a contract with the developers: it doesn't change when the styling does.
You do not have to write any of this by hand. npx playwright codegen <url> records locators as you click, and the "pick locator" button in UI mode and in the VS Code extension shows the best locator for any element you hover.
How it works
A locator is not an element. It is a recipe for finding one, and it runs again every time you use it. Creating it does nothing. click() resolves it at that moment, waits for actionability, and acts. That is why stale element errors don't exist in Playwright: there is no stored element to go stale.
Locators are also strict. If an action's locator matches two elements, Playwright refuses to guess and fails with a list of what matched. The failure is a feature. It tells you the description is ambiguous before the ambiguity picks the wrong button in production data.
The theatre version: you cast by role, never by costume. "Whoever is playing the doctor" still works when the costume department changes the coat. "The person in the white coat" does not.
A welcome side effect: getByRole reads the accessibility tree. When a test can't find the "Continue" button by role and name, it is often because the button has no accessible name, and that is a real accessibility bug.
What Tomas learned
- A locator is a lazy, strict recipe. It re-resolves on every use, and it complains when it matches more than one element.
- Role and accessible name come first, then label and text, and test ids only when there is nothing to perceive.
- Ambiguity is fixed by scoping and filtering, never by picking the nth match.
…but
Tomas's twelve tests now find the right elements and survive the redesign. They also only run on Tomas's laptop. The app has to be started by hand first, the URLs are hardcoded, and nobody has tried a browser other than Chrome. Good tests that only one person can run are not yet a suite.
Chapter 3 · Part IDay One#
The problem
Tomas now has twelve good tests, and a teammate tries to run them. It goes badly:
- The app has to be started by hand in another terminal first.
- Half the tests hardcode http://localhost:3000 and the other half https://staging.kestrel.test.
- They only ever ran in Chrome.
The first CI run reports 0 failed, 1 passed. Someone left a test.only in a file, and the runner quietly ran that one test instead of the suite.
The old Selenium suite solved these problems with a homegrown framework: a DriverFactory, a properties file, a base class that started the app. Tomas is about to start writing one. Tomas doesn't need to.
The idea
Playwright Test is the runner, the assertion library, the parallel executor, the reporter and the browser manager, all in one package. The framework you would write is one file: playwright.config.ts.
npm init playwright@latest # TypeScript, a tests/ folder, a GitHub Actions workflow
npx playwright install --with-deps
// file: playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
forbidOnly: !!process.env.CI, // a committed test.only fails the CI run
retries: process.env.CI ? 2 : 0,
reporter: [['html', { open: 'never' }]],
use: {
baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
trace: 'on-first-retry',
testIdAttribute: 'data-testid',
},
webServer: { // start the app before the tests, stop it after
command: 'npm run start',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
],
});
With baseURL set, tests say page.goto('/quote'), and one environment variable points the whole suite at staging. webServer starts the app and waits until it answers. Projects run every test once per entry. Here that means three browsers, and in chapter 9 it means four locales. use sets defaults for everything a test receives: viewport, locale, trace settings, storage state.
Remember it as: The config is your framework. Before writing a helper class, check whether it is a config option.
The day-to-day commands:
| Command | Does |
|---|---|
npx playwright test |
Run everything, headless, in parallel |
npx playwright test --ui |
UI mode: watch, filter, step through, time-travel each action |
npx playwright test -g "premium" --project=webkit |
One test title, one project |
npx playwright test --headed --debug |
Headed, paused, with the inspector |
npx playwright show-report |
The HTML report, including traces |
Install the VS Code extension too. It runs and debugs single tests from the gutter, picks locators, and records new tests into the file you have open.
Finally, lint. Several traps in this book are mechanical enough for a machine to catch: missing await, sleeps, forced clicks, committed test.only. eslint-plugin-playwright plus one TypeScript rule catch most of them:
// eslint.config.mjs
import { defineConfig } from 'eslint/config';
import tseslint from 'typescript-eslint';
import playwright from 'eslint-plugin-playwright';
export default defineConfig({
files: ['tests/**/*.ts'],
extends: [playwright.configs['flat/recommended']],
plugins: { '@typescript-eslint': tseslint.plugin },
languageOptions: { parser: tseslint.parser, parserOptions: { projectService: true } },
rules: {
'@typescript-eslint/no-floating-promises': 'error',
'playwright/no-raw-locators': 'warn',
'playwright/no-nth-methods': 'warn',
},
});
no-floating-promises is the most valuable line in the file. The epilogue has the full baseline, with every rule mentioned in the traps.
How it works
When you run npx playwright test, the runner:
- reads the config;
- starts the webServer if there is one;
- collects every test in every file, once per project;
- hands the tests to worker processes.
Each worker is a separate Node process with its own browser. Each test gets a fresh context inside that browser.
In theatre terms, the config is the running order pinned by the stage door: which venues, which casts, when the doors open. Nobody reinvents it for each performance.
fullyParallel: true lets tests in the same file run in parallel too, not just separate files. That forces good habits early: a test that only passes after another test is broken already, and parallelism simply shows it sooner.
What Tomas learned
- Environment, browsers, retries and starting the app all belong in one config file, not in a framework.
forbidOnly,no-floating-promisesand the Playwright lint rules catch mechanical mistakes before review.- UI mode and the VS Code extension replace most "add a sleep and a screenshot to see what's happening" debugging.
…but
The suite runs anywhere now. It also has forty copies of the same card-filtering chain and twenty copies of "fill the postcode, click Find, pick the address". The next redesign will cost as much as the last one, just in tidier files.
Chapter 4 · Part IThe Page Object That Knew Too Much#
The problem
Tomas does what every Selenium book recommends and introduces page objects. Two weeks later QuotePage.ts is 1,400 lines long. It has a method for every screen of the journey, isPremiumShown(): Promise<boolean> and forty other boolean getters, and a waitForPageToLoad() that nobody dares to delete.
The UK postcode lookup appears on two screens, the driver's address and the registered keeper's address. It has been written twice, slightly differently. When the lookup's "Find address" button becomes "Search", one copy is fixed and the other breaks three weeks later.
The idea
Keep page objects, but make them small and make them about the user's intent.
- One object per screen or reusable component, never one per application.
- Locators are
readonlyfields or small methods. Locators are lazy, so defining them in the constructor costs nothing and finds nothing. - Methods are intents in domain language:
choose('Breakdown cover'), notclickSecondButtonInThirdCard(). - Tests assert; page objects expose. A page object hands out locators, and the test puts them in
expect(), so the failure points at the test that cares.card()usesdescribe()so traces say "Breakdown cover card" instead of a filter chain.
// file: tests/pages/add-ons.ts
import type { Locator, Page } from '@playwright/test';
export class AddOnsScreen {
readonly heading: Locator;
readonly total: Locator;
constructor(private readonly page: Page) {
this.heading = page.getByRole('heading', { name: 'Add extras to your cover' });
this.total = page.getByTestId('premium');
}
card(name: string): Locator {
return this.page.getByRole('listitem').filter({ hasText: name }).describe(`${name} card`);
}
async choose(name: string) {
await this.card(name).getByRole('button', { name: 'Choose' }).click();
}
}
A widget that appears in several places becomes a component object that takes a root locator instead of the whole page. Everything inside it is scoped to that root, so two address lookups on one screen can't interfere:
// file: tests/components/address-lookup.ts
import type { Locator } from '@playwright/test';
export class AddressLookup {
constructor(private readonly root: Locator) {}
async find(postcode: string, line1: string) {
await this.root.getByLabel('Postcode').fill(postcode);
await this.root.getByRole('button', { name: 'Find address' }).click();
await this.root.getByLabel('Select your address').selectOption(line1);
}
}
Screens compose components. Tests read like the requirement:
// file: tests/specs/add-ons.spec.ts
import { test, expect } from '@playwright/test';
import { AddOnsScreen } from '../pages/add-ons';
test('breakdown cover is added to the premium', async ({ page }) => {
const addOns = new AddOnsScreen(page);
await page.goto('/quote/resume/Q-1042?at=add-ons');
await addOns.choose('Breakdown cover');
await expect(addOns.card('Breakdown cover')).toContainText('Added');
await expect(addOns.total).toHaveText('£447.50');
});
Remember it as: Page objects know where things are and how to do things; tests decide what should be true.
How it works
Nothing magic happens here, and that is the point. A page object in Playwright is a plain class holding Locators. Because locators are lazy and re-resolve on every use, the page object never goes stale, needs no PageFactory, and needs no "wait until loaded" step. The first action or assertion that touches a locator does the waiting.
Theatre terms: each role's lines are written once, in the prompt book. Performances don't improvise them, and when a line changes, it changes in one place.
Two design rules hold the size down:
- Split by what changes together. The add-ons screen and the address lookup change for different reasons and at different times, so they belong in different files. Forty screens can reasonably mean forty small files plus a handful of components.
- Return locators, not booleans.
isPremiumShown(): Promise<boolean>forces the caller intoexpect(await …).toBe(true), which reads once and never retries (chapter 1). Exposepremium: Locatorand let the test writeawait expect(screen.premium).toBeVisible(), a web-first assertion.
What Tomas learned
- Page objects stay small: one per screen, component objects scoped to a root locator for anything reused.
- Locators as
readonlyfields cost nothing and never go stale, so there is no init step and no "wait for load". - Page objects expose locators and intents, and the test owns the assertions.
…but
Every test now starts with the same five lines. Create the screen objects, navigate, dismiss the cookie banner, set the market. The old suite hid those in a BaseTest class, with @BeforeMethod hooks and a protected field for everything. Tomas can feel that class wanting to be written.
Chapter 5 · Part IThe BaseTest Nobody Could Change#
The problem
The old suite's BaseTest is 900 lines long. Its @BeforeMethod hooks:
- start a browser;
- log in;
- create a quote through the API;
- set the market from a system property;
- open the start page;
- dismiss the cookie banner;
- build eleven page objects.
Every test inherits all of it, including the 40% of tests that need none of it. Some tests override one hook, and a few override the override. Changing the login step once broke 300 tests at the same moment. Nobody wanted to be the second person to try.
Tomas's new tests are drifting the same way: a beforeEach that grows a line every week.
The idea
Playwright replaces setup inheritance with fixtures. A fixture is a named thing a test can ask for by putting it in its argument list. page, context, browser, request and baseURL are built-in fixtures. You add your own with test.extend, and the test gets exactly what it names, no more:
// file: tests/fixtures.ts
import { test as base, expect } from '@playwright/test';
import type { MarketId } from './kestrel/types';
import { AddOnsScreen } from './pages/add-ons';
export type TestOptions = { marketId: MarketId };
type Fixtures = { addOns: AddOnsScreen; failOnPageErrors: void };
export const test = base.extend<TestOptions & Fixtures>({
// an option: a value the config or test.use() can set per project, file or describe
marketId: ['gb', { option: true }],
// a test-scoped fixture: built for each test that names it
addOns: async ({ page }, use) => {
await use(new AddOnsScreen(page));
},
// an automatic fixture: runs for every test, whether it asks or not
failOnPageErrors: [async ({ page }, use) => {
const errors: Error[] = [];
page.on('pageerror', error => errors.push(error));
await use();
expect(errors, 'uncaught exceptions in the page').toEqual([]);
}, { auto: true }],
});
export { expect };
Everything before await use(…) is setup, everything after it is teardown, and the value passed to use is what the test receives. Specs import test from this file instead of from @playwright/test:
// file: tests/specs/add-ons.spec.ts
import { test, expect } from '../fixtures';
test('legal expenses can be added', async ({ page, addOns }) => {
await page.goto('/quote/resume/Q-1042?at=add-ons');
await addOns.choose('Legal expenses');
await expect(addOns.card('Legal expenses')).toContainText('Added');
});
test.describe('in Germany', () => {
test.use({ marketId: 'de' });
test('legal expenses are called Rechtsschutz', async ({ page, addOns }) => {
await page.goto('/quote/resume/Q-2077?at=add-ons');
await expect(addOns.card('Rechtsschutz')).toBeVisible();
});
});
A test that doesn't name addOns never builds one. A test that doesn't need a quote never creates one. Nothing is inherited.
Remember it as: Tests ask for what they need by name; fixtures build it, hand it over, and clean it up. Nothing is inherited.
Four more tools finish the kit:
- Worker scope.
{ scope: 'worker' }builds a fixture once per worker process and shares it across that worker's tests. Use it for expensive things that are safe to share: an API client, a seeded reference-data set, a test account per worker (chapter 11). - Overrides. You can redefine a built-in fixture. Overriding
page, for example, can make every test start on the quote page. The new definition receives the original, so overrides compose. - Options.
{ option: true }makes a fixture configurable fromplaywright.config.tsand fromtest.use(). That is how one project becomes "the German suite" in chapter 9. mergeTests. This combines fixture sets written in separate modules, such as your Kestrel fixtures and a shared accessibility fixture, into onetest.
// file: tests/worker-fixtures.ts
import { test as base, mergeTests, type APIRequestContext } from '@playwright/test';
import { test as kestrel } from './fixtures';
const api = base.extend<{}, { pricingApi: APIRequestContext }>({
pricingApi: [async ({ playwright }, use) => {
const context = await playwright.request.newContext({ baseURL: process.env.PRICING_URL });
await use(context);
await context.dispose();
}, { scope: 'worker' }],
});
export const test = mergeTests(kestrel, api);
How it works
Fixtures form a dependency graph. Before a test runs, Playwright looks at the names the test asks for. It builds those fixtures and whatever they depend on, in dependency order, sharing worker-scoped ones. After the test, it tears everything down in reverse. Fixtures nobody asked for are never built, except auto ones.
Theatre terms: fixtures are the stagehands. Each performance calls only for the props its scenes need. The stagehands set them before the curtain and strike them after, and no two performances share a prop unless it is marked as shared for the whole run.
Because setup lives in the graph rather than in a class hierarchy, you can: - add a fixture without touching any test that doesn't use it; - change a fixture and know exactly which tests depend on it (the ones that name it); - see everything a test depends on in its signature.
What Tomas learned
- Fixtures replace
BaseTest: each test names what it needs, and gets exactly that, set up and torn down around it. - Options plus
test.use()switch behaviour per project, file or describe block, which is the hook the locale chapters rely on. - Worker scope is for expensive, shareable things. Everything else is per test.
…but
Tomas now has clean tests, and far too many of them. There are fourteen tests for the claims screen that differ only in the number of claims and the expected premium. And the same test exists four times, once per market, with four hardcoded prices. The code is DRY. The data is copied everywhere.
Chapter 6 · Part IIForty Tests, One Difference#
The problem
Tomas counts the claims tests: fourteen of them, one for each combination somebody once cared about. No claims, one claim at fault, one not at fault, two claims, a claim last month, a claim five years and a day ago… Each is a copy of the one before, with a different number of claims and a different expected premium typed in by hand. Then there are the markets: the same fourteen tests exist again for the US, Germany and Japan, with prices in dollars, euros and yen.
A teammate tried to fix this with one test that loops over an array of cases. That test is now red for the whole release whenever any case fails, and the report can't say which one.
The idea
Separate what varies from what is done. The variation is data: what the customer answers, and what Kestrel should reply. The procedure is code, written once. Playwright then generates one test per row.
First, type the inputs. Every Kestrel test answers the same questionnaire, so one Answers type describes it all: vehicle, driver, address, claims, cover, payment. A persona is a complete, realistic set of answers. Every variation is "a persona, except…":
// file: tests/data/young-driver.ts
import { merge } from './merge';
import { baseline } from './personas';
// baseline is a careful 35-year-old commuter; this is the same person at 19
export const youngDriver = merge(baseline, {
driver: { dob: '2007-01-20', licenceYears: 1 },
noClaimsYears: 0,
});
merge combines objects key by key and replaces arrays whole. A two-claims driver is { claims: [...] } and nothing else. Scenarios stay three lines long, even though the questionnaire has forty screens.
Second, put the expected outputs in the same row as the inputs. A scenario is a small typed record:
// file: tests/data/claims-scenarios.ts
import type { Scenario } from './scenarios';
export const claimsScenarios: Scenario[] = [
{
name: 'no claims',
expect: { premium: { gb: 412.5, us: 1286, de: 389.9, jp: 58400 } },
},
{
name: 'one claim at fault',
answers: { claims: [{ date: '2024-05-10', amount: 1800, atFault: true }] },
expect: { premium: { gb: 588, us: 1840, de: 541.2, jp: 80100 } },
},
{
name: 'one claim, not at fault',
answers: { claims: [{ date: '2024-05-10', amount: 1800, atFault: false }] },
expect: { premium: { gb: 431.75, us: 1350, de: 402.5, jp: 60300 } },
},
];
The expected premium is a number per market, not a formatted string. How 588 should look in each market (£588.00, 588,00 €, ¥80,100) is derived at assertion time by the market's own formatter, which chapter 9 explains. The data says what is true, and the code says how it looks.
Third, generate the tests:
// file: tests/specs/claims.spec.ts
import { test, expect } from '../fixtures';
import { answersFor } from '../data/scenarios';
import { claimsScenarios } from '../data/claims-scenarios';
for (const scenario of claimsScenarios) {
test(`premium: ${scenario.name}`, async ({ quote, market }, testInfo) => {
await testInfo.attach('answers', { body: JSON.stringify(answersFor(scenario), null, 2), contentType: 'application/json' });
await quote.resume(answersFor(scenario), 'quote-summary');
const premium = scenario.expect.premium?.[market.id];
if (premium === undefined) throw new Error(`${scenario.name}: no expected premium for ${market.id}`);
await expect(quote.premium).toHaveText(market.money(premium));
});
}
quote.resume asks a test-only endpoint to build the quote up to a given screen. Chapter 8 explains why that is allowed. Each row becomes a real test, with its own name in the report, its own retry, its own trace and its own parallel slot. Adding a case means adding a row. The answers are attached to the report, so a failure shows exactly what was submitted.
Remember it as: A scenario is a row: what the customer says and what Kestrel must answer. The code is written once; the tests are generated.
Outputs that are more than one value. The quote summary shows a dozen facts, so check them all and see every mismatch in one run, using expect.soft:
const summary = page.getByRole('table', { name: 'Your answers' });
await expect.soft(summary.getByRole('row', { name: /Cover level/ })).toContainText('Comprehensive');
await expect.soft(summary.getByRole('row', { name: /Voluntary excess/ })).toContainText('£250');
await expect.soft(summary.getByRole('row', { name: /Named drivers/ })).toContainText('1');
When the output is a structure, snapshot it instead of listing it:
- toMatchAriaSnapshot compares the accessibility tree (roles, names, order) against a YAML description. It is robust to styling and exact about content.
- toHaveScreenshot compares pixels, for the few places where appearance is the requirement.
await expect(page.getByRole('table', { name: 'Your answers' })).toMatchAriaSnapshot(`
- table "Your answers":
- row /Cover level Comprehensive/
- row /Voluntary excess £250/
`);
When a domain check repeats, give it a name with a custom matcher (expect.extend). Chapter 7 adds one that compares journeys.
How it works
test() calls register tests while the file is being loaded, before anything runs. A for loop around test() therefore produces N independent tests, exactly as if you had typed them out. The runner can shard, parallelise, retry and filter them individually. The loop runs at collection time, so the data must be available synchronously: import it, or read a file with fs.readFileSync.
In theatre terms, the scenario table is the cast list for a run of performances. The play is the same every night, and only the names on the programme change.
You can also wrap a test.describe() in a loop to parameterise a whole group, and data can come from CSV or JSON files when the business owns it. Prefer TypeScript data where you can: a typo in a field name then fails the compile, not the test run.
What Tomas learned
- Inputs are typed personas plus small deltas, and expected outputs sit in the same row as the inputs.
- The loop goes outside
test(): one row, one test, one line in the report. - Expected values are stored as facts (numbers, ids) and formatted per market at assertion time.
…but
The scenarios jump straight to the quote summary through a test-only endpoint. That checks pricing, but not the journey. Which screens appear for a driver with two claims, and in what order? The old suite had a 180-line test for that, which clicked through forty screens in a fixed order. It broke whenever the product team added a question.
Chapter 7 · Part IIThe Wizard with Forty Screens#
The problem
The old suite's journey test is 180 lines of click, fill, click, fill, in the exact order the screens appeared in 2023. There are nine copies of it: with claims, with an extra driver, with a modified car, monthly payment, and so on. Each copy differs from the others in a few lines in the middle.
Last sprint the product team added one question, "Where is the car kept overnight?", after annual mileage. All nine copies broke at the same line. Each one was fixed by hand, and two were fixed wrongly.
Tomas wants to test which screens appear, because that is where Kestrel's bugs live. But writing out the route in advance is the thing that keeps breaking.
The idea
Stop scripting the route. Let the app choose it, and answer whatever screen appears.
Four pieces make this work.
1. Every screen says who it is. Kestrel renders <main data-screen="claims" data-step="17">. The id names the screen, and the step number increases with every screen shown, including repeats. If your app has nothing like this, ask for it. It is one attribute, and it is the most valuable testability change you will ever request. The fallback is the screen's heading.
2. A registry knows how to answer every screen. Most screens are two or three labelled questions and a Continue button, so describe them as data with formScreen. Write code only for the odd ones:
// file: tests/screens/some-screens.ts
import type { ScreenId } from '../kestrel/types';
import { formScreen } from './form-screen';
import type { ScreenHandler } from './screen';
export const someScreens: Partial<Record<ScreenId, ScreenHandler>> = {
claims: formScreen({ yesNo: 'claims.question', value: a => a.claims.length > 0 }),
// shown once per claim; `visit` says which one this is
'claim-detail': formScreen(
{ fill: 'claims.date', value: (a, i) => a.claims[i].date },
{ fill: 'claims.amount', value: (a, i) => a.claims[i].amount },
{ yesNo: 'claims.atFault', value: (a, i) => a.claims[i].atFault },
),
// the odd one: look the car up, then continue on the next screen
'vehicle-lookup': async ({ page, form, market, answers }) => {
await form.fill('vehicle.registration', answers.vehicle.registration);
await page.getByRole('button', { name: market.t('vehicle.find') }).click();
},
};
The real registry is a Record<ScreenId, ScreenHandler>, not a Partial. When someone adds 'overnight-parking' to the list of screen ids, the compiler refuses to build until the registry says how to answer it.
3. A driver loop connects them. Read the screen id, look up its handler, answer from the scenario's data, wait for the step number to change, and repeat. Record every screen on the way:
// file: tests/journeys/drive.ts
import { expect, test, type Page } from '@playwright/test';
import type { Answers, ScreenId } from '../kestrel/types';
import type { Market } from '../markets/market';
import { defaultScreens } from '../screens/registry';
import { Form } from '../components/form';
export async function drive(page: Page, market: Market, answers: Answers, last: ScreenId) {
const screen = page.locator('main[data-screen]');
const form = new Form(page, market);
const visited: ScreenId[] = [];
while (visited.length < 80) {
await expect(screen).toBeVisible();
const id = (await screen.getAttribute('data-screen')) as ScreenId;
visited.push(id);
if (id === last) return visited;
const visit = visited.filter(v => v === id).length - 1;
const answer = market.screens[id] ?? defaultScreens[id];
const step = (await screen.getAttribute('data-step')) ?? '';
await test.step(id, () => answer({ page, market, answers, form, visit }));
await expect(screen).not.toHaveAttribute('data-step', step); // the next screen has arrived
}
throw new Error(`never reached ${last}: ${visited.join(' → ')}`);
}
The quote fixture's answerUntil() is this loop, plus a friendlier error for unknown screen ids.
4. The visited path is an expected output. Each scenario says which screens must appear, in order, and which must not:
// file: tests/specs/journey.spec.ts
import { test, expect } from '../fixtures';
import { scenarios, answersFor, marketTags } from '../data/scenarios';
for (const s of scenarios) {
test(`journey: ${s.name}`, { tag: marketTags(s) }, async ({ quote }) => {
test.slow(); // forty screens: triple the timeout for this test only
await quote.start();
const visited = await quote.answerUntil(answersFor(s), 'quote-summary');
if (s.expect.visits) expect(visited).toVisitInOrder(s.expect.visits);
for (const id of s.expect.skips ?? []) expect(visited).not.toContain(id);
});
}
test.slow() triples the timeout for the long walk only. toVisitInOrder is a custom matcher defined next to the fixtures. It passes if the expected screens appear in that order, with anything in between. When a branch is wrong, the failure reads like a bug report:
expected to visit claims → claim-detail → claim-detail
stopped matching at "claim-detail"
visited: cover-start → vehicle-lookup → … → claims → claim-detail → convictions → …
Remember it as: The app picks the route; the test answers whatever screen appears and checks the path it was given.
Now the new "overnight parking" screen costs one registry entry. No scenario changes, and no route anywhere needs editing. The nine copies become nine rows in a table.
How it works
In theatre terms, the old suite's script had a fixed running order. The driver loop is an understudy who knows every scene in the repertoire. They step on stage, see which scene is playing, and play it, and they take notes on the order the scenes came in.
Three details make the loop dependable:
- Waiting on
data-step, not ondata-screen. A driver with two claims seesclaim-detailtwice in a row. The id doesn't change, but the step does. test.stepper screen. The trace viewer and the HTML report show a named step for every screen, so a failure on screen 31 is found by name, not by scrolling through 200 actions. Page object methods can get the same treatment with a decorator:
// file: tests/step.ts
import { test } from '@playwright/test';
/** Report a page object method as one boxed step: "AddOnsScreen.choose". */
export function step<This extends object, Args extends unknown[], R>(
target: (this: This, ...args: Args) => Promise<R>,
context: ClassMethodDecoratorContext<This>,
) {
return function (this: This, ...args: Args): Promise<R> {
return test.step(`${this.constructor.name}.${String(context.name)}`, () => target.call(this, ...args), { box: true });
};
}
- A hard stop. Eighty screens without reaching the target means a loop in the app or a handler that never submits. Fail with the path so far, instead of hanging until the timeout.
What Tomas learned
- A long branching journey is a loop, a registry and a scenario, not a script.
- The visited path is an output like any other: record it and assert it.
data-screenanddata-stepare the cheapest, most valuable testability hooks to ask the developers for.
…but
The full walk is thorough and slow: about 40 seconds per scenario per market. The pricing team wants the claims branch tested with fifteen claim combinations, and the add-ons screen with every add-on combination. At 40 seconds each, in four markets, that is most of an hour spent re-answering the same thirty screens to reach the one that matters.
Chapter 8 · Part IIThe Branch Nobody Reached#
The problem
The journey tests work, and there aren't nearly enough of them. The claims branch alone has fifteen interesting combinations. The add-ons screen has three kinds of breakdown cover times legal expenses times courtesy car. Multiply by payment options and four markets and you get over a thousand walks of forty screens each.
The walks also get interrupted. A "How are we doing?" survey appears on a random screen for one visitor in ten. After four minutes, a "Still there?" dialog asks whether the session should stay open. A teammate has "fixed" the survey inside the driver loop:
if (await page.getByRole('dialog', { name: 'How are we doing?' }).isVisible()) {
await page.getByRole('button', { name: 'No thanks' }).click();
}
It helps most of the time.
The idea
Three decisions tame the combinations, and one API tames the interruptions.
Start where the branch starts. A test about the claims branch doesn't need to prove that the vehicle lookup works. Twenty other tests do that. Kestrel's non-production builds expose a test-only endpoint that creates a quote with given answers, positioned at a given screen, and returns a URL to resume it. The quote fixture's resume() wraps it:
test('resume a quote at the claims screen', async ({ page, request }) => {
const response = await request.post('/api/test/quotes', {
data: { market: 'gb', at: 'claims', answers: { driver: { licenceYears: 3 } } },
});
await expect(response).toBeOK();
const { resumeUrl } = await response.json();
await page.goto(resumeUrl);
await expect(page.locator('main')).toHaveAttribute('data-screen', 'claims');
});
Branch tests then resume and let the driver loop do the rest:
// file: tests/specs/claims-branch.spec.ts
import { test, expect } from '../fixtures';
import { scenarios, answersFor, marketTags } from '../data/scenarios';
for (const s of scenarios) {
test(`from claims: ${s.name}`, { tag: marketTags(s) }, async ({ quote }) => {
const answers = answersFor(s);
await quote.resume(answers, 'claims'); // screens 1–21 built server-side
const visited = await quote.answerUntil(answers, 'quote-summary');
if (s.expect.visits) expect(visited).toVisitInOrder(s.expect.visits);
});
}
A 40-second walk becomes an 8-second one. The same scenarios table drives both. The resume point is the only difference.
No seeding endpoint? Look for deep links that restore saved quotes, or save the browser state of a quote in progress once per worker and load it into new contexts (chapter 11 shows the mechanism). And ask for the endpoint anyway: it pays for itself in a week.
Keep a few full walks. Seeding skips the real screens 1–21, so something must still cross them. One or two complete journeys per market, run on every change, prove that the route holds together. Everything else starts mid-flow.
Choose combinations deliberately. Testing every combination is rarely possible and rarely useful: most bugs involve one value or the interaction of two. Pairwise selection picks a small set of rows in which every pair of values appears at least once:
claims drivers payment breakdown
0 0 annual none
0 1 annual national
0 2 monthly european
1 0 monthly national
1 1 annual european
1 2 annual none
2 0 annual european
2 1 monthly none
2 2 annual national
Nine rows instead of 54, and every pair of values is still covered. Tools such as Microsoft's PICT generate these tables. Paste the output into the scenario file as data. Then add the boundaries the business cares about by hand, such as the claim that is exactly five years old.
Handle interruptions once, globally. For things that can appear at any moment and aren't what the test is about, register a locator handler. Before every action and every web-first check, Playwright looks for the overlay and runs the handler if it is there:
const survey = page.getByRole('dialog', { name: 'How are we doing?' });
await page.addLocatorHandler(survey, async () => {
await survey.getByRole('button', { name: 'No thanks' }).click();
});
Put it in an auto fixture and no test or handler ever mentions the survey again. The Kestrel fixtures already do this for the cookie banner.
Remember it as: Start each test where its branch starts, choose the combinations on purpose, and handle interruptions once, globally.
Sometimes what you need to wait for isn't on the page, such as a quote record reaching priced in the back end. Use expect.poll to retry any async value, or toPass to retry a block:
test('the quote is priced in the back end', async ({ request }) => {
await expect.poll(async () => {
const response = await request.get('/api/quotes/Q-1042');
return ((await response.json()) as { status: string }).status;
}, { timeout: 15_000 }).toBe('priced');
await expect(async () => {
const response = await request.get('/api/quotes/Q-1042/documents');
expect(response.status()).toBe(200);
}).toPass({ intervals: [1_000, 2_000, 5_000] });
});
How it works
addLocatorHandler hooks into the same actionability machinery as chapter 1. Before an action or assertion runs, Playwright checks whether any registered handler locator is visible. If one is, it runs the handler, waits for the overlay to disappear, and then carries on with the original action. The handler runs inside Playwright's waiting logic, so there is no race of the kind the teammate's isVisible() check had.
In theatre terms, you don't run the whole play from Act 1 to rehearse one scene in Act 3. The director says "from scene 21" and the company starts there. The dress rehearsals, the full walks, still run from the top, just not every time.
What Tomas learned
- Branch tests start mid-flow through seeding, and a few full walks per market keep the route honest.
- Pairwise rows plus deliberate boundaries beat "every combination" on both speed and signal.
- Random overlays are handled once, globally, with
addLocatorHandler, never withisVisible()checks.
…but
Everything so far has run in English. Tomas switches the project to de-DE and nothing passes. The locators say 'Continue' and the button says Weiter. The premium is asserted as £412.50 and the page shows 389,90 €. And the calendar on the start-date screen thinks today is yesterday.
Chapter 9 · Part IIIThe Quote That Said 1.234,50#
The problem
The old suite handled Germany by copying the UK tests into a de folder and translating the strings by hand. The German copies were eleven months out of date, and Japan had never been attempted.
Tomas's new suite fails in German in three different ways:
- Text: every locator says 'Continue', and the button says Weiter.
- Numbers: the premium is asserted as '£389.90', and the page shows 389,90 €.
- Time: the cover-start screen rejects "today" as a past date. The CI machine runs in UTC, the test browser uses the machine's timezone, and at 00:30 in Berlin it is still yesterday in UTC.
The idea
Treat the locale as configuration, and derive everything that depends on it.
1. Each market is a project. A project sets the browser's locale and timezoneId, and the marketId option from chapter 5, all from the same market profile:
// file: playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
import type { TestOptions } from './tests/fixtures';
import type { MarketId } from './tests/kestrel/types';
import { markets } from './tests/markets';
const market = (id: MarketId) => ({
name: markets[id].locale,
grep: new RegExp(`@${id}\\b`),
use: {
...devices['Desktop Chrome'],
marketId: id,
locale: markets[id].locale, // navigator.language, Accept-Language, Intl in the page
timezoneId: markets[id].timezoneId, // the browser's clock, whatever the CI machine uses
},
});
export default defineConfig<TestOptions>({
testDir: './tests',
use: { baseURL: process.env.BASE_URL ?? 'http://localhost:3000' },
projects: [market('gb'), market('us'), market('de'), market('jp')],
});
npx playwright test --project=de-DE runs the German suite. Without a flag, all four run, in parallel, from the same test files. The grep line is explained in chapter 10.
2. Text comes from the app's own message catalog. Kestrel's front end already has messages/de-DE.json, the file its translators edit. The market profile loads it, and every locator asks the market for its words:
// file: tests/specs/words.spec.ts
import { test } from '@playwright/test';
import { markets } from '../markets';
const { t } = markets.de;
test('continue is in the market language', async ({ page }) => {
await page.getByRole('button', { name: t('nav.continue') }).click(); // "Weiter"
await page.getByLabel(t('address.postcode')).fill('10115'); // "Postleitzahl"
});
In practice nobody writes t(…) in a test. The Form component and the screen registry take message keys ('claims.question'), and the market turns them into words. Because the catalog is imported as typed JSON, a key that doesn't exist is a compile error, and so is a German catalog that is missing a key the English one has.
3. Formatted values are derived, never typed. Expected data stores facts: 389.9, '2026-10-01'. The market formats them with the same Intl APIs the front end uses:
// file: tests/markets/format.ts
export function formatters(locale: string, currency: string) {
const money = new Intl.NumberFormat(locale, { style: 'currency', currency });
const date = new Intl.DateTimeFormat(locale, { dateStyle: 'medium', timeZone: 'UTC' });
return {
money: (amount: number) => money.format(amount),
date: (iso: string) => date.format(new Date(iso)),
};
}
| Market | money(412.5) |
date('2026-10-01') |
|---|---|---|
| en-GB | £412.50 | 1 Oct 2026 |
| en-US | $412.50 | Oct 1, 2026 |
| de-DE | 412,50 € | 01.10.2026 |
| ja-JP | ¥413 (yen has no minor unit) | 2026/10/01 |
4. Time is pinned. The project fixes the timezone, and the clock API fixes the moment. Tests about "today", "tomorrow" or "renewal in 30 days" set the time explicitly, including the awkward times:
// file: tests/specs/formats.spec.ts
import { test, expect } from '../fixtures';
import { baseline } from '../data/personas';
import { scenarios, marketTags } from '../data/scenarios';
const clean = scenarios[0];
test('premium and start date use the market format', { tag: marketTags(clean) }, async ({ page, quote, market }) => {
// 23:30 UTC on 30 September: already 1 October in Berlin and Tokyo
await page.clock.setFixedTime(new Date('2026-09-30T23:30:00Z'));
await quote.resume(baseline, 'quote-summary');
await expect(quote.premium).toHaveText(market.money(clean.expect.premium![market.id]!));
await expect(page.getByTestId('cover-start')).toHaveText(market.date(baseline.cover.startDate));
});
One test, four markets, no strings typed by hand.
Remember it as: The locale is configuration. Words come from the app's catalog, formats from Intl, and time from a pinned clock.
How it works
The theatre version: the same play tours Berlin and Tokyo. The lines are spoken in German and Japanese from the official translation, prices are in the local currency, and the performance starts at local time. Nobody rewrites the play for each city.
locale changes more than language. It sets navigator.language and the Accept-Language header, and it changes how Intl behaves inside the page. timezoneId changes what the page's Date thinks the time is. Both belong to the browser context, which is why they are set per project, not per test.
Formatted text often contains non-breaking spaces (412,50 € has one before the euro sign). toHaveText collapses all whitespace on both sides before comparing, so a string built by Intl in Node matches the page even when the two use different space characters.
What Tomas learned
- A market is a project:
locale,timezoneIdand the market option, all set from one profile. - Words come from the app's catalog by key, and formats come from
Intl, so expected data stays locale-free. - Time is part of the input: pin it, and test the midnight edges on purpose.
…but
Text and formats now work in four languages. The flow doesn't. The German run fails at screen two: there is no registration lookup in Germany, but a vehicle-key screen that the UK has never had. The US asks for a state and shows a disclosures screen. Japan asks for names family-name-first, then again in katakana. The obvious fix is if (market.id === 'de'), and it is about to appear in every file.
Chapter 10 · Part IIIThe Screen Only Germany Has#
The problem
Tomas runs the journey tests in all four markets. The UK passes. The other three fail within five screens: - Germany has no registration lookup. It asks for a vehicle key (HSN/TSN), then asks for the no-claims class (SF-Klasse) from a dropdown, and later for a SEPA mandate instead of a card. - The US asks for a state, and then shows a disclosures screen that depends on it. - Japan asks for the family name first, then asks for both names again in katakana. Its postal code fills in the prefecture by itself.
A teammate's pull request fixes Germany with eleven if (market.id === 'de') blocks spread over five files. Two more markets would make it forty.
The idea
Most "flow differences" turn out to be one of three kinds, and each kind has a home that isn't a spec.
1. Different screens. The app routes, and the loop answers. Germany's vehicle-key screen, the US disclosures screen and Japan's katakana screen are just more screen ids. The driver loop from chapter 7 already asks the page which screen it is on, and the registry already knows how to answer vehicle-hsn-tsn or name-reading. Nothing in any test changes. This covers most of the difference between Kestrel's markets.
2. The same screen, answered differently: a market override. The driver-address screen exists everywhere, but in the UK it is a postcode lookup and in Japan the postal code fills the prefecture. The market profile overrides just that screen:
// file: tests/markets/gb.ts
import messages from '../../app/messages/en-GB.json';
import { defineMarket } from './market';
export const gb = defineMarket({
id: 'gb', locale: 'en-GB', timezoneId: 'Europe/London', currency: 'GBP', messages,
screens: {
// Postcode first, then pick the address the lookup finds.
'driver-address': async ({ page, form, market, answers }) => {
await form.fill('address.postcode', answers.address.postcode);
await page.getByRole('button', { name: market.t('address.find') }).click();
await form.select('address.pick', answers.address.line1);
await form.continue();
},
},
});
The loop picks market.screens[id] ?? defaultScreens[id]. Overrides are rare and obvious: Kestrel has four across four markets.
3. Different data: local answers. A German customer doesn't live at 12 Harbour Road, Bristol, and a Japanese customer's name is family name first, with a reading. Local details sit in one table, merged between the persona and the scenario's own delta:
// file: tests/data/local-answers.ts
import type { Answers, DeepPartial, MarketId } from '../kestrel/types';
export const localAnswers: Record<MarketId, DeepPartial<Answers>> = {
gb: {},
us: {
address: { postcode: '10001', line1: '350 Fifth Avenue', city: 'New York', state: 'NY' },
contact: { phone: '(212) 555-0142' },
},
de: {
address: { postcode: '10115', line1: 'Invalidenstraße 117', city: 'Berlin' },
contact: { phone: '030 901820' },
},
jp: {
driver: { firstName: '健', lastName: '佐藤', reading: { firstName: 'ケン', lastName: 'サトウ' } },
address: { postcode: '100-0005', line1: '丸の内1-9-1', city: '千代田区' },
contact: { phone: '03-1234-5678' },
},
};
answersFor(scenario, market.id) returns persona → local answers → scenario delta. The scenario still says only what it is about.
That leaves one decision: which scenarios apply to which market. A scenario about the UK credit agreement means nothing in Germany. Scenarios list their markets, and the tests carry them as tags. Each project's grep (chapter 9) selects its own:
// file: tests/specs/quote.spec.ts
import { test, expect } from '../fixtures';
import { scenarios, answersFor, marketTags } from '../data/scenarios';
for (const s of scenarios) {
test(s.name, { tag: marketTags(s) }, async ({ quote, market }) => {
await quote.start();
const visited = await quote.answerUntil(answersFor(s, market.id), 'quote-summary');
if (s.expect.visits) expect(visited).toVisitInOrder(s.expect.visits);
for (const id of s.expect.skips ?? []) expect(visited).not.toContain(id);
const premium = s.expect.premium?.[market.id];
if (premium !== undefined) await expect(quote.premium).toHaveText(market.money(premium));
if (s.expect.declined) await expect(quote.screen).toContainText(market.t(s.expect.declined));
});
}
The German SEPA scenario carries markets: ['de'], so it is tagged @de and only the de-DE project runs it. The UK project never sees it. There are no skips in the report and no if in the spec. The two ifs that remain test which expectations a row declares, never which market is running.
Remember it as: New screens are the app's business, reshaped screens are overrides, local details are data. Specs never ask which market they are in.
Adding a fifth market. Kestrel launches in Ireland. Here is the whole change:
app/messages/en-IE.json already written by the translators
tests/kestrel/types.ts 'ie' added to MarketId (the compiler lists what's missing next)
tests/markets/ie.ts defineMarket({ id: 'ie', locale: 'en-IE', currency: 'EUR', … })
+ an override for driver-address (Eircode lookup)
tests/data/local-answers.ts an Irish address and phone number
tests/data/scenarios.ts ie premiums added to rows that assert premiums
playwright.config.ts market('ie')
No spec file changes. The Record<MarketId, …> types make the compiler produce the to-do list.
How it works
In theatre terms, the touring company plays Berlin with a scene the London production never had. The local director stages the address scene differently, the local cast has local names, and the play itself stays the same.
The market profile is the single place to read "how is Germany different?" It holds the formats, the catalog and the overrides, and the rest of the differences live in the app's routing and in the local answers.
What Tomas learned
- Most flow differences are new screens. The app routes to them, and the registry answers them without touching a test.
- Market overrides handle the same screen answered differently, and local answers hold local details, so specs stay free of
ifs. - Market tags plus project
grepchoose which scenarios run where, with no skips.
…but
Four markets times the scenario table is a lot of tests, and every one of them starts by signing in as a returning customer through the login screen. That is two-factor code and all, 400 times per run. The identity provider's rate limiter has noticed.
Chapter 11 · Part IVLogging In Four Hundred Times#
The problem
Returning customers get a different quote journey: their details are pre-filled, and saved quotes appear on a dashboard. So most scenarios start signed in. Tomas's suite signs in through the login screen at the start of every test: email, password, a one-time code from the test mailbox, then "Remember this device?". That is 12 seconds per test, 400 times per run.
Then two things break at once. The identity provider starts rate-limiting the CI machines. And because every worker signs in as the same test.customer@kestrel.test, one test's saved quote appears on another test's dashboard, and the "you have no saved quotes" check fails whenever tests happen to overlap.
The idea
Sign in once, reuse the session. A browser's signed-in state is just cookies and local storage. Playwright can save it to a file, storageState, and start new contexts with it already loaded. A setup project does the signing in, and every other project depends on it:
// file: tests/auth.setup.ts
import { test as setup, expect } from '@playwright/test';
setup('sign in as a returning customer', async ({ page }) => {
await page.goto('/sign-in');
await page.getByLabel('Email address').fill(process.env.CUSTOMER_EMAIL!);
await page.getByLabel('Password').fill(process.env.CUSTOMER_PASSWORD!);
await page.getByRole('button', { name: 'Sign in' }).click();
await expect(page.getByRole('heading', { name: 'Your quotes' })).toBeVisible();
await page.context().storageState({ path: '.auth/customer.json' });
});
// file: playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
projects: [
{ name: 'setup', testMatch: /.*\.setup\.ts/ },
{
name: 'en-GB',
dependencies: ['setup'], // runs first, even with --project=en-GB
use: { ...devices['Desktop Chrome'], storageState: '.auth/customer.json' },
},
],
});
The setup runs once per test run, as a real test: traced, retried and reported. Every test in en-GB then starts signed in, in a fresh context, without touching the login screen. Add .auth/ to .gitignore, because those files are live sessions.
One account per worker, when tests change account state. Signing in once is enough for read-only checks. Kestrel's tests save quotes, though, so parallel tests on one account interfere. Give each worker its own account with a worker-scoped fixture that overrides storageState:
// file: tests/account-fixtures.ts
import path from 'node:path';
import { test as base, expect } from '@playwright/test';
export const test = base.extend<{}, { workerStorageState: string }>({
storageState: ({ workerStorageState }, use) => use(workerStorageState),
workerStorageState: [async ({ playwright }, use) => {
const id = test.info().parallelIndex; // 0, 1, 2… stable per worker slot
const file = path.join(test.info().project.outputDir, `.auth/customer-${id}.json`);
const request = await playwright.request.newContext({ baseURL: process.env.BASE_URL });
const response = await request.post('/api/test/sessions', {
data: { persona: 'returning-customer', account: `worker-${id}` },
});
await expect(response).toBeOK();
await request.storageState({ path: file });
await request.dispose();
await use(file);
}, { scope: 'worker' }],
});
Four workers, four accounts, and no two tests ever share a dashboard. The session comes from a test-only endpoint instead of the UI. One setup test still signs in through the real screen, so the login itself stays covered.
Remember it as: Sign in once and save the session; give each worker its own account when tests change what the account owns.
Talk to the API directly when the UI isn't the point. The built-in request fixture is an HTTP client that shares the test's baseURL. page.request also shares the page's cookies, so it is already signed in. Use it for API tests in their own right, and to arrange or verify state around UI tests:
// file: tests/specs/pricing-api.spec.ts
import { test, expect } from '@playwright/test';
test('the pricing API rejects drivers under 17', async ({ request }) => {
const response = await request.post('/api/quotes', {
data: { market: 'gb', driver: { dob: '2011-03-01' } },
});
expect(response.status()).toBe(422);
expect(await response.json()).toMatchObject({ errors: [{ field: 'driver.dob' }] });
});
More than one person in a test. Some Kestrel stories involve two people. A quote is referred to an underwriter, who approves it while the customer waits. Each person is a separate context with its own storage state, in the same browser:
// file: tests/specs/referral.spec.ts
import { test, expect } from '@playwright/test';
test('an underwriter approves a referred quote', async ({ browser }) => {
const customer = await (await browser.newContext({ storageState: '.auth/customer.json' })).newPage();
const underwriter = await (await browser.newContext({ storageState: '.auth/underwriter.json' })).newPage();
await customer.goto('/quotes/Q-3310');
await underwriter.goto('/referrals/Q-3310');
await underwriter.getByRole('button', { name: 'Approve' }).click();
await expect(customer.getByRole('status')).toHaveText('Your quote has been approved');
});
How it works
A storage state file is JSON: cookies, plus local storage (and optionally IndexedDB) per origin. Loading it into a new context is instant, and the context is still fresh and isolated. Only the credentials are shared.
dependencies do more than fix the order. They make the setup part of the run: it appears in the report, it gets a trace when it fails, and fixtures work inside it. That is why a setup project beats the older globalSetup hook for authentication.
In theatre terms, cast members collect their backstage pass once at the stage door. Nobody checks their ID again before every scene, and each understudy has a pass of their own.
What Tomas learned
- A setup project signs in once, and
storageStatemakes every later context start signed in. - Tests that change account data need an account per worker, and a worker-scoped fixture provides it.
- The
requestfixture and extra contexts cover API checks and multi-user stories in the same test.
…but
The dashboard is fast now, and the payment step fails one run in five. Kestrel's card form is an iframe from a payment provider whose sandbox is slow and sometimes down. The address lookup calls a third-party API with a daily quota, which the suite burns through by lunchtime.
Chapter 12 · Part IVThe Payment Provider Was Down#
The problem
Three of Kestrel's screens depend on services Kestrel doesn't run: - Address lookup calls a postcode API with a daily quota. - Vehicle lookup calls a registration database that answers in anything from 200 ms to 6 seconds. - Payment is an iframe from a payment provider whose sandbox goes down for maintenance on Tuesday afternoons.
Tomas's suite fails whenever any of the three has a bad day, and none of those failures is Kestrel's fault. Meanwhile the one error screen that matters most, "we couldn't price your quote, please try again", has never been tested, because nobody knows how to make the pricing service fail on demand.
The idea
Playwright sits between the page and the network, so a test can decide what any request receives. page.route() intercepts requests matching a URL pattern. The handler can fulfil the request with your own response, let it continue, abort it, or fetch the real response and change it:
// file: tests/mocks/boundaries.ts
import type { Page } from '@playwright/test';
/** The postcode API: quota-limited, not ours. Always answer from data. */
export async function mockAddressLookup(page: Page, addresses: string[]) {
await page.route('**/api/address-lookup?*', route => route.fulfill({ json: { addresses } }));
}
/** Make our own pricing service fail, to test the error screen. */
export async function failPricing(page: Page, status = 503) {
await page.route('**/api/price', route => route.fulfill({ status, json: { error: 'unavailable' } }));
}
/** Keep the real response but change one field: an edge case the back end rarely produces. */
export async function withPremium(page: Page, premium: number) {
await page.route('**/api/price', async route => {
const response = await route.fetch();
const body = await response.json();
await route.fulfill({ response, json: { ...body, premium } });
});
}
Record slow or complex responses once and replay them with routeFromHAR. Run with update: true against the real service to refresh the recording, and commit the .har file:
await page.routeFromHAR('tests/hars/vehicle-lookup.har', { url: '**/api/vehicles/**', update: false });
WebSockets can be intercepted the same way, with page.routeWebSocket(). The handler can answer on the server's behalf, or connect to the real server and edit messages in flight.
Check that the mock was used. A mock the page never calls is a test that proves nothing. Wait for the request and assert what was sent. Start waiting before the action that triggers it:
const lookup = page.waitForRequest('**/api/address-lookup?*');
await page.getByRole('button', { name: 'Find address' }).click();
expect(new URL((await lookup).url()).searchParams.get('postcode')).toBe('BS1 4DJ');
Remember it as: Mock at the boundaries you don't own, fail your own services on purpose, and prove that each mock was actually called.
The rest of the browser is as reachable as the network. Kestrel's payment iframe, policy documents in new tabs, certificate downloads and the "leave your quote?" dialog look like this (see frames, pages, downloads and dialogs):
// file: tests/specs/documents.spec.ts
import { test, expect } from '@playwright/test';
test('policy documents after purchase', async ({ page }) => {
await page.goto('/quote/resume/Q-1042?at=payment-method');
// iframes: locate inside them like anywhere else
const card = page.frameLocator('iframe[title="Secure card payment"]');
await card.getByLabel('Card number').fill('4242 4242 4242 4242');
await page.getByRole('button', { name: 'Pay now' }).click();
// popups: start waiting before the click that opens them
const popupPromise = page.waitForEvent('popup');
await page.getByRole('link', { name: 'Policy wording (PDF)' }).click();
await expect(await popupPromise).toHaveURL(/policy-wording\.pdf$/);
// downloads
const downloadPromise = page.waitForEvent('download');
await page.getByRole('button', { name: 'Download certificate' }).click();
const download = await downloadPromise;
expect(download.suggestedFilename()).toBe('certificate-Q-1042.pdf');
// dialogs are dismissed automatically unless you say otherwise
page.once('dialog', dialog => dialog.accept());
await page.getByRole('link', { name: 'Start a new quote' }).click();
});
How it works
Routes belong to a page or a context and are checked in reverse order of registration, so a later, more specific route wins. A handler that calls route.fallback() passes the request to the next matching route. route.continue() sends it to the network.
In theatre terms, some scenes need a messenger from outside the company. You don't hire a real courier for every rehearsal. A stagehand delivers the letter, word-perfect, every night. Now and then, on purpose, the letter arrives with bad news, to see how the cast copes.
What to mock is a policy decision: - Always mock third parties: they are slow, rate-limited, cost money, and aren't under test. - Mock your own services only to produce states that are otherwise hard to reach, such as errors, timeouts and edge-case values. - Keep a few unmocked walks in a staging environment. They are the only thing that proves the real integrations still fit together.
What Tomas learned
- Third parties are mocked at the network boundary, and the suite no longer fails when they do.
- The same routing makes Kestrel's own error states testable on demand.
- Frames, popups, downloads and dialogs are plain API calls, and the rule for events is always: start waiting before the action.
…but
The suite is now 1,100 tests across four markets. It is reliable on Tomas's machine and takes 25 minutes on one CI runner. Three tests fail about once a day on CI and never locally. The team's response so far has been to click "re-run".
Chapter 13 · Part IVGreen Locally, Red in CI#
The problem
The suite has 1,100 tests in four markets and takes 25 minutes on one CI runner. Developers have stopped waiting for it. They merge and check later, if at all.
Three tests fail about once a day, only on CI: - "add-ons: breakdown and legal" - "journey: two claims (de-DE)" - "formats (ja-JP)"
Each time, someone clicks re-run, the test passes, and the team moves on. Last week one of those "flaky" tests was a real bug: a race in the add-ons screen that customers also hit. It had been re-run into silence for a month.
The idea
Make the suite fast by spreading it out, and make it honest by treating flakiness as information.
Parallelise, then shard. fullyParallel: true runs every test independently across the workers on one machine. --shard=1/4 splits the whole run across four machines. Each shard writes a blob report, and a final job merges them into one HTML report:
# .github/workflows/e2e.yml (excerpt)
jobs:
test:
strategy: { fail-fast: false, matrix: { shard: [1/4, 2/4, 3/4, 4/4] } }
runs-on: ubuntu-latest
container: mcr.microsoft.com/playwright:v1.63.0-noble
steps:
- uses: actions/checkout@v4
- run: npm ci
- run: npx playwright test --shard=${{ matrix.shard }}
- uses: actions/upload-artifact@v4
if: ${{ !cancelled() }}
with: { name: blob-${{ strategy.job-index }}, path: blob-report }
report:
needs: test
if: ${{ !cancelled() }}
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npm ci
- uses: actions/download-artifact@v4
with: { pattern: blob-*, path: all-blob-reports, merge-multiple: true }
- run: npx playwright merge-reports --reporter html ./all-blob-reports
Twenty-five minutes on one machine becomes about seven on four. Sharding and merging are free, built into the runner. The job runs in the official Docker image, whose version matches the package, so browsers and fonts are identical on every run.
Record evidence on the retry. The CI section of the config:
// file: playwright.config.ts
import { defineConfig } from '@playwright/test';
export default defineConfig({
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 4 : undefined,
reporter: process.env.CI ? [['blob'], ['github']] : [['html', { open: 'never' }]],
use: {
trace: 'on-first-retry', // a full recording whenever a test needed a second go
screenshot: 'only-on-failure',
video: 'retain-on-failure',
},
});
A test that fails and then passes on retry is reported as flaky, not passed. The trace from its retry is attached to the report. Open it with npx playwright show-trace, or drop it onto trace.playwright.dev. It shows every action, with a DOM snapshot before and after, the network, the console, and the source line.
Remember it as: Spread the run across machines, keep the trace from every retry, and treat "flaky" as a bug report, not a pass.
Triage flaky tests like bugs. Kestrel's three turned out to be one of each of the usual causes:
| Cause | Kestrel's example | The trace showed |
|---|---|---|
| A race in the test | Asserting the total before the add-on request finished | The assertion ran against the old total; a waitForResponse was missing |
| A race in the product | Double-clicking "Choose" added breakdown cover twice | Two POSTs 40 ms apart, which is a real bug and was filed |
| Environment | Japanese fonts missing on one runner image | Screenshot diff with tofu boxes; fixed by using the official image |
Once the known causes are fixed, set failOnFlakyTests: true in CI so the next one fails the build instead of hiding.
Slice the suite. Tags and annotations make the same suite serve several purposes:
// file: tests/specs/smoke.spec.ts
import { test, expect } from '../fixtures';
test('a returning customer can resume a saved quote', {
tag: ['@smoke', '@gb', '@us', '@de', '@jp'],
annotation: { type: 'issue', description: 'https://kestrel.example/browse/QA-311' },
}, async ({ page, market }) => {
await page.goto('/quotes');
await page.getByRole('link', { name: 'Q-1042' }).click();
await expect(page.getByRole('heading', { name: market.t('summary.heading') })).toBeVisible();
});
test.fixme('renewal quotes keep the named drivers', async () => {}); // known gap, tracked, visible in the report
The commands that go with them:
npx playwright test --grep @smoke # the 3-minute pre-merge run
npx playwright test -G @slow # everything except slow tests (--grep-invert)
npx playwright test --last-failed # only what failed last time
npx playwright test --only-changed=main # tests affected by files changed since main
npx playwright test --repeat-each=20 -g "add-ons: breakdown" # reproduce a flake locally
How it works
Shards split the test list deterministically, so every machine knows its share without talking to the others. fullyParallel lets that split happen per test rather than per file, which keeps shards evenly sized. The merged report keeps each test's retries, attachments and traces, and its timeline and "Speedboard" tabs show where the minutes go.
In theatre terms, the trace is the rehearsal recording. When a scene goes wrong one night in ten, nobody argues about what happened. They watch the tape.
What Tomas learned
- Workers, shards and merged blob reports turn a 25-minute run into a 7-minute one, without a grid.
- A flaky test is a bug report with a trace attached. Triage it: test race, product race, or environment.
- Tags,
--grep,--last-failedand--only-changedslice one suite for different moments.
…but
The suite is fast, honest and green, and a new manager asks the question every team hears eventually: "Wouldn't Cypress have been simpler? And the architecture group still says Selenium is the standard." Tomas has opinions. They should be fair ones.
Chapter 14 · Part VWhy Not Cypress or Selenium#
The problem
Kestrel's architecture group lists Selenium as the company standard. The mobile team uses Cypress and likes it. Tomas's new manager has been asked to justify a third tool, and asks Tomas for one page, "not a sales pitch".
The honest answer depends on knowing what each tool is, not just how its API looks. All three drive browsers. They do it in fundamentally different ways, and most of their strengths and limits follow from that one design decision.
The idea
Compare the architectures first. The feature lists follow from them.
- Selenium (since 2004) implements WebDriver, a W3C standard since 2018. Your test sends each command as an HTTP request to a driver process (chromedriver, geckodriver, safaridriver), which controls the browser. It is language-neutral and vendor-backed, and one round trip per command made waiting the test author's problem. Selenium is now moving to WebDriver BiDi, a two-way protocol built with the browser vendors, which narrows the gap.
- Cypress (open-sourced 2017) runs your test code inside the browser, in the same run loop as your app. That gives it direct access to the app's objects and an excellent live command log. The same choice limits it to JavaScript, to one browser tab, and to one origin per test without
cy.origin. - Playwright (Microsoft, 2020) runs your test in Node, outside the browser. It holds one persistent connection per browser and can open many isolated contexts, pages and origins inside it. Because the automation layer is close to the browser, it can do the waiting there. Playwright ships its own builds of Chromium, Firefox and WebKit, so all three engines behave the same on every machine.
Remember it as: Selenium standardises, Cypress lives inside the page, Playwright drives from outside. Their strengths and limits follow from where they sit.
| Selenium | Cypress | Playwright | |
|---|---|---|---|
| Waiting | Explicit waits you write | Built-in retries on commands and assertions | Auto-wait on actions, retrying web-first assertions |
| Isolation | A new session per test is slow, so often shared | Clears state between tests in one browser | A new context per test, in milliseconds |
| Tabs, origins, users | Yes, via window handles | One tab; cy.origin for cross-origin |
Many pages, origins and contexts in one test |
| Engines | Real Chrome, Firefox, Edge, Safari | Chrome-family, Firefox; WebKit experimental | Chromium, Firefox, WebKit builds, plus branded Chrome/Edge |
| Languages | Java, Python, C#, Ruby, JS, Kotlin | JavaScript/TypeScript only | TS/JS, Python, Java, .NET (richest runner in Node) |
| Test code style | Imperative, blocking or async per language | Chained commands, queued (not promises) | Plain async/await |
| Parallel across machines | Grid (self-hosted or vendor) | Recording to Cypress Cloud | --shard and merge-reports, built in |
| Network and WebSockets | BiDi network interception, maturing | cy.intercept for HTTP; no WebSocket frames |
route, routeFromHAR, routeWebSocket |
| Debugging | Logs, screenshots, vendor dashboards | Time-travel command log, in the browser | UI mode, trace viewer, inspector |
| Component testing | Not its job | Mature, popular | Reworked in 1.62–1.63; still maturing |
How it works
The table is a consequence of position. Cypress runs inside the page, so it sees the app's objects directly, and it can't be in two tabs at once. Selenium speaks a standard protocol through a separate driver, so every browser vendor can implement it, and every command is a round trip. Playwright sits outside the browser but holds a rich, persistent channel into it, so it can do the waiting close to the page and open as many isolated contexts as a test needs.
In theatre terms, Selenium is the touring agreement every venue signs, so the show can play anywhere, with a formal message for every cue. Cypress puts the director on stage among the actors: a superb view of one scene, and only one scene at a time. Playwright's director sits in the control booth with a headset to every stage in the building.
Where Selenium is the better choice: - Your team writes Java, C# or Ruby, and wants tests in the same language as the product. - You need real Safari or real mobile devices through Appium and device clouds. Playwright's WebKit is the engine, not Safari itself. - You depend on a standard that browser vendors implement themselves.
Selenium Manager (since 4.6) has also removed the old pain of installing drivers by hand.
Where Cypress is the better choice: - The team lives in a single-page JavaScript app and values the in-browser command log. - You want to stub the app's own objects directly. - Component testing is central to your strategy.
Where Playwright is the better choice: - Long, branching, multi-locale journeys like Kestrel's, which need cheap isolation, several contexts and origins, and real WebKit and Firefox coverage. - Teams that want the runner, fixtures, projects, sharding, network control and tracing in one free package.
Migrating. If you are moving an existing suite, don't port it test by test. Write the scenario tables and the screen registry, generate the new suite from them, run the old and new suites side by side for a few weeks, and delete old tests as each scenario row replaces them. If you already run a Selenium Grid, Playwright can connect to it (experimental) while you move. The design in this book is the main part of the migration.
What Tomas learned
- The three tools differ mainly in where they sit relative to the browser, and their strengths and limits follow from that.
- Selenium wins on languages, standards and real devices, and Cypress on in-browser debugging and component testing. Playwright wins for complex multi-context journeys and an all-in-one runner.
- A migration is a redesign. Porting line by line keeps the old problems.
EpilogueEpilogue: The Reference#
Four appendices to keep open while you work:
- A. Concept Atlas: every Playwright feature worth knowing, in one line each, with a link to the docs.
- B. Anti-pattern Catalog: all 43 traps from the chapters, for code review.
- C. Baseline configuration: a playwright.config.ts and an ESLint setup to start from.
- D. Glossary for Selenium users: old words, new words.
A. Concept Atlas
The chapter column says where a concept is taught in depth. A dash means this row is the whole story, and the docs link has the rest.
| Area | Concept | What it is and when to use it | Ch. | Docs |
|---|---|---|---|---|
| Core | Auto-waiting and actionability | Actions wait until the element is visible, stable, enabled and not covered | 1 | actionability |
| Core | Web-first assertions | await expect(locator).toX() retries until true or timeout |
1 | test-assertions |
| Core | Browser contexts | Isolated, cheap browser profiles; one per test by default | 1 | browser-contexts |
| Core | Timeouts | Test 30 s, expect 5 s, actions unbounded; override per call or test.slow() |
1 | test-timeouts |
| Locators | Role, label, text, placeholder | getByRole, getByLabel, getByText, getByPlaceholder: locate as users perceive |
2 | locators |
| Locators | Alt text and title | getByAltText for images, getByTitle for title attributes |
– | locators |
| Locators | Test ids | getByTestId; attribute set by testIdAttribute |
2 | locators |
| Locators | Filtering and chaining | filter({ hasText, has, hasNot, visible }), .and(), .or(), chaining from a container |
2 | locators |
| Locators | locator.visible() |
Only visible matches, as a method (1.63) | – | class-locator |
| Locators | Strictness | Actions fail if a locator matches more than one element | 2 | locators |
| Locators | locator.describe() |
Human name for a locator in traces and reports | 2 | class-locator |
| Locators | Frames | frameLocator() or locator.contentFrame() to reach inside iframes |
12 | frames |
| Assertions | Soft assertions | expect.soft records a failure and continues |
6 | soft assertions |
| Assertions | expect.poll and toPass |
Retry any async value, or a block of code | 8 | expect.poll |
| Assertions | Custom matchers | expect.extend({ … }) for domain checks such as toVisitInOrder |
6, 7 | custom matchers |
| Assertions | expect.configure |
A pre-configured expect, e.g. expect.configure({ soft: true, timeout: 10_000 }) |
– | expect.configure |
| Assertions | Aria snapshots | toMatchAriaSnapshot compares the accessibility tree to YAML |
6 | aria-snapshots |
| Assertions | Visual comparisons | toHaveScreenshot pixel diffs; generate baselines in the Docker image |
6, 13 | test-snapshots |
| Assertions | Page assertions | toHaveURL, toHaveTitle; toHaveURL accepts a predicate |
– | page assertions |
| Structure | Page objects | Small classes of readonly locators and intent methods | 4 | pom |
| Structure | Fixtures | test.extend: named, composable setup and teardown |
5 | test-fixtures |
| Structure | Worker fixtures | { scope: 'worker' }: built once per worker process |
5, 11 | worker fixtures |
| Structure | Option fixtures | { option: true }, set from config or test.use() |
5, 9 | test-use-options |
| Structure | Auto fixtures | { auto: true }: run for every test |
5 | automatic fixtures |
| Structure | mergeTests |
Combine fixture sets from several modules | 5 | combine fixtures |
| Structure | Parameterised tests | A loop around test() or test.describe() |
6 | test-parameterize |
| Structure | Steps | test.step() names phases in reports; box, timeout, test.step.skip |
7 | test.step |
| Structure | Annotations | test.skip, test.fixme, test.fail, test.slow, custom annotation |
13 | test-annotations |
| Structure | Tags | { tag: '@smoke' }; select with --grep or project grep |
10, 13 | tag tests |
| Structure | Serial mode | test.describe.configure({ mode: 'serial' }): dependent tests; avoid |
8 | test-parallel |
| Structure | Test locks | test('…', { lock: 'name' }) stops tests that share a resource from running together (1.63) |
– | release notes |
| Config | Projects | One run per browser, locale or environment; dependencies order them |
3, 9 | test-projects |
| Config | webServer |
Start the app before tests; wait regex for readiness |
3 | test-webserver |
| Config | Global setup and teardown | globalSetup / globalTeardown scripts; prefer setup projects |
11 | global setup |
| Config | Emulation: devices | devices['iPhone 15'] sets viewport, user agent, touch |
– | emulation |
| Config | Emulation: locale and timezone | locale, timezoneId per project or context |
9 | emulation |
| Config | Emulation: geolocation and permissions | geolocation, permissions: ['geolocation'] |
– | emulation |
| Config | Emulation: colour scheme and media | colorScheme: 'dark', reducedMotion, contrast |
– | emulation |
| Auth | Storage state | Save and load cookies, local storage and IndexedDB | 11 | auth |
| Auth | Web storage APIs | page.localStorage / page.sessionStorage read and write directly (1.61) |
– | release notes |
| Auth | Passkeys | browserContext.credentials: a virtual WebAuthn authenticator (1.61–1.62) |
– | release notes |
| Network | Routing | page.route() / context.route(): fulfill, continue, abort, fetch-and-modify |
12 | network |
| Network | HAR | routeFromHAR to replay; tracing.startHar() to record |
12 | mock with HAR |
| Network | WebSockets | page.routeWebSocket() to mock or edit messages |
12 | routeWebSocket |
| Network | Waiting for traffic | waitForRequest, waitForResponse: create the promise before the action |
12 | network events |
| Network | API testing | The request fixture: HTTP calls sharing baseURL |
11 | api-testing |
| Browser | Pages and popups | context.newPage(), page.waitForEvent('popup') |
12 | pages |
| Browser | Downloads and uploads | waitForEvent('download'); setInputFiles() for uploads, locator.drop() for drag-in files |
12 | downloads |
| Browser | Dialogs | Auto-dismissed; page.on('dialog', d => d.accept()) to handle |
12 | dialogs |
| Browser | Overlay handlers | page.addLocatorHandler() for popups that appear at random |
8 | addLocatorHandler |
| Browser | Clock | page.clock.install(), setFixedTime(), runFor() to control time |
9 | clock |
| Browser | Cancellation | signal option with an AbortSignal on most operations (1.62) |
– | release notes |
| Browser | Console and errors | page.on('pageerror'), page.consoleMessages(), page.pageErrors() |
5 | class-page |
| Running | Parallelism | fullyParallel, workers, per-project workers |
3, 13 | test-parallel |
| Running | Sharding | --shard=1/4 plus blob reports and merge-reports |
13 | test-sharding |
| Running | Retries and flakiness | retries, flaky reporting, failOnFlakyTests, retryStrategy |
13 | test-retries |
| Running | Reporters | list, html, blob, github, junit, json, custom; --add-reporter |
13 | test-reporters |
| Running | CLI filters | --grep, -G, --last-failed, --only-changed, --repeat-each, --project |
13 | test-cli |
| Running | Docker image | mcr.microsoft.com/playwright:v1.63.0-noble, version-matched |
13 | docker |
| Running | CI setup | Generated workflows for GitHub Actions and others | 13 | ci |
| Tools | UI mode | --ui: watch, filter, time-travel through each action |
3 | test-ui-mode |
| Tools | Trace viewer | Actions, DOM snapshots, network and console for one run | 7, 13 | trace-viewer |
| Tools | Codegen and pick locator | Record tests and find locators by clicking | 2 | codegen |
| Tools | Inspector and debugging | --debug, page.pause(), VS Code breakpoints |
3 | debug |
| Tools | VS Code extension | Run, debug, record and pick locators from the editor | 3 | getting-started-vscode |
| Tools | Video and screencast | video: 'retain-on-failure'; page.screencast for annotated recordings (1.59) |
13 | videos |
| Tools | Accessibility testing | @axe-core/playwright scans a page for WCAG issues inside a test |
– | accessibility-testing |
| Tools | Component testing | Mount components in a real browser; the stories model replaced the experimental packages in 1.62–1.63 | 14 | test-components |
| Tools | Test agents | Planner, generator and healer agents for LLM-assisted test writing (npx playwright init-agents) |
– | test-agents |
| Tools | MCP | npx playwright mcp exposes a browser to AI assistants over the Model Context Protocol |
– | release notes |
B. Anti-pattern Catalog
Generated from the chapters' traps by scripts/traps-checklist.py. Each name links back to the full explanation.
Waiting, locators and structure (chapters 1–5)
| Trap (ch.) | Tell-tale sign | Do instead | Lint |
|---|---|---|---|
| Sleeping for the UI (1) | page.waitForTimeout() anywhere in a test. |
Wait for the thing you actually need, with a web-first assertion or page.waitForResponse() for a specific request. |
playwright/no-wait-for-timeout |
| Waiting for network idle (1) | waitForLoadState('networkidle') or waitUntil: 'networkidle'. |
Assert on the element the user would look at: await expect(heading).toBeVisible(). |
playwright/no-networkidle |
| Asserting a snapshot of the page (1) | expect(await locator.isVisible()).toBe(true) or expect(await locator.textContent()).toBe(…). |
Put the locator in expect() and the await outside it: await expect(locator).toBeVisible(). |
playwright/prefer-web-first-assertions |
| The forgotten await (1) | An expect(locator)… or locator.click() line with no await in front. |
Turn on the lint rules. | playwright/missing-playwright-await, @typescript-eslint/no-floating-promises |
| Selectors that mirror the DOM (2) | page.locator('div.card > .btn:nth-child(2)'), or any XPath with an index in it. |
Describe the element by role, label or text; scope with a container locator and filter() instead of structural paths. |
playwright/no-raw-locators |
| Dodging strictness with first (2) | .first(), .last() or .nth(n) added right after a strictness error. |
Make the locator unique: scope it to its container, filter by text or by a child, or add a name. | playwright/no-nth-methods |
| Holding element handles (2) | page.$(), page.$$(), elementHandle, or waitForSelector returning something you keep. |
Use locators everywhere; locator.all() if you truly need to iterate over matches. |
playwright/no-element-handle, playwright/no-wait-for-selector |
| Forcing the click (2) | { force: true } on a click or fill, usually with a comment like "flaky otherwise". |
Find out what fails the check. | playwright/no-force-option |
| Turning up the timeout (3) | timeout: 120_000 in the config, or expect.timeout raised globally "because CI is slow". |
Keep the defaults. | none |
| Home-made retry wrappers (3) | Helpers like clickWithRetry(), safeFill() or a retry(3, () => …) around actions, carried over from the Selenium framework. |
Delete the wrapper. | none |
| Committed focus (3) | test.only, test.describe.only or page.pause() in a pull request. |
Set forbidOnly: !!process.env.CI in the config, and lint for both. |
playwright/no-focused-test, playwright/no-page-pause |
| The god page object (4) | One class per application, or any page object over a few hundred lines, with methods for several screens. | One object per screen, one component object per reusable widget, each scoped to a root locator. | none |
| Page objects that return booleans (4) | Methods named isVisible…, hasError… or isLoaded returning Promise<boolean>, used inside expect(). |
Expose locators, such as readonly errorSummary: Locator, and assert with await expect(screen.errorSummary).toBeVisible(). |
playwright/prefer-web-first-assertions catches the call site |
| The god beforeEach (5) | A beforeEach that creates page objects, data and navigation for every test in the file, whether they need it or not. |
Move each piece into a fixture that tests name when they need it; keep hooks for truly universal, trivial setup, or make that an auto fixture. |
playwright/no-hooks |
| Inheritance trees (5) | class CheckoutTest extends QuoteTest extends BaseTest, or page objects extending a BasePage full of utilities. |
Compose. | none |
| Module-level shared state (5) | let quoteId: string; at the top of a spec file, set in one test and read in another. |
Put per-test state in a test-scoped fixture, and share only through worker-scoped fixtures or storage files designed for it (chapter 11). | none |
Data and long flows (chapters 6–8)
| Trap (ch.) | Tell-tale sign | Do instead | Lint |
|---|---|---|---|
| Copy-paste variants (6) | Several tests whose bodies are identical except for literals. | Move the literals into a typed scenario table and generate the tests from it. | none |
| One test, many rows (6) | A single test with for (const c of cases) { … expect … } inside it. |
Put the loop outside test(), so each row is its own test with its own name. |
playwright/no-conditional-in-test flags many of these |
| Try-catch around assertions (6) | try { await expect(…) } catch { … } to "check several things" or to tolerate a known failure. |
Use expect.soft to collect several failures, and test.fail() to mark a known bug honestly. |
playwright/no-conditional-expect |
| The fifty-step script (7) | A test that clicks through a long journey in a hardcoded order, often duplicated with small changes per branch. | Drive the journey with a loop over a screen registry and assert the visited path against the scenario's expectations. | none |
| Guessing the screen (7) | Code that probes several locators with isVisible() to work out which screen is showing. |
Read a screen identifier the app renders, such as data-screen, after a web-first wait for it to change. |
playwright/no-conditional-in-test when it happens in the spec |
| If it is there, click it (8) | if (await something.isVisible()) { … } in a test, a handler or a page object. |
Make the state deterministic (seed it, disable it through a flag or cookie), or handle it with page.addLocatorHandler. |
playwright/no-conditional-in-test |
| Walking the whole journey for every branch (8) | Every branch test starts at the first screen. | Seed and resume at the branch; keep one or two full walks per market to cover the route itself. | none |
| Chained serial tests (8) | test.describe.configure({ mode: 'serial' }) with test 2 continuing where test 1 left off. |
Make each test create its own starting state through seeding or fixtures; use test.step for readable phases within one test. |
none |
| Testing every combination (8) | Nested loops over every dimension, generating hundreds of near-identical scenarios. | Generate pairwise rows, then add the boundaries and the historically buggy cases by hand. | none |
Locales (chapters 9–10)
| Trap (ch.) | Tell-tale sign | Do instead | Lint |
|---|---|---|---|
| Hardcoded English in locators (9) | getByRole('button', { name: 'Continue' }) in a suite that runs in more than one language. |
Resolve names through the market's message catalog (market.t('nav.continue')), or use test ids for controls when there is no catalog to share. |
none |
| Hand-typed formatted values (9) | Expected strings like '£412.50', '01.10.2026' or '$1,286.00' in tests or data. |
Store the fact (412.5, an ISO date) and format it at assertion time with the market's Intl formatter. |
none |
| Trusting two ICUs to agree (9) | Expected values formatted in Node fail only for one locale, or only after a Node or browser upgrade. | Pin Node alongside the browser version, and where they still disagree, format the expected value in the browser with page.evaluate(n => new Intl.NumberFormat(…).format(n), amount). |
none |
| Locale ifs in specs (10) | if (market.id === 'de') (or locale.startsWith('ja')) in a test, a handler or a page object. |
Put the difference where it belongs: the app's routing plus a registry entry, a market override, or a local-answers row. | playwright/no-conditional-in-test for the specs |
| One spec file per market (10) | quote.gb.spec.ts, quote.de.spec.ts… with nearly identical bodies. |
One spec generated from scenarios, run by one project per market. | none |
| Skipping instead of not generating (10) | test.skip(market.id !== 'gb', 'UK only') at the top of many tests. |
Tag scenarios with the markets they apply to and let each project's grep select them. |
playwright/no-skipped-test |
Auth, network and CI (chapters 11–13)
| Trap (ch.) | Tell-tale sign | Do instead | Lint |
|---|---|---|---|
| Logging in through the UI every time (11) | Tests or beforeEach hooks that fill in the login form. |
Sign in once in a setup project, save storageState, and load it in the projects that need it. |
none |
| One account for every worker (11) | All parallel tests share one set of credentials, and failures depend on which tests happened to run together. | Create or pick an account per worker in a worker-scoped fixture, keyed on test.info().parallelIndex. |
none |
| Global setup for auth (11) | Signing in inside a globalSetup script. |
Use a setup project with dependencies; it is a real test. |
none |
| Driving someone else’s UI (12) | Tests that click through the payment provider's hosted pages, a social login, or an embedded map in every run. | Mock the boundary or use the provider's test hooks in most tests; keep one tagged end-to-end walk against the real sandbox. | none |
| Mocking everything (12) | Every API call in the suite goes to a route handler, including the app's own back end. | Mock third parties always and your own services only for hard-to-reach states; keep unmocked journeys in a real environment. | none |
| Routes nobody checks (12) | A page.route() mock with no assertion that the request happened or carried the right data. |
Pair each important mock with page.waitForRequest() or record requests in the handler, and assert on what was sent. |
none |
| Retries as a cure (13) | retries: 3 added after a flaky week, and nobody looks at the flaky tests in the report. |
Keep retries low, record traces on retry, review the flaky list like a bug queue, and turn on failOnFlakyTests once it is under control. |
none |
| A screenshot for every step (13) | page.screenshot() calls sprinkled through tests and page objects "for debugging". |
Use trace: 'on-first-retry' (or 'retain-on-failure') and screenshot: 'only-on-failure' in the config. |
none |
| Snapshots from the wrong OS (13) | toHaveScreenshot baselines generated on a developer's Mac and compared on Linux CI. |
Generate and compare baselines in the same Playwright Docker image, locally and in CI; keep visual tests few and targeted. | none |
| Debugging with console.log (13) | console.log(await page.content()) or logged locator counts left in tests. |
Use UI mode locally, --debug to step through with the inspector, and the trace viewer for CI failures. |
none |
Migrating (chapter 14)
| Trap (ch.) | Tell-tale sign | Do instead | Lint |
|---|---|---|---|
| Line-by-line migration (14) | A ticket per old test, "port to Playwright", with the old waits, helpers and inheritance translated faithfully. | Rebuild around scenarios, fixtures and the screen registry, and retire old tests as scenario rows cover them. | none |
| Wrapping Playwright in the old framework (14) | A DriverWrapper, WebActions or BaseTest layer that hides Playwright behind the old framework's method names, "so tests don't have to change". |
Let tests use Playwright's API directly, and put your abstraction at the domain level (screens, journeys, scenarios), not the browser level. | none |
C. Baseline configuration
A starting point for a suite like Kestrel's: four market projects behind a setup project, CI settings, and evidence on retry.
// file: playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
import type { TestOptions } from './tests/fixtures';
import type { MarketId } from './tests/kestrel/types';
import { markets } from './tests/markets';
const market = (id: MarketId) => ({
name: markets[id].locale,
grep: new RegExp(`@${id}\\b`),
dependencies: ['setup'],
use: {
...devices['Desktop Chrome'],
marketId: id,
locale: markets[id].locale,
timezoneId: markets[id].timezoneId,
storageState: '.auth/customer.json',
},
});
export default defineConfig<TestOptions>({
testDir: './tests',
fullyParallel: true,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
workers: process.env.CI ? 4 : undefined,
reporter: process.env.CI ? [['blob'], ['github']] : [['html', { open: 'never' }]],
use: {
baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
trace: 'on-first-retry',
screenshot: 'only-on-failure',
video: 'retain-on-failure',
},
webServer: process.env.BASE_URL ? undefined : {
command: 'npm run start',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
projects: [
{ name: 'setup', testMatch: /.*\.setup\.ts/ },
market('gb'), market('us'), market('de'), market('jp'),
],
});
And the lint rules that catch the traps mechanically:
// eslint.config.mjs
import { defineConfig } from 'eslint/config';
import tseslint from 'typescript-eslint';
import playwright from 'eslint-plugin-playwright';
export default defineConfig({
files: ['tests/**/*.ts'],
extends: [playwright.configs['flat/recommended']],
plugins: { '@typescript-eslint': tseslint.plugin },
languageOptions: { parser: tseslint.parser, parserOptions: { projectService: true } },
rules: {
'@typescript-eslint/no-floating-promises': 'error', // the forgotten await; use() without await
'playwright/missing-playwright-await': 'error',
'playwright/no-wait-for-timeout': 'error',
'playwright/no-networkidle': 'error',
'playwright/prefer-web-first-assertions': 'error',
'playwright/no-element-handle': 'error',
'playwright/no-wait-for-selector': 'error',
'playwright/no-force-option': 'error',
'playwright/no-focused-test': 'error',
'playwright/no-page-pause': 'error',
'playwright/no-conditional-in-test': 'error',
'playwright/no-conditional-expect': 'error',
'playwright/no-skipped-test': 'warn',
'playwright/no-raw-locators': 'warn',
'playwright/no-nth-methods': 'warn',
'playwright/no-hooks': 'warn',
},
});
D. Glossary for Selenium users
| Selenium | Playwright |
|---|---|
WebDriver / driver session |
Browser, and a BrowserContext per test |
| Window / tab handle | Page |
WebElement |
Locator (lazy, strict, never stale) |
By.css / By.xpath |
getByRole, getByLabel, getByText, getByTestId; locator() as a last resort |
findElements().get(n) |
filter() or chaining; nth() only when order is the requirement |
WebDriverWait / ExpectedConditions |
Auto-waiting actions and web-first assertions |
| Implicit wait | None needed |
StaleElementReferenceException |
Doesn't exist |
switchTo().frame() |
frameLocator() |
switchTo().alert() |
page.on('dialog') |
Actions (hover, drag) |
locator.hover(), locator.dragTo() |
DesiredCapabilities / Options |
use in the config, per project |
@BeforeMethod / BaseTest |
Fixtures |
@DataProvider |
A loop around test() |
| TestNG groups | Tags and --grep |
| Selenium Grid | Workers and --shard |
| BrowserMob / proxy | page.route() |
| Screenshot on failure listener | screenshot: 'only-on-failure' and traces |
driver.manage().getCookies() |
context.cookies(), storageState |