Skip to the content

Test strategy

The kinds of tests in the project, where they operate, what each kind shows, and how the results get to this site.

Each claim of the thesis has a type of test that can show it. A rule with an exact value gets a unit test with fake time. A fact about a browser gets a test in that browser. A claim of agreement with web-vitals gets a test that operates the two libraries on the same page. The results page shows the latest run of all tests.

The kinds of tests

KindWhere it operatesWhat it showsCommand
Unit testsNode, with the fake timers of VitestThe rules of each part, with exact valuespnpm test
Property testsNode, with fast-checkThe INP and CLS calculators agree with a reference for all sequences of entriespnpm test
Browser testsChromium, Firefox, WebKit and Chrome, through PlaywrightThe monitors with real timers, real workers, real input and real paintspnpm test:browser
Safari testsSafari on macOS, through safaridriver (WebDriver)The same browser test files in the real Safaripnpm --filter @lag/integration-tests test:safari
iOS testsSafari on iOS in the iOS Simulator, through safaridriverThe same browser test files in Safari on iOSpnpm --filter @lag/integration-tests test:ios, with a booted simulator
web-vitals oracleThe four enginesThe vitals agree with web-vitals 6 on the same pagepnpm test:browser
CDP testsChromium in the new headless modeHidden and frozen pages, CPU throttling and compute pressurepnpm --filter @lag/integration-tests test:cdp
Back/forward cache testsChromium with the back/forward cacheA page with the library stays in the cache while another page of the origin sends messagespnpm --filter @lag/integration-tests test:bfcache
Cross-origin-isolated testsChromium, Firefox and WebKitShared memory, the fine clock and measureUserAgentSpecificMemory()pnpm --filter @lag/integration-tests test:coi
Overhead benchmarkChromium in the new headless modeThe cost of all monitors on an idle page, against budgetspnpm test:overhead
Soak testChromium, for 3 minutes or moreThe heap does not grow without limit, and stop() releases each timerpnpm test:soak
E2E testsChromium and the Grafana stack in DockerThe export of the metrics to Mimirpnpm test:e2e
CoverageNodeThe parts of @mark1russell7/lag that the unit tests usepnpm coverage
Mutation testingNode, with StrykerThe changes to @mark1russell7/lag that the unit tests findpnpm mutation
Writing checksNodeThe STE rules for the README, the site and the comments. The metric table of the README agrees with the catalog.pnpm lint:ste, pnpm readme:metrics:check

Use pnpm test for the unit tests, and pnpm test:chromium for the browser tests in Chromium only. These commands do not need Docker.

Unit tests

Each monitor gets its browser APIs from outside, as duck-typed dependencies (refer to the architecture). Thus a unit test gives each monitor fake timers, a fake clock and fake observers. The tests are fast, and they give the same result each time.

The fakes copy the behavior of the browsers that the research found. For example, the fake PerformanceObserver keeps entries for takeRecords(), and the fake lifecycle sends the events in the sequence of Chromium. Each test examines one behavior, and its name states that behavior.

An adversarial review of the library wrote a proof test for each defect that it found. The unit tests include these tests. Each one first failed, and passes after its fix. Examples are a second report of a page view after pagehide, and a hang that two tabs reported two times.

Browser tests

The browser tests use Vitest browser mode with Playwright. Each test file operates in Chromium, Firefox, WebKit and Chrome. On macOS, the same files also operate in Safari, through safaridriver (the browser (safari) project, with LAG_SAFARI=1). They also operate in Safari in the iOS Simulator (the browser (ios) project, with LAG_IOS=1). LAG_IOS_UDID selects the booted simulator. The tests examine the facts that fake time cannot show:

  • browser-facts.test.ts: the facts of the research about each engine. These are the supported entry types, and the order of the listeners at window (experiment E1). The test fails if an engine changes, and then the documents need an update.
  • soft-navigation.test.ts (Chrome): a real soft navigation in a top-level page. The command clickInTopLevelPage opens the page and clicks with Playwright, because Chrome detects soft navigations only in the top-level frame.
  • drift-accuracy.test.ts: the median lag of DriftLag on an idle page is less than 5 ms, or 10 ms in a window without focus. A 300 ms block gives more than 300 ms minus one step and 10 ms (20 ms without focus), and less than 340 ms.
  • lifecycle.test.ts: the lifecycle machine notifies its subscribers before a listener that the page added earlier. It uses the time of the event, and it removes its listeners at the end.
  • lag-monitors.test.ts: the monitors on a real page, and the INP of a real click on a button that blocks for 120 ms.
  • worker.test.ts and hang-journal.test.ts: the worker heartbeat, a hang of 6.5 s, and the record of the hang in IndexedDB.
  • worker-io.test.ts: the IndexedDB writes and the fetches of a worker during a block of 2 s. WebKit completes them only after the block, thus hang-journal.test.ts skips itself there after a probe (experiment E4). The test also measures OPFS, the Cache API, XMLHttpRequest and a WebSocket of a worker, with the expectations of each engine (experiment E6). A WebSocket server in Node (commands/probe-socket.ts) records when the message arrives.
  • peer-tab.test.ts (experiment E7) and peer-hang-watch.test.ts: a second top-level page that blocks its main thread, and that the test closes during the block. The commands of commands/peer-page.ts open the page with Playwright, or as a new window with WebdriverIO in Safari. The test page watches the other page through BroadcastChannel and the Web Locks API, and the peer hang watch of the library reports the abandoned hang.
  • stress.test.ts: the workload profiles of @lag/load (refer to stress tests).
  • browser-apis.test.ts: the monitors that need one browser API, with the real API. Examples are the layout shifts, the real CPU pressure source, the visibility-state entries, the idle periods and performance.memory.

A test that needs an API of one engine skips itself in the other engines. The skip reason comes from features.ts, for example "Long Animation Frames (the long-animation-frame entry type) are Chromium-only." The results page shows each skip with its reason.

NoteOne file at a time in each engine

The timing tests measure milliseconds. When 28 test files operated at the same time in four engines, the CPU was overloaded. A 300 ms block then measured 245 ms. Thus each engine operates its files one after the other, and the four engines operate at the same time.

The web-vitals oracle

web-vitals-oracle.test.ts operates web-vitals 6.2.3 and the page-view vitals on the same page. The two libraries read the same entries with the same rules. Thus the values must be the same, with no tolerance:

  • FCP, LCP and TTFB of the load.
  • CLS in Chromium. In Firefox and WebKit, the browser has no layout-shift entries, and the two libraries report no CLS.
  • INP of real clicks on buttons that block for 40, 100 and 180 ms, with the target, the type and the three phases of the attribution.

The oracle found a bug. The library reported a CLS of 0 in browsers without layout-shift entries. Those values were false "good" values in the CLS histogram. After the fix, the library reports no CLS there, as web-vitals does.

CDP tests

The CDP tests use the Chrome DevTools Protocol to make page states that a user makes:

  • A hidden page: the timer monitors pause, and they continue when the page is visible.
  • A frozen page: the timer monitors pause before the freeze, record no stall for it, and continue after the resume.
  • CPU throttling: 4× throttling of the same work increases the drift and the worker lag.
  • A virtual compute-pressure source: the monitor records each state of the source, as the ordinals 0 to 3.
  • A checkpoint: the page-view vitals record one value for each vital, and the events, when the page becomes hidden.

CautionThe CDP tests need Chromium in the new headless mode

Only that mode changes visibilityState when another page comes to the front. The tests also turn off the focus emulation of Playwright, through an internal API of Playwright 1.58. A new version of Playwright can break these tests.

CDP hides a page before it freezes it. Thus the monitors pause at the hidden transition, and no sample gets the discard reason frozen in these tests.

The CDP tests found a second bug. In Chromium on Windows, the 5 ms timer steps changed from 5.5 ms to 15.6 ms after a page was shown again. DriftLag then reported windows of 160 ms of lag on an idle page, and the worker saw less than 2 ms. After the fix, DriftLag follows a change of the timer granularity (refer to DriftLag).

A browser test found a third bug. DriftLag accepted a sustained load of equal tasks as a new timer granularity. Thus the lag of the load decreased to less than 15% within 4 s (experiment E5). The fix uses a probe of message tasks.

The browser test drift-load.test.ts then found two defects of the fix, with all four browsers at the same time. In Chrome, a probed step was shorter than an idle step. In Firefox, the steps were irregular after a load. The unit tests of DriftLag.load.test.ts simulate these behaviors on a simulated thread.

Overhead and soak

The overhead benchmark measures an idle page with all monitors and without them. It uses Performance.getMetrics of CDP, in 3 rounds of 10 s. The budgets come from these measurements and from the research:

BudgetLimitThe reason for the limit
CPU of all monitors3% of one core (30 ms each second)The largest measured round, plus 20%
Timer callbacks210 each secondA 5 ms chain fires 200 times each second at most. The other monitors add fewer than 10.
Wake-ups340 each second210 timers, 60 animation frames, 60 idle callbacks and 10 messages

The budgets depend on the machine. Thus the benchmark operates in pnpm results, but not in the CI workflow.

The soak test operates all monitors under a mixed workload for 3 minutes, or for LAG_SOAK_MS milliseconds. It measures the heap after a full garbage collection. A growth of more than 100 kB each minute after the warm-up is a leak. After stop(), no timer, animation frame, idle callback or worker heartbeat must stay.

Coverage and mutation testing

pnpm coverage measures the coverage of @mark1russell7/lag with V8. The command fails below the thresholds in packages/lag/vitest.config.ts.

pnpm mutation uses Stryker. Stryker makes many small changes to the code, for example - to +. It operates the unit tests on each change. A change that no test finds is a surviving mutant: a behavior that no test pins down. The command fails below the break threshold in packages/lag/stryker.config.json. The incremental mode tests again only the mutants whose code or tests changed.

Stress tests and workload profiles

The stress tests use the workload profiles of @lag/load to make main-thread load. A profile is a list of lag specs. Each spec has a weight and a distribution of durations.

Lag event durations of the heavy profile

heavy-cpuspikegc-pressurelayout-thrashloaf-long2004006008001,000Lag event duration (ms) →
Show the data as a table
Samples of each lag spec of the heavy profile (400 each, seed 42)
Lag specWeightp50p95Maximum
heavy-cpu596 ms189 ms244 ms
spike179 ms313 ms1 s
gc-pressure241 ms67 ms81 ms
layout-thrash261 ms99 ms145 ms
loaf-long1120 ms120 ms120 ms
Samples of the duration distribution of each lag spec in the heavy profile of @lag/load, with seed 42. The page only draws the numbers. It makes no load.

How the results get to this site

pnpm results collects the results:

  1. It builds all packages.
  2. It operates the unit tests with coverage, the browser tests, the overhead benchmark and the site tests. A reporter writes each test with its project and its engine.
  3. It reads the coverage summary, if the unit tests of this run wrote it. With --skip-unit, the run has no coverage. It also reads the newest Stryker report if one exists, with the commit and the time of its Stryker run.
  4. It writes one run to packages/site/public/data/results/. Git ignores this folder.

Some failures have no failed test. A beforeAll hook that fails skips the tests of its suite. An afterAll hook that fails keeps them passed, and an error at the import of a file gives no tests. Thus the reporter adds a failed test with the name Errors outside the tests for each file or suite with such errors. The command fails when a test or a budget failed, or when a step gave an exit code that is not 0.

The results page reads this folder. It shows the tests of each engine, the skips with their reasons, the coverage, the surviving mutants, the measurements and the budgets. A package with only skipped tests, or without tests, does not get the passed icon. A run with a failed budget shows as failed in the list of runs and on the home page, as in the collector.

The mutation report comes from the weekly mutation workflow, thus it is usually older than the run. The mutation view shows the commit and the date of the Stryker run. The Pages workflow also takes a run whose score is below the break threshold, thus a score that falls shows on the site.

Continuous integration

WorkflowWhenWhat
ci.ymlEach push and each pull requestThe build, the type check, the unit tests, the coverage, the browser tests in each engine, Safari on a macOS runner, Safari in the iOS Simulator, and the tests and the build of the site
e2e.ymlEach week, or manuallyThe Grafana stack in Docker, and the E2E tests
mutation.ymlEach week, or manuallyStryker. A manual run uses the incremental file of its branch, or of main, from the cache. The weekly run tests all mutants again (--force).
pages.ymlEach push to main, or manuallySafari and Safari in the iOS Simulator on a macOS runner, then pnpm results with their reports and the mutation report of the newest completed mutation.yml run of main, the build of this site and its deployment to GitHub Pages

Limits

  • Safari on iOS. The tests operate in Safari on iOS only in the iOS Simulator of a macOS runner. A simulator has the CPU and the power of the Mac, not of a phone. To open the first session, safaridriver needed up to 131 s and more than one try. Thus the iOS project has a connect timeout of 600 s.

    In the simulator, the click of safaridriver gives no INP, also in web-vitals. Thus the two INP tests skip themselves there.

  • Low Power Mode. A macOS runner refuses Low Power Mode on AC power. Thus no browser test operates in Low Power Mode. DriftLag.power.test.ts uses the timer rule of the WebKit source instead: a grid of 30 ms for nested timers.

  • Safari on a shared runner. Safari on a macOS runner of GitHub sometimes slows the timers of its window, for example by 9 s. Then a timing test fails for a reason that is not in the code. Thus the Safari job of CI reports its results, but it does not block a pull request.

    A stress profile that the browser slowed skips itself with the reason. The profile is slow if it took more than 1.5 times its time. It is also slow if a timer delay was more than 2 s while the worker saw less than 100 ms of block. A third case is a time without animation frames that is more than 2 s longer than the longest block. Safari on a runner sometimes renders no frames for its window. Then a load generator that waits for a frame stops the profile.

  • Safari and the collector. Only macOS has Safari. Thus a macOS job exports its reports (--export-reports), and the Linux job of the site imports them (--import-reports). The converter removes the checkout folder of the macOS runner from the file names, thus a test file has the same name in each engine.

  • The E2E tests need a free port. The Grafana stack uses port 3000 by default. Set GRAFANA_PORT when another program uses that port.

  • Mimir labels. Mimir puts service.name into the job label. The E2E query accepts job or service_name, and a native or a classic histogram.

  • Mimir and CORS. Mimir sends no CORS headers, thus a test page cannot read the answer of a query. A Vitest command sends the query from Node (commands/mimir.ts).

  • Shared machines. The CI runners of GitHub are shared machines. In the first CI operation, on 8 October 2026, a runner stopped a page for 490 ms. Thus a test of an idle page or of a light load examines a median or a percentile, not the largest single window.

lag: Main-thread responsiveness monitoring for browser apps, exported as OpenTelemetry metrics.

To change a page, edit its file in packages/site/content/. The writing style guide tells you how.

An AI model (Claude, from Anthropic) wrote most of the text and the code of this site and of the library, under the direction of the author. The tests and the STE linter examine them. The writing standard gives the reason for this note.