Skip to the content

Thesis

Why the usual measurements of main-thread lag in browsers are incorrect in the field, and the claims, the method, the evidence and the limits of a calibrated monitor.

Summary

  • Problem: The usual measurements of main-thread lag in browsers are incorrect in the field. Timer probes measure the timer granularity of the operating system. Clocks stop and jump. Hidden and frozen pages give samples that are not valid. The worst sessions report least.
  • Claim: A monitor can measure main-thread lag accurately in real browsers. It must calibrate each probe against its environment, validate each sample against the page lifecycle and the clocks, and measure from outside the main thread.
  • Method: We built @mark1russell7/lag and tested each part. We used unit tests, property tests, real-browser tests in three engines, a comparison with web-vitals, mutation testing and experiments on a real machine.
  • Result: On an idle page, the old timer probe reported 16 ms of lag for each 100 ms window in Chromium and 211 ms in Firefox and WebKit. The calibrated probe reports 0.1 ms, −0.3 ms and −0.4 ms. It measures a 300 ms block as 283 ms to 297 ms.

This page gives the argument. The concepts explain each mechanism, and the research gives the sources.

The problem

What lag is

A browser page has one main thread. The main thread does the JavaScript of the page, the style and layout calculations, the paint steps, and the input handlers. When one task continues for a long time, all other work waits. A click waits for its handler. A frame waits for its paint. The user sees a page that does not respond.

"Lag" is not one quantity. This project measures four different quantities:

QuantityQuestionMonitor
Blocked time in a windowHow much of the last 100 ms was the main thread busy?DriftLag
Queueing delayHow long does a task that arrives at a random time wait?Worker heartbeat
Frame delayHow late are the frames, and which script blocked them?Frame timing, LoAF
Input latencyHow long does an interaction wait for the next paint?Event Timing, Web Vitals

Why the usual measurements are not sufficient

The Interaction to Next Paint (INP) metric measures input latency. It shows nothing when the page blocks and nobody interacts. Long Animation Frames (LoAF) give accurate attribution, but only Chromium has them (browser support).

The classic "event-loop lag" probe is a chain of short timers. The probe measures how late the chain ends. This probe has five problems in the field:

  1. Timer granularity. A timer does not fire at the requested time. We measured a chained 5 ms timer on an idle page on Windows. It took 5.7 ms in Chromium and 16 ms in Firefox and WebKit. Those engines use the 15.6 ms timer tick of the system (experiment E2). A probe that thinks that each step takes 5 ms reports lag that does not exist.
  2. Clocks. performance.now() stops while the device sleeps, except on Windows. Safari calculates performance.timeOrigin again from the wall clock at each read. Date.now() can jump in the two directions (clocks and timers).
  3. Page state. Browsers throttle timers in hidden pages and stop all work in frozen pages. A sample that includes such an interval measures the browser policy, not the page.
  4. Coordinated omission. A chain of timers takes one sample for each block, also for a very long block. A 30 s block becomes one sample among hundreds, but it covers half of the time (statistics).
  5. Survivorship. A page that closes or crashes during a hang does not report the hang. Thus the worst hangs are the ones that a monitor loses (survivorship).

The export of the data also has a problem. Many browsers that write the same metric series without a writer identity break the metric store silently. The store rejects samples, or it mixes the totals of different browsers (OpenTelemetry metrics in the browser).

Claims

  1. Calibration. A timer probe is accurate only when it subtracts the idle behavior of its environment. The environment is the browser, the operating system and the power state. A sustained load must not change the calibration (refer to E5).
  2. Validation. Each sample must agree with the page lifecycle and with the clocks. A monitor must discard each sample that overlaps a hidden, frozen or suspended interval, and it must say why.
  3. An outside observer. A heartbeat from a Web Worker is the most robust estimator of the queueing delay in all engines. It is open-loop. It does not use the timers of the main thread. It can detect a hang while the hang continues. In Chromium and Firefox, it can also report the hang during the hang.
  4. The page view as the unit. Lag must have the same unit of analysis as the Core Web Vitals: the page view. Each event carries the ID of its page view.
  5. Correct aggregation. Field metrics must be counters and histograms with one writer identity for each SDK instance. Details with many values (selectors, URLs, IDs) go into events, not into metric attributes.
  6. Less survivorship bias. A monitor can report hangs that the page did not survive. In Chromium and Firefox, a journal that the worker keeps has them. In WebKit and Safari, a worker cannot write or send during a hang (E4, E6). There, another open page of the origin or the hung page itself at its close reports the hang (E7).

Method

Design

The library has one rule for each part: a part gets its dependencies from outside, and it uses only the dependencies that it needs. Each monitor is a class with duck-typed dependencies, for example a setTimeoutFn or a PerformanceObserver constructor. Each monitor has one factory that connects it to the metric catalog. The measurement conditions are one shared service. Refer to the architecture.

This design makes each part testable with fake time and fake browser APIs. It also makes the browser adapter (createBrowserDeps) the only place that reads the browser globals.

Evidence

We used eight types of evidence:

EvidenceWhat it showsWhere
Unit tests with fake timeThe rules of each part, with exact valuespackages/lag/src/*.test.ts
Property testsThe INP and CLS calculators for all sequences of entriesInpCalculator.test.ts, ClsCalculator.test.ts
Real-browser testsThe behavior in Chromium, Firefox, WebKit and Chrome, with real timers, real workers and real inputpackages/lag-integration-tests
Oracle testsThe agreement of the vitals with web-vitals 6 on the same pageweb-vitals-oracle.test.ts
CDP testsHidden and frozen pages, CPU throttling and compute pressurepackages/lag-integration-tests/src/cdp
Mutation testingThe changes to the code that no test findspnpm mutation
ExperimentsFacts about the engines that no document givesLocal experiments
ResearchThe specifications, the engine source code and the field data of othersResearch

The test strategy lists each test type. The results show the latest run.

Evidence

E1: listeners that run first

An exporter flushes when the page becomes hidden. The monitors must record their final values before that flush. At the target of an event, the DOM standard puts the capture listeners before the other listeners. We measured this in three engines.

Firefox and WebKit start the capture listeners first on document and on window. Chromium starts them first on document, but not on window. There, Chromium starts the listener that the page added first. pagehide has window as its target.

Result: the order of listeners cannot protect the final values. The library has an explicit hook, AllMonitorHandles.flush(). The OpenTelemetry setup uses it before each flush, through onBeforeFlush. Refer to page lifecycle.

E2: the timer granularity

We measured a chain of setTimeout(5) on an idle page, in the three engines from Playwright on Windows 11.

MeasurementChromium 145Firefox 146WebKit
Median step5.7 ms16 ms15 ms
Lag of the old probe, idle page, median15.8 ms211 ms211 ms
setTimeout(0) from the eighth repeat of setInterval5 ms16 ms16 ms
setTimeout(0) from a message task0 ms0 ms15 ms

Result: the old probe was incorrect by 16 ms to 211 ms on each window of an idle page. Also, a setTimeout(0) from a timer callback measures the clamp of nested timers, not the queue.

E3: the calibrated probe

DriftLag measures each step. It calculates the idle duration of one step (the baseline) from the recent steps. The baseline is the mean of the steps that are not longer than the median plus max(4 ms, half of the median). The lag of a window is its duration minus the number of steps multiplied by the baseline.

MeasurementChromium 145Firefox 146WebKit
Baseline step5.7 ms15.5 ms15.6 ms
Lag on an idle page, median0.1 ms−0.3 ms−0.4 ms
Lag after a 300 ms block297.4 ms282.7 ms294.6 ms

Result: the idle error went from 16–211 ms to less than 0.5 ms. A block is measured within one step of its duration. A browser test (drift-accuracy.test.ts) checks the two properties. The probes that use setTimeout(0) start from a message task at this time.

The clocks

The research on the engine source code (clocks and timers) gave four rules. The library obeys them:

  • Each context reads performance.timeOrigin one time (createAbsoluteClock). With one read, the absolute time of each context moves only with its monotonic clock, also in Safari.
  • The clock-drift monitor compares the wall clock with the absolute clock each second, with paired reads. It flags a change above max(50 ms, 3% of the interval). It classifies a forward change of 1 s or more as a suspend, and all other changes as a step. The lateness of its timer does not change the classification, because a hidden page can delay the timer by a minute.
  • The worker clock sync uses 8 exchanges. It keeps the exchange with the shortest round trip of all synchronizations, because the offset is constant when each context reads its origin one time.
  • A sample that overlaps a suspend is not valid. The worker gives the evidence: when its own timer is late by 5 s or more, the whole system stopped.

The worker observer

The worker sends a heartbeat from its own timer. The main thread acknowledges it. The delay until the main thread handles a heartbeat is the time that the main thread could not handle messages. In the browser tests, a 500 ms block gives heartbeat delays of more than 300 ms, while the timer of the worker stays on time.

When a heartbeat of the worker waits 5 s for its acknowledgement from the main thread, the worker reports a hang itself, through fetch with keepalive. It also writes the hang to a journal in IndexedDB, and writes it again each second.

A browser test blocks the main thread for 6.5 s. It stops the worker with terminate() before the hang ends, and it finds the record of the hang in IndexedDB. The next page of the origin reports such a record as a hang with the outcome abandoned. Two pages of the origin can start at the same time. Only one of them gets the record, because the journal removes it in one IndexedDB transaction.

E4: the input and output of a worker during a hang

The hang report and the journal need a worker that can send and write while the main thread is blocked. A browser test measures this in each engine. The worker starts an IndexedDB write or a fetch with keepalive, and the main thread blocks for 2 s. The table gives the time of the completion after the start of the block, or the number of operations that completed during the block. Safari 26.6.2 operated on a macOS runner of GitHub, on 8 October 2026. The other engines operated on Windows 11, on 7 and 8 October 2026.

OperationChromiumFirefoxWebKit (Playwright)Safari 26.6.2
IndexedDB write of a dedicated worker7 ms36 ms2015 msAfter the block
fetch of a dedicated worker19 ms14 ms2013 msAfter the block
Operations of a shared worker during the block18 of 2816 of 260 of 260
Operations of a service worker during the block18 of 2817 of 270 of 290 of 27

Result: WebKit and Safari complete the IndexedDB requests and the network requests of workers on the main thread of the page. A shared worker and a service worker do not help. Thus in WebKit and Safari the worker detects a hang, but it can send and write only after the hang. A page that does not survive the hang leaves no record. The test of the journal skips itself in these browsers, after the same measurement.

E5: a sustained load

A load can make all timer steps longer. On a thread that is busy with tasks of the same length, the steps agree with each other. Thus they look like a new timer granularity. With a load of 100% for 4 s, the calibrated probe of E3 reported less than 15% lag in each engine. In three of six cases, it reported approximately 0.

Thus the probe accepts a longer baseline only after a probe of message tasks shows an idle thread. On an idle thread, a message starts at once. On a busy thread, it waits for the current task.

The tests of this change found three behaviors of the engines. A probe changes the timing of the step and of the next step in Chromium. A message waits for the timer tick in the WebKit build for Windows. The steps of Firefox are irregular for some seconds after a load.

Load of 3 s, lag in the second halfChromium 145Firefox 146
Message tasks of 20 ms, before the changeApproximately 0 after 2.5 s0 after 2 s
Message tasks of 20 ms, after the change80% to 86%24% to 29%
Timer tasks of 30 ms, before the changeApproximately 012% after 3.5 s
Timer tasks of 30 ms, after the change78% to 85%53% to 72%

Result: the calibration follows the timer granularity, but not a sustained load. In CI, the change also works in Safari 26.6.2 on macOS (62% and 77%) and in WebKit on Linux (60% and 79%). In the WebKit build for Windows, the probe cannot see an idle thread, thus the probe there operates as before the change. A browser test (drift-load.test.ts) checks the result.

E6: the other output APIs of a worker

E4 left one question: can a worker of a blocked page write or send through another API? A browser test measures OPFS (three operations), Cache.put(), an XMLHttpRequest (asynchronous and synchronous) and a WebSocket message. The method is the same as in E4: the worker prepares the operation, and the main thread blocks for 2 s.

Operation, complete afterChromium and ChromeFirefoxWebKit (Playwright)Safari 26.6.2
OPFS write and flush through a sync access handle6 ms to 45 ms2 ms to 40 msNo OPFS in the Windows build2001 ms to 2002 ms
Cache.put()29 ms to 98 ms28 ms to 75 ms2011 ms to 2023 ms2004 ms to 2009 ms
XMLHttpRequest30 ms to 103 ms2073 ms2014 ms2005 ms
WebSocket message, at the server0 ms to 3 ms2000 ms to 2002 ms2016 ms to 2048 ms2001 ms

Result: in WebKit and Safari, each output API of a worker waits for the main thread of the page. In the source of WebKit, the worker side of each API posts its work to the main thread. Thus no worker API can keep a hang there. Firefox operates the XMLHttpRequest and the WebSocket of a worker through the main thread, but not its storage APIs or fetch.

E7: another page of the origin

Another page of the origin has its own main thread. A browser test opens a second top-level page. That page holds a Web Lock and sends heartbeats. Then it blocks its main thread, and the test closes it during the block.

MeasurementChromiumFirefoxWebKit (Playwright)Safari 26.6.2, macOSSafari 26.5, iOS Simulator
The other page operates, and the lock stays held during the blockYesYesYesYesYes
The lock comes free after the close0.5 s to 8 s later14 ms to 2.8 s laterWindows: less than 60 ms. Linux: 3 sAt the end of the block, also after 90 s5.3 s later
pagehide of the closed pageNoYes, at onceWindows: no. Linux: at the end of the blockYes, at the end of the blockNo

Result: another page sees the silence of a hung page, and the lock of the hung page shows that the page still exists.

The engines end a closed hung page in three ways. Chromium, Safari on iOS and the WebKit build for Windows stop it. Then another page can report the hang. Firefox stops the blocked script, and Safari on macOS lets the script continue to its end. Then the page closes normally, and it can report its own hang.

The peer hang watch uses both ways. A browser test in two pages expects the right report in each engine, also in Safari on macOS. Safari on iOS operates only the visible tab, thus the test cannot make a hung visible tab and a second tab that operates there.

The Web Vitals

The page-view vitals obey the rules of web-vitals 6.2. A browser test clicks a button. For 120 ms, the click handler blocks the main thread. The monitor gives an INP of 120 ms or more (119.99999999999955 ms one time in WebKit, whose clock moves in steps of 1 ms) and a processing phase of 110 ms or more. It gives the target #slow-button and the type pointer.

That test found an error in the first version of the attribution. In a real click, pointerdown is the longest entry, but the handler of click does the work. At this time, the attribution uses the frame groups of web-vitals. The processing phase spans all the events that the browser presented in the frame of the longest entry.

A second browser test makes a real soft navigation in Chrome 153, in a top-level page. The click that changes the URL gives the INP of the first page view. The second page view gets its LCP from its interaction-contentful-paint entry, a TTFB of 0 and no INP. On Windows and on Linux, Chrome delivered the entries of the click before the soft-navigation entry.

The oracle: agreement with web-vitals

A browser test operates web-vitals 6.2.3 and the page-view vitals on the same page, in Chromium, Firefox, WebKit, Chrome and Safari. The two libraries read the same entries. Thus the values must be the same, with no tolerance. This table gives one run in each browser, on 7 and 8 October 2026:

EngineFCPLCPTTFBINPCLS
Chromium108 = 108 ms108 = 108 ms3.5 = 3.5 ms184 = 184 ms0.018435 = 0.018435
Chrome148 = 148 ms148 = 148 ms4 = 4 ms192 = 192 ms0.018435 = 0.018435
Firefox312 = 312 ms312 = 312 ms9 = 9 ms184 = 184 msNo value in the two libraries
WebKit433 = 433 ms433 = 433 ms4 = 4 ms184 = 184 msNo value in the two libraries
Safari 26.6.2400 = 400 ms400 = 400 ms22 = 22 ms184 = 184 msNo value in the two libraries

The three phases of the INP attribution also agree exactly in the five browsers. The first version of the oracle found two differences. The library reported a CLS of 0 where the browser has no layout-shift entries. Also, its attribution used only the entries of the interaction, not all the events of the frame. Both differences are fixed.

Defects that the tests found

The tests found defects that the unit tests with fake time did not show:

  • An adversarial review. A review of the code wrote one proof test for each defect that it found, 29 tests in total. Examples are a second report of a page view after pagehide and a hang that two tabs reported two times. Another example is a hang that a stop lost. Each test failed first. At this time, the unit tests include all of them.
  • CLS in Firefox and WebKit. The oracle found false CLS values of 0 (refer to the oracle above).
  • A change of the timer granularity. In Chromium on Windows, the 5 ms steps changed from 5.5 ms to 15.6 ms after a page was shown again. DriftLag then reported windows of 160 ms of lag on an idle page, while the worker saw less than 2 ms. After the fix, DriftLag follows such a change. A CDP test found this defect.
  • A sustained load. A browser test found that DriftLag accepted a load of equal tasks as a new timer granularity (E5). The first version of the fix used the duration of a probed step as an idle duration. In Chrome, this gave 48 ms of lag in each window of an idle page. A second version lowered the confirmed value after one short window. In Firefox, this gave 6 ms to 10 ms of lag in each window.

The cost

The overhead benchmark measures an idle page in Chromium through CDP, with all monitors and without them:

MeasurementValue
CPU time of all monitors24.4 ms to 25.9 ms each second: 2.4% to 2.6% of one core
Script time10.4 ms to 11.3 ms each second
Timer callbacks177 to 179 each second
Wake-ups (timers, animation frames, idle callbacks, messages)Approximately 298 each second
DriftLag alone11.3 ms each second
Frame timing alone13.6 ms each second. Its animation-frame loop makes an idle page render each frame.

The budgets are 3% of one core, 210 timer callbacks and 340 wake-ups each second. All runs of the test agent were in the budgets.

The cost depends on the state of the machine. Later on the same day, the same benchmark measured 7.2% of one core, with the old code and with the new code. Then Chromium had the 15.6 ms timer tick of Windows. Each callback took approximately 4 times longer, possibly because of the power management of Windows. When the machine was in its usual state again, the final code measured 2.3% of one core and 179 timer callbacks each second.

A soak test of 10 minutes kept the heap after garbage collection at 9.50 MB to 9.58 MB. The growth was 10 kB each minute, and the limit for a leak is 100 kB each minute. After stop(), no timer, animation frame, idle callback or worker heartbeat stayed.

The quality of the tests

MeasurementFirst measurementAfter the new tests
Unit tests of @mark1russell7/lag458895
Statement coverage of @mark1russell7/lag96.1%99.3%
Branch coverage of @mark1russell7/lag89.0%98.2%
Mutation score of @mark1russell7/lag77.8% of 4095 mutants96.9% of 4690 mutants

The export works from end to end. On 8 October 2026, the E2E tests (20 tests) found the metrics of the test pages in Mimir, as native histograms. The pipeline check of grafana-infra passed 297 of 297 checks with its sample data.

A surviving mutant is a change of the code that no test finds. The first mutation run showed weak tests, for example for the sign of the delta of a web-vital event. The new tests find those mutants, and they found two more defects. The event catalog did not list lag.page_view.id for four events. Of two layout shifts with the same score, the library named a different shift than web-vitals.

A later run found five surviving mutants in the listener that makes the LCP final at a key press or a click. They were gaps in the tests, and new tests find them.

Each of the 141 mutants that survive has no effect that a user of the metrics can see (the mutation run of CI, 8 October 2026). Examples are a second guard that repeats a first one, and code that a browser cannot reach. The test strategy describes each kind of test, and the results give the newest run.

Limits

  • Engines. LoAF, CLS, visibility-state entries, the freeze event and Compute Pressure exist only in Chromium. Safari has no requestIdleCallback. The library uses each API only where it exists.
  • Test environment. The browser tests operate in the WebKit build of Playwright and in Safari 26.6.2 on a macOS runner of GitHub. They also operate in Safari 26.5 in the iOS Simulator. A simulator is not a phone: it has the CPU and the power of a Mac. A runner cannot turn on Low Power Mode, which aligns nested timers to 30 ms. Thus a unit test uses the timer rule of the WebKit source instead.
  • Sustained load. DriftLag needs a probe of message tasks to keep its baseline during a sustained load. Before the first check of the probe, a sustained load increases the baseline. This is also true in a browser in which a message waits on an idle thread. Then DriftLag reports less lag. The worker heartbeat measures that case correctly.
  • Resolution. The resolution of DriftLag is one step: 5.7 ms in Chromium and 16 ms in Firefox and WebKit on Windows.
  • Sampling. The worker sends one heartbeat each second by default. A block of B milliseconds delays a heartbeat with a probability of approximately min(1, B / 1000).
  • Hidden pages. The library discards samples from hidden pages. Thus it has no data on background work.
  • Clock steps. A forward step of the system clock of 1 s or more looks like a suspend.
  • Abandoned hangs. The journal works only for pages of the same origin, only where IndexedDB is available, and only after 30 s without updates. In WebKit, the journal and the hang report of the worker do not work for a page that does not survive its hang (E4, E6). There, the peer hang watch needs another open page on iOS. On macOS, it misses a hung script that does not end (E7).
  • A busy machine. The timing tests measure the page and the machine together. In one run of all browsers at the same time, on a busy machine, an idle Chromium page had a median lag of 5.0 ms. The same test alone gave less than 0.3 ms three times.
  • Cost. The monitors use approximately 2.5% of one core on an idle page, with approximately 300 wake-ups each second. The DriftLag chain and the animation-frame loop of the frame timing cost the most. This cost is real, especially on battery power.
  • CDP tests. The tests of hidden pages need Chromium in the new headless mode and an internal API of Playwright.

Future work

  • Duty cycles for the probes, so that the high-frequency probes run only some of the time.
  • An adaptive step size for DriftLag where the timer granularity is coarse.
  • Soft navigations in browsers other than Chromium, through the Navigation API.
  • A hang report from a worker in WebKit and Safari. The web APIs that we tested cannot do it (E4, E6), thus the change must come from WebKit. The peer hang watch covers most cases, but not a single page on iOS, or a hung script that does not end on macOS.
  • Tests on real devices: an iPhone, and a Mac on battery power in Low Power Mode.
  • Metrics from events at the collector (for example the signaltometrics component of Alloy), so that the histograms of each page view can change after the first report.
  • Segments by device class, for example from navigator.deviceMemory and navigator.cpuPerformance as resource attributes.

To change this page, edit packages/site/content/thesis/index.mdx.

lag: Main-thread responsiveness monitoring for browser apps, exported as OpenTelemetry metrics.

To change a page, edit its file in packages/site/content/. The writing style guide tells you how.

An AI model (Claude, from Anthropic) wrote most of the text and the code of this site and of the library, under the direction of the author. The tests and the STE linter examine them. The writing standard gives the reason for this note.