Skip to the content

Survivorship bias

Why the worst sessions report least, how the library reports a hang that the page does not survive, and the limits of these reports.

Note

This page is a draft. The content is not complete and can change.

A monitor on the main thread can send data only while the main thread operates. A hang stops the main thread. If the page closes during the hang, or the browser stops the page, the monitor does not send the samples of the hang. Thus the longest hangs are the ones that such a monitor loses. This page tells you how the library reports these hangs, and where it cannot.

Why the worst sessions report least

Two effects remove the worst sessions from the data (clocks and timers):

  • Abandonment. Users leave slow pages. In data of Chrome on Android, each second more of First Contentful Paint added 2.6 percentage points of abandonment, to 17.3% at 4 s. The visitors with the worst experiences are "phantom bounces" that do not report.
  • Lost beacons. A page that sends its data at the end can lose it. In a study of approximately 2 million Chrome page views (May 2023), sendBeacon at pagehide or visibilitychange was 95.8% reliable. Beacons at load, pagehide and visibilitychange together got to 98%, and an XMLHttpRequest got to 84.9%.

The usual monitors learn about a block only after its end. PerformanceObserver callbacks, timers and animation frames all start on the main thread, thus they wait for the end of the block. A page that does not recover, or a tab that the user closes during a hang, sends nothing. In the survey of RUM tools, only the server-side crash report of Chrome for an unresponsive page sees this case (RUM landscape).

What the library does

The library has four parts that decrease this bias. The worker reports a hang while the hang continues. The hang journal lets the next page report a hang that the page did not survive. In WebKit and Safari, these two parts do not operate during a hang (refer to the limits). The crash-report context connects the crash reports of Chromium with the page view. The exporter flushes each time that the page becomes hidden.

The worker reports a hang while it continues

The worker monitor sends a heartbeat from its own timer, and the main thread acknowledges each heartbeat. When no acknowledgement comes for 5 s, the worker starts a hang itself. The start of the hang is the time of the last acknowledgement.

Show the diagram source
sequenceDiagram
    participant M as Main thread
    participant W as Worker
    participant J as Hang journal
    participant C as Collector
    W->>M: heartbeat
    M-->>W: ack
    Note over M: a long task starts
    W->>M: heartbeats wait in the queue
    Note over W: 5 s without an ack
    W->>C: hang started (fetch with keepalive)
    W->>J: put the record
    loop each second while the hang continues
        W->>J: put the record again
    end
    Note over M: the long task stops
    M-->>W: ack, or stop when the page became hidden
    W->>J: remove the record
    W->>C: hang ended (fetch with keepalive)
    W->>M: hang-ended
    Note over M: lag_main_thread_hangs, outcome ended
The worker reports the start of the hang while the main thread is blocked. The first message of the main thread after the block stops the hang.

The bundled worker of @mark1russell7/lag/worker sends each hang report as an OTLP/HTTP JSON log record, with the event name lag.main_thread.hang. The record has the context of the page, the phase (started or ended) and duration_ms. A worker has no navigator.sendBeacon(), thus the worker uses fetch() with keepalive: true. With keepalive, the request can complete after the page closes. In WebKit and Safari, the request completes only when the main thread operates again (experiment E4). The Fetch standard limits the body of all keepalive requests in progress to 64 KiB (RUM landscape).

The worker sends the reports only when you give the target, with workerHangReport:

import { createInstanceId, init } from "@mark1russell7/otel-ts";
import { createBrowserDeps, createOtelLoggerAdapter, setupAllMonitors } from "@mark1russell7/lag";
import { createLagWorker } from "@mark1russell7/lag/worker";

const endpoint = "http://localhost:4318";
const serviceInstanceId = createInstanceId();
const otel = init({ serviceName : "shop", serviceInstanceId, endpoint });

const monitors = setupAllMonitors(createBrowserDeps(window, {
    meter : otel.getMeter("lag"),
    logger : createOtelLoggerAdapter(otel.getLogger("lag")),
    worker : createLagWorker(),
    // The worker sends its hang reports to this OTLP endpoint, with the identity of the page
    workerHangReport : {
        url : `${endpoint}/v1/logs`,
        resource : { "service.name" : "shop", "service.instance.id" : serviceInstanceId },
    },
    // More attributes for each hang report, for example the session ID
    pageContext : () => ({ "session.id" : otel.getSessionId() }),
}));
otel.onBeforeFlush(() => monitors.flush());

The worker sends the reports without the OpenTelemetry SDK of the page. Give workerHangReport.resource the same service.instance.id as the page. Then a query can connect the reports with the metrics of the page. The worker adds the attributes of pageContext to each report, and the ID of the current page view. Refer to OpenTelemetry setup.

When the main thread operates again, it acknowledges the heartbeats that waited. The first acknowledgement stops the hang. The worker removes the journal record, sends the report of the phase ended, and tells the main thread. The main thread then records the hang with the outcome ended, and it sends the event lag.main_thread.hang.

A stop or a start message also stops the hang, because the main thread sent it. For example, the main thread stops the worker monitor when the page becomes hidden. That can occur before it handles the heartbeats that waited. While the monitor is stopped, the main thread still handles hang-ended, thus the hang counts.

When the timer of the worker itself is late by 5 s or more, the worker did not operate. Then the worker resets the time of the last acknowledgement, and it does not start a hang. A system suspend thus does not give a hang. During a hang, the worker also moves the start of the hang forward by that time. Thus a sleep on Windows, in which performance.now() continues, does not count as time of the hang. Refer to measurement validity.

A unit test of setupAllMonitors() blocks the main thread for 8 s. It makes sure that the worker reports the start of the hang during the block. After the block, the main thread records one hang of 5 s or more.

The hang journal and the outcome abandoned

The worker keeps a record of each hang in progress in IndexedDB, in the hang journal. A record has these properties:

  • pageId: the random ID of the page instance whose main thread hung.
  • startedAt: the time of the last acknowledgement before the hang.
  • lastSeenAt: the last time that the worker saw the hang continue.
  • attributes: the context of the page at the time of the hang, for example lag.page_view.id.

The times are wall-clock times (Date.now()). Other pages compare them with their own time, and the wall clock is the only clock that all pages share. The monotonic clock of a page stops while the device sleeps, except on Windows.

The worker writes the record when the hang starts, writes it again each second while the hang continues, and removes it when the hang stops. A stop of the monitor also stops the hang, thus it also removes the record. A failed operation of the journal does not stop the heartbeats.

Show the diagram source
sequenceDiagram
    participant W as Worker of page 1
    participant J as Hang journal
    participant P as Page 2 of the same origin
    W->>J: put the record of the hang
    Note over W: page 1 closes during the hang
    Note over J: no update for 30 s or more
    P->>J: list all records, one time at the start
    J-->>P: the record of page 1
    P->>J: take the record, only if it is still stale
    J-->>P: the record, to one page only
    Note over P: lag_main_thread_hangs, outcome abandoned
The next page of the same origin takes the record that nobody updated for 30 s, and reports it as abandoned. Only one page can take a record.

The pages and the workers of one origin share the IndexedDB database. When a new page starts the worker monitor, the monitor reads the journal one time. It selects the records of other pages that nobody updated for 30 s or more. For each of these records, it does these steps:

  1. It takes the record with take(pageId, latestSeenAt). The method removes the record and gives it, but only if the record is still stale. IndexedDB does the read and the removal in one read-write transaction. Thus only one page gets the record, also when two pages start at the same time, for example at a session restore. The page skips a record that it did not get.
  2. It records the hang with the outcome abandoned. The duration is the time from startedAt to lastSeenAt, thus the time until the worker saw the hang for the last time.
  3. It sends the event lag.main_thread.hang with the phase abandoned. The event has the attributes of the record, lag.hang.page_id, and lag.hang.source: "journal".

If the monitor stops after a take, the page puts the record back for the next page. A unit test starts two pages at the same time. It makes sure that the two pages report the hang one time, with a journal in memory and with IndexedDB.

The peer hang watch: another page, or the page at its close

The journal needs a worker that can write during the hang. WebKit and Safari do not let a worker write or send during a hang. The peer hang watch uses two other ways. The engines end a hung page that the user closes in different ways, and each way leaves one of them (experiment E7, local experiments):

  • The browser stops the page: Chromium and Safari on iOS. Then another open page of the origin reports the hang. A visible page holds a Web Lock and sends a heartbeat each second through a BroadcastChannel. Each other page asks for that lock, and the browser releases it when it stops the page. A lock that comes free after 5 s or more without a heartbeat gives a report with lag.hang.source: "peer".
  • The page operates again before it closes: Firefox stops the blocked script, and Safari on macOS lets the script continue to its end. Then the page itself reports its hang at its close, with lag.hang.source: "self". Its heartbeats show the hang: none for 5 s or more before the close.

The report of another page comes at once, not at the start of a later page. In Chromium, that page also takes the journal record of the hung page, thus the next page does not report the hang again. The own report goes out in the last export of the page.

In Firefox, the hung page reports the hang itself, and the other page reports nothing. The worker of the hung page wrote a record during the hang, and Firefox can stop the worker before it removes the record. The page cannot wait for IndexedDB at its close. Thus it tries to remove the record, and it also writes a mark in localStorage (key lag-hang-reported:<page ID>). The journal reader of a later page takes the record of a marked page, but it does not count the hang again. A peer report without the record also writes a mark, for example when the take fails.

MetricKindUnitAttributesDescription
lag_main_thread_hangsCounter{hang}outcomeThe number of main-thread hangs that the worker detected. In a hang, the main thread does not acknowledge heartbeats. The outcome `abandoned` means that the page closed or crashed during the hang. The next page of the origin reports it from the hang journal. Another open page of the origin, or the page itself at its close, can also report it (PeerHangWatch).
lag_main_thread_hang_duration_histogramHistogrammsoutcomeThe duration of each main-thread hang. For an abandoned hang, the duration until the worker saw the hang for the last time. Without a record of the worker, it is the duration from the last heartbeat of the page to its end.

createBrowserDeps() gives the journal when the page has IndexedDB, and a worker or the peer hang watch. Only the worker monitor reads the journal at its start. The peer hang watch only takes the record of a hung page. With hangJournal: false, the monitor does not read the journal, and it sends no page ID to the worker. Without a page ID, the worker writes no records.

A browser test blocks the main thread for 6.5 s and stops the worker before the end of the block, as when the page closes. Then it finds one record in IndexedDB, with a duration of 5 s or more and the ID of the page view. The test skips itself where a probe finds that the IndexedDB requests of a worker wait for the main thread, as in WebKit and Safari. The skip comes from that measurement, not from the name of the browser. The browser test worker-io.test.ts measures the behavior of each engine, also of a service worker, and it fails if an engine changes.

The crash-report context

Chromium can stop an unresponsive page and send a crash report to the reporting endpoint of the page. From Chrome 145, the page can add keys and values to such reports with window.crashReport. The page-view context sets the key lag.page_view.id there at each new page view. Thus the server can connect a crash report with the events and the Web Vitals of the same page view. Refer to page views.

Flush on hide

otel-ts exports all telemetry each time the page becomes hidden. At pagehide, it exports again: a page that goes into the back/forward cache keeps its providers, and a page that unloads stops them. The otel-ts library flushes in the export phase of the lifecycle tracker that the page shares. Thus the monitors record each transition before the export. With onBeforeFlush(() => monitors.flush()), these exports also contain the latest values of the page view. Refer to the flush hook.

The library does not use unload. From Chrome 154, Chrome does not start unload listeners (browser support). The README of web-vitals gives the same advice: send the data when the page goes into the background or unloads. Refer to Web Vitals algorithms.

The metrics use cumulative temporality, thus an export that is lost does not lose data. The next export contains the totals. Only the data after the last successful export is lost (OpenTelemetry metrics in the browser).

Limits

  • WebKit and Safari. WebKit and Safari 26.6.2 complete the IndexedDB requests and the fetch requests of a worker on the main thread of the page (experiment E4). The other output of a worker also waits: OPFS, the Cache API, XMLHttpRequest, WebSocket and BroadcastChannel (experiments E6 and E7). The worker detects the hang, because its timer continues. But the hang report and the journal record complete only when the main thread operates again. A service worker also waits, thus it cannot keep the journal either.

    Thus there, only the peer hang watch keeps a hang that the page does not survive. On macOS, the page reports the hang itself when its script ends. On iOS, another open page of the origin must report it, and Safari there operates only the visible tab. With one page on iOS, or with a script that does not end on macOS, the hang is lost.

  • Firefox closes a hung page normally. When the user closes a hung page, Firefox stops the blocked script, and the page closes normally (experiment E7). Then the page reports its own hang at its close. The worker monitor can also count the hang as ended before the close. Then the page does not report it again.

  • A frozen tab. If the browser freezes the page and its worker during a hang, the worker cannot report or update the record. Chrome on Android freezes background pages and their workers after 1 minute, from Chrome 139 (browser support). After 30 s without an update, the next page reports the hang as abandoned, also if the frozen page continues later.

  • The stale time of 30 s. A record counts as abandoned only after 30 s without an update. The monitor reads the journal only one time, at its start. Thus a page that starts less than 30 s after the hung page closed does not report the record, and a later page must report it. If no page of the origin starts the worker monitor again, nobody reports the hang.

  • The same origin only. IndexedDB belongs to one origin. Only a later page of the same origin, with the worker monitor and the journal, can report the hang.

  • No localStorage. Without the marks, a later page can count a hang from the journal that the hung page reported already. For example, a sandboxed frame cannot use localStorage.

  • Private windows. The journal needs IndexedDB. The code comments name some private windows as an example of a browser that does not permit IndexedDB. Then each operation of the journal fails, the worker continues without the journal, and the monitor logs a warning. We did not examine this case in each browser.

  • Short hangs. A hang starts when a heartbeat waits 5 s for its acknowledgement. A page that closes during a block of less than 5 s gives no report and no record.

  • The duration of an abandoned hang. The duration stops at the last write of the record. The worker writes the record at a heartbeat, at most one time each second. Thus the duration is too short by up to 1 s, or by up to one heartbeat interval if the interval is longer.

  • The wall clock. The stale test compares the wall-clock time of the new page with the times that the worker of the earlier page wrote. A step of the system clock between the write and the read changes the age of a record (clocks and time). We did not test this case.

  • No target. Without workerHangReport, the worker does not send reports during the hang. The library then reports a hang only when it stops, or from the journal.

  • Crash reports. Only Chromium sends crash reports, and only to the reporting endpoint that the page declares. A script cannot read them (browser support).

lag: Main-thread responsiveness monitoring for browser apps, exported as OpenTelemetry metrics.

To change a page, edit its file in packages/site/content/. The writing style guide tells you how.

An AI model (Claude, from Anthropic) wrote most of the text and the code of this site and of the library, under the direction of the author. The tests and the STE linter examine them. The writing standard gives the reason for this note.