Survivorship bias
Why the worst sessions report least, how the library reports a hang that the page does not survive, and the limits of these reports.
Note
A monitor on the main thread can send data only while the main thread operates. A hang stops the main thread. If the page closes during the hang, or the browser stops the page, the monitor does not send the samples of the hang. Thus the longest hangs are the ones that such a monitor loses. This page tells you how the library reports these hangs, and where it cannot.
Why the worst sessions report least
Two effects remove the worst sessions from the data (clocks and timers):
- Abandonment. Users leave slow pages. In data of Chrome on Android, each second more of First Contentful Paint added 2.6 percentage points of abandonment, to 17.3% at 4 s. The visitors with the worst experiences are "phantom bounces" that do not report.
- Lost beacons. A page that sends its data at the end can lose it. In a study of approximately 2 million Chrome page views (May 2023),
sendBeaconatpagehideorvisibilitychangewas 95.8% reliable. Beacons atload,pagehideandvisibilitychangetogether got to 98%, and anXMLHttpRequestgot to 84.9%.
The usual monitors learn about a block only after its end. PerformanceObserver callbacks, timers and animation frames all start on the main thread, thus they wait for the end of the block. A page that does not recover, or a tab that the user closes during a hang, sends nothing. In the survey of RUM tools, only the server-side crash report of Chrome for an unresponsive page sees this case (RUM landscape).
What the library does
The library has four parts that decrease this bias. The worker reports a hang while the hang continues. The hang journal lets the next page report a hang that the page did not survive. In WebKit and Safari, these two parts do not operate during a hang (refer to the limits). The crash-report context connects the crash reports of Chromium with the page view. The exporter flushes each time that the page becomes hidden.
The worker reports a hang while it continues
The worker monitor sends a heartbeat from its own timer, and the main thread acknowledges each heartbeat. When no acknowledgement comes for 5 s, the worker starts a hang itself. The start of the hang is the time of the last acknowledgement.
Show the diagram source
sequenceDiagram
participant M as Main thread
participant W as Worker
participant J as Hang journal
participant C as Collector
W->>M: heartbeat
M-->>W: ack
Note over M: a long task starts
W->>M: heartbeats wait in the queue
Note over W: 5 s without an ack
W->>C: hang started (fetch with keepalive)
W->>J: put the record
loop each second while the hang continues
W->>J: put the record again
end
Note over M: the long task stops
M-->>W: ack, or stop when the page became hidden
W->>J: remove the record
W->>C: hang ended (fetch with keepalive)
W->>M: hang-ended
Note over M: lag_main_thread_hangs, outcome endedThe bundled worker of @mark1russell7/lag/worker sends each hang report as an OTLP/HTTP JSON log record, with the event name lag.main_thread.hang. The record has the context of the page, the phase (started or ended) and duration_ms. A worker has no navigator.sendBeacon(), thus the worker uses fetch() with keepalive: true. With keepalive, the request can complete after the page closes. In WebKit and Safari, the request completes only when the main thread operates again (experiment E4). The Fetch standard limits the body of all keepalive requests in progress to 64 KiB (RUM landscape).
The worker sends the reports only when you give the target, with workerHangReport:
import { createInstanceId, init } from "@mark1russell7/otel-ts";
import { createBrowserDeps, createOtelLoggerAdapter, setupAllMonitors } from "@mark1russell7/lag";
import { createLagWorker } from "@mark1russell7/lag/worker";
const endpoint = "http://localhost:4318";
const serviceInstanceId = createInstanceId();
const otel = init({ serviceName : "shop", serviceInstanceId, endpoint });
const monitors = setupAllMonitors(createBrowserDeps(window, {
meter : otel.getMeter("lag"),
logger : createOtelLoggerAdapter(otel.getLogger("lag")),
worker : createLagWorker(),
// The worker sends its hang reports to this OTLP endpoint, with the identity of the page
workerHangReport : {
url : `${endpoint}/v1/logs`,
resource : { "service.name" : "shop", "service.instance.id" : serviceInstanceId },
},
// More attributes for each hang report, for example the session ID
pageContext : () => ({ "session.id" : otel.getSessionId() }),
}));
otel.onBeforeFlush(() => monitors.flush());
The worker sends the reports without the OpenTelemetry SDK of the page. Give workerHangReport.resource the same service.instance.id as the page. Then a query can connect the reports with the metrics of the page. The worker adds the attributes of pageContext to each report, and the ID of the current page view. Refer to OpenTelemetry setup.
When the main thread operates again, it acknowledges the heartbeats that waited. The first acknowledgement stops the hang. The worker removes the journal record, sends the report of the phase ended, and tells the main thread. The main thread then records the hang with the outcome ended, and it sends the event lag.main_thread.hang.
A stop or a start message also stops the hang, because the main thread sent it. For example, the main thread stops the worker monitor when the page becomes hidden. That can occur before it handles the heartbeats that waited. While the monitor is stopped, the main thread still handles hang-ended, thus the hang counts.
When the timer of the worker itself is late by 5 s or more, the worker did not operate. Then the worker resets the time of the last acknowledgement, and it does not start a hang. A system suspend thus does not give a hang. During a hang, the worker also moves the start of the hang forward by that time. Thus a sleep on Windows, in which performance.now() continues, does not count as time of the hang. Refer to measurement validity.
A unit test of setupAllMonitors() blocks the main thread for 8 s. It makes sure that the worker reports the start of the hang during the block. After the block, the main thread records one hang of 5 s or more.
The hang journal and the outcome abandoned
The worker keeps a record of each hang in progress in IndexedDB, in the hang journal. A record has these properties:
pageId: the random ID of the page instance whose main thread hung.startedAt: the time of the last acknowledgement before the hang.lastSeenAt: the last time that the worker saw the hang continue.attributes: the context of the page at the time of the hang, for examplelag.page_view.id.
The times are wall-clock times (Date.now()). Other pages compare them with their own time, and the wall clock is the only clock that all pages share. The monotonic clock of a page stops while the device sleeps, except on Windows.
The worker writes the record when the hang starts, writes it again each second while the hang continues, and removes it when the hang stops. A stop of the monitor also stops the hang, thus it also removes the record. A failed operation of the journal does not stop the heartbeats.
Show the diagram source
sequenceDiagram
participant W as Worker of page 1
participant J as Hang journal
participant P as Page 2 of the same origin
W->>J: put the record of the hang
Note over W: page 1 closes during the hang
Note over J: no update for 30 s or more
P->>J: list all records, one time at the start
J-->>P: the record of page 1
P->>J: take the record, only if it is still stale
J-->>P: the record, to one page only
Note over P: lag_main_thread_hangs, outcome abandonedThe pages and the workers of one origin share the IndexedDB database. When a new page starts the worker monitor, the monitor reads the journal one time. It selects the records of other pages that nobody updated for 30 s or more. For each of these records, it does these steps:
- It takes the record with
take(pageId, latestSeenAt). The method removes the record and gives it, but only if the record is still stale. IndexedDB does the read and the removal in one read-write transaction. Thus only one page gets the record, also when two pages start at the same time, for example at a session restore. The page skips a record that it did not get. - It records the hang with the outcome
abandoned. The duration is the time fromstartedAttolastSeenAt, thus the time until the worker saw the hang for the last time. - It sends the event
lag.main_thread.hangwith the phaseabandoned. The event has the attributes of the record,lag.hang.page_id, andlag.hang.source: "journal".
If the monitor stops after a take, the page puts the record back for the next page. A unit test starts two pages at the same time. It makes sure that the two pages report the hang one time, with a journal in memory and with IndexedDB.
The peer hang watch: another page, or the page at its close
The journal needs a worker that can write during the hang. WebKit and Safari do not let a worker write or send during a hang. The peer hang watch uses two other ways. The engines end a hung page that the user closes in different ways, and each way leaves one of them (experiment E7, local experiments):
- The browser stops the page: Chromium and Safari on iOS. Then another open page of the origin reports the hang. A visible page holds a Web Lock and sends a heartbeat each second through a
BroadcastChannel. Each other page asks for that lock, and the browser releases it when it stops the page. A lock that comes free after 5 s or more without a heartbeat gives a report withlag.hang.source: "peer". - The page operates again before it closes: Firefox stops the blocked script, and Safari on macOS lets the script continue to its end. Then the page itself reports its hang at its close, with
lag.hang.source: "self". Its heartbeats show the hang: none for 5 s or more before the close.
The report of another page comes at once, not at the start of a later page. In Chromium, that page also takes the journal record of the hung page, thus the next page does not report the hang again. The own report goes out in the last export of the page.
In Firefox, the hung page reports the hang itself, and the other page reports nothing. The worker of the hung page wrote a record during the hang, and Firefox can stop the worker before it removes the record. The page cannot wait for IndexedDB at its close. Thus it tries to remove the record, and it also writes a mark in localStorage (key lag-hang-reported:<page ID>). The journal reader of a later page takes the record of a marked page, but it does not count the hang again. A peer report without the record also writes a mark, for example when the take fails.
| Metric | Kind | Unit | Attributes | Description |
|---|---|---|---|---|
lag_main_thread_hangs | Counter | {hang} | outcome | The number of main-thread hangs that the worker detected. In a hang, the main thread does not acknowledge heartbeats. The outcome `abandoned` means that the page closed or crashed during the hang. The next page of the origin reports it from the hang journal. Another open page of the origin, or the page itself at its close, can also report it (PeerHangWatch). |
lag_main_thread_hang_duration_histogram | Histogram | ms | outcome | The duration of each main-thread hang. For an abandoned hang, the duration until the worker saw the hang for the last time. Without a record of the worker, it is the duration from the last heartbeat of the page to its end. |
createBrowserDeps() gives the journal when the page has IndexedDB, and a worker or the peer hang watch. Only the worker monitor reads the journal at its start. The peer hang watch only takes the record of a hung page. With hangJournal: false, the monitor does not read the journal, and it sends no page ID to the worker. Without a page ID, the worker writes no records.
A browser test blocks the main thread for 6.5 s and stops the worker before the end of the block, as when the page closes. Then it finds one record in IndexedDB, with a duration of 5 s or more and the ID of the page view. The test skips itself where a probe finds that the IndexedDB requests of a worker wait for the main thread, as in WebKit and Safari. The skip comes from that measurement, not from the name of the browser. The browser test worker-io.test.ts measures the behavior of each engine, also of a service worker, and it fails if an engine changes.
The crash-report context
Chromium can stop an unresponsive page and send a crash report to the reporting endpoint of the page. From Chrome 145, the page can add keys and values to such reports with window.crashReport. The page-view context sets the key lag.page_view.id there at each new page view. Thus the server can connect a crash report with the events and the Web Vitals of the same page view. Refer to page views.
Flush on hide
otel-ts exports all telemetry each time the page becomes hidden. At pagehide, it exports again: a page that goes into the back/forward cache keeps its providers, and a page that unloads stops them. The otel-ts library flushes in the export phase of the lifecycle tracker that the page shares. Thus the monitors record each transition before the export. With onBeforeFlush(() => monitors.flush()), these exports also contain the latest values of the page view. Refer to the flush hook.
The library does not use unload. From Chrome 154, Chrome does not start unload listeners (browser support). The README of web-vitals gives the same advice: send the data when the page goes into the background or unloads. Refer to Web Vitals algorithms.
The metrics use cumulative temporality, thus an export that is lost does not lose data. The next export contains the totals. Only the data after the last successful export is lost (OpenTelemetry metrics in the browser).
Limits
-
WebKit and Safari. WebKit and Safari 26.6.2 complete the IndexedDB requests and the
fetchrequests of a worker on the main thread of the page (experiment E4). The other output of a worker also waits: OPFS, the Cache API,XMLHttpRequest, WebSocket and BroadcastChannel (experiments E6 and E7). The worker detects the hang, because its timer continues. But the hang report and the journal record complete only when the main thread operates again. A service worker also waits, thus it cannot keep the journal either.Thus there, only the peer hang watch keeps a hang that the page does not survive. On macOS, the page reports the hang itself when its script ends. On iOS, another open page of the origin must report it, and Safari there operates only the visible tab. With one page on iOS, or with a script that does not end on macOS, the hang is lost.
-
Firefox closes a hung page normally. When the user closes a hung page, Firefox stops the blocked script, and the page closes normally (experiment E7). Then the page reports its own hang at its close. The worker monitor can also count the hang as ended before the close. Then the page does not report it again.
-
A frozen tab. If the browser freezes the page and its worker during a hang, the worker cannot report or update the record. Chrome on Android freezes background pages and their workers after 1 minute, from Chrome 139 (browser support). After 30 s without an update, the next page reports the hang as abandoned, also if the frozen page continues later.
-
The stale time of 30 s. A record counts as abandoned only after 30 s without an update. The monitor reads the journal only one time, at its start. Thus a page that starts less than 30 s after the hung page closed does not report the record, and a later page must report it. If no page of the origin starts the worker monitor again, nobody reports the hang.
-
The same origin only. IndexedDB belongs to one origin. Only a later page of the same origin, with the worker monitor and the journal, can report the hang.
-
No
localStorage. Without the marks, a later page can count a hang from the journal that the hung page reported already. For example, a sandboxed frame cannot uselocalStorage. -
Private windows. The journal needs IndexedDB. The code comments name some private windows as an example of a browser that does not permit IndexedDB. Then each operation of the journal fails, the worker continues without the journal, and the monitor logs a warning. We did not examine this case in each browser.
-
Short hangs. A hang starts when a heartbeat waits 5 s for its acknowledgement. A page that closes during a block of less than 5 s gives no report and no record.
-
The duration of an abandoned hang. The duration stops at the last write of the record. The worker writes the record at a heartbeat, at most one time each second. Thus the duration is too short by up to 1 s, or by up to one heartbeat interval if the interval is longer.
-
The wall clock. The stale test compares the wall-clock time of the new page with the times that the worker of the earlier page wrote. A step of the system clock between the write and the read changes the age of a record (clocks and time). We did not test this case.
-
No target. Without
workerHangReport, the worker does not send reports during the hang. The library then reports a hang only when it stops, or from the journal. -
Crash reports. Only Chromium sends crash reports, and only to the reporting endpoint that the page declares. A script cannot read them (browser support).