Skip to the content

Metrics model

Why the monitors record only counters and histograms, how the metric catalog defines them, and how to aggregate the series of many browsers.

The monitors record OpenTelemetry counters and histograms through a Meter that the caller gives. They send the details with many different values as events. Many browsers export the same metrics at the same time, and each page load is a separate writer. This page tells you the rules that keep these metrics correct when the backend aggregates them. The research note OpenTelemetry metrics in the browser gives the sources.

Counters and histograms only

The library records only two types of instrument:

  • A counter adds values, for example the number of discarded samples.
  • A histogram records a distribution of values, for example the lag of each window.

The values of counters and histograms aggregate correctly across many browsers. The sum of the rates of all pages is the rate of the fleet, and the buckets of the histograms of all pages add together.

The library records no gauges, for two causes:

  • No meaning for a fleet. A gauge gives the last value of one browser. The last value of one browser has no meaning for many browsers.
  • Stale values. With cumulative temporality, the OpenTelemetry SDK for JavaScript keeps each attribute set of an observable gauge that it observed one time. It exports the last value of the set in each later export. The issue opentelemetry-js #4096 is open since 2023-08-30. Its fix (#6884) was not merged at its last update, on 2026-09-24.

The metric catalog

METRIC_CATALOG in metric-catalog.ts is the single source of truth for the metrics. Each definition has these properties:

  • name: the instrument name, for example lag_drift_histogram.
  • kind: counter or histogram.
  • unit: for example ms, By, 1, or an annotation in braces, for example {sample}.
  • monitor: the monitor that records the metric.
  • description: the description that the instrument and the documentation show.
  • attributes: each attribute name with its permitted values.
  • advice: only for a histogram, its bucket boundaries.

The factories make their instruments only with createHistogram(meter, definition) and createCounter(meter, definition). These functions take the name, the unit, the description and the advice from the catalog. If the definition has a different kind, they do not make the instrument, and they give an error. The tables of metrics on this site also come from the catalog, for example on each page of the monitors. Thus the code, the instruments and the documentation agree.

Tests make sure of this agreement:

  • The naming rules. Each name starts with lag_ and has only lowercase letters, digits and underscores. A histogram name has the suffix _histogram, and a counter name does not. Each metric has a description, a unit and a closed set of values for each attribute.
  • The conformance test. A unit test of setupAllMonitors() starts all monitors with fake browser APIs, and it makes activity. Then it examines each instrument that the monitors made. Each instrument must be in the catalog, with the same kind and unit, and each attribute value must be a permitted value of the catalog.
  • The coverage test. The same test file makes sure that the normal activity records a value for each metric of the catalog. The exceptions are the metrics that need a special condition, for example a stall or a clock jump. Other tests cover them.

Attributes with a small number of values

The backend makes one series for each combination of attribute values. Thus each attribute of a metric has a small, fixed set of permitted values, for example reason: hidden, frozen or suspend. A measured value, a timestamp or an ID is not a metric attribute at any time.

The details with many different values go into events: CSS selectors, URLs, script names and IDs. The library removes the query string and the fragment from each URL. They can contain tokens or personal data, and they make each value unique.

Events for the details

An event is an OpenTelemetry log record with an event name. createOtelEventSink() sends each event in this way, with the severity INFO, and it removes each attribute that has no value. In Loki, the attributes of an event become structured metadata, and the event name becomes event_name.

The body of the record is the line of the event (formatEventLine()): the event name, and then the attributes as sorted key=value pairs. A value goes into JSON quotes when it is empty or has a space, a quote or =. For example, lag.stall with the attributes kind: "hang" and duration_ms: 6200 gives the line lag.stall duration_ms=6200 kind=hang. Loki drops an entry when the previous entry of its stream has the same timestamp and the same line. The five Web Vitals of one report can have the same millisecond, thus the attributes make each line different. The worker uses the same line for its hang reports.

The events of the monitors, from the event catalog
EventMonitorAttributes
browser.web_vitalPageViewVitalsbrowser.web_vital.name, browser.web_vital.value, browser.web_vital.delta, browser.web_vital.id, browser.web_vital.rating, browser.web_vital.navigation_type, lag.page_view.id, lag.page_view.url, lag.web_vital.*
lag.main_thread.hangWorkerLagMonitorphase, duration_ms, lag.hang.page_id, lag.hang.source, lag.page_view.id
lag.clock.jumpClockDriftMonitordirection, kind, magnitude_ms, skew_ms, lateness_ms, lag.page_view.id
lag.long_animation_frameLongAnimationFrameMonitorduration_ms, blocking_duration_ms, script.invoker, script.invoker_type, script.source_url, script.duration_ms, lag.page_view.id
lag.browser_reportBrowserReportMonitortype, id, message, source_file, line_number, lag.page_view.id
lag.stallMeasurementConditionskind, duration_ms, lag.page_view.id
lag.lifecycle.transitionLifecycleStateMachinefrom, to, trigger, lag.page_view.id
lag.page_view.startPageViewVitalsnavigation_type, lag.page_view.id, lag.page_view.url, lag.page_view.previous_id, lag.page_view.trace_id, lag.page_view.span_id
lag.pressure.changeComputePressureMonitorsource, state, previous_state, lag.page_view.id

The time of an event

Each event has the time of its occurrence (EventOptions.time, in Unix milliseconds). For an event with a duration, for example a hang, a stall or a long animation frame, it is the start. The time of a Web Vital is the time of the occurrence that gave the value (page-view vitals). The monitors calculate the times from the absolute clock of the page (createAbsoluteClock). The OpenTelemetry SDK gives the metric points times on the same clock. Thus an event and a metric point of the same time are at the same place on a chart.

The OTel event sink puts the time of the occurrence into the timestamp of the record. The observed time of the record is the time of the call. Loki keeps the events of one name in one stream for all pages of a service. It rejects a record that is older than the newest record of its stream by more than max_chunk_age / 2: 1 hour by default.

Some occurrences are reported much later. Examples are an LCP that the page reports when it becomes hidden after some hours, and an abandoned hang that the next page finds. Thus the sink uses the time of the occurrence only when it is no more than 30 minutes from the time of the call (MAX_EVENT_TIME_OFFSET_MS). Otherwise, the record gets the time of the call, and the attribute lag.event.time keeps the time of the occurrence. The hang reports of the worker use the same rule.

The monitors keep the volume of events small. For example, the long-animation-frame monitor sends an event only for a frame that blocks for 150 ms or more, and at most 10 events each minute.

The page-view ID on each event

setupAllMonitors() wraps the event sink with withEventContext(). When the page-view vitals operate, each event gets the attribute lag.page_view.id, with the ID of the current page view:

import { withEventContext, type EventSink } from "@mark1russell7/lag";

const sink : EventSink = { emit : (name, attributes) => console.log(name, attributes) };
let viewId = "view-1";

// The wrapper reads the context again for each event
const events = withEventContext(sink, () => ({ "lag.page_view.id" : viewId }));
events.emit("lag.stall", { kind : "hang", duration_ms : 6200 });
viewId = "view-2";
events.emit("lag.stall", { kind : "suspend", duration_ms : 60000 });

The wrapper uses the context function again for each event, thus each event gets the page view at the time of the event. An attribute of the event replaces an attribute of the context with the same name. The ID connects an event with the Web Vitals of its page view. Refer to page views.

Spans for the periods

A hang, a stall, a long animation frame and a hidden period have a start and an end. An event gives only the start and the duration. With a span sink (spans), the monitors also send each of these periods as a span, with its real start and end. A trace viewer, for example Grafana Tempo, then shows each page view as a timeline.

  • One trace for each page view. The span lag.page_view starts at the start of the view. It is the root of the trace of the view. The other spans of the monitors are in it.
  • The end of the span of a view. The span ends when the page is hidden for the first time in the view, or at the final report of the view. The browser can discard a hidden page without an event, and an exporter sends only the spans that ended. The span gets the values of the Web Vitals at its end as attributes (lag.web_vital.lcp, for example). A value that changes after that time is in the metrics and in the events.
  • A hang that another page reports. The page-view context of a page has the identity of the span of its view (lag.page_view.trace_id and lag.page_view.span_id). The hang journal and the peer hang watch keep this context. Thus the span of an abandoned hang is in the trace of the page view that hung. It has a link to the page view that reported it.
  • Sampling. The sampler of the SDK decides for the span of a view. The spans in a view are sampled as the view. A page whose view was not sampled puts no span identity into its context.
The spans of the monitors, from the span catalog
SpanMonitorPeriodAttributes
lag.page_viewPageViewVitalsOne page view, from its start to the first time that the page is hidden in the view, or to its final report. It has the values of the Web Vitals at its end. The root of the trace of the view.lag.page_view.id, navigation_type, lag.page_view.url, lag.web_vital.*
lag.main_thread.hangWorkerLagMonitorA main-thread hang, from its start to its end. An abandoned hang is in the page view of the page that hung, when its record has the identity of that span. Then it has a link to the page view of the page that reported it.phase, duration_ms, lag.hang.page_id, lag.hang.source
lag.stallMeasurementConditionsOne stall episode, from the start of its first window to the end of its longest sample.kind, duration_ms
lag.long_animation_frameLongAnimationFrameMonitorA long animation frame above the attribution threshold, with the script that blocked it most.duration_ms, blocking_duration_ms, script.invoker, script.invoker_type, script.source_url, script.duration_ms
lag.page.hiddenLifecycleStateMachineA period in which the page was hidden, from the transition to hidden to the next transition.trigger
lag.page.frozenLifecycleStateMachineA period in which the browser froze the page or kept it in the back/forward cache, to the next transition.trigger

The event of a period stays. The dashboards and the alerts use the events and the metrics. The spans are for the investigation of one page view. The playground records the spans and the events of its own page, and it shows them on a timeline with the metrics.

Writer identity

Each SDK instance must write its own series. Without a writer identity, all browsers of a service write to the same series. Then Mimir rejects samples with the errors new-value-for-timestamp and sample-out-of-order. The samples that it accepts mix the running totals of different browsers. PromQL sees a counter reset at almost each sample, thus rate(), increase() and histogram_quantile() give incorrect values.

For example, the counter of the first browser reads 100, 101 and 102. The counter of the second browser reads 5, 6 and 7, at times between those of the first browser. The mixed series is 100, 5, 101, 6, 102, 7. increase() sees a reset at each decrease, and it gives 210. The correct increase is 4.

In the OpenTelemetry data model, more than one writer for a metric stream is an error state. Mimir makes the instance label from the resource attribute service.instance.id, and the job label from service.name. Thus each SDK instance needs its own service.instance.id:

  • otel-ts sets a new random UUID each time that init() starts, and it does not keep the value (OpenTelemetry setup). Each page load and each worker gets a different value.
  • The browser build of the OpenTelemetry SDK does not set service.instance.id. A setup without otel-ts must set it.
  • session.id is not a writer identity. A duplicated tab copies the session, a reload keeps it, and a worker shares it.

Cumulative temporality

The exporter of otel-ts requests cumulative temporality by default. Each export contains the running totals since the start of the SDK instance. In its default configuration, Mimir rejects delta sums and delta histograms. It answers with the HTTP status 400, also when it accepted the other parts of the request. The delta ingestion of Mimir is experimental.

Cumulative data tolerates a lost export: the next export contains the totals. Only the data after the last successful export is lost. A short page load has a different problem. rate() and increase() use only the increases between the samples of a series, thus they lose the first export of each page load. The Mimir limit otel_created_timestamp_zero_ingestion_enabled adds a sample of 0 at the start time of each series. This limit is experimental, and the Grafana stack of this project sets it.

Native histograms

With histogramAggregation: "exponential", otel-ts records each histogram as a base-2 exponential histogram. Mimir keeps it as a native histogram: one series for each attribute set. An explicit-bucket histogram needs one series for each bucket, one for +Inf, and one each for _sum and _count. Mimir must accept native histograms (native_histograms_ingestion_enabled). The default of otel-ts is "explicit".

With the default settings of Mimir, the series of a native histogram has the name of the instrument, without a suffix. The PromQL functions histogram_quantile(), histogram_sum(), histogram_count() and histogram_fraction() read it.

Bucket boundaries

With the explicit-bucket aggregation, the SDK puts each value into a bucket with fixed boundaries. The default boundaries of the SDK go from 0 to 10000, for durations in milliseconds. With them, all scores and ratios go into the bucket from 0 to 5, and all byte counts go into the last bucket. Thus each histogram of the catalog has its own boundaries in advice.explicitBucketBoundaries. createHistogram() gives this advice to the Meter. An exponential histogram does not use the advice, and a view of the SDK can replace it.

The catalog has one set of boundaries for each type of value (HISTOGRAM_BOUNDARIES):

  • Durations (ms): from 0 to 60000 ms. The boundaries below 1 ms separate the clock resolutions of the browsers. The boundaries 16 and 33 are the frame budgets at 60 Hz and 30 Hz. The boundary 50 is the limit of a long task, and 5000 is the limit of a stall.
  • Scores and ratios (1): from 0 to 1, for layout shifts, CLS and the memory ratio. The CLS thresholds 0.1 and 0.25 are boundaries.
  • The pressure state: 0, 1, 2 and 3. Each state has its own bucket.
  • Bytes (By): the powers of 2 from 1 MiB to 16 GiB.

Each bucket includes its upper boundary. The thresholds of the Web Vitals are boundaries, thus the buckets give the exact number of good values. For example, the _bucket series of lag_web_vital_cls_histogram with le="0.1" counts the page views with a good CLS. The boundaries of two library versions must agree: a different set of boundaries gives different _bucket series.

The bucket boundaries of the histograms, from the metric catalog
UnitBoundariesHistograms
ms0, 0.01, 0.05, 0.1, 0.5, 1, 2, 5, 10, 16, 25, 33, 50, 75, 100, 150, 200, 300, 500, 800, 1000, 1800, 2500, 3000, 4000, 5000, 10000, 30000, 60000All 27 histograms with the unit ms
10, 0.001, 0.01, 0.025, 0.05, 0.1, 0.15, 0.25, 0.5, 0.75, 0.9, 1lag_layout_shift_histogram, lag_web_vital_cls_histogram, lag_memory_usage_ratio_histogram
By1 MiB, 2 MiB, 4 MiB, 8 MiB, 16 MiB, 32 MiB, 64 MiB, 128 MiB, 256 MiB, 512 MiB, 1 GiB, 2 GiB, 4 GiB, 8 GiB, 16 GiBlag_memory_used_bytes_histogram
10, 1, 2, 3lag_pressure_state_histogram

Aggregation in PromQL

Sum the rates of all writers first. Then calculate the quantile. A quantile of each page and then a mean of those quantiles gives a value that no sample has.

For example, page A has 100 samples of 10 ms. Page B has 85 samples of 10 ms and 15 samples of 500 ms. The 95th percentile of A is 10 ms, and the 95th percentile of B is 500 ms. The mean of the two is 255 ms, but no sample has this value. The 95th percentile of the 200 samples together is 500 ms.

With native histograms:

histogram_quantile(0.95, sum by (service_name) (rate(lag_drift_histogram[5m])))

With explicit-bucket histograms:

histogram_quantile(0.95, sum by (service_name, le) (rate(lag_drift_histogram_bucket[5m])))

For a counter:

sum by (reason) (rate(lag_samples_discarded[5m]))

These rules also help:

  • Do not sum _bucket series without rate().
  • Set the range of a rate to at least 4 times the export interval. otel-ts exports each 15 s by default.
  • Use recording rules for long ranges. The Grafana stack of this project has recording rules for the most expensive queries of its dashboard.
  • Do not use a metric attribute to find one page. The instance label and the events in Loki connect a series with its page load.

The label service_name in these examples comes from the resource attribute service.name, which the Mimir configuration of the Grafana stack promotes to each series. Refer to the Grafana stack.

All metrics

All metrics of setupAllMonitors()
MetricKindUnitAttributesDescription
lag_drift_histogramHistogrammsNoneThe lag of one window (approximately 100 ms) of chained timeouts: its duration minus the idle duration of its steps. Each block of the main thread in the window adds to the lag.
lag_drift_baseline_histogramHistogrammsNoneThe idle duration of one timer step: the mean of the recent steps that are not blocks. An increase needs a probe that shows an idle thread. It is the timer granularity of the browser and the operating system. DriftLag subtracts it.
lag_macrotask_histogramHistogrammsNoneThe time that a zero-delay timeout waits in the task queue. The monitor measures one sample every 5 seconds.
lag_samples_discardedCounter{sample}reasonThe number of samples that a monitor did not record because the measurement window was not valid.
lag_stallsCounter{stall}kindThe number of stall episodes: very long samples (5000 ms or more) of all monitors whose windows overlap count as one episode. A hang has no evidence of a suspend. A suspend has evidence that the system stopped.
lag_stall_duration_histogramHistogrammskindThe duration of each stall episode: its longest sample.
lag_worker_main_block_histogramHistogrammsNoneThe time that a worker heartbeat waited for the main thread. This is main-thread blocking, measured from outside the main thread.
lag_worker_self_lag_histogramHistogrammsNoneThe lateness of the heartbeat timer of the worker. A high value shows that the worker itself did not operate.
lag_worker_clock_offset_histogramHistogrammsNoneThe absolute offset between the worker clock and the main-thread clock, from the clock synchronization exchange.
lag_main_thread_hangsCounter{hang}outcomeThe number of main-thread hangs that the worker detected. In a hang, the main thread does not acknowledge heartbeats. The outcome `abandoned` means that the page closed or crashed during the hang. The next page of the origin reports it from the hang journal. Another open page of the origin, or the page itself at its close, can also report it (PeerHangWatch).
lag_main_thread_hang_duration_histogramHistogrammsoutcomeThe duration of each main-thread hang. For an abandoned hang, the duration until the worker saw the hang for the last time. Without a record of the worker, it is the duration from the last heartbeat of the page to its end.
lag_loaf_blocking_histogramHistogrammsNoneThe blocking duration of each long animation frame.
lag_loaf_duration_histogramHistogrammsNoneThe total duration of each long animation frame.
lag_event_duration_histogramHistogrammsinteractionThe duration of each interaction event of 16 ms or more, from input to the next paint.
lag_event_input_delay_histogramHistogrammsinteractionThe time from the input to the start of the event handlers.
lag_event_processing_histogramHistogrammsinteractionThe time that the event handlers used to process the event.
lag_event_presentation_delay_histogramHistogrammsinteractionThe time from the end of the event handlers to the next paint.
lag_layout_shift_histogramHistogram1NoneThe score of each layout shift that did not follow user input.
lag_web_vital_inp_histogramHistogrammsnavigation_typeInteraction to Next Paint (INP) for each page view.
lag_web_vital_cls_histogramHistogram1navigation_typeCumulative Layout Shift (CLS) for each page view, in browsers that have layout-shift entries.
lag_web_vital_lcp_histogramHistogrammsnavigation_typeLargest Contentful Paint (LCP) for each page view.
lag_web_vital_fcp_histogramHistogrammsnavigation_typeFirst Contentful Paint (FCP) for each page view.
lag_web_vital_ttfb_histogramHistogrammsnavigation_typeTime to First Byte (TTFB) for each page view. A restore from the back/forward cache and a soft navigation have no network response and get 0, as in web-vitals. A page without a navigation entry gets no value.
lag_frame_delta_histogramHistogrammsNoneThe time between two animation frame callbacks.
lag_framesCounter{frame}outcomeThe number of delivered frames and the estimated number of dropped frames.
lag_idle_time_remaining_histogramHistogrammsNoneThe idle time that was available when an idle callback started.
lag_idle_gap_histogramHistogrammsNoneThe time between two idle callbacks.
lag_idle_callbacksCounter{callback}timed_outThe number of idle callbacks. A callback that timed out started because no idle period came before its timeout.
lag_scheduling_microtask_histogramHistogrammsNoneThe latency of a queueMicrotask callback. This value stays near 0 and is a baseline.
lag_scheduling_macrotask_histogramHistogrammsNoneThe latency of a zero-delay timeout.
lag_scheduling_message_channel_histogramHistogrammsNoneThe latency of a MessageChannel message.
lag_memory_used_bytes_histogramHistogramBysourceThe used heap memory of each sample.
lag_memory_usage_ratio_histogramHistogram1NoneThe used heap divided by the heap limit. Only the legacy source supplies the limit.
lag_pressure_state_histogramHistogram1sourceThe compute pressure state of each record: 0 nominal, 1 fair, 2 serious, 3 critical.
lag_gc_eventsCounter{gc}NoneThe number of garbage collections that the detector saw.
lag_lifecycle_transitionsCounter{transition}from, to, triggerThe number of page lifecycle transitions. A restore from the back/forward cache always counts, with the trigger pageshow, also when the state does not change (Chromium makes the page visible before pageshow).
lag_timer_calibrationsCounter{calibration}throttledThe number of timer calibration rounds. A throttled round shows that the browser slowed the timers.
lag_clock_resolution_histogramHistogrammsNoneThe resolution of performance.now(). The checker measures it one time for each page.
lag_clock_skew_histogramHistogrammsNoneThe absolute difference between Date.now() and the absolute monotonic clock (timeOrigin from the start plus performance.now()).
lag_clock_jumpsCounter{jump}direction, kindThe number of discontinuities between the wall clock and the monotonic clock: a suspend (the monotonic clock stopped while the device slept) or a step of the system clock.
lag_browser_reportsCounter{report}typeThe number of reports from the Reporting API, for example interventions and deprecations.
lag_liveness_block_histogramHistogrammsNoneThe duration of each main-thread block that a worker saw through shared memory.

lag: Main-thread responsiveness monitoring for browser apps, exported as OpenTelemetry metrics.

To change a page, edit its file in packages/site/content/. The writing style guide tells you how.

An AI model (Claude, from Anthropic) wrote most of the text and the code of this site and of the library, under the direction of the author. The tests and the STE linter examine them. The writing standard gives the reason for this note.