Skip to the content

Dashboards

The Grafana dashboards of the stack, the query of each panel, and how to add a panel.

Note

This page is a draft. The content is not complete and can change.

The Grafana stack provisions two dashboards from the grafana-infra repository. The Lag Monitor dashboard shows the metrics of the monitors in Mimir and their events in Loki. The Frontend Observability dashboard shows span metrics, traces and logs. This page describes each panel and the query that it uses.

NoteThe version

This page describes the dashboards of grafana-infra at commit 6ebe0dc, the main branch on 2026-10-09, after pull request 5 ("annotations"). The lockfile of this repository has this commit, and pnpm test:e2e starts its stack.

The dashboards

DashboardUIDFileContent
Lag Monitorlag-monitorlag-monitor.json72 panels in 21 rows, and 8 annotation layers: the metrics of the monitors in Mimir, and their events in Loki. The script scripts/build-dashboards.mjs writes the file.
Frontend Observabilityfrontend-observabilityfrontend.json11 panels in 5 rows: the span metrics of Tempo, the traces in Tempo, and the logs in Loki. It does not use the lag metrics.

The files are in config/grafana/provisioning/dashboards of grafana-infra. The file provider Default of Grafana loads them. Grafana reads the files again in 30 seconds or less (updateIntervalSeconds: 30). By default, both dashboards show the last hour, and they refresh each 30 seconds.

The verification script of grafana-infra loads one dashboard through the Grafana API and sends each panel query and the query of each annotation layer. The option --dashboard selects the dashboard, and the default is lag-monitor. Refer to the verification workflow.

Variables

The Lag Monitor dashboard has four variables:

VariableLabelTypeQuery or value
service_nameServiceQuery, Mimirlabel_values(lag_drift_histogram, service_name)
instancePage loadQuery, Mimirlabel_values(lag_drift_histogram{service_name=~"$service_name"}, instance)
session_idSessionText boxThe start of a session.id. The default is empty.
navigation_typeNavigation typeQuery, Mimirlabel_values(lag_web_vital_lcp_histogram{service_name=~"$service_name"}, navigation_type)

The default value of each query variable is All, and All sends the regular expression .+. Thus each query uses the operator =~, for example service_name=~"$service_name". Grafana gets the values again when the time range changes (refresh: 2). You can select more than one value of Service and of Navigation type. Page load has one value or All.

The service_name label comes from the resource attribute service.name. Mimir promotes it to a label, and Loki indexes it. The Service list shows only the services that send lag_drift_histogram as a native histogram. A service with explicit-bucket histograms has only _bucket, _sum and _count series, thus the list does not show it.

The instance variable selects one page load: one SDK instance. Mimir makes the instance label from the resource attribute service.instance.id. In Loki, the same value is in the structured metadata service_instance_id. Each Mimir query has the matcher instance=~"$instance", and each Loki query has the filter service_instance_id=~"$instance". Thus, with one page load, each panel and each annotation layer shows only that page load. The row Long range is different: its recording rules sum all page loads, thus the variable does not apply to it.

The Page load list has the page loads that sent lag_drift_histogram in the time range. The list is in alphabetical order, and each value is a UUID. To find a page load, use the Page loads table. A click on a page load in the table sets the variable.

With All, the matcher is instance=~".+". This matcher does not find a series without the instance label. Thus the panels do not show the data of an SDK that does not set service.instance.id. In the same way, the Loki panels do not show an event without service_instance_id.

The session_id variable filters only the Page loads table. The table uses session_id=~"${session_id:regex}.*". Grafana escapes the special characters of the text, and .* lets the text match the start of a session ID.

The navigation_type variable filters the Web Vitals panels, the INP p75 panel of the Overview row, and the recorded Web Vitals panel.

Annotation layers

Each kind of lag event has an annotation layer. A layer shows a mark on each time series panel at the time of each event. The top of the dashboard has a switch for each layer.

LayerEventOn by defaultColorTitle of a markText of a markTags
Hangslag.main_thread.hangYesRedHang {{phase}}{{duration_ms}} ms, source {{lag_hang_source}}phase, lag_hang_source
Stallslag.stallYesOrangeStall: {{kind}}{{duration_ms}} mskind
Page viewslag.page_view.startNoBluePage view: {{navigation_type}}{{lag_page_view_url}}navigation_type
Lifecyclelag.lifecycle.transitionNoPurple{{from}} to {{to}}{{trigger}}trigger
Compute pressurelag.pressure.changeNoYellowPressure {{source}}: {{state}}from {{previous_state}}source, state
Clock jumpslag.clock.jumpNoGray (#8f8f8f)Clock {{kind}} {{direction}}{{magnitude_ms}} mskind
Long animation frameslag.long_animation_frameNoPurple (#b877d9)Long animation frame{{blocking_duration_ms}} ms blocking, {{script_invoker}}script_invoker_type
Browser reportslag.browser_reportNoBlue (#5794f2){{type}}: {{id}}{{message}}type

The marks show these items:

  • Hangs: the phase of the hang (started, ended or abandoned) and its duration. For an abandoned hang, the source tells which page reported it: journal (the next page of the origin), peer (another open page of the origin) or self (the page itself at its close). The other phases have no source. Refer to Hangs.
  • Stalls: the kind of the stall episode (hang or suspend) and its duration.
  • Page views: the start of each page view, with its navigation type and its URL.
  • Lifecycle: each transition of the lifecycle state, with the browser event that caused it.
  • Compute pressure: each change of the pressure state of a source, with the state before it. The first record of a source has no state before it.
  • Clock jumps: the kind, the direction and the magnitude of each jump.
  • Long animation frames: the blocking duration of each frame that the monitor sent as an event, and the script that blocked the frame most.
  • Browser reports: the type, the ID and the message of each report.

The Hangs and Stalls layers are on by default. The other layers can have many events, thus they are off by default. With all page loads, the marks of all pages are on the panels. Thus, first select one page load, then turn on the layers. A layer gets at most 500 events (maxLines: 500).

Each layer follows the Service and Page load variables. It does not follow Navigation type or Session. The query of a layer has this form:

{service_name=~"$service_name", event_name="lag.main_thread.hang"} | service_instance_id=~"$instance" | logfmt

The line of an event in Loki is the event name, then the attributes as sorted key=value pairs. Refer to events in Loki. The parser logfmt gets the attributes from the line. Thus the title, the text and the tags of a mark can use each attribute. The names have underscores in place of dots, for example lag_hang_source for lag.hang.source.

The time of a mark is the time of the log record. The lag library gives each event the time of its occurrence. For an event with a duration, for example a hang or a stall, it is the start. Thus a mark is at the time of the metric values that it explains.

An occurrence can be more than 30 minutes before or after the call. Then the record gets the time of the call, and the attribute lag.event.time has the time of the occurrence. Refer to the time of an event.

NoteThe query of a layer in Grafana 13

The Loki data source of Grafana 13 reads the query of an annotation layer from the layer itself (annotationQuery), not from its target. Thus each layer has the query in the top-level property expr, with maxLines, instant, titleFormat, textFormat and tagKeys. The target of the layer has the same query, for the query editor and for the verification script. With the query only in target, Grafana sent no query for the layers (commit 91a5cd8 of grafana-infra).

The verification script makes sure that the query of each layer finds events in the sample data. It does not examine the title and the text of the marks. The screenshots of the verification workflow show the marks.

Examine one page load

  1. Open the Lag Monitor dashboard.
  2. Go to the row Events (Loki).
  3. To find the page loads of one session, write the start of its session ID in Session.
  4. In the Page loads table, click the ID of a page load. If Grafana shows a menu, click "Show only this page load".
  5. At the top of the dashboard, turn on the annotation layers that you want to see.

The link of step 4 opens the dashboard with the Page load variable, the Service value and the time range in the URL. To see all page loads again, set Page load to All.

Read a panel

The queries of the Lag Monitor dashboard obey these rules:

  • A histogram query first sums the rates of all page loads, then it takes the quantile, for example histogram_quantile(0.95, sum(rate(x[$__rate_interval]))). Thus a p95 line is the p95 of all samples of the selected page loads. It is not a mean of the quantiles of each page load. Refer to aggregation in PromQL.
  • A counter query uses rate() or increase(), not the cumulative value. A Mimir panel with "per minute" in its title multiplies the rate in each second by 60.
  • Each query filters by the Service and the Page load variables. Only the queries of the row Long range do not filter by Page load.
  • No Mimir query groups by instance or by session. Only the Page loads table groups by page load and by session. It counts the events of each page load in Loki.
  • A rate in a time series uses the range $__rate_interval. The Mimir data source sets timeInterval to 15s, the export interval. Then $__rate_interval is 60 seconds or more.
  • A stat panel with "(time range)" in its title uses an instant query with increase(...[$__range]). The panels INP p75, Hangs, Stalls and performance.now() resolution also use such a query. These panels give one value for the full time range of the dashboard.
  • A mean line uses histogram_avg(). A count of observations uses histogram_count().
  • The Loki tables use instant queries over $__range, or they show the latest log lines.

The panels use these units:

UnitMeaning
msMilliseconds.
bytesBytes.
noneA number without a unit, for example a count or a score.
percentunitA ratio from 0 to 1. Grafana shows it as a percentage.

The Web Vitals panels show the good and the poor thresholds of each vital. The color changes to orange at the good threshold and to red at the poor threshold. The time series show the thresholds as dashed lines.

VitalGoodPoorUnit
INP200 or lessMore than 500ms
CLS0.1 or lessMore than 0.25none
LCP2500 or lessMore than 4000ms
FCP1800 or lessMore than 3000ms
TTFB800 or lessMore than 1800ms

To see the metrics and the events of one page load, set the Page load variable. Refer to examine one page load. In Mimir, the ID of the page load is the instance label. In Loki, the same value is in service_instance_id.

The metric names are a contract

The panels, the variables and the recording rules use the metric names of the catalog, packages/lag/src/metric-catalog.ts. Mimir keeps these names, because it adds no unit suffix and no _total suffix. Thus the metric names are a contract between the catalog and the dashboard. A change of a name in the catalog without the same change in grafana-infra gives empty panels.

The names are in these files:

RepositoryFileWhat it has
lagpackages/lag/src/metric-catalog.tsThe catalog: the name, kind, unit and permitted attribute values of each metric, and the events with their attributes.
grafana-infrascripts/lib/lag-catalog.mjsA copy of the catalog. The dashboard script, the sample-data script and the verification script use it.
grafana-infrascripts/build-dashboards.mjsThe panel queries, the annotation layers and the variables. The script writes lag-monitor.json.
grafana-infraconfig/mimir/rules/anonymous/lag.yamlThe recording rules of 11 histograms.
grafana-infraconfig/grafana/provisioning/dashboards/lag-monitor.jsonThe dashboard that the script writes.

The attribute names are also part of the contract. The panels group by attributes, for example outcome, kind and navigation_type. The Loki panels use the event names and the event attributes, for example blocking_duration_ms. The annotation layers use the event names, and their marks use the event attributes, for example lag_hang_source.

On 2026-10-09, the copy of commit 6ebe0dc agrees with the catalog on the main branch of this repository. Both have 42 metrics (32 histograms and 10 counters) and 9 events, with the same names, kinds, units, attributes and permitted values. The 9 events are browser.web_vital, lag.main_thread.hang, lag.clock.jump, lag.long_animation_frame, lag.browser_report, lag.stall, lag.lifecycle.transition, lag.page_view.start and lag.pressure.change. The check verify --catalog of grafana-infra compares the two catalogs, and it found no difference.

The copy also has a list of optional attributes (optional) for 3 events. Some events of these names do not have these attributes:

EventOptional attributes
lag.main_thread.hanglag.hang.page_id, lag.hang.source
lag.page_view.startlag.page_view.url, lag.page_view.previous_id
lag.pressure.changeprevious_state

The catalog of this repository has no such list. The verification script does not expect the optional attributes in the structured metadata of each event.

The dashboard uses all 42 metric names, and no other metric name that starts with lag_. The other lag_ names in the dashboard are the structured metadata of the events: lag_page_view_id, lag_page_view_url, lag_page_view_previous_id, lag_hang_page_id and lag_hang_source.

Examine a change of a name

  1. Change the name in metric-catalog.ts.

  2. In grafana-infra, change the name in scripts/lib/lag-catalog.mjs.

  3. Change the name in scripts/build-dashboards.mjs. The variables use lag_drift_histogram and lag_web_vital_lcp_histogram.

  4. If the metric has recording rules, change the name in config/mimir/rules/anonymous/lag.yaml.

  5. Write the dashboard again:

    pnpm dashboards
    

    The script stops with the message "Metrics without a panel" if a metric of the copy is in no query.

  6. Search lag-monitor.json and lag.yaml for the old name. Make sure that the search finds nothing.

  7. Start the stack.

  8. Send the sample data:

    pnpm sample-data --summary sample-summary.json
    
  9. Start the verification script with the catalog of this repository:

    pnpm verify --summary sample-summary.json --catalog ../lag/packages/lag/src/metric-catalog.ts
    

With --catalog, the verification script compares the copy with the catalog. It compares the names, the kinds, the units, the attribute names and the attributes of each event. It does not compare the permitted values of the attributes. It also examines if Mimir stores each metric under its catalog name, if each panel query gives data, and if each annotation layer finds events. The path in the example is correct when lag and grafana-infra are in the same directory.

The comparison of the catalogs is the first part of the output, and it does not use the stack. Without the stack, the next checks fail, and the script stops at the CORS check with the error fetch failed. The CI workflow of grafana-infra does not use --catalog, because it does not have a copy of this repository.

Add a panel

The script scripts/build-dashboards.mjs writes the Lag Monitor dashboard. Do not edit lag-monitor.json. To add a panel, do these steps:

  1. Open scripts/build-dashboards.mjs in grafana-infra.

  2. Find the row() call of the row for the new panel.

  3. After the last panel of that row, add a call to timeseries(), stat(), metricTable() or eventTable().

  4. Give the targets with prom(), promInstant(), loki() or lokiInstant().

  5. Write each PromQL query with the helpers, for example quantile(), sumRate(), perMinute() and observationsPerMinute(). The helpers add the matchers service_name=~"$service_name" and instance=~"$instance".

  6. Start each LogQL query with the helper events(), for example events('="lag.stall"'). It adds the Service matcher and the Page load filter.

  7. Make sure that the query obeys the rules in Read a panel.

  8. Write the dashboard again:

    pnpm dashboards
    
  9. Look at each panel and its queries:

    node scripts/build-dashboards.mjs --list
    
  10. Wait for 30 seconds.

  11. Load the dashboard again in Grafana.

  12. Send sample data.

  13. Start the verification script. Make sure that the query of the new panel gives data.

  14. Commit scripts/build-dashboards.mjs and lag-monitor.json together. The CI workflow writes the dashboard again, and it fails if the file is different.

This example adds the p75 of the drift lag to the drift row:

row("Drift (DriftLag)");
// The panels of the row
timeseries("Drift lag p75",
    "The 75th percentile of the drift lag of all page loads.",
    [prom(quantile(0.75, "lag_drift_histogram"), "p75")], { w : 12 });

By default, a time series panel has the unit ms, a width of 12 and a height of 8. A stat panel has the unit ms, a width of 4 and a height of 4. The layout puts the panels from left to right. It starts a new line when a panel does not fit in the 24 columns of the grid.

A panel for a new metric needs more steps. First add the metric to scripts/lib/lag-catalog.mjs. Then record the metric in scripts/send-sample-data.mjs. That script stops with "Metrics without values" if a metric of the copy has no values.

Add an annotation layer

The list ANNOTATIONS in scripts/build-dashboards.mjs has the layers. Each item is a call to annotation(), with the name of the layer, the event name, and the options color, enable, title, text and tags. The function writes the query with events() and | logfmt. It puts the query in expr and in target. This example is the Stalls layer:

annotation("Stalls", "lag.stall", { color : "orange", enable : true, title : "Stall: {{kind}}", text : "{{duration_ms}} ms", tags : "kind" }),

A layer for a new event needs more steps. First add the event to EVENTS in scripts/lib/lag-catalog.mjs. Then send the event in scripts/send-sample-data.mjs. Without events in the sample data, the verification script fails for the layer.

The verification workflow

The CI workflow .github/workflows/verify.yml of grafana-infra examines the stack, the sample data and the dashboard from end to end. It starts for each pull request, for each push to main, and by hand (workflow_dispatch). The job uses Node.js 24 and the pnpm version of packageManager in package.json. It stops after 30 minutes.

StepWhat it does
The dashboard file agrees with scripts/build-dashboards.mjsWrites the dashboard again. The step fails if git diff --exit-code finds a change in config/grafana/provisioning/dashboards/.
Start the stackdocker compose up -d. Then it waits up to 5 minutes for Alloy, Mimir, Loki and Grafana.
Send sample datanode scripts/send-sample-data.mjs --summary sample-summary.json
Verify the pipeline and the dashboardnode scripts/verify-pipeline.mjs --summary sample-summary.json, without --catalog
Screenshots of the dashboardInstalls Chromium for Playwright, then node scripts/screenshot-dashboard.mjs --summary sample-summary.json --out screenshots.
Upload the screenshotsUploads the directory screenshots/ as the artifact dashboard-screenshots.
Stop the stackdocker compose down -v. When a step fails, the workflow first writes the last 200 log lines of the containers.

The screenshots, the upload and the stop of the stack also occur when an earlier step fails.

The sample data

scripts/send-sample-data.mjs simulates six page loads of two services, with the OpenTelemetry JS SDK. It sends each event of the catalog copy, also lag.page_view.start, lag.lifecycle.transition and lag.pressure.change. Each event has the line of the event as its body, in the form of formatEventLine of the lag library. Thus the lines in Loki, which the annotation layers parse, have the same form as the lines of the library. A hang event and its stall event have the start of the hang as their time.

The checks of the dashboard

scripts/verify-pipeline.mjs examines the dashboard through the Grafana API:

  • Grafana loads the dashboard.
  • Each label_values() variable gives values. The Session text box has no query.
  • Each panel query gives data. The script sends the query to /api/ds/query. It replaces $service_name with the services of the sample, $navigation_type and $instance with .+, and ${session_id:regex} with an empty text.
  • The query of each annotation layer finds events. The script sends the query of the target of the layer.

Before these checks, the script compares the number of events of each name in Loki with the number that the sample sent. It waits up to 30 seconds for the last events of the sample. Then it examines the structured metadata of one event of each name, without the optional attributes of the catalog copy.

On 2026-10-09, the workflow passed on the merge commit 6ebe0dc (workflow run 37920009653). All 134 panel queries gave data, each of the 8 annotation layers found events, and the script passed 269 of 269 checks.

The screenshots

scripts/screenshot-dashboard.mjs takes screenshots of the dashboard with Playwright and Chromium, after the sample data. It reads the dashboard from Grafana and saves a copy with all annotation layers on: the UID lag-monitor-review and the title "Lag Monitor (all annotations)". A URL cannot turn on a layer, thus the script needs this copy. The time range is from 1 minute before the start of the sample to 1 minute after its end. The browser window is 1600 by 1200 pixels, and Grafana shows the dashboard in kiosk mode.

The script writes these files into the directory of --out. The default directory is screenshots.

FileContent
fleet.pngThe full dashboard as provisioned, for all page loads.
table-<title>.pngOne screenshot for each table of events: Page loads, Recent page views and lifecycle transitions, Recent hangs and stalls, Recent clock jumps, and LoAF attribution.
tables.jsonFor each table: the title, the file and the first 400 characters of its text, or the error.
page-<index>.pngThe review copy for one page load of the sample, with all annotation layers on. The option --pages selects the page loads. The default is 0,1.
annotations.jsonThe annotation queries that Grafana sent, with the number of frames and rows of each result. Also the console errors of the browser, and the fields of the log frames of the Loki panels.

To take the screenshots on your computer, do these steps:

  1. Send the sample data with --summary sample-summary.json.

  2. Install Chromium for Playwright:

    pnpm exec playwright install chromium
    
  3. Take the screenshots:

    node scripts/screenshot-dashboard.mjs --summary sample-summary.json
    
  4. Open the directory screenshots.

The script sends no credentials to Grafana. It saves the dashboard lag-monitor-review in your Grafana, and it replaces an earlier copy. Git ignores the directory screenshots/.

The Lag Monitor panels

Each row of the dashboard has a section here. The table gives each panel, and the code block gives its queries from lag-monitor.json.

Overview

Six stat panels give the state of the selected services and page loads.

PanelTypeUnitWhat it shows
Active page loadsStatnoneThe number of page loads that sent drift samples in the minute before each point. Each page load has its own service.instance.id, thus its own series. The panel shows the last value.
Drift p95StatmsThe 95th percentile of the drift lag of all selected page loads. The panel shows the last value that is not null.
Heartbeat delay p95StatmsThe 95th percentile of the delivery delay of the worker heartbeats: the lag that an event at a random time sees.
INP p75StatmsThe 75th percentile of INP of the page views in the time range. The color changes to orange at 200 ms and to red at 500 ms.
HangsStatnoneThe main-thread hangs that the workers detected in the time range, with the outcomes ended and abandoned.
StallsStatnoneThe stall episodes in the time range: very long samples from hangs and from suspends.
# Active page loads
count(count_over_time(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[1m])) or vector(0)
# Drift p95
histogram_quantile(0.95, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Heartbeat delay p95
histogram_quantile(0.95, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# INP p75
histogram_quantile(0.75, sum(increase(lag_web_vital_inp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# Hangs
sum(increase(lag_main_thread_hangs{service_name=~"$service_name", instance=~"$instance"}[$__range]))
# Stalls
sum(increase(lag_stalls{service_name=~"$service_name", instance=~"$instance"}[$__range]))

Drift (DriftLag)

Refer to DriftLag.

PanelTypeUnitWhat it shows
Drift lag p50 / p95 / p99Time seriesmsThe quantiles and the mean of the lag of each window of chained timeouts (approximately 100 ms). Each block of the main thread in a window adds to the lag of the window.
Drift baseline: timer granularityTime seriesmsThe p50 and the p95 of the idle duration of one timer step: the mean of the recent steps that are not blocks. An increase needs a probe that shows an idle thread, thus a sustained load does not change it. It is the timer granularity of the browser and the operating system. DriftLag subtracts it.
Drift windows per secondTime seriesnoneThe number of windows that DriftLag recorded each second, for all page loads. The monitors discard the windows of hidden and frozen pages.
# Drift lag p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_avg(sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Drift baseline: timer granularity
histogram_quantile(0.5, sum (rate(lag_drift_baseline_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_drift_baseline_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Drift windows per second
histogram_count(sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Macrotask (MacrotaskLag)

Refer to MacrotaskLag.

PanelTypeUnitWhat it shows
Macrotask queue delay p50 / p95 / p99Time seriesmsThe quantiles and the mean of the time that a zero-delay timeout waits in the task queue. Each page load measures one sample every 5 seconds.
# Macrotask queue delay p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_avg(sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Scheduling fairness (SchedulingFairnessMonitor)

Refer to scheduling fairness.

PanelTypeUnitWhat it shows
Scheduling latency p50, by primitiveTime seriesmsThe median latency of a microtask, a zero-delay timeout (setTimeout 0) and a MessageChannel message. The microtask latency stays near 0 and is the baseline.
Scheduling latency p99, by primitiveTime seriesmsThe 99th percentile of the same three latencies. A high macrotask or MessageChannel value shows a busy task queue.
# Scheduling latency p50, by primitive
histogram_quantile(0.5, sum (rate(lag_scheduling_microtask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.5, sum (rate(lag_scheduling_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.5, sum (rate(lag_scheduling_message_channel_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Scheduling latency p99, by primitive
histogram_quantile(0.99, sum (rate(lag_scheduling_microtask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_scheduling_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_scheduling_message_channel_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Worker heartbeat (WorkerLagMonitor)

Refer to worker lag.

PanelTypeUnitWhat it shows
Heartbeat delivery delay: the lag that a random event seesTime seriesmsThe quantiles and the mean of the time that each worker heartbeat waited for the main thread. The worker sends a heartbeat each second from its own thread. Thus the heartbeats arrive at random times, as user events do. This is the primary lag estimator.
Worker self-lagTime seriesmsThe lateness of the heartbeat timer of the worker. A high value shows that the worker itself did not operate. Then its heartbeat delays are not reliable.
Worker clock offsetTime seriesmsThe absolute offset between the worker clock and the main-thread clock, from the clock synchronization.
# Heartbeat delivery delay: the lag that a random event sees
histogram_quantile(0.5, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_avg(sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Worker self-lag
histogram_quantile(0.5, sum (rate(lag_worker_self_lag_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_worker_self_lag_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_worker_self_lag_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Worker clock offset
histogram_quantile(0.5, sum (rate(lag_worker_clock_offset_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_worker_clock_offset_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_worker_clock_offset_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Hangs (WorkerLagMonitor)

Refer to worker lag.

PanelTypeUnitWhat it shows
Hangs per minute, by outcomeTime series, barsnoneThe hangs that the workers detected, each minute, for each outcome. In a hang, the main thread does not acknowledge the heartbeats. abandoned: the page closed or crashed during the hang. The next page of the origin, another open page of the origin, or the page itself at its close reported it. The attribute lag.hang.source of the event tells which (journal, peer or self).
Hang duration p50 / p95, by outcomeTime seriesmsThe duration of each hang. For an abandoned hang, it is the duration until the worker saw the hang for the last time, or until the end of the page.
# Hangs per minute, by outcome
60 * sum by (outcome) (rate(lag_main_thread_hangs{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Hang duration p50 / p95, by outcome
histogram_quantile(0.5, sum by (outcome) (rate(lag_main_thread_hang_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum by (outcome) (rate(lag_main_thread_hang_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

The annotation layer Hangs shows each hang event on the charts, with its phase, its duration and its source.

Measurement conditions (stalls and discarded samples)

Refer to measurement validity.

PanelTypeUnitWhat it shows
Stalls per minute, by kindTime series, barsnoneThe stall episodes each minute, for each kind. hang: no evidence of a suspend. suspend: evidence that the system stopped, for example a sleep of the device.
Stall duration p50 / p95, by kindTime seriesmsThe duration of each stall episode: its longest sample.
Samples discarded per minute, by reasonTime series, stackednoneThe samples that the monitors did not record, for each reason: the page was hidden or frozen, or the system was suspended during the measurement window.
# Stalls per minute, by kind
60 * sum by (kind) (rate(lag_stalls{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Stall duration p50 / p95, by kind
histogram_quantile(0.5, sum by (kind) (rate(lag_stall_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum by (kind) (rate(lag_stall_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Samples discarded per minute, by reason
60 * sum by (reason) (rate(lag_samples_discarded{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))

The annotation layer Stalls shows each lag.stall event on the charts, at the start of the stall episode.

Clock (ClockDriftMonitor and ClockReliabilityChecker)

Refer to clock drift and clock reliability.

PanelTypeUnitWhat it shows
Clock jumps per minute, by kind and directionTime series, barsnoneThe discontinuities between the wall clock and the monotonic clock each minute, for each kind and direction. suspend: the monotonic clock stopped while the device slept. step: the system clock changed.
Clock skew p50 / p95 / p99Time seriesmsThe absolute difference between Date.now() and the absolute monotonic clock (timeOrigin plus performance.now()).
performance.now() resolutionStatmsThe p50 and the p95 of the resolution of performance.now() of the page loads in the time range, with 3 decimals. Each page load measures it one time.
# Clock jumps per minute, by kind and direction
60 * sum by (kind, direction) (rate(lag_clock_jumps{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Clock skew p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_clock_skew_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_clock_skew_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_clock_skew_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# performance.now() resolution
histogram_quantile(0.5, sum(increase(lag_clock_resolution_histogram{service_name=~"$service_name", instance=~"$instance"}[$__range])))
histogram_quantile(0.95, sum(increase(lag_clock_resolution_histogram{service_name=~"$service_name", instance=~"$instance"}[$__range])))

The annotation layer Clock jumps shows each lag.clock.jump event on the charts. It is off by default.

Long animation frames (LongAnimationFrameMonitor)

Refer to long animation frames.

PanelTypeUnitWhat it shows
LoAF blocking duration p50 / p95 / p99Time seriesmsThe blocking duration of each long animation frame.
LoAF duration p50 / p95 / p99Time seriesmsThe total duration of each long animation frame.
Long animation frames per minuteTime seriesnoneThe number of long animation frames of all page loads, each minute: the count of the duration histogram.
LoAF attribution: blocking time by script (events)TableBlocking time: ms, Frames: noneOne row for each script, with the columns Invoker type, Invoker and Script URL. The columns Blocking time and Frames give the sum of the blocking time and the number of frames in the time range. The data comes from the lag.long_animation_frame events. The monitor sends this event only for a frame that blocks for 150 ms or more. The event names the script that blocked the frame most. The table sorts by the blocking time. Two instant queries give the values. The transformations put each label in a column and merge the two queries into one row for each script.
# LoAF blocking duration p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_loaf_blocking_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_loaf_blocking_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_loaf_blocking_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# LoAF duration p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Long animation frames per minute
60 * histogram_count(sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# LoAF attribution: blocking time by script (events)
sum by (script_invoker_type, script_invoker, script_source_url) (sum_over_time({service_name=~"$service_name", event_name="lag.long_animation_frame"} | service_instance_id=~"$instance" | keep script_invoker_type, script_invoker, script_source_url, blocking_duration_ms | unwrap blocking_duration_ms [$__range]))
sum by (script_invoker_type, script_invoker, script_source_url) (count_over_time({service_name=~"$service_name", event_name="lag.long_animation_frame"} | service_instance_id=~"$instance" | keep script_invoker_type, script_invoker, script_source_url [$__range]))

The annotation layer Long animation frames shows each lag.long_animation_frame event on the charts, at the start of the frame. It is off by default.

Event timing (EventTimingMonitor)

Refer to Event Timing.

PanelTypeUnitWhat it shows
Event duration p95, by interaction typeTime seriesmsThe 95th percentile of the duration of the interaction events of 16 ms or more, from the input to the next paint, for each interaction value.
Event phases p95Time seriesmsWhere the time of an interaction goes: the input delay before the handlers, the processing time of the handlers, and the presentation delay to the next paint.
Interactions per minute, by interaction typeTime series, stackednoneThe interaction events of 16 ms or more, each minute, for each interaction value.
# Event duration p95, by interaction type
histogram_quantile(0.95, sum by (interaction) (rate(lag_event_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Event phases p95
histogram_quantile(0.95, sum (rate(lag_event_input_delay_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_event_processing_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_event_presentation_delay_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Interactions per minute, by interaction type
60 * histogram_count(sum by (interaction) (rate(lag_event_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Layout shift (LayoutShiftMonitor)

Refer to layout shift.

PanelTypeUnitWhat it shows
Layout shift score p50 / p95 / p99Time seriesnoneThe score of each layout shift that did not follow user input, with 3 decimals.
Layout shifts per minuteTime seriesnoneThe layout shifts without user input, each minute.
# Layout shift score p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Layout shifts per minute
60 * histogram_count(sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Web Vitals (PageViewVitals)

Refer to page-view vitals. The histogram of a vital gets one value for each page view, usually when the page becomes hidden for the first time. Later changes go only into the browser.web_vital events.

PanelTypeUnitWhat it shows
INP p75 (time range)StatmsThe p75 of INP of the page views in the time range, for each navigation type. The color changes to orange at 200 ms and to red at 500 ms.
CLS p75 (time range)StatnoneThe p75 of CLS, with 3 decimals. The color changes to orange at 0.1 and to red at 0.25.
LCP p75 (time range)StatmsThe p75 of LCP. The color changes to orange at 2500 ms and to red at 4000 ms.
FCP p75 (time range)StatmsThe p75 of FCP. The color changes to orange at 1800 ms and to red at 3000 ms.
TTFB p75 (time range)StatmsThe p75 of TTFB. The color changes to orange at 800 ms and to red at 1800 ms.
Page views (time range)StatnoneThe page views that reported LCP in the time range: the sample size of the vitals. The count is approximate, because increase() extrapolates.
INP p75, by navigation typeTime seriesmsThe p75 of INP of the page views that reported in each interval, for each navigation type. Dashed lines show the thresholds.
CLS p75, by navigation typeTime seriesnoneThe same for CLS.
LCP p75, by navigation typeTime seriesmsThe same for LCP.
FCP p75, by navigation typeTime seriesmsThe same for FCP.
TTFB p75, by navigation typeTime seriesmsThe same for TTFB.
Web Vitals p75 from events, by navigation typeTableEvents: noneThe p75 and the number of events of each vital and navigation type, from the browser.web_vital events. A page view sends a new event at each change of a vital. Thus this p75 also includes the earlier values.
Web Vitals p75 from eventsTime seriesmsThe p75 of browser_web_vital_value for each vital, from the events in a 5-minute window. CLS uses the right axis.
# INP p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_inp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# CLS p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_cls_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# LCP p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_lcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# FCP p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_fcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# TTFB p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_ttfb_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# Page views (time range)
histogram_count(sum(increase(lag_web_vital_lcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# INP p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_inp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# CLS p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_cls_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# LCP p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_lcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# FCP p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_fcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# TTFB p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_ttfb_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# Web Vitals p75 from events, by navigation type
quantile_over_time(0.75, {service_name=~"$service_name", event_name="browser.web_vital"} | service_instance_id=~"$instance" | browser_web_vital_navigation_type=~"$navigation_type" | keep browser_web_vital_name, browser_web_vital_navigation_type, browser_web_vital_value | unwrap browser_web_vital_value [$__range]) by (browser_web_vital_name, browser_web_vital_navigation_type)
sum by (browser_web_vital_name, browser_web_vital_navigation_type) (count_over_time({service_name=~"$service_name", event_name="browser.web_vital"} | service_instance_id=~"$instance" | browser_web_vital_navigation_type=~"$navigation_type" | keep browser_web_vital_name, browser_web_vital_navigation_type [$__range]))
# Web Vitals p75 from events
quantile_over_time(0.75, {service_name=~"$service_name", event_name="browser.web_vital"} | service_instance_id=~"$instance" | browser_web_vital_navigation_type=~"$navigation_type" | keep browser_web_vital_name, browser_web_vital_value | unwrap browser_web_vital_value [5m]) by (browser_web_vital_name)

The annotation layer Page views shows the start of each page view (lag.page_view.start) on the charts. It is off by default.

Frames (FrameTimingMonitor)

Refer to frame timing.

PanelTypeUnitWhat it shows
Frames per second, delivered and droppedTime series, stackednoneThe animation frames of all visible page loads each second, for each outcome. The monitor estimates the dropped frames from long gaps between frames.
Dropped frame ratioTime seriespercentunitThe dropped frames divided by all frames.
Frame delta p50 / p95 / p99Time seriesmsThe time between two animation frame callbacks. 16.7 ms is 60 frames per second.
# Frames per second, delivered and dropped
sum by (outcome) (rate(lag_frames{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Dropped frame ratio
sum(rate(lag_frames{service_name=~"$service_name", instance=~"$instance", outcome="dropped"}[$__rate_interval])) / sum(rate(lag_frames{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Frame delta p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_frame_delta_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_frame_delta_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_frame_delta_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Idle time (IdleAvailabilityMonitor)

Refer to idle availability.

PanelTypeUnitWhat it shows
Idle time remaining p10 / p50 / p90Time seriesmsThe idle time that was available when an idle callback started. Low values show a busy main thread.
Idle gap p50 / p95 / p99Time seriesmsThe time between two idle callbacks.
Idle callbacks per minute, by timed_outTime series, stackednoneThe idle callbacks each minute, for each timed_out value. A callback that timed out started because no idle period came before its timeout.
Idle callbacks that timed outTime seriespercentunitThe idle callbacks that timed out, divided by all idle callbacks.
# Idle time remaining p10 / p50 / p90
histogram_quantile(0.1, sum (rate(lag_idle_time_remaining_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.5, sum (rate(lag_idle_time_remaining_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.9, sum (rate(lag_idle_time_remaining_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Idle gap p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_idle_gap_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_idle_gap_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_idle_gap_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Idle callbacks per minute, by timed_out
60 * sum by (timed_out) (rate(lag_idle_callbacks{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Idle callbacks that timed out
sum(rate(lag_idle_callbacks{service_name=~"$service_name", instance=~"$instance", timed_out="true"}[$__rate_interval])) / sum(rate(lag_idle_callbacks{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))

Memory (MemoryMonitor)

Refer to memory.

PanelTypeUnitWhat it shows
Used JS heap p50 / p95, by sourceTime seriesbytesThe used heap memory of each sample. modern: performance.measureUserAgentSpecificMemory(). legacy: performance.memory.
Heap usage ratio p50 / p95Time seriespercentunitThe used heap divided by the heap limit. Only the legacy source gives the limit.
# Used JS heap p50 / p95, by source
histogram_quantile(0.5, sum by (source) (rate(lag_memory_used_bytes_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum by (source) (rate(lag_memory_used_bytes_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Heap usage ratio p50 / p95
histogram_quantile(0.5, sum (rate(lag_memory_usage_ratio_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_memory_usage_ratio_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Compute pressure (ComputePressureMonitor)

Refer to compute pressure.

PanelTypeUnitWhat it shows
Mean pressure state, by sourceTime seriesnoneThe mean compute pressure state of the records: 0 nominal, 1 fair, 2 serious, 3 critical. The axis stops at 3.
Records at serious or critical pressure, by sourceTime seriespercentunitThe records with the state serious (2) or critical (3), divided by all records.
# Mean pressure state, by source
histogram_avg(sum by (source) (rate(lag_pressure_state_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Records at serious or critical pressure, by source
1 - histogram_fraction(-Inf, 1.5, sum by (source) (rate(lag_pressure_state_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

The annotation layer Compute pressure shows each change of the pressure state of a source (lag.pressure.change) on the charts. It is off by default.

GC, page lifecycle and timer throttling

Refer to GC signal, lifecycle and timer throttle.

PanelTypeUnitWhat it shows
GC events per minuteTime seriesnoneThe garbage collections that the GC signal detector saw, each minute, for all page loads.
Lifecycle transitions per minuteTime series, barsnoneThe page lifecycle transitions each minute, for each from, to and trigger. The legend is a table.
Throttled timer calibrationsTime seriespercentunitThe timer calibration rounds that found throttled timers, divided by all rounds.
# GC events per minute
60 * sum (rate(lag_gc_events{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Lifecycle transitions per minute
60 * sum by (from, to, trigger) (rate(lag_lifecycle_transitions{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Throttled timer calibrations
sum(rate(lag_timer_calibrations{service_name=~"$service_name", instance=~"$instance", throttled="true"}[$__rate_interval])) / sum(rate(lag_timer_calibrations{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))

The annotation layer Lifecycle shows each lag.lifecycle.transition event on the charts. It is off by default. The table Recent page views and lifecycle transitions gives the same events.

Browser reports (BrowserReportMonitor)

Refer to browser reports.

PanelTypeUnitWhat it shows
Browser reports per minute, by typeTime series, barsnoneThe reports from the Reporting API each minute, for each type: intervention or deprecation.
Recent browser reports (events)Table–The latest lag.browser_report events (at most 200 lines): the service, the type, the ID, the message, the source file and the line. A Grafana log frame has its own id field, thus the query writes the report ID to report_id.
# Browser reports per minute, by type
60 * sum by (type) (rate(lag_browser_reports{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Recent browser reports (events)
{service_name=~"$service_name", event_name="lag.browser_report"} | service_instance_id=~"$instance" | label_format report_id=id

The annotation layer Browser reports shows each lag.browser_report event on the charts. It is off by default.

Shared-memory liveness (SharedLivenessMonitor)

Refer to shared-memory liveness.

PanelTypeUnitWhat it shows
Liveness block duration p50 / p95 / p99Time seriesmsThe duration of each main-thread block that a worker saw through shared memory. Only cross-origin isolated pages have this monitor.
Liveness blocks per minuteTime seriesnoneThe main-thread blocks that the shared-memory watcher saw, each minute.
# Liveness block duration p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Liveness blocks per minute
60 * histogram_count(sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))

Events (Loki)

The panels of this row read the events in Loki. Refer to events in Loki.

PanelTypeUnitWhat it shows
Events per minute, by event nameTime seriesnoneThe lag events in Loki in the minute before each point, for each event_name. The panel has an interval of 15 s.
Page loadsTable–The page loads (SDK instances) with lag events in the time range. Each row has the columns Page load (service_instance_id), Session (session_id) and Events, the number of events in the time range. The table has at most 50 rows (topk(50, ...)), and it sorts by Events. The Session variable filters the rows. Each page load ID has the link "Show only this page load", which sets the Page load variable. The query also filters by Page load, thus with one page load the table has only that page load.
Recent page views and lifecycle transitionsTable–The latest lag.page_view.start and lag.lifecycle.transition events (at most 200 lines), with the columns Time, Event, Navigation, From, To, Trigger, URL, Page view, Previous view, Session and Instance. A page view start fills Navigation, URL and Previous view. The first page view of a page load has no previous view. A transition fills From, To and Trigger. The annotation layers Page views and Lifecycle show the same events on the charts.
Recent hangs and stallsTableDuration: msThe latest lag.main_thread.hang and lag.stall events (at most 200 lines): the time, the service, the event, the phase, the kind, the duration, the sender (scope_name), the page view, the hung page, the session and the instance. The worker sends the start of a hang itself, with the scope @mark1russell7/lag/worker. An abandoned hang comes from the next page of the origin (journal), from another open page of the origin (peer), or from the page itself at its close (self). The table has no column for this source, lag_hang_source. The Hangs annotation layer shows it.
Recent clock jumpsTableMagnitude: ms, Skew: ms, Lateness: msThe latest lag.clock.jump events (at most 200 lines): the time, the service, the kind, the direction, the magnitude, the skew, the lateness, the page view and the session.

The Page loads table uses the transformations of the LoAF attribution table: one column for each label, then one row for each series. The value of its instant query is Value or Value #A, by the version of Grafana, and the table shows it as Events. The other tables get their columns from the structured metadata of each log line (extractFields), and they sort by the time.

# Events per minute, by event name
sum by (event_name) (count_over_time({service_name=~"$service_name", event_name=~".+"} | service_instance_id=~"$instance" [1m]))
# Page loads
topk(50, sum by (service_instance_id, session_id) (count_over_time({service_name=~"$service_name", event_name=~".+"} | service_instance_id=~"$instance" | session_id=~"${session_id:regex}.*" [$__range])))
# Recent page views and lifecycle transitions
{service_name=~"$service_name", event_name=~"lag.page_view.start|lag.lifecycle.transition"} | service_instance_id=~"$instance"
# Recent hangs and stalls
{service_name=~"$service_name", event_name=~"lag.main_thread.hang|lag.stall"} | service_instance_id=~"$instance"
# Recent clock jumps
{service_name=~"$service_name", event_name="lag.clock.jump"} | service_instance_id=~"$instance"

Long range (recording rules, all page loads)

The panels of this row read the series of the recording rules, not the raw series. The rules use 5-minute windows, and Mimir evaluates them each minute. Use these panels for long time ranges. The six histograms are the drift, the macrotask delay, the heartbeat delay, the LoAF blocking duration, the event duration and the frame delta. The rules sum all page loads of a service, thus the Page load variable does not apply to this row.

PanelTypeUnitWhat it shows
p95 by service (recorded)Time seriesmsservice_name:<metric>:p95_rate5m of the six histograms, for each service. The legend is a table.
p99 by service (recorded)Time seriesmsservice_name:<metric>:p99_rate5m of the same six histograms.
Fleet p95 from recorded ratesTime seriesmsThe p95 of all selected services, from the recorded native-histogram rates service_name:<metric>:rate5m. A sum of recorded rates is correct. A sum of recorded quantiles is not correct.
Web Vitals p75 from recorded ratesTime seriesmsThe p75 of each vital, from service_name_navigation_type:<vital>:rate5m. CLS uses the right axis.
# p95 by service (recorded)
service_name:lag_drift_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_macrotask_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_worker_main_block_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_loaf_blocking_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_event_duration_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_frame_delta_histogram:p95_rate5m{service_name=~"$service_name"}
# p99 by service (recorded)
service_name:lag_drift_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_macrotask_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_worker_main_block_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_loaf_blocking_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_event_duration_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_frame_delta_histogram:p99_rate5m{service_name=~"$service_name"}
# Fleet p95 from recorded rates
histogram_quantile(0.95, sum(service_name:lag_drift_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_macrotask_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_worker_main_block_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_loaf_blocking_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_event_duration_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_frame_delta_histogram:rate5m{service_name=~"$service_name"}))
# Web Vitals p75 from recorded rates
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_inp_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_cls_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_lcp_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_fcp_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_ttfb_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))

The Frontend Observability panels

The Frontend Observability dashboard is from the first commit of grafana-infra. The only later change is the query of the Log Entries panel. It uses the span metrics that the metrics generator of Tempo writes to Mimir (traces_spanmetrics_*), the traces in Tempo, and the logs in Loki. It has the variable service (label_values(traces_spanmetrics_calls_total, service)), but no panel query uses it.

The verification script does not examine this dashboard, thus its panels are not verified. The Loki configuration of the stack makes only service_name and event_name index labels. Thus the selector {exporter="OTLP"} of the Log Stream panel finds no stream (not verified).

RowPanelTypeUnitWhat it shows
OverviewTotal Traces (1h)StatshortThe increase of traces_spanmetrics_calls_total in the last hour.
OverviewErrors (1h)StatshortThe same for the spans with the status code STATUS_CODE_ERROR.
OverviewAvg Span DurationStatmsThe mean duration of the spans in 5-minute windows.
OverviewLog Entries (1h)StatshortThe number of log lines in Loki in the last hour, of all services. Loki does not accept the empty selector {}.
Page Load PerformancePage Load Duration (documentLoad spans)Time seriesmsThe p50, p95 and p99 duration of the documentLoad spans, from explicit buckets.
Page Load PerformanceSpan Rate by TypeTime series, stacked barsshortThe spans each second, for each span name.
HTTP RequestsHTTP Request Duration (fetch spans)Time seriesmsThe p50 and the p95 duration of the HTTP spans: span names that start with HTTP, fetch or an HTTP method.
HTTP RequestsHTTP Request RateTime seriesreqpsThe HTTP spans each second, and the HTTP spans with an error.
Traces & LogsRecent TracesTraces–A TraceQL search in Tempo: at most 20 traces with the status error.
Traces & LogsLog StreamLogs–The log lines of the stream selector {exporter="OTLP"}.
Service MapService GraphNode graph–The service map of Tempo.
# Total Traces (1h)
sum(increase(traces_spanmetrics_calls_total[1h]))
# Errors (1h)
sum(increase(traces_spanmetrics_calls_total{status_code="STATUS_CODE_ERROR"}[1h]))
# Avg Span Duration
sum(rate(traces_spanmetrics_latency_sum[5m])) / sum(rate(traces_spanmetrics_latency_count[5m])) * 1000
# Page Load Duration (documentLoad spans)
histogram_quantile(0.50, sum(rate(traces_spanmetrics_latency_bucket{span_name="documentLoad"}[5m])) by (le))
histogram_quantile(0.95, sum(rate(traces_spanmetrics_latency_bucket{span_name="documentLoad"}[5m])) by (le))
histogram_quantile(0.99, sum(rate(traces_spanmetrics_latency_bucket{span_name="documentLoad"}[5m])) by (le))
# Span Rate by Type
sum by (span_name) (rate(traces_spanmetrics_calls_total[5m]))
# HTTP Request Duration (fetch spans)
histogram_quantile(0.50, sum(rate(traces_spanmetrics_latency_bucket{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*"}[5m])) by (le))
histogram_quantile(0.95, sum(rate(traces_spanmetrics_latency_bucket{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*"}[5m])) by (le))
# HTTP Request Rate
sum(rate(traces_spanmetrics_calls_total{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*"}[5m]))
sum(rate(traces_spanmetrics_calls_total{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*", status_code="STATUS_CODE_ERROR"}[5m]))
# Log Entries (1h)
sum(count_over_time({service_name=~".+"} [1h]))
# Log Stream
{exporter="OTLP"}

Source

  • grafana-infra on GitHub.
  • scripts/build-dashboards.mjs: the panels, the annotation layers and the variables of the Lag Monitor dashboard, and the query rules.
  • scripts/lib/lag-catalog.mjs: the copy of the metric catalog.
  • config/grafana/provisioning/dashboards: lag-monitor.json, frontend.json and the dashboard provider.
  • scripts/send-sample-data.mjs: the sample data.
  • scripts/verify-pipeline.mjs: the verification of each panel query and each annotation layer, and the comparison of the catalogs.
  • scripts/screenshot-dashboard.mjs: the screenshots of the dashboard.
  • .github/workflows/verify.yml: the CI workflow.
  • packages/lag/src/metric-catalog.ts in this repository: the metric catalog.
  • The Grafana documentation of provisioning.

lag: Main-thread responsiveness monitoring for browser apps, exported as OpenTelemetry metrics.

To change a page, edit its file in packages/site/content/. The writing style guide tells you how.

An AI model (Claude, from Anthropic) wrote most of the text and the code of this site and of the library, under the direction of the author. The tests and the STE linter examine them. The writing standard gives the reason for this note.