Dashboards
The Grafana dashboards of the stack, the query of each panel, and how to add a panel.
Note
The Grafana stack provisions two dashboards from the grafana-infra repository. The Lag Monitor dashboard shows the metrics of the monitors in Mimir and their events in Loki. The Frontend Observability dashboard shows span metrics, traces and logs. This page describes each panel and the query that it uses.
NoteThe version
This page describes the dashboards of grafana-infra at commit 6ebe0dc, the main branch on 2026-10-09, after pull request 5 ("annotations"). The lockfile of this repository has this commit, and pnpm test:e2e starts its stack.
The dashboards
| Dashboard | UID | File | Content |
|---|---|---|---|
| Lag Monitor | lag-monitor | lag-monitor.json | 72 panels in 21 rows, and 8 annotation layers: the metrics of the monitors in Mimir, and their events in Loki. The script scripts/build-dashboards.mjs writes the file. |
| Frontend Observability | frontend-observability | frontend.json | 11 panels in 5 rows: the span metrics of Tempo, the traces in Tempo, and the logs in Loki. It does not use the lag metrics. |
The files are in config/grafana/provisioning/dashboards of grafana-infra. The file provider Default of Grafana loads them. Grafana reads the files again in 30 seconds or less (updateIntervalSeconds: 30). By default, both dashboards show the last hour, and they refresh each 30 seconds.
The verification script of grafana-infra loads one dashboard through the Grafana API and sends each panel query and the query of each annotation layer. The option --dashboard selects the dashboard, and the default is lag-monitor. Refer to the verification workflow.
Variables
The Lag Monitor dashboard has four variables:
| Variable | Label | Type | Query or value |
|---|---|---|---|
service_name | Service | Query, Mimir | label_values(lag_drift_histogram, service_name) |
instance | Page load | Query, Mimir | label_values(lag_drift_histogram{service_name=~"$service_name"}, instance) |
session_id | Session | Text box | The start of a session.id. The default is empty. |
navigation_type | Navigation type | Query, Mimir | label_values(lag_web_vital_lcp_histogram{service_name=~"$service_name"}, navigation_type) |
The default value of each query variable is All, and All sends the regular expression .+. Thus each query uses the operator =~, for example service_name=~"$service_name". Grafana gets the values again when the time range changes (refresh: 2). You can select more than one value of Service and of Navigation type. Page load has one value or All.
The service_name label comes from the resource attribute service.name. Mimir promotes it to a label, and Loki indexes it. The Service list shows only the services that send lag_drift_histogram as a native histogram. A service with explicit-bucket histograms has only _bucket, _sum and _count series, thus the list does not show it.
The instance variable selects one page load: one SDK instance. Mimir makes the instance label from the resource attribute service.instance.id. In Loki, the same value is in the structured metadata service_instance_id. Each Mimir query has the matcher instance=~"$instance", and each Loki query has the filter service_instance_id=~"$instance". Thus, with one page load, each panel and each annotation layer shows only that page load. The row Long range is different: its recording rules sum all page loads, thus the variable does not apply to it.
The Page load list has the page loads that sent lag_drift_histogram in the time range. The list is in alphabetical order, and each value is a UUID. To find a page load, use the Page loads table. A click on a page load in the table sets the variable.
With All, the matcher is instance=~".+". This matcher does not find a series without the instance label. Thus the panels do not show the data of an SDK that does not set service.instance.id. In the same way, the Loki panels do not show an event without service_instance_id.
The session_id variable filters only the Page loads table. The table uses session_id=~"${session_id:regex}.*". Grafana escapes the special characters of the text, and .* lets the text match the start of a session ID.
The navigation_type variable filters the Web Vitals panels, the INP p75 panel of the Overview row, and the recorded Web Vitals panel.
Annotation layers
Each kind of lag event has an annotation layer. A layer shows a mark on each time series panel at the time of each event. The top of the dashboard has a switch for each layer.
| Layer | Event | On by default | Color | Title of a mark | Text of a mark | Tags |
|---|---|---|---|---|---|---|
| Hangs | lag.main_thread.hang | Yes | Red | Hang {{phase}} | {{duration_ms}} ms, source {{lag_hang_source}} | phase, lag_hang_source |
| Stalls | lag.stall | Yes | Orange | Stall: {{kind}} | {{duration_ms}} ms | kind |
| Page views | lag.page_view.start | No | Blue | Page view: {{navigation_type}} | {{lag_page_view_url}} | navigation_type |
| Lifecycle | lag.lifecycle.transition | No | Purple | {{from}} to {{to}} | {{trigger}} | trigger |
| Compute pressure | lag.pressure.change | No | Yellow | Pressure {{source}}: {{state}} | from {{previous_state}} | source, state |
| Clock jumps | lag.clock.jump | No | Gray (#8f8f8f) | Clock {{kind}} {{direction}} | {{magnitude_ms}} ms | kind |
| Long animation frames | lag.long_animation_frame | No | Purple (#b877d9) | Long animation frame | {{blocking_duration_ms}} ms blocking, {{script_invoker}} | script_invoker_type |
| Browser reports | lag.browser_report | No | Blue (#5794f2) | {{type}}: {{id}} | {{message}} | type |
The marks show these items:
- Hangs: the phase of the hang (
started,endedorabandoned) and its duration. For an abandoned hang, the source tells which page reported it:journal(the next page of the origin),peer(another open page of the origin) orself(the page itself at its close). The other phases have no source. Refer to Hangs. - Stalls: the kind of the stall episode (
hangorsuspend) and its duration. - Page views: the start of each page view, with its navigation type and its URL.
- Lifecycle: each transition of the lifecycle state, with the browser event that caused it.
- Compute pressure: each change of the pressure state of a source, with the state before it. The first record of a source has no state before it.
- Clock jumps: the kind, the direction and the magnitude of each jump.
- Long animation frames: the blocking duration of each frame that the monitor sent as an event, and the script that blocked the frame most.
- Browser reports: the type, the ID and the message of each report.
The Hangs and Stalls layers are on by default. The other layers can have many events, thus they are off by default. With all page loads, the marks of all pages are on the panels. Thus, first select one page load, then turn on the layers. A layer gets at most 500 events (maxLines: 500).
Each layer follows the Service and Page load variables. It does not follow Navigation type or Session. The query of a layer has this form:
{service_name=~"$service_name", event_name="lag.main_thread.hang"} | service_instance_id=~"$instance" | logfmt
The line of an event in Loki is the event name, then the attributes as sorted key=value pairs. Refer to events in Loki. The parser logfmt gets the attributes from the line. Thus the title, the text and the tags of a mark can use each attribute. The names have underscores in place of dots, for example lag_hang_source for lag.hang.source.
The time of a mark is the time of the log record. The lag library gives each event the time of its occurrence. For an event with a duration, for example a hang or a stall, it is the start. Thus a mark is at the time of the metric values that it explains.
An occurrence can be more than 30 minutes before or after the call. Then the record gets the time of the call, and the attribute lag.event.time has the time of the occurrence. Refer to the time of an event.
NoteThe query of a layer in Grafana 13
The Loki data source of Grafana 13 reads the query of an annotation layer from the layer itself (annotationQuery), not from its target. Thus each layer has the query in the top-level property expr, with maxLines, instant, titleFormat, textFormat and tagKeys. The target of the layer has the same query, for the query editor and for the verification script. With the query only in target, Grafana sent no query for the layers (commit 91a5cd8 of grafana-infra).
The verification script makes sure that the query of each layer finds events in the sample data. It does not examine the title and the text of the marks. The screenshots of the verification workflow show the marks.
Examine one page load
- Open the Lag Monitor dashboard.
- Go to the row Events (Loki).
- To find the page loads of one session, write the start of its session ID in Session.
- In the Page loads table, click the ID of a page load. If Grafana shows a menu, click "Show only this page load".
- At the top of the dashboard, turn on the annotation layers that you want to see.
The link of step 4 opens the dashboard with the Page load variable, the Service value and the time range in the URL. To see all page loads again, set Page load to All.
Read a panel
The queries of the Lag Monitor dashboard obey these rules:
- A histogram query first sums the rates of all page loads, then it takes the quantile, for example
histogram_quantile(0.95, sum(rate(x[$__rate_interval]))). Thus a p95 line is the p95 of all samples of the selected page loads. It is not a mean of the quantiles of each page load. Refer to aggregation in PromQL. - A counter query uses
rate()orincrease(), not the cumulative value. A Mimir panel with "per minute" in its title multiplies the rate in each second by 60. - Each query filters by the Service and the Page load variables. Only the queries of the row Long range do not filter by Page load.
- No Mimir query groups by
instanceor by session. Only the Page loads table groups by page load and by session. It counts the events of each page load in Loki. - A rate in a time series uses the range
$__rate_interval. The Mimir data source setstimeIntervalto15s, the export interval. Then$__rate_intervalis 60 seconds or more. - A stat panel with "(time range)" in its title uses an instant query with
increase(...[$__range]). The panels INP p75, Hangs, Stalls andperformance.now()resolution also use such a query. These panels give one value for the full time range of the dashboard. - A mean line uses
histogram_avg(). A count of observations useshistogram_count(). - The Loki tables use instant queries over
$__range, or they show the latest log lines.
The panels use these units:
| Unit | Meaning |
|---|---|
ms | Milliseconds. |
bytes | Bytes. |
none | A number without a unit, for example a count or a score. |
percentunit | A ratio from 0 to 1. Grafana shows it as a percentage. |
The Web Vitals panels show the good and the poor thresholds of each vital. The color changes to orange at the good threshold and to red at the poor threshold. The time series show the thresholds as dashed lines.
| Vital | Good | Poor | Unit |
|---|---|---|---|
| INP | 200 or less | More than 500 | ms |
| CLS | 0.1 or less | More than 0.25 | none |
| LCP | 2500 or less | More than 4000 | ms |
| FCP | 1800 or less | More than 3000 | ms |
| TTFB | 800 or less | More than 1800 | ms |
To see the metrics and the events of one page load, set the Page load variable. Refer to examine one page load. In Mimir, the ID of the page load is the instance label. In Loki, the same value is in service_instance_id.
The metric names are a contract
The panels, the variables and the recording rules use the metric names of the catalog, packages/lag/src/metric-catalog.ts. Mimir keeps these names, because it adds no unit suffix and no _total suffix. Thus the metric names are a contract between the catalog and the dashboard. A change of a name in the catalog without the same change in grafana-infra gives empty panels.
The names are in these files:
| Repository | File | What it has |
|---|---|---|
| lag | packages/lag/src/metric-catalog.ts | The catalog: the name, kind, unit and permitted attribute values of each metric, and the events with their attributes. |
| grafana-infra | scripts/lib/lag-catalog.mjs | A copy of the catalog. The dashboard script, the sample-data script and the verification script use it. |
| grafana-infra | scripts/build-dashboards.mjs | The panel queries, the annotation layers and the variables. The script writes lag-monitor.json. |
| grafana-infra | config/mimir/rules/anonymous/lag.yaml | The recording rules of 11 histograms. |
| grafana-infra | config/grafana/provisioning/dashboards/lag-monitor.json | The dashboard that the script writes. |
The attribute names are also part of the contract. The panels group by attributes, for example outcome, kind and navigation_type. The Loki panels use the event names and the event attributes, for example blocking_duration_ms. The annotation layers use the event names, and their marks use the event attributes, for example lag_hang_source.
On 2026-10-09, the copy of commit 6ebe0dc agrees with the catalog on the main branch of this repository. Both have 42 metrics (32 histograms and 10 counters) and 9 events, with the same names, kinds, units, attributes and permitted values. The 9 events are browser.web_vital, lag.main_thread.hang, lag.clock.jump, lag.long_animation_frame, lag.browser_report, lag.stall, lag.lifecycle.transition, lag.page_view.start and lag.pressure.change. The check verify --catalog of grafana-infra compares the two catalogs, and it found no difference.
The copy also has a list of optional attributes (optional) for 3 events. Some events of these names do not have these attributes:
| Event | Optional attributes |
|---|---|
lag.main_thread.hang | lag.hang.page_id, lag.hang.source |
lag.page_view.start | lag.page_view.url, lag.page_view.previous_id |
lag.pressure.change | previous_state |
The catalog of this repository has no such list. The verification script does not expect the optional attributes in the structured metadata of each event.
The dashboard uses all 42 metric names, and no other metric name that starts with lag_. The other lag_ names in the dashboard are the structured metadata of the events: lag_page_view_id, lag_page_view_url, lag_page_view_previous_id, lag_hang_page_id and lag_hang_source.
Examine a change of a name
-
Change the name in
metric-catalog.ts. -
In grafana-infra, change the name in
scripts/lib/lag-catalog.mjs. -
Change the name in
scripts/build-dashboards.mjs. The variables uselag_drift_histogramandlag_web_vital_lcp_histogram. -
If the metric has recording rules, change the name in
config/mimir/rules/anonymous/lag.yaml. -
Write the dashboard again:
pnpm dashboardsThe script stops with the message "Metrics without a panel" if a metric of the copy is in no query.
-
Search
lag-monitor.jsonandlag.yamlfor the old name. Make sure that the search finds nothing. -
Start the stack.
-
Send the sample data:
pnpm sample-data --summary sample-summary.json -
Start the verification script with the catalog of this repository:
pnpm verify --summary sample-summary.json --catalog ../lag/packages/lag/src/metric-catalog.ts
With --catalog, the verification script compares the copy with the catalog. It compares the names, the kinds, the units, the attribute names and the attributes of each event. It does not compare the permitted values of the attributes. It also examines if Mimir stores each metric under its catalog name, if each panel query gives data, and if each annotation layer finds events. The path in the example is correct when lag and grafana-infra are in the same directory.
The comparison of the catalogs is the first part of the output, and it does not use the stack. Without the stack, the next checks fail, and the script stops at the CORS check with the error fetch failed. The CI workflow of grafana-infra does not use --catalog, because it does not have a copy of this repository.
Add a panel
The script scripts/build-dashboards.mjs writes the Lag Monitor dashboard. Do not edit lag-monitor.json. To add a panel, do these steps:
-
Open
scripts/build-dashboards.mjsin grafana-infra. -
Find the
row()call of the row for the new panel. -
After the last panel of that row, add a call to
timeseries(),stat(),metricTable()oreventTable(). -
Give the targets with
prom(),promInstant(),loki()orlokiInstant(). -
Write each PromQL query with the helpers, for example
quantile(),sumRate(),perMinute()andobservationsPerMinute(). The helpers add the matchersservice_name=~"$service_name"andinstance=~"$instance". -
Start each LogQL query with the helper
events(), for exampleevents('="lag.stall"'). It adds the Service matcher and the Page load filter. -
Make sure that the query obeys the rules in Read a panel.
-
Write the dashboard again:
pnpm dashboards -
Look at each panel and its queries:
node scripts/build-dashboards.mjs --list -
Wait for 30 seconds.
-
Load the dashboard again in Grafana.
-
Send sample data.
-
Start the verification script. Make sure that the query of the new panel gives data.
-
Commit
scripts/build-dashboards.mjsandlag-monitor.jsontogether. The CI workflow writes the dashboard again, and it fails if the file is different.
This example adds the p75 of the drift lag to the drift row:
row("Drift (DriftLag)");
// The panels of the row
timeseries("Drift lag p75",
"The 75th percentile of the drift lag of all page loads.",
[prom(quantile(0.75, "lag_drift_histogram"), "p75")], { w : 12 });
By default, a time series panel has the unit ms, a width of 12 and a height of 8. A stat panel has the unit ms, a width of 4 and a height of 4. The layout puts the panels from left to right. It starts a new line when a panel does not fit in the 24 columns of the grid.
A panel for a new metric needs more steps. First add the metric to scripts/lib/lag-catalog.mjs. Then record the metric in scripts/send-sample-data.mjs. That script stops with "Metrics without values" if a metric of the copy has no values.
Add an annotation layer
The list ANNOTATIONS in scripts/build-dashboards.mjs has the layers. Each item is a call to annotation(), with the name of the layer, the event name, and the options color, enable, title, text and tags. The function writes the query with events() and | logfmt. It puts the query in expr and in target. This example is the Stalls layer:
annotation("Stalls", "lag.stall", { color : "orange", enable : true, title : "Stall: {{kind}}", text : "{{duration_ms}} ms", tags : "kind" }),
A layer for a new event needs more steps. First add the event to EVENTS in scripts/lib/lag-catalog.mjs. Then send the event in scripts/send-sample-data.mjs. Without events in the sample data, the verification script fails for the layer.
The verification workflow
The CI workflow .github/workflows/verify.yml of grafana-infra examines the stack, the sample data and the dashboard from end to end. It starts for each pull request, for each push to main, and by hand (workflow_dispatch). The job uses Node.js 24 and the pnpm version of packageManager in package.json. It stops after 30 minutes.
| Step | What it does |
|---|---|
The dashboard file agrees with scripts/build-dashboards.mjs | Writes the dashboard again. The step fails if git diff --exit-code finds a change in config/grafana/provisioning/dashboards/. |
| Start the stack | docker compose up -d. Then it waits up to 5 minutes for Alloy, Mimir, Loki and Grafana. |
| Send sample data | node scripts/send-sample-data.mjs --summary sample-summary.json |
| Verify the pipeline and the dashboard | node scripts/verify-pipeline.mjs --summary sample-summary.json, without --catalog |
| Screenshots of the dashboard | Installs Chromium for Playwright, then node scripts/screenshot-dashboard.mjs --summary sample-summary.json --out screenshots. |
| Upload the screenshots | Uploads the directory screenshots/ as the artifact dashboard-screenshots. |
| Stop the stack | docker compose down -v. When a step fails, the workflow first writes the last 200 log lines of the containers. |
The screenshots, the upload and the stop of the stack also occur when an earlier step fails.
The sample data
scripts/send-sample-data.mjs simulates six page loads of two services, with the OpenTelemetry JS SDK. It sends each event of the catalog copy, also lag.page_view.start, lag.lifecycle.transition and lag.pressure.change. Each event has the line of the event as its body, in the form of formatEventLine of the lag library. Thus the lines in Loki, which the annotation layers parse, have the same form as the lines of the library. A hang event and its stall event have the start of the hang as their time.
The checks of the dashboard
scripts/verify-pipeline.mjs examines the dashboard through the Grafana API:
- Grafana loads the dashboard.
- Each
label_values()variable gives values. The Session text box has no query. - Each panel query gives data. The script sends the query to
/api/ds/query. It replaces$service_namewith the services of the sample,$navigation_typeand$instancewith.+, and${session_id:regex}with an empty text. - The query of each annotation layer finds events. The script sends the query of the
targetof the layer.
Before these checks, the script compares the number of events of each name in Loki with the number that the sample sent. It waits up to 30 seconds for the last events of the sample. Then it examines the structured metadata of one event of each name, without the optional attributes of the catalog copy.
On 2026-10-09, the workflow passed on the merge commit 6ebe0dc (workflow run 37920009653). All 134 panel queries gave data, each of the 8 annotation layers found events, and the script passed 269 of 269 checks.
The screenshots
scripts/screenshot-dashboard.mjs takes screenshots of the dashboard with Playwright and Chromium, after the sample data. It reads the dashboard from Grafana and saves a copy with all annotation layers on: the UID lag-monitor-review and the title "Lag Monitor (all annotations)". A URL cannot turn on a layer, thus the script needs this copy. The time range is from 1 minute before the start of the sample to 1 minute after its end. The browser window is 1600 by 1200 pixels, and Grafana shows the dashboard in kiosk mode.
The script writes these files into the directory of --out. The default directory is screenshots.
| File | Content |
|---|---|
fleet.png | The full dashboard as provisioned, for all page loads. |
table-<title>.png | One screenshot for each table of events: Page loads, Recent page views and lifecycle transitions, Recent hangs and stalls, Recent clock jumps, and LoAF attribution. |
tables.json | For each table: the title, the file and the first 400 characters of its text, or the error. |
page-<index>.png | The review copy for one page load of the sample, with all annotation layers on. The option --pages selects the page loads. The default is 0,1. |
annotations.json | The annotation queries that Grafana sent, with the number of frames and rows of each result. Also the console errors of the browser, and the fields of the log frames of the Loki panels. |
To take the screenshots on your computer, do these steps:
-
Send the sample data with
--summary sample-summary.json. -
Install Chromium for Playwright:
pnpm exec playwright install chromium -
Take the screenshots:
node scripts/screenshot-dashboard.mjs --summary sample-summary.json -
Open the directory
screenshots.
The script sends no credentials to Grafana. It saves the dashboard lag-monitor-review in your Grafana, and it replaces an earlier copy. Git ignores the directory screenshots/.
The Lag Monitor panels
Each row of the dashboard has a section here. The table gives each panel, and the code block gives its queries from lag-monitor.json.
Overview
Six stat panels give the state of the selected services and page loads.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Active page loads | Stat | none | The number of page loads that sent drift samples in the minute before each point. Each page load has its own service.instance.id, thus its own series. The panel shows the last value. |
| Drift p95 | Stat | ms | The 95th percentile of the drift lag of all selected page loads. The panel shows the last value that is not null. |
| Heartbeat delay p95 | Stat | ms | The 95th percentile of the delivery delay of the worker heartbeats: the lag that an event at a random time sees. |
| INP p75 | Stat | ms | The 75th percentile of INP of the page views in the time range. The color changes to orange at 200 ms and to red at 500 ms. |
| Hangs | Stat | none | The main-thread hangs that the workers detected in the time range, with the outcomes ended and abandoned. |
| Stalls | Stat | none | The stall episodes in the time range: very long samples from hangs and from suspends. |
# Active page loads
count(count_over_time(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[1m])) or vector(0)
# Drift p95
histogram_quantile(0.95, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Heartbeat delay p95
histogram_quantile(0.95, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# INP p75
histogram_quantile(0.75, sum(increase(lag_web_vital_inp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# Hangs
sum(increase(lag_main_thread_hangs{service_name=~"$service_name", instance=~"$instance"}[$__range]))
# Stalls
sum(increase(lag_stalls{service_name=~"$service_name", instance=~"$instance"}[$__range]))
Drift (DriftLag)
Refer to DriftLag.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Drift lag p50 / p95 / p99 | Time series | ms | The quantiles and the mean of the lag of each window of chained timeouts (approximately 100 ms). Each block of the main thread in a window adds to the lag of the window. |
| Drift baseline: timer granularity | Time series | ms | The p50 and the p95 of the idle duration of one timer step: the mean of the recent steps that are not blocks. An increase needs a probe that shows an idle thread, thus a sustained load does not change it. It is the timer granularity of the browser and the operating system. DriftLag subtracts it. |
| Drift windows per second | Time series | none | The number of windows that DriftLag recorded each second, for all page loads. The monitors discard the windows of hidden and frozen pages. |
# Drift lag p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_avg(sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Drift baseline: timer granularity
histogram_quantile(0.5, sum (rate(lag_drift_baseline_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_drift_baseline_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Drift windows per second
histogram_count(sum (rate(lag_drift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Macrotask (MacrotaskLag)
Refer to MacrotaskLag.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Macrotask queue delay p50 / p95 / p99 | Time series | ms | The quantiles and the mean of the time that a zero-delay timeout waits in the task queue. Each page load measures one sample every 5 seconds. |
# Macrotask queue delay p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_avg(sum (rate(lag_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Scheduling fairness (SchedulingFairnessMonitor)
Refer to scheduling fairness.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Scheduling latency p50, by primitive | Time series | ms | The median latency of a microtask, a zero-delay timeout (setTimeout 0) and a MessageChannel message. The microtask latency stays near 0 and is the baseline. |
| Scheduling latency p99, by primitive | Time series | ms | The 99th percentile of the same three latencies. A high macrotask or MessageChannel value shows a busy task queue. |
# Scheduling latency p50, by primitive
histogram_quantile(0.5, sum (rate(lag_scheduling_microtask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.5, sum (rate(lag_scheduling_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.5, sum (rate(lag_scheduling_message_channel_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Scheduling latency p99, by primitive
histogram_quantile(0.99, sum (rate(lag_scheduling_microtask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_scheduling_macrotask_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_scheduling_message_channel_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Worker heartbeat (WorkerLagMonitor)
Refer to worker lag.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Heartbeat delivery delay: the lag that a random event sees | Time series | ms | The quantiles and the mean of the time that each worker heartbeat waited for the main thread. The worker sends a heartbeat each second from its own thread. Thus the heartbeats arrive at random times, as user events do. This is the primary lag estimator. |
| Worker self-lag | Time series | ms | The lateness of the heartbeat timer of the worker. A high value shows that the worker itself did not operate. Then its heartbeat delays are not reliable. |
| Worker clock offset | Time series | ms | The absolute offset between the worker clock and the main-thread clock, from the clock synchronization. |
# Heartbeat delivery delay: the lag that a random event sees
histogram_quantile(0.5, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_avg(sum (rate(lag_worker_main_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Worker self-lag
histogram_quantile(0.5, sum (rate(lag_worker_self_lag_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_worker_self_lag_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_worker_self_lag_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Worker clock offset
histogram_quantile(0.5, sum (rate(lag_worker_clock_offset_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_worker_clock_offset_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_worker_clock_offset_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Hangs (WorkerLagMonitor)
Refer to worker lag.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Hangs per minute, by outcome | Time series, bars | none | The hangs that the workers detected, each minute, for each outcome. In a hang, the main thread does not acknowledge the heartbeats. abandoned: the page closed or crashed during the hang. The next page of the origin, another open page of the origin, or the page itself at its close reported it. The attribute lag.hang.source of the event tells which (journal, peer or self). |
| Hang duration p50 / p95, by outcome | Time series | ms | The duration of each hang. For an abandoned hang, it is the duration until the worker saw the hang for the last time, or until the end of the page. |
# Hangs per minute, by outcome
60 * sum by (outcome) (rate(lag_main_thread_hangs{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Hang duration p50 / p95, by outcome
histogram_quantile(0.5, sum by (outcome) (rate(lag_main_thread_hang_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum by (outcome) (rate(lag_main_thread_hang_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
The annotation layer Hangs shows each hang event on the charts, with its phase, its duration and its source.
Measurement conditions (stalls and discarded samples)
Refer to measurement validity.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Stalls per minute, by kind | Time series, bars | none | The stall episodes each minute, for each kind. hang: no evidence of a suspend. suspend: evidence that the system stopped, for example a sleep of the device. |
| Stall duration p50 / p95, by kind | Time series | ms | The duration of each stall episode: its longest sample. |
| Samples discarded per minute, by reason | Time series, stacked | none | The samples that the monitors did not record, for each reason: the page was hidden or frozen, or the system was suspended during the measurement window. |
# Stalls per minute, by kind
60 * sum by (kind) (rate(lag_stalls{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Stall duration p50 / p95, by kind
histogram_quantile(0.5, sum by (kind) (rate(lag_stall_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum by (kind) (rate(lag_stall_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Samples discarded per minute, by reason
60 * sum by (reason) (rate(lag_samples_discarded{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
The annotation layer Stalls shows each lag.stall event on the charts, at the start of the stall episode.
Clock (ClockDriftMonitor and ClockReliabilityChecker)
Refer to clock drift and clock reliability.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Clock jumps per minute, by kind and direction | Time series, bars | none | The discontinuities between the wall clock and the monotonic clock each minute, for each kind and direction. suspend: the monotonic clock stopped while the device slept. step: the system clock changed. |
| Clock skew p50 / p95 / p99 | Time series | ms | The absolute difference between Date.now() and the absolute monotonic clock (timeOrigin plus performance.now()). |
performance.now() resolution | Stat | ms | The p50 and the p95 of the resolution of performance.now() of the page loads in the time range, with 3 decimals. Each page load measures it one time. |
# Clock jumps per minute, by kind and direction
60 * sum by (kind, direction) (rate(lag_clock_jumps{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Clock skew p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_clock_skew_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_clock_skew_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_clock_skew_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# performance.now() resolution
histogram_quantile(0.5, sum(increase(lag_clock_resolution_histogram{service_name=~"$service_name", instance=~"$instance"}[$__range])))
histogram_quantile(0.95, sum(increase(lag_clock_resolution_histogram{service_name=~"$service_name", instance=~"$instance"}[$__range])))
The annotation layer Clock jumps shows each lag.clock.jump event on the charts. It is off by default.
Long animation frames (LongAnimationFrameMonitor)
Refer to long animation frames.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| LoAF blocking duration p50 / p95 / p99 | Time series | ms | The blocking duration of each long animation frame. |
| LoAF duration p50 / p95 / p99 | Time series | ms | The total duration of each long animation frame. |
| Long animation frames per minute | Time series | none | The number of long animation frames of all page loads, each minute: the count of the duration histogram. |
| LoAF attribution: blocking time by script (events) | Table | Blocking time: ms, Frames: none | One row for each script, with the columns Invoker type, Invoker and Script URL. The columns Blocking time and Frames give the sum of the blocking time and the number of frames in the time range. The data comes from the lag.long_animation_frame events. The monitor sends this event only for a frame that blocks for 150 ms or more. The event names the script that blocked the frame most. The table sorts by the blocking time. Two instant queries give the values. The transformations put each label in a column and merge the two queries into one row for each script. |
# LoAF blocking duration p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_loaf_blocking_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_loaf_blocking_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_loaf_blocking_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# LoAF duration p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Long animation frames per minute
60 * histogram_count(sum (rate(lag_loaf_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# LoAF attribution: blocking time by script (events)
sum by (script_invoker_type, script_invoker, script_source_url) (sum_over_time({service_name=~"$service_name", event_name="lag.long_animation_frame"} | service_instance_id=~"$instance" | keep script_invoker_type, script_invoker, script_source_url, blocking_duration_ms | unwrap blocking_duration_ms [$__range]))
sum by (script_invoker_type, script_invoker, script_source_url) (count_over_time({service_name=~"$service_name", event_name="lag.long_animation_frame"} | service_instance_id=~"$instance" | keep script_invoker_type, script_invoker, script_source_url [$__range]))
The annotation layer Long animation frames shows each lag.long_animation_frame event on the charts, at the start of the frame. It is off by default.
Event timing (EventTimingMonitor)
Refer to Event Timing.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Event duration p95, by interaction type | Time series | ms | The 95th percentile of the duration of the interaction events of 16 ms or more, from the input to the next paint, for each interaction value. |
| Event phases p95 | Time series | ms | Where the time of an interaction goes: the input delay before the handlers, the processing time of the handlers, and the presentation delay to the next paint. |
| Interactions per minute, by interaction type | Time series, stacked | none | The interaction events of 16 ms or more, each minute, for each interaction value. |
# Event duration p95, by interaction type
histogram_quantile(0.95, sum by (interaction) (rate(lag_event_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Event phases p95
histogram_quantile(0.95, sum (rate(lag_event_input_delay_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_event_processing_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_event_presentation_delay_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Interactions per minute, by interaction type
60 * histogram_count(sum by (interaction) (rate(lag_event_duration_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Layout shift (LayoutShiftMonitor)
Refer to layout shift.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Layout shift score p50 / p95 / p99 | Time series | none | The score of each layout shift that did not follow user input, with 3 decimals. |
| Layout shifts per minute | Time series | none | The layout shifts without user input, each minute. |
# Layout shift score p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Layout shifts per minute
60 * histogram_count(sum (rate(lag_layout_shift_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Web Vitals (PageViewVitals)
Refer to page-view vitals. The histogram of a vital gets one value for each page view, usually when the page becomes hidden for the first time. Later changes go only into the browser.web_vital events.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| INP p75 (time range) | Stat | ms | The p75 of INP of the page views in the time range, for each navigation type. The color changes to orange at 200 ms and to red at 500 ms. |
| CLS p75 (time range) | Stat | none | The p75 of CLS, with 3 decimals. The color changes to orange at 0.1 and to red at 0.25. |
| LCP p75 (time range) | Stat | ms | The p75 of LCP. The color changes to orange at 2500 ms and to red at 4000 ms. |
| FCP p75 (time range) | Stat | ms | The p75 of FCP. The color changes to orange at 1800 ms and to red at 3000 ms. |
| TTFB p75 (time range) | Stat | ms | The p75 of TTFB. The color changes to orange at 800 ms and to red at 1800 ms. |
| Page views (time range) | Stat | none | The page views that reported LCP in the time range: the sample size of the vitals. The count is approximate, because increase() extrapolates. |
| INP p75, by navigation type | Time series | ms | The p75 of INP of the page views that reported in each interval, for each navigation type. Dashed lines show the thresholds. |
| CLS p75, by navigation type | Time series | none | The same for CLS. |
| LCP p75, by navigation type | Time series | ms | The same for LCP. |
| FCP p75, by navigation type | Time series | ms | The same for FCP. |
| TTFB p75, by navigation type | Time series | ms | The same for TTFB. |
| Web Vitals p75 from events, by navigation type | Table | Events: none | The p75 and the number of events of each vital and navigation type, from the browser.web_vital events. A page view sends a new event at each change of a vital. Thus this p75 also includes the earlier values. |
| Web Vitals p75 from events | Time series | ms | The p75 of browser_web_vital_value for each vital, from the events in a 5-minute window. CLS uses the right axis. |
# INP p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_inp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# CLS p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_cls_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# LCP p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_lcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# FCP p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_fcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# TTFB p75 (time range)
histogram_quantile(0.75, sum by (navigation_type) (increase(lag_web_vital_ttfb_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# Page views (time range)
histogram_count(sum(increase(lag_web_vital_lcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__range])))
# INP p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_inp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# CLS p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_cls_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# LCP p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_lcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# FCP p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_fcp_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# TTFB p75, by navigation type
histogram_quantile(0.75, sum by (navigation_type) (rate(lag_web_vital_ttfb_histogram{service_name=~"$service_name", instance=~"$instance", navigation_type=~"$navigation_type"}[$__rate_interval])))
# Web Vitals p75 from events, by navigation type
quantile_over_time(0.75, {service_name=~"$service_name", event_name="browser.web_vital"} | service_instance_id=~"$instance" | browser_web_vital_navigation_type=~"$navigation_type" | keep browser_web_vital_name, browser_web_vital_navigation_type, browser_web_vital_value | unwrap browser_web_vital_value [$__range]) by (browser_web_vital_name, browser_web_vital_navigation_type)
sum by (browser_web_vital_name, browser_web_vital_navigation_type) (count_over_time({service_name=~"$service_name", event_name="browser.web_vital"} | service_instance_id=~"$instance" | browser_web_vital_navigation_type=~"$navigation_type" | keep browser_web_vital_name, browser_web_vital_navigation_type [$__range]))
# Web Vitals p75 from events
quantile_over_time(0.75, {service_name=~"$service_name", event_name="browser.web_vital"} | service_instance_id=~"$instance" | browser_web_vital_navigation_type=~"$navigation_type" | keep browser_web_vital_name, browser_web_vital_value | unwrap browser_web_vital_value [5m]) by (browser_web_vital_name)
The annotation layer Page views shows the start of each page view (lag.page_view.start) on the charts. It is off by default.
Frames (FrameTimingMonitor)
Refer to frame timing.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Frames per second, delivered and dropped | Time series, stacked | none | The animation frames of all visible page loads each second, for each outcome. The monitor estimates the dropped frames from long gaps between frames. |
| Dropped frame ratio | Time series | percentunit | The dropped frames divided by all frames. |
| Frame delta p50 / p95 / p99 | Time series | ms | The time between two animation frame callbacks. 16.7 ms is 60 frames per second. |
# Frames per second, delivered and dropped
sum by (outcome) (rate(lag_frames{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Dropped frame ratio
sum(rate(lag_frames{service_name=~"$service_name", instance=~"$instance", outcome="dropped"}[$__rate_interval])) / sum(rate(lag_frames{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Frame delta p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_frame_delta_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_frame_delta_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_frame_delta_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Idle time (IdleAvailabilityMonitor)
Refer to idle availability.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Idle time remaining p10 / p50 / p90 | Time series | ms | The idle time that was available when an idle callback started. Low values show a busy main thread. |
| Idle gap p50 / p95 / p99 | Time series | ms | The time between two idle callbacks. |
| Idle callbacks per minute, by timed_out | Time series, stacked | none | The idle callbacks each minute, for each timed_out value. A callback that timed out started because no idle period came before its timeout. |
| Idle callbacks that timed out | Time series | percentunit | The idle callbacks that timed out, divided by all idle callbacks. |
# Idle time remaining p10 / p50 / p90
histogram_quantile(0.1, sum (rate(lag_idle_time_remaining_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.5, sum (rate(lag_idle_time_remaining_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.9, sum (rate(lag_idle_time_remaining_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Idle gap p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_idle_gap_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_idle_gap_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_idle_gap_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Idle callbacks per minute, by timed_out
60 * sum by (timed_out) (rate(lag_idle_callbacks{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Idle callbacks that timed out
sum(rate(lag_idle_callbacks{service_name=~"$service_name", instance=~"$instance", timed_out="true"}[$__rate_interval])) / sum(rate(lag_idle_callbacks{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
Memory (MemoryMonitor)
Refer to memory.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Used JS heap p50 / p95, by source | Time series | bytes | The used heap memory of each sample. modern: performance.measureUserAgentSpecificMemory(). legacy: performance.memory. |
| Heap usage ratio p50 / p95 | Time series | percentunit | The used heap divided by the heap limit. Only the legacy source gives the limit. |
# Used JS heap p50 / p95, by source
histogram_quantile(0.5, sum by (source) (rate(lag_memory_used_bytes_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum by (source) (rate(lag_memory_used_bytes_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Heap usage ratio p50 / p95
histogram_quantile(0.5, sum (rate(lag_memory_usage_ratio_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_memory_usage_ratio_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Compute pressure (ComputePressureMonitor)
Refer to compute pressure.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Mean pressure state, by source | Time series | none | The mean compute pressure state of the records: 0 nominal, 1 fair, 2 serious, 3 critical. The axis stops at 3. |
| Records at serious or critical pressure, by source | Time series | percentunit | The records with the state serious (2) or critical (3), divided by all records. |
# Mean pressure state, by source
histogram_avg(sum by (source) (rate(lag_pressure_state_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Records at serious or critical pressure, by source
1 - histogram_fraction(-Inf, 1.5, sum by (source) (rate(lag_pressure_state_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
The annotation layer Compute pressure shows each change of the pressure state of a source (lag.pressure.change) on the charts. It is off by default.
GC, page lifecycle and timer throttling
Refer to GC signal, lifecycle and timer throttle.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| GC events per minute | Time series | none | The garbage collections that the GC signal detector saw, each minute, for all page loads. |
| Lifecycle transitions per minute | Time series, bars | none | The page lifecycle transitions each minute, for each from, to and trigger. The legend is a table. |
| Throttled timer calibrations | Time series | percentunit | The timer calibration rounds that found throttled timers, divided by all rounds. |
# GC events per minute
60 * sum (rate(lag_gc_events{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Lifecycle transitions per minute
60 * sum by (from, to, trigger) (rate(lag_lifecycle_transitions{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Throttled timer calibrations
sum(rate(lag_timer_calibrations{service_name=~"$service_name", instance=~"$instance", throttled="true"}[$__rate_interval])) / sum(rate(lag_timer_calibrations{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
The annotation layer Lifecycle shows each lag.lifecycle.transition event on the charts. It is off by default. The table Recent page views and lifecycle transitions gives the same events.
Browser reports (BrowserReportMonitor)
Refer to browser reports.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Browser reports per minute, by type | Time series, bars | none | The reports from the Reporting API each minute, for each type: intervention or deprecation. |
| Recent browser reports (events) | Table | – | The latest lag.browser_report events (at most 200 lines): the service, the type, the ID, the message, the source file and the line. A Grafana log frame has its own id field, thus the query writes the report ID to report_id. |
# Browser reports per minute, by type
60 * sum by (type) (rate(lag_browser_reports{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval]))
# Recent browser reports (events)
{service_name=~"$service_name", event_name="lag.browser_report"} | service_instance_id=~"$instance" | label_format report_id=id
The annotation layer Browser reports shows each lag.browser_report event on the charts. It is off by default.
Shared-memory liveness (SharedLivenessMonitor)
Refer to shared-memory liveness.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Liveness block duration p50 / p95 / p99 | Time series | ms | The duration of each main-thread block that a worker saw through shared memory. Only cross-origin isolated pages have this monitor. |
| Liveness blocks per minute | Time series | none | The main-thread blocks that the shared-memory watcher saw, each minute. |
# Liveness block duration p50 / p95 / p99
histogram_quantile(0.5, sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.95, sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
histogram_quantile(0.99, sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
# Liveness blocks per minute
60 * histogram_count(sum (rate(lag_liveness_block_histogram{service_name=~"$service_name", instance=~"$instance"}[$__rate_interval])))
Events (Loki)
The panels of this row read the events in Loki. Refer to events in Loki.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| Events per minute, by event name | Time series | none | The lag events in Loki in the minute before each point, for each event_name. The panel has an interval of 15 s. |
| Page loads | Table | – | The page loads (SDK instances) with lag events in the time range. Each row has the columns Page load (service_instance_id), Session (session_id) and Events, the number of events in the time range. The table has at most 50 rows (topk(50, ...)), and it sorts by Events. The Session variable filters the rows. Each page load ID has the link "Show only this page load", which sets the Page load variable. The query also filters by Page load, thus with one page load the table has only that page load. |
| Recent page views and lifecycle transitions | Table | – | The latest lag.page_view.start and lag.lifecycle.transition events (at most 200 lines), with the columns Time, Event, Navigation, From, To, Trigger, URL, Page view, Previous view, Session and Instance. A page view start fills Navigation, URL and Previous view. The first page view of a page load has no previous view. A transition fills From, To and Trigger. The annotation layers Page views and Lifecycle show the same events on the charts. |
| Recent hangs and stalls | Table | Duration: ms | The latest lag.main_thread.hang and lag.stall events (at most 200 lines): the time, the service, the event, the phase, the kind, the duration, the sender (scope_name), the page view, the hung page, the session and the instance. The worker sends the start of a hang itself, with the scope @mark1russell7/lag/worker. An abandoned hang comes from the next page of the origin (journal), from another open page of the origin (peer), or from the page itself at its close (self). The table has no column for this source, lag_hang_source. The Hangs annotation layer shows it. |
| Recent clock jumps | Table | Magnitude: ms, Skew: ms, Lateness: ms | The latest lag.clock.jump events (at most 200 lines): the time, the service, the kind, the direction, the magnitude, the skew, the lateness, the page view and the session. |
The Page loads table uses the transformations of the LoAF attribution table: one column for each label, then one row for each series. The value of its instant query is Value or Value #A, by the version of Grafana, and the table shows it as Events. The other tables get their columns from the structured metadata of each log line (extractFields), and they sort by the time.
# Events per minute, by event name
sum by (event_name) (count_over_time({service_name=~"$service_name", event_name=~".+"} | service_instance_id=~"$instance" [1m]))
# Page loads
topk(50, sum by (service_instance_id, session_id) (count_over_time({service_name=~"$service_name", event_name=~".+"} | service_instance_id=~"$instance" | session_id=~"${session_id:regex}.*" [$__range])))
# Recent page views and lifecycle transitions
{service_name=~"$service_name", event_name=~"lag.page_view.start|lag.lifecycle.transition"} | service_instance_id=~"$instance"
# Recent hangs and stalls
{service_name=~"$service_name", event_name=~"lag.main_thread.hang|lag.stall"} | service_instance_id=~"$instance"
# Recent clock jumps
{service_name=~"$service_name", event_name="lag.clock.jump"} | service_instance_id=~"$instance"
Long range (recording rules, all page loads)
The panels of this row read the series of the recording rules, not the raw series. The rules use 5-minute windows, and Mimir evaluates them each minute. Use these panels for long time ranges. The six histograms are the drift, the macrotask delay, the heartbeat delay, the LoAF blocking duration, the event duration and the frame delta. The rules sum all page loads of a service, thus the Page load variable does not apply to this row.
| Panel | Type | Unit | What it shows |
|---|---|---|---|
| p95 by service (recorded) | Time series | ms | service_name:<metric>:p95_rate5m of the six histograms, for each service. The legend is a table. |
| p99 by service (recorded) | Time series | ms | service_name:<metric>:p99_rate5m of the same six histograms. |
| Fleet p95 from recorded rates | Time series | ms | The p95 of all selected services, from the recorded native-histogram rates service_name:<metric>:rate5m. A sum of recorded rates is correct. A sum of recorded quantiles is not correct. |
| Web Vitals p75 from recorded rates | Time series | ms | The p75 of each vital, from service_name_navigation_type:<vital>:rate5m. CLS uses the right axis. |
# p95 by service (recorded)
service_name:lag_drift_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_macrotask_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_worker_main_block_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_loaf_blocking_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_event_duration_histogram:p95_rate5m{service_name=~"$service_name"}
service_name:lag_frame_delta_histogram:p95_rate5m{service_name=~"$service_name"}
# p99 by service (recorded)
service_name:lag_drift_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_macrotask_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_worker_main_block_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_loaf_blocking_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_event_duration_histogram:p99_rate5m{service_name=~"$service_name"}
service_name:lag_frame_delta_histogram:p99_rate5m{service_name=~"$service_name"}
# Fleet p95 from recorded rates
histogram_quantile(0.95, sum(service_name:lag_drift_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_macrotask_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_worker_main_block_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_loaf_blocking_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_event_duration_histogram:rate5m{service_name=~"$service_name"}))
histogram_quantile(0.95, sum(service_name:lag_frame_delta_histogram:rate5m{service_name=~"$service_name"}))
# Web Vitals p75 from recorded rates
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_inp_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_cls_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_lcp_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_fcp_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
histogram_quantile(0.75, sum(service_name_navigation_type:lag_web_vital_ttfb_histogram:rate5m{service_name=~"$service_name", navigation_type=~"$navigation_type"}))
The Frontend Observability panels
The Frontend Observability dashboard is from the first commit of grafana-infra. The only later change is the query of the Log Entries panel. It uses the span metrics that the metrics generator of Tempo writes to Mimir (traces_spanmetrics_*), the traces in Tempo, and the logs in Loki. It has the variable service (label_values(traces_spanmetrics_calls_total, service)), but no panel query uses it.
The verification script does not examine this dashboard, thus its panels are not verified. The Loki configuration of the stack makes only service_name and event_name index labels. Thus the selector {exporter="OTLP"} of the Log Stream panel finds no stream (not verified).
| Row | Panel | Type | Unit | What it shows |
|---|---|---|---|---|
| Overview | Total Traces (1h) | Stat | short | The increase of traces_spanmetrics_calls_total in the last hour. |
| Overview | Errors (1h) | Stat | short | The same for the spans with the status code STATUS_CODE_ERROR. |
| Overview | Avg Span Duration | Stat | ms | The mean duration of the spans in 5-minute windows. |
| Overview | Log Entries (1h) | Stat | short | The number of log lines in Loki in the last hour, of all services. Loki does not accept the empty selector {}. |
| Page Load Performance | Page Load Duration (documentLoad spans) | Time series | ms | The p50, p95 and p99 duration of the documentLoad spans, from explicit buckets. |
| Page Load Performance | Span Rate by Type | Time series, stacked bars | short | The spans each second, for each span name. |
| HTTP Requests | HTTP Request Duration (fetch spans) | Time series | ms | The p50 and the p95 duration of the HTTP spans: span names that start with HTTP, fetch or an HTTP method. |
| HTTP Requests | HTTP Request Rate | Time series | reqps | The HTTP spans each second, and the HTTP spans with an error. |
| Traces & Logs | Recent Traces | Traces | – | A TraceQL search in Tempo: at most 20 traces with the status error. |
| Traces & Logs | Log Stream | Logs | – | The log lines of the stream selector {exporter="OTLP"}. |
| Service Map | Service Graph | Node graph | – | The service map of Tempo. |
# Total Traces (1h)
sum(increase(traces_spanmetrics_calls_total[1h]))
# Errors (1h)
sum(increase(traces_spanmetrics_calls_total{status_code="STATUS_CODE_ERROR"}[1h]))
# Avg Span Duration
sum(rate(traces_spanmetrics_latency_sum[5m])) / sum(rate(traces_spanmetrics_latency_count[5m])) * 1000
# Page Load Duration (documentLoad spans)
histogram_quantile(0.50, sum(rate(traces_spanmetrics_latency_bucket{span_name="documentLoad"}[5m])) by (le))
histogram_quantile(0.95, sum(rate(traces_spanmetrics_latency_bucket{span_name="documentLoad"}[5m])) by (le))
histogram_quantile(0.99, sum(rate(traces_spanmetrics_latency_bucket{span_name="documentLoad"}[5m])) by (le))
# Span Rate by Type
sum by (span_name) (rate(traces_spanmetrics_calls_total[5m]))
# HTTP Request Duration (fetch spans)
histogram_quantile(0.50, sum(rate(traces_spanmetrics_latency_bucket{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*"}[5m])) by (le))
histogram_quantile(0.95, sum(rate(traces_spanmetrics_latency_bucket{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*"}[5m])) by (le))
# HTTP Request Rate
sum(rate(traces_spanmetrics_calls_total{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*"}[5m]))
sum(rate(traces_spanmetrics_calls_total{span_name=~"HTTP.*|fetch.*|GET.*|POST.*|PUT.*|DELETE.*|PATCH.*", status_code="STATUS_CODE_ERROR"}[5m]))
# Log Entries (1h)
sum(count_over_time({service_name=~".+"} [1h]))
# Log Stream
{exporter="OTLP"}
Source
- grafana-infra on GitHub.
scripts/build-dashboards.mjs: the panels, the annotation layers and the variables of the Lag Monitor dashboard, and the query rules.scripts/lib/lag-catalog.mjs: the copy of the metric catalog.config/grafana/provisioning/dashboards:lag-monitor.json,frontend.jsonand the dashboard provider.scripts/send-sample-data.mjs: the sample data.scripts/verify-pipeline.mjs: the verification of each panel query and each annotation layer, and the comparison of the catalogs.scripts/screenshot-dashboard.mjs: the screenshots of the dashboard..github/workflows/verify.yml: the CI workflow.packages/lag/src/metric-catalog.tsin this repository: the metric catalog.- The Grafana documentation of provisioning.