Grafana stack
The Docker Compose backend that receives the data from the browser, with Grafana Alloy, Mimir, Loki, Tempo and Grafana.
Note
The grafana-infra repository has a Docker Compose stack for development and for the integration tests. The browser exports the metrics and the events with OTLP over HTTP. Grafana Alloy receives the data and sends the metrics to Mimir, the logs to Loki and the traces to Tempo. Grafana shows the dashboards. The page OpenTelemetry setup tells you how to configure the SDK in the page.
NoteThe versions
This page describes grafana-infra at commit 6ebe0dc, the main branch on 2026-10-09. The lockfile of this repository has this commit. The scripts infra:up and test:e2e of @lag/integration-tests start the stack from node_modules/@mark1russell7/grafana-infra. The settings of OpenTelemetry setup need otel-ts at commit a4a7722 or later, and the lockfile has that commit too.
The services
| Service | Image | What it does in this pipeline |
|---|---|---|
| Alloy | grafana/alloy:v1.20.1 | Receives OTLP from the pages and the workers. Prepares the log records for Loki. Sends each signal to its store. |
| Mimir | grafana/mimir:3.2.2 | Stores the metrics as Prometheus series, with native histograms. Evaluates the recording rules. |
| Loki | grafana/loki:3.6.0 | Stores the events and the logs. |
| Tempo | grafana/tempo:2.7.2 | Stores the traces. Its metrics generator writes span metrics and service graphs to Mimir. |
| Grafana | grafana/grafana:13.0.10 | Shows the dashboards. It has a data source for Mimir, Loki, Tempo and Pyroscope. |
| Pyroscope | grafana/pyroscope:latest | Stores profiles. The lag monitors do not send profiles. |
All images except Pyroscope have fixed versions. The Mimir configuration uses experimental options, and a new Mimir version can change their names. Docker starts a stopped container again, but not a container that you stopped (restart: unless-stopped). The data stays in the Docker volumes mimir_data, loki_data, tempo_data, grafana_data and pyroscope_data.
The data path
Show the diagram source
flowchart LR
page["Page: OpenTelemetry SDK"] -- "OTLP/HTTP, port 4318" --> receiver["Alloy: OTLP receiver"]
worker["Worker: hang reports"] -- "OTLP/HTTP JSON, /v1/logs" --> receiver
receiver -- "metrics" --> batch["Alloy: batch"]
receiver -- "logs" --> transform["Alloy: transform<br/>event name and line"]
transform --> batch
receiver -- "traces" --> batch
batch -- "/otlp/v1/metrics" --> mimir["Mimir"]
batch -- "/otlp/v1/logs" --> loki["Loki"]
batch -- "OTLP/gRPC" --> tempo["Tempo"]
tempo -- "span metrics" --> mimir
mimir --> grafana["Grafana"]
loki --> grafana
tempo --> grafanaThe Alloy configuration is in config/alloy/config.alloy. It has these components:
| Component | Settings | What it does |
|---|---|---|
otelcol.receiver.otlp "default" | gRPC on 0.0.0.0:4317, HTTP on 0.0.0.0:4318 | Receives OTLP/gRPC, and OTLP/HTTP with JSON or protobuf. Sends the metrics and the traces to batch, and the logs to transform. |
otelcol.processor.transform "event_name" | error_mode = "ignore" | Prepares the log records for Loki. Refer to events in Loki. |
otelcol.processor.batch "default" | timeout = "1s", send_batch_size = 1000 | Puts the data into batches. The short timeout gives fast feedback in development. |
otelcol.exporter.otlphttp "mimir" | http://mimir:9009/otlp | Sends the metrics to the OTLP endpoint of Mimir. |
otelcol.exporter.otlphttp "loki" | http://loki:3100/otlp | Sends the logs to the OTLP endpoint of Loki. |
otelcol.exporter.otlp "tempo" | tempo:4317, insecure = true | Sends the traces to Tempo with OTLP/gRPC, without TLS. |
Alloy keeps no state. A restart of Alloy loses no metric data, because browsers export cumulative values.
Send data from a page
A page sends its data to the OTLP/HTTP receiver of Alloy on port 4318:
| Data | URL | Store |
|---|---|---|
| Metrics | http://localhost:4318/v1/metrics | Mimir |
| Logs and events | http://localhost:4318/v1/logs | Loki |
The traces go to the same receiver, and Alloy sends them to Tempo. With otel-ts, set endpoint: "http://localhost:4318". The worker of @mark1russell7/lag/worker sends its hang reports itself, as OTLP/HTTP JSON. For this stack, set workerHangReport.url to http://localhost:4318/v1/logs. Refer to the hang reports of the worker.
CORS
The origin of a page on a development server is not the origin of the receiver. Thus the receiver must answer the CORS requests of the browser. The cors block of the HTTP receiver has these settings:
| Setting | Value |
|---|---|
allowed_origins | http://localhost, http://localhost:*, https://localhost, https://localhost:*, http://127.0.0.1, http://127.0.0.1:*, https://127.0.0.1, https://127.0.0.1:* |
allowed_headers | ["*"] |
max_age | 7200, the value of the Access-Control-Max-Age header |
Thus pages on localhost and 127.0.0.1 can post, on all ports, through HTTP or HTTPS. A request with a JSON body (Content-Type: application/json) gets a CORS preflight. The receiver answers the preflight, and the answer has the origin of the page in Access-Control-Allow-Origin. The OTLP exporters of the page and the hang reports of the worker (fetch with keepalive) send JSON. Thus both get a preflight, and the receiver accepts both.
The Alloy configuration says that the answer also has Access-Control-Allow-Credentials: true. Thus navigator.sendBeacon with a JSON body also works. The verification script shows this header, but it does not examine it.
To accept a page from a different origin, do these steps:
-
Add the origin to
allowed_originsinconfig/alloy/config.alloy. -
Restart Alloy:
docker compose restart alloy
A collector in production must also accept the CORS requests of the page. Refer to OpenTelemetry setup.
Metrics in Mimir
Mimir operates in monolithic mode, without multitenancy. Thus the tenant is anonymous. The configuration is in config/mimir/mimir.yaml. Its comments say that the option names agree with Mimir 3.2.2.
From OTLP to Prometheus series
Mimir translates each OTLP metric to Prometheus series:
service.namebecomes thejoblabel. Withservice.namespace, thejoblabel is<namespace>/<name>.service.instance.idbecomes theinstancelabel. Thus each page load writes its own series.promote_otel_resource_attributescopiesservice.name,service.namespace,service.version,deployment.environment.name,browser.platformandbrowser.mobileto labels on each series, for exampleservice_name. These attributes are constant for an SDK instance, so they add no series.- The other resource attributes go to the
target_infoseries. - The metric names do not change. Mimir adds no unit suffix and no
_totalsuffix, becauseotel_metric_suffixes_enabledis false. For example,lag_drift_histogram(unitms) stayslag_drift_histogram, and the counterlag_stallsstayslag_stalls.
PromQL shows an info notice for rate() on a counter without the _total suffix. The notice has no effect on the result.
Cautionsession.id
Do not add session.id to promote_otel_resource_attributes. Each session then gets its own series. Also, session.id is not a correct writer identity.
Native histograms
The OpenTelemetry setup of this project records each histogram as an exponential histogram (histogramAggregation: "exponential"). With native_histograms_ingestion_enabled: true, Mimir stores each exponential histogram as one native histogram series, without _bucket, _sum and _count series. Mimir 3.2 enables this setting by default.
An explicit-bucket histogram becomes _bucket, _sum and _count series. The Lag Monitor dashboard does not use these series. The PromQL functions histogram_quantile(), histogram_count(), histogram_avg() and histogram_fraction() read a native histogram. Refer to aggregation in PromQL.
Other settings
| Setting | Value | What it does |
|---|---|---|
native_histograms_ingestion_enabled | true | Stores exponential histograms as native histograms. |
otel_created_timestamp_zero_ingestion_enabled (experimental) | true | Adds a zero sample at the OTLP start time of each new series. Then rate() and increase() also count the first export of a page load. Mimir uses the start time only when it is at most 5 minutes before the sample. |
promote_otel_resource_attributes (experimental) | Six attributes | Copies these resource attributes to labels on each series. |
out_of_order_time_window | 10m | Accepts samples up to 10 minutes older than the newest sample of a series, for example from retries and final exports. It does not merge writers that share a series. |
max_global_series_per_user | 500000 | The series budget of the tenant. |
max_global_series_per_metric | 100000 | The series budget of one metric. |
query_scheduler.max_outstanding_requests_per_tenant | 4096 | Lets the Lag Monitor dashboard send all its queries at the same time. |
The Lag Monitor dashboard sends approximately 135 queries when it loads: 134 panel queries, and one query for each annotation layer that is on. Query sharding splits each query into up to 16 requests. With the default limit of 100, Mimir answers 429 "too many outstanding requests" to some panels.
The series limits use the rule in-memory series = R × 2.5 h × S. S is the number of series of one page load, and the configuration uses S = 100. R is the number of page loads in one hour, and the stack plans for R = 2,000. Thus 2,000 × 2.5 × 100 = 500,000. The series of a closed page stay in memory until the next head compaction, thus the rule uses 2.5 h. Refer to the number of series.
Do not enable otel_native_delta_ingestion. Browsers export cumulative temporality, and the delta ingestion is experimental.
Recording rules
The Mimir ruler loads config/mimir/rules/anonymous/lag.yaml. Docker Compose mounts config/mimir/rules at /etc/mimir/rules, read-only. The ruler storage reads the files of each tenant from <directory>/<tenant>/. Thus the directory name anonymous is the tenant, and the file name is the rule namespace.
The quantile rules first sum the rates of all page loads, then they take the quantile. The rate5m rules store the summed rate as a native histogram. A panel can take any quantile of it, or sum it across services. The file has two rule groups, and Mimir evaluates each group every minute (interval: 1m):
| Group | Rules |
|---|---|
lag-latency-5m | 18 rules: three rules for each of six latency histograms |
lag-web-vitals-5m | 5 rules: one rule for each Web Vital |
| Rule | Expression |
|---|---|
service_name:<metric>:rate5m | sum by (service_name) (rate(<metric>[5m])), a native histogram |
service_name:<metric>:p95_rate5m | The p95 of the same rate |
service_name:<metric>:p99_rate5m | The p99 of the same rate |
service_name_navigation_type:<vital>:rate5m | sum by (service_name, navigation_type) (rate(<vital>[5m])) |
The latency rules use these histograms:
lag_drift_histogramlag_macrotask_histogramlag_worker_main_block_histogramlag_loaf_blocking_histogramlag_event_duration_histogramlag_frame_delta_histogram
The vital rules use the five lag_web_vital_*_histogram metrics. Mimir accepts at most 20 rules in one group (ruler_max_rules_per_rule_group).
Use the recorded series for long time ranges. To get a quantile of many services, sum the rate5m series first. Do not sum or average the recorded quantiles. For example, the Lag Monitor dashboard gets the p95 of all selected services with this query:
histogram_quantile(0.95, sum(service_name:lag_drift_histogram:rate5m{service_name=~"$service_name"}))
The long range row of the dashboard uses the rules. The rules sum all page loads, thus the Page load variable of the dashboard does not apply to them. Mimir reads the rule file again every 10 minutes. To load a change at once, restart Mimir:
docker compose restart mimir
Events in Loki
The monitors send each event as an OpenTelemetry log record with an event name (createOtelEventSink). The worker sends its hang reports as OTLP/HTTP JSON log records. The transform processor of Alloy changes each log record with these two statements:
set(log.body, ToKeyValueString(log.attributes, "=", " ", true)) where log.body == nil or log.body == ""
set(log.attributes["event.name"], log.event_name) where log.event_name != ""
- A record without a body gets its attributes as the line, as sorted
key=valuepairs. Loki drops an entry when the previous entry of its stream has the same timestamp and the same line. A browser can send several events in the same millisecond, for example the five Web Vitals of one report. The library gives each of its events a body with the attributes (refer to the next table). The statement applies to other senders that send no body. The sample data of grafana-infra sends a body in the form of the library. - The Alloy configuration says that Loki 3.6 and 3.7 ignore the
EventNamefield of an OTLP log record. Thus Alloy copies the event name to the log attributeevent.name. Loki makes the index labelevent_namefrom this attribute.
Labels and structured metadata
The setting limits_config.otlp_config in config/loki/loki.yaml sets the labels:
- The index labels are
service_name, from the resource attributeservice.name, andevent_name, from the log attributeevent.name. Both have few values. ignore_defaults: trueremoves the default list of index labels. That list makesservice.instance.idan index label, but each page load has a new value.- All other attributes of the resource and of the record go to the structured metadata. Loki changes the dots in the attribute names to underscores. For example,
service.instance.idbecomesservice_instance_id,session.idbecomessession_id, andbrowser.web_vital.valuebecomesbrowser_web_vital_value. - The structured metadata also has the scope name in
scope_name, for example@mark1russell7/lag/workerfor the hang reports of the worker. - A log record without an event name, for example from
createOtelLoggerAdapter, goes to a stream without the labelevent_name.
The line of an entry is the body of the record:
| Sender | Body, and the line in Loki |
|---|---|
The event sink of the monitors (createOtelEventSink) | The event name, then the attributes as sorted key=value pairs (formatEventLine). For example, lag.stall duration_ms=6200 kind=hang lag.page_view.id=lag-1760000000000-4821733201941. |
| The hang reports of the worker | The same form, with the name lag.main_thread.hang, duration_ms, phase and the context of the page |
The logger adapter (createOtelLoggerAdapter) | The log message |
| A record without a body | No body. Alloy writes the attributes as the line. |
NoteEvents of the same millisecond
The attributes in the line make the lines of two events different, also when the two events have the same name and the same timestamp. Thus the five Web Vitals of one report have five different lines. A unit test of the event sink makes sure that such events get different bodies. Two events with the same name, the same attributes and the same timestamp still have the same line.
The sample data of grafana-infra makes its bodies with a copy of the rules of formatEventLine. The verification script makes sure that Loki has each event that the sample data sent. It does not examine the bodies of the library itself.
The annotation layers of the Lag Monitor dashboard parse the line with logfmt. Thus the title and the text of a mark can use each attribute of the event. Refer to annotation layers.
This LogQL query gives the p75 of each Web Vital in 5-minute windows:
quantile_over_time(0.75,
{service_name="shop", event_name="browser.web_vital"}
| keep browser_web_vital_name, browser_web_vital_value
| unwrap browser_web_vital_value [5m]) by (browser_web_vital_name)
Loki has no authentication in this stack (auth_enabled: false). It stores the data in the local file system, with the tsdb index and the schema v13. For more information, refer to OTLP ingestion in the Loki documentation.
Traces in Tempo
Alloy sends the traces to Tempo with OTLP/gRPC, on port 4317 in the Docker network. Tempo stores them in the local file system. The lag monitors send no traces. The traces come from the tracing of the SDK, for example otel-ts with tracing: true.
The metrics generator of Tempo makes service graphs and span metrics, with the dimension service.name. It writes them to Mimir at http://mimir:9009/api/v1/push, with exemplars. The Frontend Observability dashboard uses these span metrics.
Grafana
Grafana gets its configuration from config/grafana/grafana.ini and from the provisioning files in config/grafana/provisioning:
- The administrator is
adminwith the passwordadmin. The variablesGF_SECURITY_ADMIN_USERandGF_SECURITY_ADMIN_PASSWORDin.envset these values. grafana.inienables the anonymous access, with the roleAdmin. The README of grafana-infra says the same, and it tells you to use this setting only on your own computer. The screenshot script of grafana-infra saves a dashboard through the API without credentials, and this step passed in the CI workflow on 2026-10-09. You can also sign in asadminto edit a dashboard.grafana.inialso setsallow_embedding = trueand the dark theme.- The variable
GF_INSTALL_PLUGINSindocker-compose.ymladds the plugingrafana-pyroscope-app.
The provisioning files add these data sources:
| Name | Type | UID | URL | Settings |
|---|---|---|---|---|
| Mimir (default) | prometheus | mimir | http://mimir:9009/prometheus | httpMethod: POST and timeInterval: 15s. Exemplars link to the traces in Tempo. |
| Loki | loki | loki | http://loki:3100 | The derived field TraceID links the trace IDs in a log line to Tempo. |
| Tempo | tempo | tempo | http://tempo:3200 | Links from traces to the metrics in Mimir and to the logs in Loki (tag service.name). A node graph and a service map. |
| Pyroscope | grafana-pyroscope-datasource | – | http://pyroscope:4040 | – |
timeInterval: 15s is the export interval of the browser. Then $__rate_interval is 60 seconds or more. The file provider Default loads the dashboards in config/grafana/provisioning/dashboards and reads them again each 30 seconds. Refer to dashboards.
Ports and addresses
The file .env sets the host ports. In the Docker network, the services use the service names, for example mimir:9009.
| Service | Address on the host | Variable | What it gives |
|---|---|---|---|
| Alloy OTLP/HTTP | http://localhost:4318 | OTLP_HTTP_PORT | The endpoint for pages and workers: /v1/metrics and /v1/logs. |
| Alloy OTLP/gRPC | localhost:4317 | OTLP_GRPC_PORT | The endpoint for OTLP/gRPC senders. |
| Alloy UI | http://localhost:12345 | ALLOY_UI_PORT | The user interface of Alloy, and /-/ready. |
| Mimir | http://localhost:9009 | MIMIR_PORT | The Prometheus API at /prometheus, the OTLP endpoint at /otlp, /metrics and /ready. |
| Loki | http://localhost:3100 | LOKI_PORT | The Loki API at /loki/api/v1, the OTLP endpoint at /otlp, and /ready. |
| Tempo | http://localhost:3200 | TEMPO_PORT | The Tempo API. Tempo also receives OTLP/gRPC on tempo:4317 in the Docker network. |
| Pyroscope | http://localhost:4040 | PYROSCOPE_PORT | The Pyroscope API. |
| Grafana | http://localhost:3000 | GRAFANA_PORT | The dashboards, and /api/health. |
If a port is in use, set a different port in the environment. This example moves Grafana to port 3300:
GRAFANA_PORT=3300 docker compose up -d
$env:GRAFANA_PORT = "3300"
docker compose up -d
The containers have fixed names, for example mimir and grafana. To change a name, set its variable, for example MIMIR_CONTAINER_NAME.
Start the stack
You need Docker with Docker Compose. The scripts of grafana-infra also need Node.js and pnpm.
-
Open a terminal in the directory of grafana-infra.
-
Start the stack:
docker compose up -d -
Open Grafana at
http://localhost:3000. -
To edit a dashboard, sign in as
adminwith the passwordadmin.
When another program uses port 3000, set GRAFANA_PORT before the start, for example GRAFANA_PORT=3007. In this repository, pnpm run infra:up in packages/lag-integration-tests starts the stack from node_modules/@mark1russell7/grafana-infra, at the commit of the lockfile.
Check the stack
These addresses tell you if a service is ready:
| Service | Address | The answer contains |
|---|---|---|
| Alloy | http://localhost:12345/-/ready | ready |
| Mimir | http://localhost:9009/ready | ready |
| Loki | http://localhost:3100/ready | ready |
| Grafana | http://localhost:3000/api/health | ok |
To examine the full pipeline, send sample data and start the verification script:
-
Install the dependencies of the scripts:
pnpm install -
Start the stack.
-
Send sample data for 3 minutes:
pnpm sample-data --summary sample-summary.json -
Examine the pipeline:
pnpm verify --summary sample-summary.json
scripts/send-sample-data.mjs uses the OpenTelemetry JS SDK in the same way as a browser. It simulates six page loads of two services, lag-sample-shop and lag-sample-docs. Each page load has its own service.instance.id. The script records each metric of the catalog and sends each lag event, with the line of the event as the body. It also posts hang reports in the JSON format of the lag worker, with fetch, keepalive and the origin http://localhost:5173. The options --duration, --interval, --pages, --seed and --endpoint change the run.
scripts/verify-pipeline.mjs examines these items:
- Alloy, Mimir, Loki and Grafana are ready.
- The receiver answers the CORS preflight of
localhostpages. It does not answer the preflight of other origins, for examplehttps://example.com. - Mimir stores each metric of the catalog under its catalog name, and each histogram as a native histogram.
- The
instancelabel is theservice.instance.id. No series has asession_idlabel. - The recording rules are healthy and have data.
- Loki has only the index labels
service_nameandevent_name. It has all the lag events with their structured metadata, and the hang reports of the worker. For each event name, Loki has the number of events that the sample sent. The script waits up to 30 seconds for the last events. - Grafana loads the Lag Monitor dashboard, each panel query gives data, and each annotation layer finds events.
The exit code is 1 if a check fails. If Grafana uses a different port, add --grafana http://localhost:3300. To compare the catalog copy scripts/lib/lag-catalog.mjs with the lag catalog, add --catalog ../lag/packages/lag/src/metric-catalog.ts.
The script scripts/screenshot-dashboard.mjs takes screenshots of the dashboard after the sample data. The CI workflow .github/workflows/verify.yml of grafana-infra starts the stack, sends the sample data, starts the verification script and uploads the screenshots. Refer to the verification workflow.
The integration tests
The e2e project of @lag/integration-tests uses the stack:
pnpm test:e2estarts the stack withinfra:up. Then it startslag-monitors.test.tsandstress.test.tsin Chromium.- The tests export with otel-ts to
http://localhost:4318, each 5 s. - The tests use the default histograms of otel-ts, which are explicit-bucket histograms. Thus the histogram panels of the Lag Monitor dashboard do not show their data.
- In the
e2eproject, the tests must find the_countseries oflag_drift_histogramin Mimir. The query accepts the labelservice_nameorjob, and the unit suffix_milliseconds.
Stop the stack
-
Stop the stack:
docker compose down -
To also delete the data in the Docker volumes, add
-v:docker compose down -v
To stop the stack of node_modules, use pnpm run infra:down in packages/lag-integration-tests. To restart one service, for example after a change of its configuration, use docker compose restart <service>.
Troubleshooting
| Problem | Cause | What to do |
|---|---|---|
| The histogram panels show no data. | The app sends explicit-bucket histograms. Mimir stores them as _bucket, _sum and _count series, and the dashboard does not use them. | Configure exponential histograms. |
The series have no instance label. | The app does not set service.instance.id. Then many page loads write to the same series, and rate() gives incorrect values. The Lag Monitor dashboard also does not show these series, because its Page load variable sends instance=~".+". | Give each SDK instance a new service.instance.id. |
| Mimir rejects samples. | Mimir counts each rejection, with its reason, in cortex_discarded_samples_total. | Look at cortex_discarded_samples_total at http://localhost:9009/metrics. |
| The browser shows a CORS error. | The origin of the page is not in allowed_origins. | Add the origin to allowed_origins in config/alloy/config.alloy. Then restart Alloy. |
| A panel shows the error "too many outstanding requests". | The dashboard sends more requests than the limit of the query scheduler. | Increase max_outstanding_requests_per_tenant in config/mimir/mimir.yaml. |
Source
- grafana-infra on GitHub.
docker-compose.ymland.env: the services, the images, the ports and the volumes.config/alloy/config.alloy: the receiver, the CORS settings, thetransformprocessor and the exporters.config/mimir/mimir.yamlandconfig/mimir/rules/anonymous/lag.yaml: the Mimir settings and the recording rules.config/loki/loki.yamlandconfig/tempo/tempo.yaml: the labels of Loki, and the metrics generator of Tempo.config/grafana:grafana.ini, the data sources and the dashboard provider.scripts/send-sample-data.mjsandscripts/verify-pipeline.mjs: the sample data and the verification.scripts/screenshot-dashboard.mjsand.github/workflows/verify.yml: the screenshots of the dashboard and the CI workflow.- The Grafana documentation of Alloy, Mimir, Loki and Tempo.