Skip to the content

Grafana stack

The Docker Compose backend that receives the data from the browser, with Grafana Alloy, Mimir, Loki, Tempo and Grafana.

Note

This page is a draft. The content is not complete and can change.

The grafana-infra repository has a Docker Compose stack for development and for the integration tests. The browser exports the metrics and the events with OTLP over HTTP. Grafana Alloy receives the data and sends the metrics to Mimir, the logs to Loki and the traces to Tempo. Grafana shows the dashboards. The page OpenTelemetry setup tells you how to configure the SDK in the page.

NoteThe versions

This page describes grafana-infra at commit 6ebe0dc, the main branch on 2026-10-09. The lockfile of this repository has this commit. The scripts infra:up and test:e2e of @lag/integration-tests start the stack from node_modules/@mark1russell7/grafana-infra. The settings of OpenTelemetry setup need otel-ts at commit a4a7722 or later, and the lockfile has that commit too.

The services

ServiceImageWhat it does in this pipeline
Alloygrafana/alloy:v1.20.1Receives OTLP from the pages and the workers. Prepares the log records for Loki. Sends each signal to its store.
Mimirgrafana/mimir:3.2.2Stores the metrics as Prometheus series, with native histograms. Evaluates the recording rules.
Lokigrafana/loki:3.6.0Stores the events and the logs.
Tempografana/tempo:2.7.2Stores the traces. Its metrics generator writes span metrics and service graphs to Mimir.
Grafanagrafana/grafana:13.0.10Shows the dashboards. It has a data source for Mimir, Loki, Tempo and Pyroscope.
Pyroscopegrafana/pyroscope:latestStores profiles. The lag monitors do not send profiles.

All images except Pyroscope have fixed versions. The Mimir configuration uses experimental options, and a new Mimir version can change their names. Docker starts a stopped container again, but not a container that you stopped (restart: unless-stopped). The data stays in the Docker volumes mimir_data, loki_data, tempo_data, grafana_data and pyroscope_data.

The data path

Show the diagram source
flowchart LR
    page["Page: OpenTelemetry SDK"] -- "OTLP/HTTP, port 4318" --> receiver["Alloy: OTLP receiver"]
    worker["Worker: hang reports"] -- "OTLP/HTTP JSON, /v1/logs" --> receiver
    receiver -- "metrics" --> batch["Alloy: batch"]
    receiver -- "logs" --> transform["Alloy: transform<br/>event name and line"]
    transform --> batch
    receiver -- "traces" --> batch
    batch -- "/otlp/v1/metrics" --> mimir["Mimir"]
    batch -- "/otlp/v1/logs" --> loki["Loki"]
    batch -- "OTLP/gRPC" --> tempo["Tempo"]
    tempo -- "span metrics" --> mimir
    mimir --> grafana["Grafana"]
    loki --> grafana
    tempo --> grafana
The page and its worker post OTLP/HTTP to Alloy. Alloy sends each signal to its store, and Grafana reads the stores.

The Alloy configuration is in config/alloy/config.alloy. It has these components:

ComponentSettingsWhat it does
otelcol.receiver.otlp "default"gRPC on 0.0.0.0:4317, HTTP on 0.0.0.0:4318Receives OTLP/gRPC, and OTLP/HTTP with JSON or protobuf. Sends the metrics and the traces to batch, and the logs to transform.
otelcol.processor.transform "event_name"error_mode = "ignore"Prepares the log records for Loki. Refer to events in Loki.
otelcol.processor.batch "default"timeout = "1s", send_batch_size = 1000Puts the data into batches. The short timeout gives fast feedback in development.
otelcol.exporter.otlphttp "mimir"http://mimir:9009/otlpSends the metrics to the OTLP endpoint of Mimir.
otelcol.exporter.otlphttp "loki"http://loki:3100/otlpSends the logs to the OTLP endpoint of Loki.
otelcol.exporter.otlp "tempo"tempo:4317, insecure = trueSends the traces to Tempo with OTLP/gRPC, without TLS.

Alloy keeps no state. A restart of Alloy loses no metric data, because browsers export cumulative values.

Send data from a page

A page sends its data to the OTLP/HTTP receiver of Alloy on port 4318:

DataURLStore
Metricshttp://localhost:4318/v1/metricsMimir
Logs and eventshttp://localhost:4318/v1/logsLoki

The traces go to the same receiver, and Alloy sends them to Tempo. With otel-ts, set endpoint: "http://localhost:4318". The worker of @mark1russell7/lag/worker sends its hang reports itself, as OTLP/HTTP JSON. For this stack, set workerHangReport.url to http://localhost:4318/v1/logs. Refer to the hang reports of the worker.

CORS

The origin of a page on a development server is not the origin of the receiver. Thus the receiver must answer the CORS requests of the browser. The cors block of the HTTP receiver has these settings:

SettingValue
allowed_originshttp://localhost, http://localhost:*, https://localhost, https://localhost:*, http://127.0.0.1, http://127.0.0.1:*, https://127.0.0.1, https://127.0.0.1:*
allowed_headers["*"]
max_age7200, the value of the Access-Control-Max-Age header

Thus pages on localhost and 127.0.0.1 can post, on all ports, through HTTP or HTTPS. A request with a JSON body (Content-Type: application/json) gets a CORS preflight. The receiver answers the preflight, and the answer has the origin of the page in Access-Control-Allow-Origin. The OTLP exporters of the page and the hang reports of the worker (fetch with keepalive) send JSON. Thus both get a preflight, and the receiver accepts both.

The Alloy configuration says that the answer also has Access-Control-Allow-Credentials: true. Thus navigator.sendBeacon with a JSON body also works. The verification script shows this header, but it does not examine it.

To accept a page from a different origin, do these steps:

  1. Add the origin to allowed_origins in config/alloy/config.alloy.

  2. Restart Alloy:

    docker compose restart alloy
    

A collector in production must also accept the CORS requests of the page. Refer to OpenTelemetry setup.

Metrics in Mimir

Mimir operates in monolithic mode, without multitenancy. Thus the tenant is anonymous. The configuration is in config/mimir/mimir.yaml. Its comments say that the option names agree with Mimir 3.2.2.

From OTLP to Prometheus series

Mimir translates each OTLP metric to Prometheus series:

  • service.name becomes the job label. With service.namespace, the job label is <namespace>/<name>.
  • service.instance.id becomes the instance label. Thus each page load writes its own series.
  • promote_otel_resource_attributes copies service.name, service.namespace, service.version, deployment.environment.name, browser.platform and browser.mobile to labels on each series, for example service_name. These attributes are constant for an SDK instance, so they add no series.
  • The other resource attributes go to the target_info series.
  • The metric names do not change. Mimir adds no unit suffix and no _total suffix, because otel_metric_suffixes_enabled is false. For example, lag_drift_histogram (unit ms) stays lag_drift_histogram, and the counter lag_stalls stays lag_stalls.

PromQL shows an info notice for rate() on a counter without the _total suffix. The notice has no effect on the result.

Cautionsession.id

Do not add session.id to promote_otel_resource_attributes. Each session then gets its own series. Also, session.id is not a correct writer identity.

Native histograms

The OpenTelemetry setup of this project records each histogram as an exponential histogram (histogramAggregation: "exponential"). With native_histograms_ingestion_enabled: true, Mimir stores each exponential histogram as one native histogram series, without _bucket, _sum and _count series. Mimir 3.2 enables this setting by default.

An explicit-bucket histogram becomes _bucket, _sum and _count series. The Lag Monitor dashboard does not use these series. The PromQL functions histogram_quantile(), histogram_count(), histogram_avg() and histogram_fraction() read a native histogram. Refer to aggregation in PromQL.

Other settings

SettingValueWhat it does
native_histograms_ingestion_enabledtrueStores exponential histograms as native histograms.
otel_created_timestamp_zero_ingestion_enabled (experimental)trueAdds a zero sample at the OTLP start time of each new series. Then rate() and increase() also count the first export of a page load. Mimir uses the start time only when it is at most 5 minutes before the sample.
promote_otel_resource_attributes (experimental)Six attributesCopies these resource attributes to labels on each series.
out_of_order_time_window10mAccepts samples up to 10 minutes older than the newest sample of a series, for example from retries and final exports. It does not merge writers that share a series.
max_global_series_per_user500000The series budget of the tenant.
max_global_series_per_metric100000The series budget of one metric.
query_scheduler.max_outstanding_requests_per_tenant4096Lets the Lag Monitor dashboard send all its queries at the same time.

The Lag Monitor dashboard sends approximately 135 queries when it loads: 134 panel queries, and one query for each annotation layer that is on. Query sharding splits each query into up to 16 requests. With the default limit of 100, Mimir answers 429 "too many outstanding requests" to some panels.

The series limits use the rule in-memory series = R × 2.5 h × S. S is the number of series of one page load, and the configuration uses S = 100. R is the number of page loads in one hour, and the stack plans for R = 2,000. Thus 2,000 × 2.5 × 100 = 500,000. The series of a closed page stay in memory until the next head compaction, thus the rule uses 2.5 h. Refer to the number of series.

Do not enable otel_native_delta_ingestion. Browsers export cumulative temporality, and the delta ingestion is experimental.

Recording rules

The Mimir ruler loads config/mimir/rules/anonymous/lag.yaml. Docker Compose mounts config/mimir/rules at /etc/mimir/rules, read-only. The ruler storage reads the files of each tenant from <directory>/<tenant>/. Thus the directory name anonymous is the tenant, and the file name is the rule namespace.

The quantile rules first sum the rates of all page loads, then they take the quantile. The rate5m rules store the summed rate as a native histogram. A panel can take any quantile of it, or sum it across services. The file has two rule groups, and Mimir evaluates each group every minute (interval: 1m):

GroupRules
lag-latency-5m18 rules: three rules for each of six latency histograms
lag-web-vitals-5m5 rules: one rule for each Web Vital
RuleExpression
service_name:<metric>:rate5msum by (service_name) (rate(<metric>[5m])), a native histogram
service_name:<metric>:p95_rate5mThe p95 of the same rate
service_name:<metric>:p99_rate5mThe p99 of the same rate
service_name_navigation_type:<vital>:rate5msum by (service_name, navigation_type) (rate(<vital>[5m]))

The latency rules use these histograms:

  • lag_drift_histogram
  • lag_macrotask_histogram
  • lag_worker_main_block_histogram
  • lag_loaf_blocking_histogram
  • lag_event_duration_histogram
  • lag_frame_delta_histogram

The vital rules use the five lag_web_vital_*_histogram metrics. Mimir accepts at most 20 rules in one group (ruler_max_rules_per_rule_group).

Use the recorded series for long time ranges. To get a quantile of many services, sum the rate5m series first. Do not sum or average the recorded quantiles. For example, the Lag Monitor dashboard gets the p95 of all selected services with this query:

histogram_quantile(0.95, sum(service_name:lag_drift_histogram:rate5m{service_name=~"$service_name"}))

The long range row of the dashboard uses the rules. The rules sum all page loads, thus the Page load variable of the dashboard does not apply to them. Mimir reads the rule file again every 10 minutes. To load a change at once, restart Mimir:

docker compose restart mimir

Events in Loki

The monitors send each event as an OpenTelemetry log record with an event name (createOtelEventSink). The worker sends its hang reports as OTLP/HTTP JSON log records. The transform processor of Alloy changes each log record with these two statements:

set(log.body, ToKeyValueString(log.attributes, "=", " ", true)) where log.body == nil or log.body == ""
set(log.attributes["event.name"], log.event_name) where log.event_name != ""
  1. A record without a body gets its attributes as the line, as sorted key=value pairs. Loki drops an entry when the previous entry of its stream has the same timestamp and the same line. A browser can send several events in the same millisecond, for example the five Web Vitals of one report. The library gives each of its events a body with the attributes (refer to the next table). The statement applies to other senders that send no body. The sample data of grafana-infra sends a body in the form of the library.
  2. The Alloy configuration says that Loki 3.6 and 3.7 ignore the EventName field of an OTLP log record. Thus Alloy copies the event name to the log attribute event.name. Loki makes the index label event_name from this attribute.

Labels and structured metadata

The setting limits_config.otlp_config in config/loki/loki.yaml sets the labels:

  • The index labels are service_name, from the resource attribute service.name, and event_name, from the log attribute event.name. Both have few values.
  • ignore_defaults: true removes the default list of index labels. That list makes service.instance.id an index label, but each page load has a new value.
  • All other attributes of the resource and of the record go to the structured metadata. Loki changes the dots in the attribute names to underscores. For example, service.instance.id becomes service_instance_id, session.id becomes session_id, and browser.web_vital.value becomes browser_web_vital_value.
  • The structured metadata also has the scope name in scope_name, for example @mark1russell7/lag/worker for the hang reports of the worker.
  • A log record without an event name, for example from createOtelLoggerAdapter, goes to a stream without the label event_name.

The line of an entry is the body of the record:

SenderBody, and the line in Loki
The event sink of the monitors (createOtelEventSink)The event name, then the attributes as sorted key=value pairs (formatEventLine). For example, lag.stall duration_ms=6200 kind=hang lag.page_view.id=lag-1760000000000-4821733201941.
The hang reports of the workerThe same form, with the name lag.main_thread.hang, duration_ms, phase and the context of the page
The logger adapter (createOtelLoggerAdapter)The log message
A record without a bodyNo body. Alloy writes the attributes as the line.

NoteEvents of the same millisecond

The attributes in the line make the lines of two events different, also when the two events have the same name and the same timestamp. Thus the five Web Vitals of one report have five different lines. A unit test of the event sink makes sure that such events get different bodies. Two events with the same name, the same attributes and the same timestamp still have the same line.

The sample data of grafana-infra makes its bodies with a copy of the rules of formatEventLine. The verification script makes sure that Loki has each event that the sample data sent. It does not examine the bodies of the library itself.

The annotation layers of the Lag Monitor dashboard parse the line with logfmt. Thus the title and the text of a mark can use each attribute of the event. Refer to annotation layers.

This LogQL query gives the p75 of each Web Vital in 5-minute windows:

quantile_over_time(0.75,
  {service_name="shop", event_name="browser.web_vital"}
  | keep browser_web_vital_name, browser_web_vital_value
  | unwrap browser_web_vital_value [5m]) by (browser_web_vital_name)

Loki has no authentication in this stack (auth_enabled: false). It stores the data in the local file system, with the tsdb index and the schema v13. For more information, refer to OTLP ingestion in the Loki documentation.

Traces in Tempo

Alloy sends the traces to Tempo with OTLP/gRPC, on port 4317 in the Docker network. Tempo stores them in the local file system. The lag monitors send no traces. The traces come from the tracing of the SDK, for example otel-ts with tracing: true.

The metrics generator of Tempo makes service graphs and span metrics, with the dimension service.name. It writes them to Mimir at http://mimir:9009/api/v1/push, with exemplars. The Frontend Observability dashboard uses these span metrics.

Grafana

Grafana gets its configuration from config/grafana/grafana.ini and from the provisioning files in config/grafana/provisioning:

  • The administrator is admin with the password admin. The variables GF_SECURITY_ADMIN_USER and GF_SECURITY_ADMIN_PASSWORD in .env set these values.
  • grafana.ini enables the anonymous access, with the role Admin. The README of grafana-infra says the same, and it tells you to use this setting only on your own computer. The screenshot script of grafana-infra saves a dashboard through the API without credentials, and this step passed in the CI workflow on 2026-10-09. You can also sign in as admin to edit a dashboard.
  • grafana.ini also sets allow_embedding = true and the dark theme.
  • The variable GF_INSTALL_PLUGINS in docker-compose.yml adds the plugin grafana-pyroscope-app.

The provisioning files add these data sources:

NameTypeUIDURLSettings
Mimir (default)prometheusmimirhttp://mimir:9009/prometheushttpMethod: POST and timeInterval: 15s. Exemplars link to the traces in Tempo.
Lokilokilokihttp://loki:3100The derived field TraceID links the trace IDs in a log line to Tempo.
Tempotempotempohttp://tempo:3200Links from traces to the metrics in Mimir and to the logs in Loki (tag service.name). A node graph and a service map.
Pyroscopegrafana-pyroscope-datasource–http://pyroscope:4040–

timeInterval: 15s is the export interval of the browser. Then $__rate_interval is 60 seconds or more. The file provider Default loads the dashboards in config/grafana/provisioning/dashboards and reads them again each 30 seconds. Refer to dashboards.

Ports and addresses

The file .env sets the host ports. In the Docker network, the services use the service names, for example mimir:9009.

ServiceAddress on the hostVariableWhat it gives
Alloy OTLP/HTTPhttp://localhost:4318OTLP_HTTP_PORTThe endpoint for pages and workers: /v1/metrics and /v1/logs.
Alloy OTLP/gRPClocalhost:4317OTLP_GRPC_PORTThe endpoint for OTLP/gRPC senders.
Alloy UIhttp://localhost:12345ALLOY_UI_PORTThe user interface of Alloy, and /-/ready.
Mimirhttp://localhost:9009MIMIR_PORTThe Prometheus API at /prometheus, the OTLP endpoint at /otlp, /metrics and /ready.
Lokihttp://localhost:3100LOKI_PORTThe Loki API at /loki/api/v1, the OTLP endpoint at /otlp, and /ready.
Tempohttp://localhost:3200TEMPO_PORTThe Tempo API. Tempo also receives OTLP/gRPC on tempo:4317 in the Docker network.
Pyroscopehttp://localhost:4040PYROSCOPE_PORTThe Pyroscope API.
Grafanahttp://localhost:3000GRAFANA_PORTThe dashboards, and /api/health.

If a port is in use, set a different port in the environment. This example moves Grafana to port 3300:

GRAFANA_PORT=3300 docker compose up -d

The containers have fixed names, for example mimir and grafana. To change a name, set its variable, for example MIMIR_CONTAINER_NAME.

Start the stack

You need Docker with Docker Compose. The scripts of grafana-infra also need Node.js and pnpm.

  1. Open a terminal in the directory of grafana-infra.

  2. Start the stack:

    docker compose up -d
    
  3. Open Grafana at http://localhost:3000.

  4. To edit a dashboard, sign in as admin with the password admin.

When another program uses port 3000, set GRAFANA_PORT before the start, for example GRAFANA_PORT=3007. In this repository, pnpm run infra:up in packages/lag-integration-tests starts the stack from node_modules/@mark1russell7/grafana-infra, at the commit of the lockfile.

Check the stack

These addresses tell you if a service is ready:

ServiceAddressThe answer contains
Alloyhttp://localhost:12345/-/readyready
Mimirhttp://localhost:9009/readyready
Lokihttp://localhost:3100/readyready
Grafanahttp://localhost:3000/api/healthok

To examine the full pipeline, send sample data and start the verification script:

  1. Install the dependencies of the scripts:

    pnpm install
    
  2. Start the stack.

  3. Send sample data for 3 minutes:

    pnpm sample-data --summary sample-summary.json
    
  4. Examine the pipeline:

    pnpm verify --summary sample-summary.json
    

scripts/send-sample-data.mjs uses the OpenTelemetry JS SDK in the same way as a browser. It simulates six page loads of two services, lag-sample-shop and lag-sample-docs. Each page load has its own service.instance.id. The script records each metric of the catalog and sends each lag event, with the line of the event as the body. It also posts hang reports in the JSON format of the lag worker, with fetch, keepalive and the origin http://localhost:5173. The options --duration, --interval, --pages, --seed and --endpoint change the run.

scripts/verify-pipeline.mjs examines these items:

  • Alloy, Mimir, Loki and Grafana are ready.
  • The receiver answers the CORS preflight of localhost pages. It does not answer the preflight of other origins, for example https://example.com.
  • Mimir stores each metric of the catalog under its catalog name, and each histogram as a native histogram.
  • The instance label is the service.instance.id. No series has a session_id label.
  • The recording rules are healthy and have data.
  • Loki has only the index labels service_name and event_name. It has all the lag events with their structured metadata, and the hang reports of the worker. For each event name, Loki has the number of events that the sample sent. The script waits up to 30 seconds for the last events.
  • Grafana loads the Lag Monitor dashboard, each panel query gives data, and each annotation layer finds events.

The exit code is 1 if a check fails. If Grafana uses a different port, add --grafana http://localhost:3300. To compare the catalog copy scripts/lib/lag-catalog.mjs with the lag catalog, add --catalog ../lag/packages/lag/src/metric-catalog.ts.

The script scripts/screenshot-dashboard.mjs takes screenshots of the dashboard after the sample data. The CI workflow .github/workflows/verify.yml of grafana-infra starts the stack, sends the sample data, starts the verification script and uploads the screenshots. Refer to the verification workflow.

The integration tests

The e2e project of @lag/integration-tests uses the stack:

  • pnpm test:e2e starts the stack with infra:up. Then it starts lag-monitors.test.ts and stress.test.ts in Chromium.
  • The tests export with otel-ts to http://localhost:4318, each 5 s.
  • The tests use the default histograms of otel-ts, which are explicit-bucket histograms. Thus the histogram panels of the Lag Monitor dashboard do not show their data.
  • In the e2e project, the tests must find the _count series of lag_drift_histogram in Mimir. The query accepts the label service_name or job, and the unit suffix _milliseconds.

Stop the stack

  1. Stop the stack:

    docker compose down
    
  2. To also delete the data in the Docker volumes, add -v:

    docker compose down -v
    

To stop the stack of node_modules, use pnpm run infra:down in packages/lag-integration-tests. To restart one service, for example after a change of its configuration, use docker compose restart <service>.

Troubleshooting

ProblemCauseWhat to do
The histogram panels show no data.The app sends explicit-bucket histograms. Mimir stores them as _bucket, _sum and _count series, and the dashboard does not use them.Configure exponential histograms.
The series have no instance label.The app does not set service.instance.id. Then many page loads write to the same series, and rate() gives incorrect values. The Lag Monitor dashboard also does not show these series, because its Page load variable sends instance=~".+".Give each SDK instance a new service.instance.id.
Mimir rejects samples.Mimir counts each rejection, with its reason, in cortex_discarded_samples_total.Look at cortex_discarded_samples_total at http://localhost:9009/metrics.
The browser shows a CORS error.The origin of the page is not in allowed_origins.Add the origin to allowed_origins in config/alloy/config.alloy. Then restart Alloy.
A panel shows the error "too many outstanding requests".The dashboard sends more requests than the limit of the query scheduler.Increase max_outstanding_requests_per_tenant in config/mimir/mimir.yaml.

Source

  • grafana-infra on GitHub.
  • docker-compose.yml and .env: the services, the images, the ports and the volumes.
  • config/alloy/config.alloy: the receiver, the CORS settings, the transform processor and the exporters.
  • config/mimir/mimir.yaml and config/mimir/rules/anonymous/lag.yaml: the Mimir settings and the recording rules.
  • config/loki/loki.yaml and config/tempo/tempo.yaml: the labels of Loki, and the metrics generator of Tempo.
  • config/grafana: grafana.ini, the data sources and the dashboard provider.
  • scripts/send-sample-data.mjs and scripts/verify-pipeline.mjs: the sample data and the verification.
  • scripts/screenshot-dashboard.mjs and .github/workflows/verify.yml: the screenshots of the dashboard and the CI workflow.
  • The Grafana documentation of Alloy, Mimir, Loki and Tempo.

lag: Main-thread responsiveness monitoring for browser apps, exported as OpenTelemetry metrics.

To change a page, edit its file in packages/site/content/. The writing style guide tells you how.

An AI model (Claude, from Anthropic) wrote most of the text and the code of this site and of the library, under the direction of the author. The tests and the STE linter examine them. The writing standard gives the reason for this note.