Skip to content

Spans and metrics

Everything the OpenTelemetry exporter emits.

Names here are public API: renaming one breaks every dashboard built on it. They are defined in polars_telemetry.export.semconv, and a test asserts this page documents every one of them.

Query span

The span is named polars.collect.

Attribute Type Notes
polars.query_id str UUIDv7 from polars; time-ordered
polars.query.label str Set with polars_telemetry.label(); nested labels joined with /. Never a metric dimension
polars.plan.fingerprint str Hash of the plan shape — see below
polars.engine str streaming, or in-memory for eager operations and an explicit engine="in-memory". Absent when the query failed before planning
polars.cpu_ms float Summed node self time; exceeds wall time when parallel
polars.parallelism float cpu_ms / wall_ms
polars.parallel_efficiency float cpu_ms / wall_ms / cpu_count, 0–1
polars.cpu_count int Cores visible to the process
polars.node_count int Physical plan nodes
polars.result.rows int Rows reaching the sink, when reported

Call site

Where in your code the query ran, under OpenTelemetry's own code attributes, so a backend that already understands them links a query to its source.

Attribute Type Notes
code.file.path str Absolute path of the innermost frame outside polars
code.line.number int Line that ran the query
code.function.name str Enclosing function

Absent when the query came from code with no file on disk — exec, the REPL, or a notebook cell, whose temporary filename changes on every run — and for collect_async() and collect_batches(), which polars reports from its own threads.

This is the identity a person can act on. The fingerprint groups runs of the same plan but is a hash, and it changes whenever polars changes its optimiser; pipeline.py:142 does not. Disable with Config(call_site=False).

None of the three is a metric dimension: a line number changes whenever the file above it is edited, which would restart every series on an unrelated edit.

The fingerprint

A hash of node kinds, topology and column identity — not literal values. So amount > 10 and amount > 90 produce the same fingerprint, while a different grouping column produces a different one. A scanned file counts by its name, with numbers and dates masked: data-2024-01-01.parquet and data-2024-01-02.parquet in any directory are the same shape, orders.parquet is another. It is bounded by your code paths, which is what makes it safe as a metric dimension where polars.query_id is not.

Hot node

The single most expensive node, which is usually the whole answer.

Attribute Type Notes
polars.hot_node.kind str e.g. GroupBy, EquiJoin
polars.hot_node.cpu_ms float Its self time
polars.hot_node.share float Fraction of total CPU, 0–1

Diagnostics

Derived from counters already collected. Absent when the plan has no node of the relevant kind.

Attribute Type What it tells you
polars.filter.selectivity float Rows surviving the filter, 0–1
polars.filter.rows_dropped int Rows removed before the rest of the plan
polars.join.amplification float Rows out over probe-side rows in; above 1 is fan-out
polars.projection.efficiency float Columns read over columns in the file
polars.morsel.skew float Largest morsel over the mean; above 1 is uneven
polars.scan.predicate_pushed bool True if any scan filters inside the scan
polars.scan.row_groups_skipped bool Whether parquet row groups were skipped
polars.scan.has_statistics bool Whether the optimiser had table statistics

Plan shape

Attribute Type Notes
polars.scan.count int Number of scan nodes
polars.scan.sources str[] Paths or URIs scanned
polars.scan.predicates str[] Predicates pushed into the scan
polars.scan.columns int Columns actually read, summed across scans
polars.join.count int Number of join nodes
polars.join.types str[] e.g. INNER, LEFT
polars.join.keys str[] Left-hand join keys
polars.groupby.count int Number of group-by nodes
polars.groupby.keys str[] Grouping expressions
polars.sort.columns str[] Sort expressions

These are read from the IR plan, which keeps your own column names. The physical plan rewrites group-by keys and aggregations to _POLARS_TMP_N, so reading them from there would be useless to a human.

Data quality

Attribute Type Notes
polars.metrics.complete bool False when the closing snapshot caught unfinished nodes
polars.metrics.incomplete_nodes int How many; only set when non-zero

When polars.metrics.complete is false, every counter below is a floor, not a total — polars called close() before the engine had finished flushing.

The full plan

Attribute Type Notes
polars.plan str The whole plan and its counters as JSON

Off by default — it is kilobytes per span and identical for every run of a shape. Enable with Config(include_plan=True) when you want the topology, which nothing else carries. Contains both the physical and IR node lists with id, kind and inputs, plus every per-node counter.

Metrics

Query-level, dimensioned by polars.plan.fingerprint and polars.engine:

Instrument Type Unit
polars.query.duration histogram ms
polars.query.cpu_time histogram ms
polars.query.parallel_efficiency histogram 1

Node-level, dimensioned by polars.node.kind and polars.engine:

Instrument Type Unit
polars.node.cpu_time histogram ms
polars.node.poll_time histogram ms
polars.node.max_poll_time histogram ms
polars.node.state_update_time histogram ms
polars.node.max_state_update_time histogram ms
polars.node.largest_morsel histogram {row}
polars.node.stolen_ratio histogram 1
polars.node.io_time histogram ms
polars.node.rows_in counter {row}
polars.node.rows_out counter {row}
polars.node.morsels_in counter {morsel}
polars.node.morsels_out counter {morsel}
polars.node.polls counter {poll}
polars.node.state_updates counter {update}
polars.node.io_bytes counter By

polars.node.io_bytes and polars.node.largest_morsel carry one extra dimension, polars.direction. For bytes its values are requested, received and sent; for morsels, received and sent.

Every metric dimension is drawn from a bounded set — node kinds, io directions, engine, and the plan fingerprint. Plan literals are never metric attributes: their values are unbounded and would destroy series cardinality.

Attributes that can carry your data

These contain query content

polars.scan.sources · polars.scan.predicates · polars.join.keys · polars.groupby.keys · polars.sort.columns

polars.plan also contains plan detail, when enabled.

A filter on col("email") == "someone@example.com" arrives verbatim. Data and privacy covers how to mask it. The authoritative list is polars_telemetry.export.semconv.CARRIES_USER_DATA.