How it works¶
The hook¶
polars imports a module named polars_cloud and reads QueryCloudObserver off
it by name, then duck-types the result. It never checks the type. So this
package supplies that name.
If the real polars-cloud is installed, its factory is kept and forwarded to.
Otherwise a module is registered in sys.modules under that name — we never
publish a distribution called polars_cloud, which would collide with theirs.
The protocol, verified against polars 1.44.1 and 1.44.2:
polars_cloud.authenticate()
polars_cloud.QueryCloudObserver(workspace, organization) -> observer
observer.on_query_started(query_id: UUID)
observer.on_query_planned(query_id, handle, ir: bytes, phys: bytes) -> guard
observer.on_query_failed(...)
guard.close()
ir and phys are MessagePack plans. handle has exactly one public method,
snapshot_query_metrics(), returning MessagePack per-node counters keyed by
phys_node_key — the same ids as the physical plan, so metrics attribute to
nodes without any name matching.
Failure isolation¶
Instrumentation runs inside your data path, so nothing here may surface as an exception in your query. Every callback is wrapped: errors are counted, each distinct one logged once, and past a threshold the hook disarms itself for the rest of the process.
polars-telemetry: disabling observer after 5 errors. Queries are unaffected.
on_query_planned is a special case — polars calls close() on whatever it
returns, so it always returns a guard even when it has failed internally.
Why there are no per-node spans¶
The counters polars reports are cumulative, and none of the twenty fields is a timestamp. A node interval could therefore only be sampled, which costs 4.7% of query wall time at 25 ms and 15.1% at 5 ms — and still resolves poorly: on a 48 ms query sampled at 5 ms, eight of eleven nodes collapsed onto two identical windows.
Read once at query end, the same counters are exact and cost nothing measurable. If polars exposes per-node timestamps, node spans become exact and free, and they go back in.
Overhead¶
The instrumentation itself, measured with an exporter that does nothing, stays below measurement noise on a 3M-row join and aggregation, interleaved against an uninstrumented run on the same engine. CI fails if it exceeds 10%.
Exporters add their own cost on top, on the thread that ran the query. Each exporter's page states it.