Skip to main content
Conduit exposes operational metrics in the Prometheus exposition format — the de-facto standard scraped natively by Prometheus, Grafana Alloy, the OpenTelemetry Collector, Datadog, VictoriaMetrics, Telegraf, and the managed Prometheus services of the major clouds. The same metrics are also pushed over OTLP to every instance-level OpenTelemetry exporter configured in Settings → Instance → Telemetry (and to one configured via the standard OTEL_EXPORTER_OTLP_* environment variables), alongside the traces and events Conduit already sends — see Telemetry Export. If your observability backend ingests OTLP, you get these metrics with no extra configuration; the scrape endpoint below is for Prometheus-style pull pipelines. Metrics are aggregates over the whole instance, so they are not part of what a workspace’s own exporters receive.

Enabling the endpoint

Set CONDUIT_METRICS_ADDR to start a dedicated listener serving /metrics:
Unset (the default), no listener starts and nothing is exposed. On Kubernetes with the Helm chart, set metrics.enabled: true — it sets the variable and adds a metrics container port (9464). The port is deliberately not part of the chart’s Service or Ingress: scraping happens at the pod IP, and the endpoint can never be exposed through the public entry point by accident.

Security posture

The endpoint is unauthenticated by design — the standard posture for scrape endpoints across self-hosted software (scrapers are robots; access control is network position). It reveals operational shape (connector names, request rates, session counts) but never payloads, secrets, or per-user data. Treat it accordingly:
  • Bind a private interface, or rely on network segmentation (on Kubernetes, a NetworkPolicy restricting the port to your monitoring namespace).
  • Never route it through a public ingress or load balancer.
  • If your network is flat and you need transport auth, front the port with an authenticating proxy — Prometheus supports bearer tokens, basic auth, OAuth2, and mutual TLS in scrape_config.

Scraping

A minimal Prometheus scrape_config:
On Kubernetes, any pod-scraping mechanism works. With annotation-based discovery, add to the chart’s podAnnotations:
With the Prometheus operator, a PodMonitor selecting the release’s pod labels and the metrics port does the same job. Each replica is scraped individually and reports its own counters — that is the Prometheus model, not a limitation. Aggregate across replicas in queries: sum by (connector) (rate(conduit_mcp_tool_calls_total[5m])).

Restarts and counter resets

Metrics are aggregated in memory and reset to zero when the process restarts. Consumers are built for this: rate() and increase() detect and splice across resets, and at most the few seconds since the last scrape are lost. Durable history lives in your monitoring backend, not in Conduit — anything Conduit must count exactly (usage dashboards, the audit log) is stored in its database independently of these metrics.

Metric reference

Conventions: counters end in _total, durations are histograms in seconds (exposed as _bucket, _sum, and _count series), and every label has a small, bounded value set — metrics never carry user, client, or per-request identifiers.

Gateway

HTTP

Database

Each database metric is reported per database (db is main or usage) and per connection pool (pool is read or write):

Process

Standard Go runtime and process metrics (go_* — goroutines, memory, GC — and process_* — CPU, RSS, file descriptors, start time) are exposed alongside the application metrics.

Useful starting queries

Paste these into the Prometheus Graph tab or a Grafana panel. Most of them also make sensible first alert rules — thresholds depend on your traffic, so chart them for a while before alerting.

Tool-call error rate

The fraction of executed tool calls that failed, per connector. A single connector spiking usually means an upstream outage or an expired credential; everything spiking at once points at the gateway or its network.

Tool-call latency

The 99th-percentile duration of executed tool calls. Aggregating _bucket series with sum by (le) makes this correct across replicas; add connector to the by clause to break it down per connector.

Policy denials

Calls blocked by access policy. A steady trickle is normal (clients probe tools they can’t see); a step change after a policy edit means the edit cut off something real, and a sustained stream from one connector is a client misconfigured or probing.

Authentication failures

Rejected MCP requests by reason. Alert on sustained invalid_token volume — that’s either a fleet of clients with stale credentials or someone guessing bearer tokens.

Replica health

Prometheus synthesizes an up series per scrape target — no Conduit metric needed. This is the first alert to configure.

Database pool saturation

Time spent waiting for a database connection, per second. Zero is the healthy steady state; any sustained non-zero value means requests are queuing for the pool and latency is following.

Usage pipeline drops

Usage events dropped because the recorder’s buffer was full. Anything above zero means the in-product usage dashboards are undercounting under load.