> ## Documentation Index
> Fetch the complete documentation index at: https://pipedream.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Monitoring

> The Prometheus metrics endpoint — how to enable and scrape it, and a reference for every metric Conduit exposes.

Conduit exposes operational metrics in the Prometheus exposition format — the
de-facto standard scraped natively by Prometheus, Grafana Alloy, the
OpenTelemetry Collector, Datadog, VictoriaMetrics, Telegraf, and the managed
Prometheus services of the major clouds.

The same metrics are also **pushed over OTLP** to every instance-level
OpenTelemetry exporter configured in **Settings → Instance → Telemetry** (and
to one configured via the standard `OTEL_EXPORTER_OTLP_*` environment
variables), alongside the traces and events Conduit already sends — see
[Telemetry Export](/docs/conduit/configure/telemetry). If your observability backend
ingests OTLP, you get these metrics with no extra configuration; the scrape
endpoint below is for Prometheus-style pull pipelines. Metrics are aggregates
over the whole instance, so they are not part of what a workspace's own
exporters receive.

## Enabling the endpoint

Set `CONDUIT_METRICS_ADDR` to start a dedicated listener serving `/metrics`:

```bash theme={null}
docker run -p 7272:7272 \
  -e CONDUIT_METRICS_ADDR=0.0.0.0:9464 \
  ghcr.io/pipedreamhq/conduit
```

Unset (the default), no listener starts and nothing is exposed.

On Kubernetes with the [Helm chart](/docs/conduit/deploy/kubernetes), set `metrics.enabled: true`
— it sets the variable and adds a `metrics` container port (9464). The port is
deliberately **not** part of the chart's Service or Ingress: scraping happens
at the pod IP, and the endpoint can never be exposed through the public entry
point by accident.

### Security posture

The endpoint is unauthenticated by design — the standard posture for scrape
endpoints across self-hosted software (scrapers are robots; access control is
network position). It reveals operational shape (connector names, request
rates, session counts) but never payloads, secrets, or per-user data. Treat it
accordingly:

* Bind a private interface, or rely on network segmentation (on Kubernetes, a
  NetworkPolicy restricting the port to your monitoring namespace).
* Never route it through a public ingress or load balancer.
* If your network is flat and you need transport auth, front the port with an
  authenticating proxy — Prometheus supports bearer tokens, basic auth, OAuth2,
  and mutual TLS in `scrape_config`.

## Scraping

A minimal Prometheus `scrape_config`:

```yaml theme={null}
scrape_configs:
  - job_name: conduit
    static_configs:
      - targets: ["conduit.internal:9464"]
```

On Kubernetes, any pod-scraping mechanism works. With annotation-based
discovery, add to the chart's `podAnnotations`:

```yaml theme={null}
podAnnotations:
  prometheus.io/scrape: "true"
  prometheus.io/port: "9464"
```

With the Prometheus operator, a `PodMonitor` selecting the release's pod
labels and the `metrics` port does the same job.

Each replica is scraped individually and reports its own counters — that is
the Prometheus model, not a limitation. Aggregate across replicas in queries:
`sum by (connector) (rate(conduit_mcp_tool_calls_total[5m]))`.

### Restarts and counter resets

Metrics are aggregated in memory and reset to zero when the process restarts.
Consumers are built for this: `rate()` and `increase()` detect and splice
across resets, and at most the few seconds since the last scrape are lost.
Durable history lives in your monitoring backend, not in Conduit — anything
Conduit must count exactly (usage dashboards, the audit log) is stored in its
database independently of these metrics.

## Metric reference

Conventions: counters end in `_total`, durations are histograms in seconds
(exposed as `_bucket`, `_sum`, and `_count` series), and every label has a
small, bounded value set — metrics never carry user, client, or per-request
identifiers.

### Gateway

| Metric | Description |
| - | - |
| `conduit_mcp_tool_calls_total` | Counter of tool calls served by the MCP gateway, by `connector` and `outcome`. Outcomes: `ok`, `error` (the call ran and failed), and `denied` (blocked by access policy before running). |
| `conduit_mcp_tool_call_duration_seconds` | Histogram of executed tool-call duration, by `connector`, including the upstream round trip. Denied calls never run, so they record no duration. |
| `conduit_mcp_prompt_gets_total` | Counter of prompt fetches (`prompts/get`) served by the MCP gateway, by `connector` and `outcome` (`ok`, `error`, `denied`). |
| `conduit_mcp_resource_reads_total` | Counter of resource reads (`resources/read`) served by the MCP gateway, by `connector` and `outcome` (`ok`, `error`, `denied`). |
| `conduit_mcp_auth_failures_total` | Counter of rejected MCP requests, by `reason`: `missing_token`, `invalid_token`, `session_expired`, `org_ambiguous`, `client_revoked`, `client_not_allowed`, `sign_in_not_accepted`, `sign_in_too_old`, `sign_in_provider_gone`, `sign_in_password_off`, or `rate_limited`. A burst of the credential reasons is the signature of a client with a broken credential — or of someone guessing; `client_revoked` and `client_not_allowed` are a workspace refusing a valid token's client (a workspace revoke, or its policy against dynamically registered clients); the `sign_in_*` reasons are a workspace's sign-in policy refusing how the token's holder signed in, or the way they signed in being turned off (a provider disabled or deleted, or password sign-in turned off). |
| `conduit_mcp_sse_sessions` | Gauge of currently connected MCP SSE sessions on this replica. A sudden drop to zero while the process is up usually means a load balancer or proxy severed the streams. |

### HTTP

| Metric | Description |
| - | - |
| `http_server_request_duration_seconds` | Histogram of API, OAuth, and SCIM request duration, by `http_route` (the matched route pattern, never the raw URL), `http_request_method`, and `http_response_status_code`. Health probes, static assets, and the long-lived MCP streams are excluded — they would only skew a request-duration distribution. |

### Database

Each database metric is reported per database (`db` is `main` or `usage`) and
per connection pool (`pool` is `read` or `write`):

| Metric | Description |
| - | - |
| `conduit_db_connections_open` | Gauge of open connections. |
| `conduit_db_connections_in_use` | Gauge of connections currently executing a query. |
| `conduit_db_connections_waits_total` | Counter of times a request had to wait for a free connection — the pool-saturation signal. |
| `conduit_db_connections_wait_duration_seconds_total` | Counter of total time spent waiting for connections. Growing wait time with flat traffic means the pool is undersized for the workload. |

### Process

| Metric | Description |
| - | - |
| `conduit_build_info` | Gauge with constant value `1`, labeled with the running `version`. Join other metrics against it to see which version each replica runs — most useful mid-rollout. |
| `conduit_usage_recorder_dropped_total` | Counter of usage events dropped because the internal buffer was full. Sustained growth means the usage pipeline (the durable store behind the in-product dashboards) is shedding load. |

Standard Go runtime and process metrics (`go_*` — goroutines, memory, GC —
and `process_*` — CPU, RSS, file descriptors, start time) are exposed
alongside the application metrics.

## Useful starting queries

Paste these into the Prometheus Graph tab or a Grafana panel. Most of them
also make sensible first alert rules — thresholds depend on your traffic, so
chart them for a while before alerting.

### Tool-call error rate

The fraction of executed tool calls that failed, per connector. A single
connector spiking usually means an upstream outage or an expired credential;
everything spiking at once points at the gateway or its network.

```
sum by (connector) (rate(conduit_mcp_tool_calls_total{outcome="error"}[5m]))
  / sum by (connector) (rate(conduit_mcp_tool_calls_total[5m]))
```

### Tool-call latency

The 99th-percentile duration of executed tool calls. Aggregating `_bucket`
series with `sum by (le)` makes this correct across replicas; add `connector`
to the `by` clause to break it down per connector.

```
histogram_quantile(0.99,
  sum by (le) (rate(conduit_mcp_tool_call_duration_seconds_bucket[5m])))
```

### Policy denials

Calls blocked by access policy. A steady trickle is normal (clients probe
tools they can't see); a step change after a policy edit means the edit cut
off something real, and a sustained stream from one connector is a client
misconfigured or probing.

```
sum by (connector) (rate(conduit_mcp_tool_calls_total{outcome="denied"}[5m]))
```

### Authentication failures

Rejected MCP requests by reason. Alert on sustained `invalid_token` volume —
that's either a fleet of clients with stale credentials or someone guessing
bearer tokens.

```
sum by (reason) (rate(conduit_mcp_auth_failures_total[5m]))
```

### Replica health

Prometheus synthesizes an `up` series per scrape target — no Conduit metric
needed. This is the first alert to configure.

```
up == 0
```

### Database pool saturation

Time spent waiting for a database connection, per second. Zero is the healthy
steady state; any sustained non-zero value means requests are queuing for the
pool and latency is following.

```
rate(conduit_db_connections_wait_duration_seconds_total[5m])
```

### Usage pipeline drops

Usage events dropped because the recorder's buffer was full. Anything above
zero means the in-product usage dashboards are undercounting under load.

```
rate(conduit_usage_recorder_dropped_total[5m]) > 0
```
