Skip to main content
Conduit deployments come in three tiers. Each is a strict upgrade you can adopt when you need it — start simple, and the path up never requires re-architecting what you have.

Embedded (the default)

One container, one volume, no external services — the right choice until you have a reason otherwise. The rule that comes with it: exactly one instance per data directory. The embedded database is single-writer, and Conduit enforces the rule — an instance that detects another live instance using the same database file stops immediately rather than risk corruption. Never put the data directory on a network filesystem (NFS, EFS, Azure Files). Availability at this tier is fast recovery of the single instance, and it’s better than it sounds because restarts are cheap by design: sessions, tokens, and all state live in the database, so a restart signs nobody out, MCP clients reconnect on their own, and migrations apply at boot. In Kubernetes terms: liveness-probe restarts recover in seconds; node failure reschedules in a minute or two (bounded by volume reattach). What it can’t give you is zero-downtime upgrades or surviving an AZ outage without a restore — for that, move up a tier. Back up the volume on a schedule and keep the encryption key separate from those backups; snapshots alone can’t decrypt stored secrets.

PostgreSQL (single replica)

Set CONDUIT_DATABASE_URL to a postgres:// URL and Conduit — usage metrics included — runs on your PostgreSQL instead of the embedded database. Requirements: PostgreSQL 14+, CONDUIT_ENCRYPTION_KEY set in the environment (there is no data-directory key file to fall back on), and CONDUIT_ADMIN_PASSWORD until first-run setup completes. Run this tier when your organization already has managed PostgreSQL (RDS, Cloud SQL, …) and wants Conduit’s data under that backup, replication, and point-in-time-recovery regime — or as the stepping stone to active-active. The instance itself is now stateless: no volume, nothing to snapshot except the database and the key you already hold.

IAM authentication on AWS

On RDS or Aurora, Conduit can authenticate with IAM instead of a stored password: set CONDUIT_DATABASE_AUTH=rds-iam and leave the password out of the database URL. Each new connection authenticates with a short-lived token minted from the AWS identity the instance runs as — an EKS service-account role (IRSA or Pod Identity), an EC2 instance profile, or AWS credentials in the environment — so there is no database credential to store or rotate. Anything the AWS SDK’s default credential chain resolves works, including a cross-account role assumed via the shared AWS config file. The AWS region comes from the same source (set AWS_REGION if nothing else provides it). On the AWS side:
  1. Enable IAM database authentication on the RDS instance or Aurora cluster.
  2. Create the database user and delegate its authentication to IAM:
    Granting rds_iam makes the user IAM-only — its password (if it has one) stops working. Use a dedicated user rather than granting it to one that anything else logs into with a password. rds_iam covers authentication only — the user is created with no privileges on any data, so a sign-in that succeeds is still followed by permission denied for table … on every query. Conduit creates and migrates its own tables, so the user needs ownership-level access, not just SELECT. For a fresh install, make it the owner of Conduit’s database (RDS creates the initial database owned by the master user, so run this as that user; if the database doesn’t exist yet, CREATE DATABASE conduit OWNER conduit; does both steps):
    When switching an existing Conduit deployment over from a password user, grant that user’s role instead, so the IAM user can read and migrate the tables the old role owns:
    Ownership, rather than piecemeal grants, is also what keeps this working on PostgreSQL 15+, where the public schema is no longer writable by every user.
  3. Allow Conduit’s AWS role to connect. The policy resource names the resource ID (from the RDS console’s Configuration tab), not the database’s name or ARN — using the wrong identifier is the most common setup mistake. Which ID depends on what the URL connects to: an RDS instance’s db-…, an Aurora cluster’s cluster-…, or an RDS Proxy’s prx-… (AWS’s policy reference has the details and wildcard forms):
  4. Point Conduit at the database with TLS and no password — RDS refuses IAM logins over unencrypted connections, and while rds-iam is set any password in the URL is ignored. RDS server certificates chain to AWS-private CAs that no system trust store carries, so sslmode=verify-full also needs sslrootcert pointing at the AWS certificate bundle (https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem), downloaded to somewhere the server can read:
    Conduit enforces the TLS half of this at startup: with rds-iam set, a URL whose sslmode permits an unencrypted connection — disable, allow, or prefer, which is the default when sslmode is absent — refuses to boot. The token is a live credential sent as the connection password, and RDS rejects unencrypted IAM logins only after the token has already crossed the wire. sslmode=require is accepted but logs a warning, since it encrypts without verifying the server’s certificate. On Kubernetes the Helm chart wires all of this — IRSA annotation, env, and the bundle mount; see IAM authentication to RDS.
If sign-in to the database fails with PAM authentication failed, check the rds_iam grant, the resource ID in the policy, and that the database user in the URL exactly matches the one in the policy (it is case-sensitive). A no pg_hba.conf entry … SSL off error is different: the connection isn’t using TLS — fix sslmode in the URL. A certificate signed by unknown authority error means sslrootcert is missing or doesn’t hold the AWS bundle. Tokens are signed for the exact host and port in the URL, and RDS only accepts tokens signed for its own endpoint names — so the URL must name the real RDS, Aurora, or proxy endpoint. A custom DNS alias (a Route 53 CNAME in front of the database) breaks IAM auth even though it resolves to the same host. RDS Proxy works: the token is minted for the proxy’s endpoint, and the policy names the proxy’s prx-… resource ID. Note that IAM then authenticates the Conduit-to-proxy hop; the proxy signs in to the database behind it with its own Secrets Manager credential unless the proxy is configured for end-to-end IAM authentication.

Active-active (N replicas)

Scale the same PostgreSQL deployment to multiple replicas behind a load balancer. This is a supported, first-class mode — not replicas bolted onto a single-instance design:
  • Every instance is equal. No leader, and no session affinity. Sign-ins, OAuth flows, the setup wizard, and MCP tool calls all survive landing on different replicas mid-flow — including a tool call that asks the user a question in-band (MCP elicitation or sampling) and a request cancellation: the client’s answer or cancel may reach any replica, and it is handed to the one holding the call.
  • Policy changes propagate live. When an admin changes access policies, groups, or connectors on one replica, the others invalidate their caches within milliseconds over the database’s notification channel — and even if that channel drops, every replica independently re-checks authorization within a few seconds. Stale access is bounded by construction.
  • Abuse throttles hold cluster-wide. The MCP authentication-failure limiter counts in the shared database, so adding replicas doesn’t multiply what an attacker gets to try.
  • Maintenance runs once. Rollups, retention sweeps, and upgrade migrations coordinate through database locks; N replicas don’t duplicate or race the work.
  • Upgrades are zero-downtime. Rolling-update the deployment; draining instances finish their requests while new ones pass readiness (/healthz/ready, which verifies database connectivity) before taking traffic.
Point your load balancer’s health check at /healthz/ready so an instance that loses the database is ejected; keep container liveness on /healthz so a database blip doesn’t restart-loop the fleet. See Deploy on Kubernetes for the manifest changes — they amount to: more replicas, no volume, three secrets.
What stays per-replicaGateway log messages (notifications/message) that a client receives on its standing notification stream are pushed by the replica that handled the call. If that stream is open on a different replica, the client misses those frames for that call. They are informational only. Nothing else depends on which replica serves a request, so a plain round-robin load balancer — an AWS ALB, ingress-nginx, a cloud load balancer — needs no sticky-session configuration.
Conduit’s availability ceiling at this tier is your PostgreSQL’s: pair it with your provider’s HA database offering and the failure of any single Conduit instance, node, or zone is absorbed without operator action.