Embedded (the default)
One container, one volume, no external services — the right choice until you have a reason otherwise. The rule that comes with it: exactly one instance per data directory. The embedded database is single-writer, and Conduit enforces the rule — an instance that detects another live instance using the same database file stops immediately rather than risk corruption. Never put the data directory on a network filesystem (NFS, EFS, Azure Files). Availability at this tier is fast recovery of the single instance, and it’s better than it sounds because restarts are cheap by design: sessions, tokens, and all state live in the database, so a restart signs nobody out, MCP clients reconnect on their own, and migrations apply at boot. In Kubernetes terms: liveness-probe restarts recover in seconds; node failure reschedules in a minute or two (bounded by volume reattach). What it can’t give you is zero-downtime upgrades or surviving an AZ outage without a restore — for that, move up a tier. Back up the volume on a schedule and keep the encryption key separate from those backups; snapshots alone can’t decrypt stored secrets.PostgreSQL (single replica)
SetCONDUIT_DATABASE_URL to a postgres:// URL and Conduit — usage
metrics included — runs on your PostgreSQL instead of the embedded database.
Requirements: PostgreSQL 14+, CONDUIT_ENCRYPTION_KEY set in the
environment (there is no data-directory key file to fall back on), and
CONDUIT_ADMIN_PASSWORD until first-run setup completes.
Run this tier when your organization already has managed PostgreSQL (RDS,
Cloud SQL, …) and wants Conduit’s data under that backup, replication, and
point-in-time-recovery regime — or as the stepping stone to active-active.
The instance itself is now stateless: no volume, nothing to snapshot except
the database and the key you already hold.
IAM authentication on AWS
On RDS or Aurora, Conduit can authenticate with IAM instead of a stored password: setCONDUIT_DATABASE_AUTH=rds-iam and leave the password out of
the database URL. Each new connection authenticates with a short-lived token
minted from the AWS identity the instance runs as — an EKS service-account
role (IRSA or Pod Identity), an EC2 instance profile, or AWS credentials in
the environment — so there is no database credential to store or rotate.
Anything the AWS SDK’s default credential chain resolves works, including a
cross-account role assumed via the shared AWS config file. The AWS region
comes from the same source (set AWS_REGION if nothing else provides it).
On the AWS side:
- Enable IAM database authentication on the RDS instance or Aurora cluster.
-
Create the database user and delegate its authentication to IAM:
Granting
rds_iammakes the user IAM-only — its password (if it has one) stops working. Use a dedicated user rather than granting it to one that anything else logs into with a password.rds_iamcovers authentication only — the user is created with no privileges on any data, so a sign-in that succeeds is still followed bypermission denied for table …on every query. Conduit creates and migrates its own tables, so the user needs ownership-level access, not justSELECT. For a fresh install, make it the owner of Conduit’s database (RDS creates the initial database owned by the master user, so run this as that user; if the database doesn’t exist yet,CREATE DATABASE conduit OWNER conduit;does both steps):When switching an existing Conduit deployment over from a password user, grant that user’s role instead, so the IAM user can read and migrate the tables the old role owns:Ownership, rather than piecemeal grants, is also what keeps this working on PostgreSQL 15+, where thepublicschema is no longer writable by every user. -
Allow Conduit’s AWS role to connect. The policy resource names the
resource ID (from the RDS console’s Configuration tab), not the
database’s name or ARN — using the wrong identifier is the most common
setup mistake. Which ID depends on what the URL connects to: an RDS
instance’s
db-…, an Aurora cluster’scluster-…, or an RDS Proxy’sprx-…(AWS’s policy reference has the details and wildcard forms): -
Point Conduit at the database with TLS and no password — RDS refuses IAM
logins over unencrypted connections, and while
rds-iamis set any password in the URL is ignored. RDS server certificates chain to AWS-private CAs that no system trust store carries, sosslmode=verify-fullalso needssslrootcertpointing at the AWS certificate bundle (https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem), downloaded to somewhere the server can read:Conduit enforces the TLS half of this at startup: withrds-iamset, a URL whosesslmodepermits an unencrypted connection —disable,allow, orprefer, which is the default whensslmodeis absent — refuses to boot. The token is a live credential sent as the connection password, and RDS rejects unencrypted IAM logins only after the token has already crossed the wire.sslmode=requireis accepted but logs a warning, since it encrypts without verifying the server’s certificate. On Kubernetes the Helm chart wires all of this — IRSA annotation, env, and the bundle mount; see IAM authentication to RDS.
PAM authentication failed, check the
rds_iam grant, the resource ID in the policy, and that the database user in
the URL exactly matches the one in the policy (it is case-sensitive). A
no pg_hba.conf entry … SSL off error is different: the connection isn’t
using TLS — fix sslmode in the URL. A certificate signed by unknown authority error means sslrootcert is missing or doesn’t hold the AWS
bundle.
Tokens are signed for the exact host and port in the URL, and RDS only
accepts tokens signed for its own endpoint names — so the URL must name the
real RDS, Aurora, or proxy endpoint. A custom DNS alias (a Route 53 CNAME in
front of the database) breaks IAM auth even though it resolves to the same
host.
RDS Proxy works: the token is minted for the proxy’s endpoint, and the policy
names the proxy’s prx-… resource ID. Note that IAM then authenticates the
Conduit-to-proxy hop; the proxy signs in to the database behind it with its
own Secrets Manager credential unless the proxy is configured for end-to-end
IAM authentication.
Active-active (N replicas)
Scale the same PostgreSQL deployment to multiple replicas behind a load balancer. This is a supported, first-class mode — not replicas bolted onto a single-instance design:- Every instance is equal. No leader, and no session affinity. Sign-ins, OAuth flows, the setup wizard, and MCP tool calls all survive landing on different replicas mid-flow — including a tool call that asks the user a question in-band (MCP elicitation or sampling) and a request cancellation: the client’s answer or cancel may reach any replica, and it is handed to the one holding the call.
- Policy changes propagate live. When an admin changes access policies, groups, or connectors on one replica, the others invalidate their caches within milliseconds over the database’s notification channel — and even if that channel drops, every replica independently re-checks authorization within a few seconds. Stale access is bounded by construction.
- Abuse throttles hold cluster-wide. The MCP authentication-failure limiter counts in the shared database, so adding replicas doesn’t multiply what an attacker gets to try.
- Maintenance runs once. Rollups, retention sweeps, and upgrade migrations coordinate through database locks; N replicas don’t duplicate or race the work.
- Upgrades are zero-downtime. Rolling-update the deployment; draining
instances finish their requests while new ones pass readiness
(
/healthz/ready, which verifies database connectivity) before taking traffic.
/healthz/ready so an instance
that loses the database is ejected; keep container liveness on /healthz
so a database blip doesn’t restart-loop the fleet. See
Deploy on Kubernetes for the manifest changes — they
amount to: more replicas, no volume, three secrets.
What stays per-replicaGateway log messages (
notifications/message) that a client receives on its
standing notification stream are pushed by the replica that handled the
call. If that stream is open on a different replica, the client misses those
frames for that call. They are informational only. Nothing else depends on
which replica serves a request, so a plain round-robin load balancer — an
AWS ALB, ingress-nginx, a cloud load balancer — needs no sticky-session
configuration.