The short answerOne Conduit pod — 1 CPU / 512 MiB, the shape the
Kubernetes guide recommends — sustains
8,859 modeled active users. Most teams fit in one pod; add replicas
for headroom or high availability, not raw capacity.
What “a user” means here
A capacity number is only useful if the modeled user resembles yours. The workload model is calibrated against measured traffic from a production Conduit deployment serving agentic MCP clients (Claude Code, Cursor, and similar), and it reproduces the patterns that actually stress a gateway:- Each user holds one open connection (the SSE stream MCP clients keep for server notifications) for the whole test.
- Calls arrive in bursts, not on a timer. An agent fires several calls seconds apart, then goes quiet for minutes — matching measured inter-call gaps (median 6 s inside a burst, ~90 s on average overall).
- A few users are bots. 1% of modeled users are unattended scripts calling continuously — in production a cohort that small produced almost half of all traffic.
- Upstream tools are slow, with a long tail. Each call’s duration is drawn from the measured distribution: median 1.2 s, 5% over 12 s, capped at 50 s. Slow calls are held open by the upstream, so the gateway carries the same in-flight state a slow real tool imposes.
- Some calls fail. 8% of calls return a tool error, as measured in production (upstream timeouts, oversized responses, disconnected accounts).
- Responses vary in size — mostly a few KB, 2% around 1 MB.
- Clients churn. Users periodically reconnect (handshake + tool listing) and re-list tools between calls; 5% of calls come with a dashboard page view.
Sizing your deployment
Conduit ships in two storage shapes: the default embedded database (one pod, zero external dependencies) and PostgreSQL (required to run more than one replica). Per-pod throughput is within ~15% between the two — choose PostgreSQL for replicas and operational preference, not speed.
Two resources set the ceiling, in this order:
- Memory per connection. Every connected user holds an SSE stream (~40 KB each); the 512 MiB pod runs out of memory before it runs out of CPU under the modeled workload. More users per pod = more memory.
- CPU per request. Raw request throughput (thousands of calls per second per pod — see the table below) is far above what realistic user counts generate; it matters only for unusually hot, bot-heavy workloads.
The measured numbers
Measured 2026-08-31 on an Apple M4 Max (Docker VM: 8 CPUs / 32 GiB, all components on one host — treat results as lower bounds and validate on your own infrastructure for sizing decisions). Every Conduit pod is pinned to 1 CPU / 512 MiB. Ceilings are sustained maxima under the pass/fail bar below, found by ramping until it breaks and binary-searching to ±10%. The pass/fail bar (every run must hold all of these): p95 tool call under 15 s (the modeled upstream latency alone has a ~12 s p95), transport error rate under 0.5%, tool-error rate under 15% (8% is modeled), no shed load, healthy connection keepalives.Users per pod (the headline)
The ceiling is the pod’s 512 MiB of memory, not CPU: the last passing step
runs 5 MiB under the limit with a quarter of the CPU idle.
Request-rate ceilings (synthetic)
Constant-rate probes of one request type at a time, under the standard fixed upstream latency (calls 1 s ± 0.5, listings 250 ms ± 100). These measure per-request cost, not realistic traffic — useful for spotting regressions and for bot-heavy sizing.
The dashboard probe runs against a usage store holding ~500k recorded tool
calls (about a month of busy traffic), so its queries do realistic work.
PostgreSQL figures are from the 2026-07-22 pass series (see
history); per-pod differences between the backends have stayed
within ~15%.
Idle connections
An idle connected user costs ~2.5 goroutines and ~40 KB of server memory (measured holding 500–5,000 streams) — roughly 10–12k held streams fit in a 512 MiB pod, which is why memory, not CPU, sets the users-per-pod ceiling.Protocol revisions
Conduit serves clients on the stateless MCP revision2026-07-28 and on
the handshake revisions (see the MCP client reference).
Every scenario runs in either: -era stateless has the generators send
2026-07-28 requests (per-request metadata, server/discover in place of
the handshake) and hold subscriptions/listen streams in place of the
handshake era’s stream, with the workload otherwise identical.
Measured 2026-09-23 on a smaller host than the numbers above (Docker VM:
4 CPUs), both revisions back to back with identical parameters on a
freshly seeded pod each — compare within this table, not with the
headline:
A held
subscriptions/listen stream costs slightly more memory than the
handshake era’s stream (its request carries more headers, which the
server holds for the stream’s life), so users per pod lands one
measurement step lower; per request, 2026-07-28 costs the same or less.
The stateless run replays the handshake era’s request mix, including its
re-listing and reconnects — a 2026-07-28 client that honors the caching
hints on listings and on server/discover sends fewer requests, so its
numbers are a conservative bound.
Run it on your own infrastructure
Two images and a URL. Deploy them however you deploy anything; Conduit and the mock upstream may live in different clusters or clouds. Neither image is published: build both from the repository at the Conduit release you are measuring and push them to a registry your environments pull from. Below,<registry>/loadtest-conduit:<tag> and
<registry>/loadtest-mcp-server:<tag> are those pushes, and the example
hostnames are yours to replace.
1. Deploy the mock upstream
Image<registry>/loadtest-mcp-server:<tag>, wherever your upstream MCP
servers run. Port 8765, GET /healthz for probes, stateless
(run as many replicas as you like). Conduit must reach it over https on a
hostname: give it a certificate with TLS_CERT_FILE / TLS_KEY_FILE, or
run it plain behind whatever terminates TLS for you.
Environment, the connector envelope behind the published numbers:
2. Deploy Conduit from the load-test image
Image<registry>/loadtest-conduit:<tag>, wherever you deploy Conduit,
exactly as you would deploy conduit:<tag> — with the
Helm chart set image.repository and
image.tag; with your own manifests swap the image. It is the product
image plus the load-test CLI as entrypoint: same server binary, same
layers. Keep your normal CONDUIT_* configuration and add:
- Pin the pod shape you want to quote, requests equal to limits (the published numbers are 1 CPU / 512 MiB).
- If the mock’s certificate isn’t from a public authority, mount the CA
and point
SSL_CERT_FILEat it. - Treat the instance as a throwaway. Whoever holds the seed secret holds every seeded user’s credentials, including an instance admin’s.
auto-seed complete, and /tmp/loadtest.seeded exists in the container.
A restart re-runs the pass and changes nothing. A 401 from the
generators in step 3 means seeding hasn’t finished, or the secret or user
count differs between the two sides.
3. Run the load
From any machine that reaches Conduit’s URL, with k6 installed and this repository checked out, the same two values the deployment has:-era stateless to conduit-bench (or -e ERA=stateless to stock
k6) to drive the same workload over MCP 2026-07-28; each result records
the revision it measured.
Results land in loadtest/results/ as JSON and Markdown. Every scenario
first verifies that the server honors the generator’s client IPs and stops
with instructions if it doesn’t. For a -url run, server CPU and memory
come from your own monitoring, and the result states that the server’s
shape is unknown. Run the generator close to Conduit: a laptop over a WAN
link adds its own latency and connection limits to the numbers.
4. Tear down
Delete both deployments. To remove the seeded workspace, users and connectors from a database you keep, runconduit-bench unseed -from-env
inside the Conduit container (it is on the image’s PATH), or from anywhere
with conduit-bench unseed -database-url … -base-url ….
Variants
One Docker host, end to end.loadtest/e2e-docker.sh in the repository
runs steps 1 to 4 above literally, with docker run as the deployment
mechanism: it builds both images from the checkout, runs the mock and Conduit
as two containers on one network, waits for seeding, drives the mixed
scenario and the SSE holder from containers on the same network, removes the
seeded population, and tears everything down. Read it as the template for
your own environment’s script, with docker run swapped for however you
deploy:
kubectl access. The repository’s
loadtest/k8s/ holds a complete rig: values pinning the pod shape and
wiring the variables above, a chart for the mock, an optional in-cluster
PostgreSQL, and make targets that install everything into a fresh namespace
and wait for seeding. With this target the harness also runs the generators
as pods next to the server and collects CPU, memory and runtime signals:
USERS, BOT_FRACTION, CALL_P50_MS, CALL_ERROR_FRACTION,
…) are environment variables documented in loadtest/k6/mixed.js; set them
to match your own traffic before quoting numbers to yourself.