Skip to main content
This page answers one question: how much hardware does Conduit need for your team? The numbers are measured, not estimated — a load-testing harness ships in the repository, every result below states the exact command that produced it, and the history table at the bottom tracks how capacity has changed across versions.
The short answerOne Conduit pod — 1 CPU / 512 MiB, the shape the Kubernetes guide recommends — sustains 8,859 modeled active users. Most teams fit in one pod; add replicas for headroom or high availability, not raw capacity.

What “a user” means here

A capacity number is only useful if the modeled user resembles yours. The workload model is calibrated against measured traffic from a production Conduit deployment serving agentic MCP clients (Claude Code, Cursor, and similar), and it reproduces the patterns that actually stress a gateway:
  • Each user holds one open connection (the SSE stream MCP clients keep for server notifications) for the whole test.
  • Calls arrive in bursts, not on a timer. An agent fires several calls seconds apart, then goes quiet for minutes — matching measured inter-call gaps (median 6 s inside a burst, ~90 s on average overall).
  • A few users are bots. 1% of modeled users are unattended scripts calling continuously — in production a cohort that small produced almost half of all traffic.
  • Upstream tools are slow, with a long tail. Each call’s duration is drawn from the measured distribution: median 1.2 s, 5% over 12 s, capped at 50 s. Slow calls are held open by the upstream, so the gateway carries the same in-flight state a slow real tool imposes.
  • Some calls fail. 8% of calls return a tool error, as measured in production (upstream timeouts, oversized responses, disconnected accounts).
  • Responses vary in size — mostly a few KB, 2% around 1 MB.
  • Clients churn. Users periodically reconnect (handshake + tool listing) and re-list tools between calls; 5% of calls come with a dashboard page view.
If your users are chattier than this, divide accordingly — the harness knobs (call rate, bot share, latency, error rate) are all adjustable, and you can run the same tests against your own numbers.

Sizing your deployment

Conduit ships in two storage shapes: the default embedded database (one pod, zero external dependencies) and PostgreSQL (required to run more than one replica). Per-pod throughput is within ~15% between the two — choose PostgreSQL for replicas and operational preference, not speed. Two resources set the ceiling, in this order:
  1. Memory per connection. Every connected user holds an SSE stream (~40 KB each); the 512 MiB pod runs out of memory before it runs out of CPU under the modeled workload. More users per pod = more memory.
  2. CPU per request. Raw request throughput (thousands of calls per second per pod — see the table below) is far above what realistic user counts generate; it matters only for unusually hot, bot-heavy workloads.
Real-world latency depends on your upstream tools far more than on Conduit: the gateway adds single-digit milliseconds; production tool calls measure a 1.2 s median on their own.

The measured numbers

Measured 2026-08-31 on an Apple M4 Max (Docker VM: 8 CPUs / 32 GiB, all components on one host — treat results as lower bounds and validate on your own infrastructure for sizing decisions). Every Conduit pod is pinned to 1 CPU / 512 MiB. Ceilings are sustained maxima under the pass/fail bar below, found by ramping until it breaks and binary-searching to ±10%. The pass/fail bar (every run must hold all of these): p95 tool call under 15 s (the modeled upstream latency alone has a ~12 s p95), transport error rate under 0.5%, tool-error rate under 15% (8% is modeled), no shed load, healthy connection keepalives.

Users per pod (the headline)

The ceiling is the pod’s 512 MiB of memory, not CPU: the last passing step runs 5 MiB under the limit with a quarter of the CPU idle.

Request-rate ceilings (synthetic)

Constant-rate probes of one request type at a time, under the standard fixed upstream latency (calls 1 s ± 0.5, listings 250 ms ± 100). These measure per-request cost, not realistic traffic — useful for spotting regressions and for bot-heavy sizing.
The dashboard probe runs against a usage store holding ~500k recorded tool calls (about a month of busy traffic), so its queries do realistic work. PostgreSQL figures are from the 2026-07-22 pass series (see history); per-pod differences between the backends have stayed within ~15%.

Idle connections

An idle connected user costs ~2.5 goroutines and ~40 KB of server memory (measured holding 500–5,000 streams) — roughly 10–12k held streams fit in a 512 MiB pod, which is why memory, not CPU, sets the users-per-pod ceiling.

Protocol revisions

Conduit serves clients on the stateless MCP revision 2026-07-28 and on the handshake revisions (see the MCP client reference). Every scenario runs in either: -era stateless has the generators send 2026-07-28 requests (per-request metadata, server/discover in place of the handshake) and hold subscriptions/listen streams in place of the handshake era’s stream, with the workload otherwise identical. Measured 2026-09-23 on a smaller host than the numbers above (Docker VM: 4 CPUs), both revisions back to back with identical parameters on a freshly seeded pod each — compare within this table, not with the headline: A held subscriptions/listen stream costs slightly more memory than the handshake era’s stream (its request carries more headers, which the server holds for the stream’s life), so users per pod lands one measurement step lower; per request, 2026-07-28 costs the same or less. The stateless run replays the handshake era’s request mix, including its re-listing and reconnects — a 2026-07-28 client that honors the caching hints on listings and on server/discover sends fewer requests, so its numbers are a conservative bound.

Run it on your own infrastructure

Two images and a URL. Deploy them however you deploy anything; Conduit and the mock upstream may live in different clusters or clouds. Neither image is published: build both from the repository at the Conduit release you are measuring and push them to a registry your environments pull from. Below, <registry>/loadtest-conduit:<tag> and <registry>/loadtest-mcp-server:<tag> are those pushes, and the example hostnames are yours to replace.

1. Deploy the mock upstream

Image <registry>/loadtest-mcp-server:<tag>, wherever your upstream MCP servers run. Port 8765, GET /healthz for probes, stateless (run as many replicas as you like). Conduit must reach it over https on a hostname: give it a certificate with TLS_CERT_FILE / TLS_KEY_FILE, or run it plain behind whatever terminates TLS for you. Environment, the connector envelope behind the published numbers:

2. Deploy Conduit from the load-test image

Image <registry>/loadtest-conduit:<tag>, wherever you deploy Conduit, exactly as you would deploy conduit:<tag> — with the Helm chart set image.repository and image.tag; with your own manifests swap the image. It is the product image plus the load-test CLI as entrypoint: same server binary, same layers. Keep your normal CONDUIT_* configuration and add:
Then:
  • Pin the pod shape you want to quote, requests equal to limits (the published numbers are 1 CPU / 512 MiB).
  • If the mock’s certificate isn’t from a public authority, mount the CA and point SSL_CERT_FILE at it.
  • Treat the instance as a throwaway. Whoever holds the seed secret holds every seeded user’s credentials, including an instance admin’s.
The instance seeds itself once the server is ready. Its log prints auto-seed complete, and /tmp/loadtest.seeded exists in the container. A restart re-runs the pass and changes nothing. A 401 from the generators in step 3 means seeding hasn’t finished, or the secret or user count differs between the two sides.

3. Run the load

From any machine that reaches Conduit’s URL, with k6 installed and this repository checked out, the same two values the deployment has:
The headline measurement, ramping modeled users until the pass/fail bar breaks (one held connection per user, ~10 minutes at these settings):
One scenario at a fixed load, with stock k6:
Add -era stateless to conduit-bench (or -e ERA=stateless to stock k6) to drive the same workload over MCP 2026-07-28; each result records the revision it measured. Results land in loadtest/results/ as JSON and Markdown. Every scenario first verifies that the server honors the generator’s client IPs and stops with instructions if it doesn’t. For a -url run, server CPU and memory come from your own monitoring, and the result states that the server’s shape is unknown. Run the generator close to Conduit: a laptop over a WAN link adds its own latency and connection limits to the numbers.

4. Tear down

Delete both deployments. To remove the seeded workspace, users and connectors from a database you keep, run conduit-bench unseed -from-env inside the Conduit container (it is on the image’s PATH), or from anywhere with conduit-bench unseed -database-url … -base-url ….

Variants

One Docker host, end to end. loadtest/e2e-docker.sh in the repository runs steps 1 to 4 above literally, with docker run as the deployment mechanism: it builds both images from the checkout, runs the mock and Conduit as two containers on one network, waits for seeding, drives the mixed scenario and the SSE holder from containers on the same network, removes the seeded population, and tears everything down. Read it as the template for your own environment’s script, with docker run swapped for however you deploy:
Kubernetes, with the chart and kubectl access. The repository’s loadtest/k8s/ holds a complete rig: values pinning the pod shape and wiring the variables above, a chart for the mock, an optional in-cluster PostgreSQL, and make targets that install everything into a fresh namespace and wait for seeding. With this target the harness also runs the generators as pods next to the server and collects CPU, memory and runtime signals:
Locally. The numbers above come from pinned Docker Compose environments in the repository:
Workload knobs (USERS, BOT_FRACTION, CALL_P50_MS, CALL_ERROR_FRACTION, …) are environment variables documented in loadtest/k6/mixed.js; set them to match your own traffic before quoting numbers to yourself.

History

One row per test pass, newest first. Passes before 2026-08-31 modeled a lighter, uniform workload (a call every 30–60 s, 1 s fixed upstream latency, no bots, no errors), so their user counts are not directly comparable with today’s model — the row notes what changed.