> ## Documentation Index
> Fetch the complete documentation index at: https://pipedream.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Load Tests

> How many users one Conduit pod supports, what workload that number assumes, and how to size your deployment.

This page answers one question: **how much hardware does Conduit need for
your team?** The numbers are measured, not estimated — a load-testing
harness ships in the repository, every result below states the exact
command that produced it, and the [history](#history) table at the bottom
tracks how capacity has changed across versions.

<Note>
  **The short answer**

  One Conduit pod — **1 CPU / 512 MiB**, the shape the
  [Kubernetes guide](/docs/conduit/deploy/kubernetes) recommends — sustains
  **8,859 modeled active users**. Most teams fit in one pod; add replicas
  for headroom or high availability, not raw capacity.
</Note>

## What "a user" means here

A capacity number is only useful if the modeled user resembles yours. The
workload model is calibrated against measured traffic from a production
Conduit deployment serving agentic MCP clients (Claude Code, Cursor, and
similar), and it reproduces the patterns that actually stress a gateway:

* **Each user holds one open connection** (the SSE stream MCP clients keep
  for server notifications) for the whole test.
* **Calls arrive in bursts, not on a timer.** An agent fires several calls
  seconds apart, then goes quiet for minutes — matching measured inter-call
  gaps (median 6 s inside a burst, \~90 s on average overall).
* **A few users are bots.** 1% of modeled users are unattended scripts
  calling continuously — in production a cohort that small produced almost
  half of all traffic.
* **Upstream tools are slow, with a long tail.** Each call's duration is
  drawn from the measured distribution: median 1.2 s, 5% over 12 s, capped
  at 50 s. Slow calls are held open by the upstream, so the gateway carries
  the same in-flight state a slow real tool imposes.
* **Some calls fail.** 8% of calls return a tool error, as measured in
  production (upstream timeouts, oversized responses, disconnected
  accounts).
* **Responses vary in size** — mostly a few KB, 2% around 1 MB.
* **Clients churn.** Users periodically reconnect (handshake + tool
  listing) and re-list tools between calls; 5% of calls come with a
  dashboard page view.

If your users are chattier than this, divide accordingly — the harness
knobs (call rate, bot share, latency, error rate) are all adjustable, and
you can [run the same tests](#run-it-on-your-own-infrastructure) against
your own numbers.

## Sizing your deployment

Conduit ships in two storage shapes: the default **embedded database**
(one pod, zero external dependencies) and **PostgreSQL** (required to run
more than one replica). Per-pod throughput is within \~15% between the two —
choose PostgreSQL for replicas and operational preference, not speed.

| Your situation | Recommended shape |
| - | - |
| Up to \~8,000 active users | 1 pod, embedded database |
| More users, or zero-downtime upgrades | PostgreSQL + 2–3 pods behind a load balancer |
| Growing beyond that | Add pods — capacity is expected to scale roughly linearly with replicas (checked periodically, not every pass) |

Two resources set the ceiling, in this order:

1. **Memory per connection.** Every connected user holds an SSE stream
   (\~40 KB each); the 512 MiB pod runs out of memory before it runs out of
   CPU under the modeled workload. More users per pod = more memory.
2. **CPU per request.** Raw request throughput (thousands of calls per
   second per pod — see the table below) is far above what realistic user
   counts generate; it matters only for unusually hot, bot-heavy
   workloads.

Real-world latency depends on your upstream tools far more than on
Conduit: the gateway adds single-digit milliseconds; production tool calls
measure a 1.2 s median on their own.

## The measured numbers

Measured 2026-08-31 on an Apple M4 Max (Docker VM: 8 CPUs / 32 GiB, all
components on one host — treat results as lower bounds and validate on
your own infrastructure for sizing decisions). Every
Conduit pod is pinned to 1 CPU / 512 MiB. Ceilings are sustained maxima
under the pass/fail bar below, found by ramping until it breaks and
binary-searching to ±10%.

**The pass/fail bar** (every run must hold all of these): p95 tool call
under 15 s (the modeled upstream latency alone has a \~12 s p95), transport
error rate under 0.5%, tool-error rate under 15% (8% is modeled), no shed
load, healthy connection keepalives.

### Users per pod (the headline)

```sh theme={null}
make loadtest-tier1
go run ./cmd/conduit-bench knee -compose loadtest/compose/tier1.yml \
  -scenario mixed -ramp-var USERS -sse-per-user -start 3000 -factor 1.5 \
  -slo-call-p95-ms 15000 -env DURATION=150s
```

| Modeled users | Result | Server CPU | Server memory |
| - | - | - | - |
| 3,000 | pass — 100% checks, call p95 11.1 s | 22% | 272 MiB |
| 6,750 | pass — 100% checks, call p95 11.6 s | 46% | 482 MiB |
| **8,859** | **pass** — all streams held, call p95 11.3 s, tool errors 8.0% (exactly as modeled) | 76% | 507 MiB |
| 9,281 | breach — memory limit reached, streams dropped, 8.3% transport errors | 103% | 496 MiB |

The ceiling is the pod's 512 MiB of memory, not CPU: the last passing step
runs 5 MiB under the limit with a quarter of the CPU idle.

### Request-rate ceilings (synthetic)

Constant-rate probes of one request type at a time, under the standard
fixed upstream latency (calls 1 s ± 0.5, listings 250 ms ± 100). These
measure per-request cost, not realistic traffic — useful for spotting
regressions and for bot-heavy sizing.

```sh theme={null}
go run ./cmd/conduit-bench knee -compose loadtest/compose/tier1.yml \
  -scenario tools-call -ramp-var RATE -start 400 -factor 2 -max 6400 -env DURATION=30s,USERS=4000
go run ./cmd/conduit-bench knee -compose loadtest/compose/tier1.yml \
  -scenario tools-list -ramp-var RATE -start 100 -factor 2 -max 3200 -env DURATION=30s,USERS=3000
go run ./cmd/conduit-bench knee -compose loadtest/compose/tier1.yml \
  -scenario dashboard -ramp-var RATE -start 1600 -factor 2 -max 25600 -env DURATION=30s,USERS=1600
```

| Request type | Embedded DB (1 pod) | PostgreSQL (1 pod) |
| - | - | - |
| Tool calls /s | 3,000 | 4,000 |
| Tool listings /s (5 connectors) | 750 | 800 |
| Dashboard views /s | 3,600 | 3,000 |

The dashboard probe runs against a usage store holding \~500k recorded tool
calls (about a month of busy traffic), so its queries do realistic work.
PostgreSQL figures are from the 2026-07-22 pass series (see
[history](#history)); per-pod differences between the backends have stayed
within \~15%.

### Idle connections

An idle connected user costs \~2.5 goroutines and \~40 KB of server memory
(measured holding 500–5,000 streams) — roughly 10–12k held streams fit in
a 512 MiB pod, which is why memory, not CPU, sets the users-per-pod
ceiling.

### Protocol revisions

Conduit serves clients on the stateless MCP revision `2026-07-28` and on
the handshake revisions (see the [MCP client reference](/docs/conduit/use/mcp-reference)).
Every scenario runs in either: `-era stateless` has the generators send
2026-07-28 requests (per-request metadata, `server/discover` in place of
the handshake) and hold `subscriptions/listen` streams in place of the
handshake era's stream, with the workload otherwise identical.

Measured 2026-09-23 on a smaller host than the numbers above (Docker VM:
4 CPUs), both revisions back to back with identical parameters on a
freshly seeded pod each — compare within this table, not with the
headline:

| | Handshake revisions | `2026-07-28` |
| - | - | - |
| Modeled users per pod | 8,437 | 8,015 |
| Tool calls /s | 1,300 | 1,600 |
| Tool listings /s (5 connectors) | 750 | 750 |

A held `subscriptions/listen` stream costs slightly more memory than the
handshake era's stream (its request carries more headers, which the
server holds for the stream's life), so users per pod lands one
measurement step lower; per request, 2026-07-28 costs the same or less.
The stateless run replays the handshake era's request mix, including its
re-listing and reconnects — a 2026-07-28 client that honors the caching
hints on listings and on `server/discover` sends fewer requests, so its
numbers are a conservative bound.

```sh theme={null}
go run ./cmd/conduit-bench knee -compose loadtest/compose/tier1.yml -era stateless \
  -scenario mixed -ramp-var USERS -sse-per-user -start 3000 -factor 1.5 \
  -slo-call-p95-ms 15000 -env DURATION=150s
```

## Run it on your own infrastructure

Two images and a URL. Deploy them however you deploy anything; Conduit and
the mock upstream may live in different clusters or clouds. Neither image is
published: [build both from the repository](/docs/conduit/deploy/build-from-source#every-image-the-repository-builds)
at the Conduit release you are measuring and push them to a registry your
environments pull from. Below, `<registry>/loadtest-conduit:<tag>` and
`<registry>/loadtest-mcp-server:<tag>` are those pushes, and the example
hostnames are yours to replace.

### 1. Deploy the mock upstream

Image `<registry>/loadtest-mcp-server:<tag>`, wherever your upstream MCP
servers run. Port 8765, `GET /healthz` for probes, stateless
(run as many replicas as you like). Conduit must reach it over `https` on a
hostname: give it a certificate with `TLS_CERT_FILE` / `TLS_KEY_FILE`, or
run it plain behind whatever terminates TLS for you.

Environment, the connector envelope behind the published numbers:

```sh theme={null}
AUTH_MODE=none
LOG_REQUESTS=0
TOOL_COUNT=20
LIST_LATENCY_MS=250
LIST_JITTER_MS=100
CALL_LATENCY_MS=1000
CALL_JITTER_MS=500
RESPONSE_BYTES=1024
```

### 2. Deploy Conduit from the load-test image

Image `<registry>/loadtest-conduit:<tag>`, wherever you deploy Conduit,
exactly as you would deploy `conduit:<tag>` — with the
[Helm chart](/docs/conduit/deploy/kubernetes) set `image.repository` and
`image.tag`; with your own manifests swap the image. It is the product
image plus the load-test CLI as entrypoint: same server binary, same
layers. Keep your normal `CONDUIT_*` configuration and add:

```sh theme={null}
LOADTEST_SEED_SECRET=<openssl rand -hex 32>   # every seeded token derives from it; you need it again in step 3
LOADTEST_USERS=5000                            # virtual users to seed; at least the largest population you will model
LOADTEST_CONNECTOR_URLS=https://mock.internal.example.com:8765,https://mock.internal.example.com:8765,https://mock.internal.example.com:8765,https://mock.internal.example.com:8765,https://mock.internal.example.com:8765
                                               # one connector per entry, as reachable FROM Conduit; the published numbers use five on one mock
CONDUIT_SAFEHTTP_ALLOWED_PRIVATE_HOSTS=mock.internal.example.com   # only if the mock resolves to a private address
CONDUIT_TRUSTED_PROXIES=10.0.0.0/8             # the hop your load generator's traffic arrives through
```

Then:

* **Pin the pod shape** you want to quote, requests equal to limits (the
  published numbers are 1 CPU / 512 MiB).
* **If the mock's certificate isn't from a public authority**, mount the CA
  and point `SSL_CERT_FILE` at it.
* **Treat the instance as a throwaway.** Whoever holds the seed secret holds
  every seeded user's credentials, including an instance admin's.

The instance seeds itself once the server is ready. Its log prints
`auto-seed complete`, and `/tmp/loadtest.seeded` exists in the container.
A restart re-runs the pass and changes nothing. A `401` from the
generators in step 3 means seeding hasn't finished, or the secret or user
count differs between the two sides.

### 3. Run the load

From any machine that reaches Conduit's URL, with [k6](https://k6.io)
installed and this repository checked out, the same two values the
deployment has:

```sh theme={null}
export LOADTEST_SEED_SECRET=<the value from step 2>
export SEEDED_USERS=5000
```

The headline measurement, ramping modeled users until the pass/fail bar
breaks (one held connection per user, \~10 minutes at these settings):

```sh theme={null}
go run ./cmd/conduit-bench knee -url https://conduit.loadtest.example.com \
  -scenario mixed -ramp-var USERS -sse-per-user -start 500 -factor 1.5 \
  -slo-call-p95-ms 15000 -env DURATION=150s
```

One scenario at a fixed load, with stock k6:

```sh theme={null}
k6 run -e BASE_URL=https://conduit.loadtest.example.com \
  -e LOADTEST_SEED_SECRET=$LOADTEST_SEED_SECRET -e SEEDED_USERS=$SEEDED_USERS \
  -e USERS=2000 -e DURATION=150s loadtest/k6/mixed.js
```

Add `-era stateless` to `conduit-bench` (or `-e ERA=stateless` to stock
k6) to drive the same workload over MCP `2026-07-28`; each result records
the revision it measured.

Results land in `loadtest/results/` as JSON and Markdown. Every scenario
first verifies that the server honors the generator's client IPs and stops
with instructions if it doesn't. For a `-url` run, server CPU and memory
come from your own monitoring, and the result states that the server's
shape is unknown. Run the generator close to Conduit: a laptop over a WAN
link adds its own latency and connection limits to the numbers.

### 4. Tear down

Delete both deployments. To remove the seeded workspace, users and
connectors from a database you keep, run `conduit-bench unseed -from-env`
inside the Conduit container (it is on the image's PATH), or from anywhere
with `conduit-bench unseed -database-url … -base-url …`.

### Variants

**One Docker host, end to end.** `loadtest/e2e-docker.sh` in the repository
runs steps 1 to 4 above literally, with `docker run` as the deployment
mechanism: it builds both images from the checkout, runs the mock and Conduit
as two containers on one network, waits for seeding, drives the mixed
scenario and the SSE holder from containers on the same network, removes the
seeded population, and tears everything down. Read it as the template for
your own environment's script, with `docker run` swapped for however you
deploy:

```sh theme={null}
USERS=200 DURATION_S=60 loadtest/e2e-docker.sh    # KEEP=1 leaves the containers running
```

**Kubernetes, with the chart and `kubectl` access.** The repository's
`loadtest/k8s/` holds a complete rig: values pinning the pod shape and
wiring the variables above, a chart for the mock, an optional in-cluster
PostgreSQL, and make targets that install everything into a fresh namespace
and wait for seeding. With this target the harness also runs the generators
as pods next to the server and collects CPU, memory and runtime signals:

```sh theme={null}
make loadtest-k8s-images LT_REGISTRY=registry.example.com/conduit   # build both images here, push them
make loadtest-k8s-up LT_REGISTRY=registry.example.com/conduit       # LT_DB=postgres LT_REPLICAS=3 for the PostgreSQL tier
LOADTEST_SEED_SECRET=$(make -s loadtest-k8s-secret) \
go run ./cmd/conduit-bench knee -k8s conduit-loadtest/conduit \
  -scenario mixed -ramp-var USERS -sse-per-user -start 3000 -factor 1.5 \
  -slo-call-p95-ms 15000 -env DURATION=150s
make loadtest-k8s-down
```

**Locally.** The numbers above come from pinned Docker Compose
environments in the repository:

```sh theme={null}
make loadtest-tier1     # build + start; the pod seeds itself
go run ./cmd/conduit-bench knee -compose loadtest/compose/tier1.yml -scenario mixed ...
make loadtest-down
```

Workload knobs (`USERS`, `BOT_FRACTION`, `CALL_P50_MS`, `CALL_ERROR_FRACTION`,
…) are environment variables documented in `loadtest/k6/mixed.js`; set them
to match your own traffic before quoting numbers to yourself.

## History

One row per test pass, newest first. Passes before 2026-08-31 modeled a
lighter, uniform workload (a call every 30–60 s, 1 s fixed upstream
latency, no bots, no errors), so their user counts are not directly
comparable with today's model — the row notes what changed.

| Date | Users per pod | What changed |
| - | - | - |
| 2026-08-31 | **8,859** | Workload model rebuilt from measured production traffic (bursts, bots, heavy-tailed latency, errors, response sizes, reconnect churn) — new baseline, not comparable with earlier rows |
| 2026-07-22 (6) | 10,125 | Session validation collapsed to one database query; PostgreSQL tool calls +11% (4,000/s), dashboard +30% (5,200/s) |
| 2026-07-22 (5) | 10,125 | Database pools keep idle connections; PostgreSQL tool calls 850 → 3,600/s with zero failures |
| 2026-07-22 (4) | 10,125 | Upstream latency model raised to 1 s calls / 250 ms listings (earlier rows used 100 ms); surfaced a PostgreSQL concurrency limit, fixed in the next row |
| 2026-07-22 (3) | 10,687 | Runtime memory limit derived from the container ceiling; the previous out-of-memory failure mode became clean GC pressure (+53%) |
| 2026-07-22 (2) | 7,000 | Connecting a stream no longer performs upstream work, fixing stalls under reconnect bursts; first measurable headline |
| 2026-07-22 (1) | — | First baseline: rate ceilings (2,600 calls/s embedded) and per-connection cost measured; headline blocked on the stream-connect stall |
