Skip to content

Sizing and performance

TinyGuard is a single process. Its footprint is dominated by one thing — password hashing — and that dominates it as a transient peak, not as a resident cost. This page gives the measured numbers and the model behind them, so you can size a deployment from your own hashing settings rather than from a guess.

Container image 21 MB (distroless, uncompressed)
Idle memory ≈ 3 MB
After 44 sign-ins, including a concurrent burst ≈ 13 MB
Startup to ready ≈ 1.4 s
Throughput, cheap endpoints ≈ 7,800 requests/s

Argon2id is deliberately memory-hard. At the shipped defaults each hash asks for 19 MiB (hashing.argon2_m_cost = 19456 KiB, the OWASP parameter) and the server runs at most hashing.max_hash_threads (default 2) at once — a semaphore, so further sign-ins wait for a slot rather than allocating more.

That memory is used and released. Measured in the shipped container, counting only sign-ins that actually reached the server:

Sign-ins Memory
Idle ≈ 3 MB
1 ≈ 6 MB
8 ≈ 7.5 MB
32 ≈ 8.8 MB
44, including a 12-way burst ≈ 13 MB

The peak is bounded by max_hash_threads, and what a deployment should plan for is small:

peak ≈ base (≈ 10 MB) + max_hash_threads × argon2_m_cost

At the defaults that is roughly 50 MB, and it does not accumulate. Size the container limit for the peak with headroom — the chart’s 512Mi clears it many times over — and size the request for steady state.

Three things worth knowing before tuning:

  • This was not always true. Older builds climbed roughly one hash’s memory per sign-in and never came back: eight sign-ins reached 140 MB and thirty-eight reached 389 MB. glibc serves an allocation the size of a hash with mmap and frees it with munmap, but the first free raises its mmap threshold to the freed size, so every later hash came from the heap instead. The binary now pins that threshold at startup, which holds the same forty-four sign-ins at 13 MB. A deployment that runs the binary directly gets the fix because it is in the program; there is no environment variable to set.
  • Lowering the cost lowers the security with it. The documented floor is 19456 KiB, which is the OWASP parameter and also the shipped cost. Raising max_hash_threads raises the peak proportionally, so raise the limit with it.
  • Memory is not the bottleneck to size against. Sign-in cost is CPU: a hash is hundreds of milliseconds by design, and that is what makes a guessed password expensive.

Measured against a release-build container on one host, 16 concurrent clients, no errors:

Endpoint Requests/s
/auth/v1/ping ≈ 7,900
/auth/v1/version ≈ 7,800
/.well-known/openid-configuration ≈ 7,700

These are client-bound — the load generator and the container’s port mapping are in the path, so treat them as a floor rather than a ceiling. The useful signal is the shape: discovery is as cheap as a health check, so relying parties fetching metadata are not a load problem.

Sign-ins are a different order of magnitude by design: one costs an Argon2id hash — a few hundred milliseconds of CPU and 19 MiB by intent, because that is what makes a guessed password expensive — plus the proof-of-work gate. Size for concurrency, not for rate: the pool bounds simultaneous hashes, and requests beyond it wait.

Startup, storage, and what a restart costs

Section titled “Startup, storage, and what a restart costs”

Startup to a ready listener is about 1.4 s for the container, including storage migrations. A rolling restart is therefore cheap, and the readiness probe keeps traffic off a pod until it is serving.

Storage is small: tenant, role, group, client, session and event records, all compact JSON. 10 GB — the chart’s default volume — is far more than most deployments will use; watch it if you keep a long event archive, and see Events for retention.

A replica holds up to pool_size connections (default 10), opened as needed, so a deployment needs replicas x pool_size connections plus a few for migrations and tools. See the connection budget.

The pool is what lets a replica use them. Measured on one host against a local server, with 64 concurrent callers running a mix of mostly reads, some writes and the occasional listing:

pool_size Operations/s Relative
1 ≈ 1,300 – 2,000 1.0×
2 ≈ 3,100 – 4,200 ≈ 2.2×
4 ≈ 5,400 – 6,400 ≈ 3.7×
10 ≈ 10,300 – 12,100 ≈ 6 – 7.7×

A pool of one is how earlier builds behaved — every operation of a replica queued behind a single connection — so the last row is the improvement. Treat the absolute figures as a floor: they come from a development build and a server on the same machine, with no network between them. The relative shape is the useful part, and it is close to linear until the pool is as large as the concurrency the workload actually has.

Memory is unchanged by the pool: connections are small next to a hash.

Being explicit about the gaps, so nobody reads the numbers above as more than they are:

  • One machine, one process. The figures come from a development host (Apple silicon) running one container. No multi-tenant load and no long soak has been measured.
  • No sign-in throughput figure. Only endpoint throughput was measured; sign-in cost depends on your hashing settings and your client mix.
  • PostgreSQL over a real network. The server was on the same machine, so the round trip was as cheap as it gets. Over a network each operation pays its latency, and the pool is what hides it: size pool_size for your latency and concurrency, and check the result against your own server. Nor was a database shared with another component under its own load measured.

If you need figures for a specific deployment, the model above is what to reason from, and the configuration you would change is in the configuration reference.