Files
Rendezvous/docs/operations/capacity-and-resilience.md
T

11 KiB

Capacity, resilience, and availability gate

Tracking: #18

This gate turns the v1 budgets in ADR 0003 into a repeatable release decision. It does not turn Rendezvous into a horizontally scalable service: v1 remains one active process with bounded in-memory state. A second process may be a cold standby, but it must not accept traffic until the first process has stopped and released the public HTTP and UDP endpoints.

Launch envelope and approved core-state profile

The approved core-state profile is one Linux process limited to 2 vCPU and 2 GiB RAM. Public HTTP/UDP numbers are launch objectives that require the #23 real-network canary before they become a supported service claim:

Dimension Value Evidence status
Visible listings 25,000 Enforced and measured here
Active join attempts 10,000 Enforced and measured here
Core control path 200 operations/second; p95 at most 200 ms Measured here
Core mediation path 2,000 pairings/second; p95 at most 100 ms Measured here
Sustained HTTP demand 200 requests/second #23 launch objective; not yet a supported claim
Sustained UDP demand 2,000 datagrams/second #23 launch objective; not yet a supported claim
Public HTTP/UDP latency p95 at most 200 ms / 100 ms #23 launch objective; not yet a supported claim
Capacity-phase average CPU / peak working memory below 70% / below 1.5 GiB Measured for the core candidate
Valid in-profile monthly availability 99.5%, excluding announced maintenance Operational objective
Process-ready RTO / host-visible recovery 15 seconds / 90 seconds 15 seconds automated; 90-second deployment drill required

The proposed public-network mix is 20% registration/update, 30% lease-critical renew/delete, 30% browse, and 20% join authorization for HTTP. The UDP mix is 60% authenticated host-presence refresh, 30% attempt contributions, and 10% invalid or duplicate traffic that must be dropped early. A deployment may use a lower per-game profile, but must not claim a higher one without new versioned evidence.

The capacity harness fills the complete state ceilings, then measures registration plus presence, renewal, a 100-item compatible browse, join issuance, simultaneous two-peer pairing, principal revocation, and telemetry. It applies 200/100 ms guardrails and minimum 200 control / 2,000 mediation operations per second to the core hot path. Those measurements deliberately exclude Kestrel, LiteNetLib, TLS, JSON, socket scheduling, and the documented mixed traffic shape. The #23 real-network canary must exercise those layers, rate-shape the mix, record errors and shedding, and meet the public objectives before launch; a core result is not a public-network latency or throughput claim.

Reproduce the evidence

Every push runs the quick profile and the selected fault matrix:

./scripts/run-capacity-gate.sh

Run the production candidate on an otherwise idle Linux host and restrict the runtime to two logical CPUs. The default candidate includes a five-minute, high-intensity expiry soak; use 3,600 seconds for a release-candidate endurance run:

export RENDEZVOUS_CAPACITY_PROFILE=candidate
export RENDEZVOUS_CAPACITY_CPUSET=0,1
export RENDEZVOUS_CAPACITY_OUTPUT="$PWD/artifacts/capacity/candidate.json"
./scripts/run-capacity-gate.sh

# Release-candidate endurance override:
dotnet run --project tests/FinalFactory.Rendezvous.Capacity \
  --configuration Release --no-build -- \
  --profile candidate --soak-seconds 3600 \
  --output artifacts/capacity/candidate-endurance.json

The machine must have at least 2 GiB available to the process. For formal deployment evidence, run inside the same cgroup/container shape as production. The v2 JSON embeds the commit and tree state, command, image context, CPU model, kernel, affinity, cgroup quota/limit, collector mode, and workload seed. Supply RENDEZVOUS_EVIDENCE_IMAGE_DIGEST when running a release image. Do not compare results collected under a debugger, concurrent build, thermal throttling, or oversubscribed CI host.

The checked-in baseline is candidate-2cpu.json. It was produced on .NET 10.0.9/Linux x64 with CPU affinity restricted to two logical CPUs. It filled 25,000 listings and 10,000 attempts, peaked at about 162 MiB, and cleared all active/retained state. The five-minute baseline supersedes any earlier local probe when its timestamp and target duration differ.

Soak and bounded-state interpretation

Each soak cycle creates a listing, repeatedly renews its lease and refreshes presence, creates a join attempt, replay marker, and retained outcome, checks that scheduled expiry entries remain proportional to live keys, then advances the injected monotonic clock beyond all deadlines, and verifies that listings, attempts, replay, idempotency, and outcome state return to zero. The candidate also measures managed-memory and process handle deltas after full collection. Failure is any retained state, more than 64 MiB retained managed memory, more than eight retained handles, a working set above 1.5 GiB, an untyped capacity result, or failure to admit work after expiry.

This accelerated soak intentionally executes far more state lifecycle/cleanup events than wall-clock traffic would permit. It catches stale deadline-queue entries, cache growth, replay/idempotency retention, and cleanup cost. Because it does not open Kestrel/LiteNetLib connections, its process-handle delta is only a harness guard and is not evidence of transport stability by itself. The selected production-process gate adds a ten-second real HTTP/UDP transport soak, samples child-process handles and RSS, asserts bounded growth, then verifies a clean SIGTERM and socket release. #23 must extend that into the full rate-shaped multi-client canary while sampling queues, managed memory, and state cardinalities. A one-hour core override remains required before tagging a production release.

Fault and recovery matrix

run-capacity-gate.sh runs these deterministic production paths before the numeric profile:

Fault Required result
HTTP/UDP overload and tracker exhaustion Typed HTTP 429/CapacityExceeded, silent UDP drop, bounded tracker keys, recovery after the window
Optional traffic saturation Lease-critical renew/update/delete capacity remains available
Store/dependency unavailable Readiness fails; new authorization returns typed ServiceUnavailable; liveness remains independent
Graceful drain/SIGTERM New work returns Draining; existing pairing may finish; process exits 0 and releases TCP/UDP before the deadline
Hard restart In-flight state is lost; SDK reports typed ServiceUnavailable; a host re-registers, rebinds presence, and becomes the only browser-visible replacement
UDP listener bind/restart Readiness stays false without the required listener; rebinding the advertised port restores native LiteNetLib pairing
Wall-clock jump/skew Monotonic lease/attempt authority is neither shortened nor extended; credential skew remains capped at 30 seconds
Signing-secret rotation New key signs, overlap verifies, retired/revoked key rejects, missing material fails startup
Principal revocation Listing, presence, attempts, and outcome paths are removed atomically within the latency budget

No external database exists in v1, so “dependency/store failure” means the process-local atomic store is marked unavailable or a required listener/key is unready. The service fails closed rather than pretending a degraded writable mode exists.

Bandwidth and amplification

  • Accepted application datagrams are at most 1,200 bytes.
  • Malformed, oversized, unauthenticated, stale, replayed, wrong-role, and rate-limited traffic receives zero response bytes.
  • A completing authenticated contribution produces at most one introduction to each observed peer, and the combined response is at most 2.0 times that contribution's bytes.
  • The frozen-envelope and native LiteNetLib socket tests measure this on the real UDP listener; the hostile corpus and allocation gate exercise 10,000+ inputs without input-sized logs, tasks, or queues.

Bandwidth planning must therefore reserve ingress for the configured 2,000 datagrams/second plus edge overhead and egress for a worst-case verified 2.0 amplification. Actual successful pairs normally use two contributions and two introductions; normal gameplay leaves Rendezvous entirely.

Availability decision

Single-active remains the v1 topology. The measured core profile proves bounded state and substantial core-path headroom, while public launch capacity remains conditional on #23. The service has a bounded stop-before-start restart path. Its failure domain is deliberately one process/node/public UDP endpoint: node, kernel, host network, DNS/TLS edge, secret configuration, or operator error can remove all readiness until the cold replacement owns the same source-preserving endpoint.

The 99.5% objective permits about 216 minutes of unannounced downtime in a 30-day month. Operations must target process readiness within 15 seconds and host-visible re-registration within 90 seconds, page when no ready instance exists, and include detection plus recovery in the monthly budget. The current in-process test validates typed downtime, same-port HTTP restart, fresh registration, presence rebinding, and browser visibility in under five seconds; the production-process test separately validates graceful termination, TCP/UDP release, replacement startup on the same endpoints, UDP readiness, and the 15-second process-ready RTO. Cold-standby activation policy and the 90-second operator-to-host recovery objective still require a deployment drill before release. Rollout and rollback use the deployment runbook's drain, stop, socket-release, start, smoke sequence; never overlap old and new active processes.

Bring shared TTL/CAS state and deterministic mediator routing forward before enabling two active instances if any of these occurs:

  • one node cannot sustain 150% of the measured 30-day peak while meeting SLOs;
  • CPU stays above 70%, memory above 75%, attempt depth above 70%, or limiter drops/latency remain elevated after abusive traffic is excluded;
  • the availability target rises above 99.5% or planned maintenance must preserve listings; or
  • one region requires multiple simultaneously active mediator endpoints.

Rendezvous makes no multi-instance claim today, so a two-node atomic-pairing test is intentionally not applicable. It becomes a hard release gate with the shared-state/routing implementation; until then SingleActiveInstance=false fails production startup. Multi-region and relay remain separate evidence-driven decisions.