Files
Rendezvous/docs/operations/capacity-and-resilience.md
T

190 lines
11 KiB
Markdown

# Capacity, resilience, and availability gate
Tracking: #18
This gate turns the v1 budgets in ADR 0003 into a repeatable release decision.
It does not turn Rendezvous into a horizontally scalable service: v1 remains one
active process with bounded in-memory state. A second process may be a cold
standby, but it must not accept traffic until the first process has stopped and
released the public HTTP and UDP endpoints.
## Launch envelope and approved core-state profile
The approved core-state profile is one Linux process limited to 2 vCPU and
2 GiB RAM. Public HTTP/UDP numbers are launch objectives that require the #23
real-network canary before they become a supported service claim:
| Dimension | Value | Evidence status |
| --- | --- | --- |
| Visible listings | 25,000 | Enforced and measured here |
| Active join attempts | 10,000 | Enforced and measured here |
| Core control path | 200 operations/second; p95 at most 200 ms | Measured here |
| Core mediation path | 2,000 pairings/second; p95 at most 100 ms | Measured here |
| Sustained HTTP demand | 200 requests/second | #23 launch objective; not yet a supported claim |
| Sustained UDP demand | 2,000 datagrams/second | #23 launch objective; not yet a supported claim |
| Public HTTP/UDP latency | p95 at most 200 ms / 100 ms | #23 launch objective; not yet a supported claim |
| Capacity-phase average CPU / peak working memory | below 70% / below 1.5 GiB | Measured for the core candidate |
| Valid in-profile monthly availability | 99.5%, excluding announced maintenance | Operational objective |
| Process-ready RTO / host-visible recovery | 15 seconds / 90 seconds | 15 seconds automated; 90-second deployment drill required |
The proposed public-network mix is 20% registration/update, 30% lease-critical
renew/delete, 30% browse, and 20% join authorization for HTTP. The UDP mix is
60% authenticated host-presence refresh, 30% attempt contributions, and 10%
invalid or duplicate traffic that must be dropped early. A deployment may use a
lower per-game profile, but must not claim a higher one without new versioned
evidence.
The capacity harness fills the complete state ceilings, then measures
registration plus presence, renewal, a 100-item compatible browse, join
issuance, simultaneous two-peer pairing, principal revocation, and telemetry.
It applies 200/100 ms guardrails and minimum 200 control / 2,000 mediation
operations per second to the core hot path. Those measurements deliberately
exclude Kestrel, LiteNetLib, TLS, JSON, socket scheduling, and the documented
mixed traffic shape. The #23 real-network canary must exercise those layers,
rate-shape the mix, record errors and shedding, and meet the public objectives
before launch; a core result is not a public-network latency or throughput claim.
## Reproduce the evidence
Every push runs the quick profile and the selected fault matrix:
```bash
./scripts/run-capacity-gate.sh
```
Run the production candidate on an otherwise idle Linux host and restrict the
runtime to two logical CPUs. The default candidate includes a five-minute,
high-intensity expiry soak; use 3,600 seconds for a release-candidate endurance
run:
```bash
export RENDEZVOUS_CAPACITY_PROFILE=candidate
export RENDEZVOUS_CAPACITY_CPUSET=0,1
export RENDEZVOUS_CAPACITY_OUTPUT="$PWD/artifacts/capacity/candidate.json"
./scripts/run-capacity-gate.sh
# Release-candidate endurance override:
dotnet run --project tests/FinalFactory.Rendezvous.Capacity \
--configuration Release --no-build -- \
--profile candidate --soak-seconds 3600 \
--output artifacts/capacity/candidate-endurance.json
```
The machine must have at least 2 GiB available to the process. For formal
deployment evidence, run inside the same cgroup/container shape as production.
The v2 JSON embeds the commit and tree state, command, image context, CPU model,
kernel, affinity, cgroup quota/limit, collector mode, and workload seed. Supply
`RENDEZVOUS_EVIDENCE_IMAGE_DIGEST` when running a release image. Do not compare
results collected under a debugger,
concurrent build, thermal throttling, or oversubscribed CI host.
The checked-in baseline is
[`candidate-2cpu.json`](../evidence/capacity/v2/candidate-2cpu.json). It was
produced on .NET 10.0.9/Linux x64 with CPU affinity restricted to two logical
CPUs. It filled 25,000 listings and 10,000 attempts, peaked at about 162 MiB,
and cleared all active/retained state. The five-minute baseline supersedes any
earlier local probe when its timestamp and target duration differ.
## Soak and bounded-state interpretation
Each soak cycle creates a listing, repeatedly renews its lease and refreshes
presence, creates a join attempt, replay marker, and retained outcome, checks
that scheduled expiry entries remain proportional to live keys, then advances
the injected monotonic clock beyond all
deadlines, and verifies that listings, attempts, replay, idempotency, and outcome
state return to zero. The candidate also measures managed-memory and process
handle deltas after full collection. Failure is any retained state, more than
64 MiB retained managed memory, more than eight retained handles, a working set
above 1.5 GiB, an untyped capacity result, or failure to admit work after expiry.
This accelerated soak intentionally executes far more state lifecycle/cleanup
events than wall-clock traffic would permit. It catches stale deadline-queue
entries, cache growth, replay/idempotency retention, and cleanup cost. Because
it does not open Kestrel/LiteNetLib connections, its process-handle delta is only
a harness guard and is not evidence of transport stability by itself. The
selected production-process gate adds a ten-second real HTTP/UDP transport soak,
samples child-process handles and RSS, asserts bounded growth, then verifies a
clean SIGTERM and socket release. #23 must extend that into the full rate-shaped
multi-client canary while sampling queues, managed memory, and state
cardinalities. A one-hour core override remains required before tagging a
production release.
## Fault and recovery matrix
`run-capacity-gate.sh` runs these deterministic production paths before the
numeric profile:
| Fault | Required result |
| --- | --- |
| HTTP/UDP overload and tracker exhaustion | Typed HTTP `429`/`CapacityExceeded`, silent UDP drop, bounded tracker keys, recovery after the window |
| Optional traffic saturation | Lease-critical renew/update/delete capacity remains available |
| Store/dependency unavailable | Readiness fails; new authorization returns typed `ServiceUnavailable`; liveness remains independent |
| Graceful drain/SIGTERM | New work returns `Draining`; existing pairing may finish; process exits 0 and releases TCP/UDP before the deadline |
| Hard restart | In-flight state is lost; SDK reports typed `ServiceUnavailable`; a host re-registers, rebinds presence, and becomes the only browser-visible replacement |
| UDP listener bind/restart | Readiness stays false without the required listener; rebinding the advertised port restores native LiteNetLib pairing |
| Wall-clock jump/skew | Monotonic lease/attempt authority is neither shortened nor extended; credential skew remains capped at 30 seconds |
| Signing-secret rotation | New key signs, overlap verifies, retired/revoked key rejects, missing material fails startup |
| Principal revocation | Listing, presence, attempts, and outcome paths are removed atomically within the latency budget |
No external database exists in v1, so “dependency/store failure” means the
process-local atomic store is marked unavailable or a required listener/key is
unready. The service fails closed rather than pretending a degraded writable
mode exists.
## Bandwidth and amplification
- Accepted application datagrams are at most 1,200 bytes.
- Malformed, oversized, unauthenticated, stale, replayed, wrong-role, and
rate-limited traffic receives zero response bytes.
- A completing authenticated contribution produces at most one introduction to
each observed peer, and the combined response is at most 2.0 times that
contribution's bytes.
- The frozen-envelope and native LiteNetLib socket tests measure this on the real
UDP listener; the hostile corpus and allocation gate exercise 10,000+ inputs
without input-sized logs, tasks, or queues.
Bandwidth planning must therefore reserve ingress for the configured 2,000
datagrams/second plus edge overhead and egress for a worst-case verified 2.0
amplification. Actual successful pairs normally use two contributions and two
introductions; normal gameplay leaves Rendezvous entirely.
## Availability decision
Single-active remains the v1 topology. The measured core profile proves bounded
state and substantial core-path headroom, while public launch capacity remains
conditional on #23. The service has a bounded stop-before-start restart path.
Its failure domain is deliberately
one process/node/public UDP endpoint: node, kernel, host network, DNS/TLS edge,
secret configuration, or operator error can remove all readiness until the cold
replacement owns the same source-preserving endpoint.
The 99.5% objective permits about 216 minutes of unannounced downtime in a
30-day month. Operations must target process readiness within 15 seconds and
host-visible re-registration within 90 seconds, page when no ready instance
exists, and include detection plus recovery in the monthly budget. The current
in-process test validates typed downtime, same-port HTTP restart, fresh
registration, presence rebinding, and browser visibility in under five seconds;
the production-process test separately validates graceful termination, TCP/UDP
release, replacement startup on the same endpoints, UDP readiness, and the
15-second process-ready RTO. Cold-standby activation policy and the 90-second
operator-to-host recovery objective still require a deployment drill before
release. Rollout and rollback use the
deployment runbook's drain, stop, socket-release, start, smoke sequence; never
overlap old and new active processes.
Bring shared TTL/CAS state and deterministic mediator routing forward before
enabling two active instances if any of these occurs:
- one node cannot sustain 150% of the measured 30-day peak while meeting SLOs;
- CPU stays above 70%, memory above 75%, attempt depth above 70%, or limiter
drops/latency remain elevated after abusive traffic is excluded;
- the availability target rises above 99.5% or planned maintenance must preserve
listings; or
- one region requires multiple simultaneously active mediator endpoints.
Rendezvous makes no multi-instance claim today, so a two-node atomic-pairing
test is intentionally not applicable. It becomes a hard release gate with the
shared-state/routing implementation; until then `SingleActiveInstance=false`
fails production startup. Multi-region and relay remain separate evidence-driven
decisions.