diff --git a/README.md b/README.md index 6a93ae7..b83b61f 100644 --- a/README.md +++ b/README.md @@ -76,3 +76,6 @@ The initial service does not provide: ## Project status Rendezvous is currently in its initial design and bootstrap stage. The first implementation should establish the contracts, directory leases, LiteNetLib mediator, client SDK, thin test client, and a three-party integration test before either game depends on it for production connectivity. + +The ratified v1 boundaries, trust decisions, privacy rules, safety budgets, and +threat model are indexed in [the architecture documentation](docs/architecture/README.md). diff --git a/docs/architecture/0001-v1-control-plane-boundaries.md b/docs/architecture/0001-v1-control-plane-boundaries.md new file mode 100644 index 0000000..6800aac --- /dev/null +++ b/docs/architecture/0001-v1-control-plane-boundaries.md @@ -0,0 +1,113 @@ +# ADR 0001: v1 control-plane boundaries and domain + +- Status: Accepted +- Date: 2026-07-16 +- Tracking: #2 + +## Context + +Rendezvous must help two game peers discover and attempt an authenticated direct +connection without becoming a game server, an identity provider, or a gameplay +traffic service. The HTTP API and UDP mediator share short-lived state and must +agree on authorization, endpoint freshness, and tenant scope. + +## Decision + +V1 is one ASP.NET Core deployable with separable directory, join-authorization, +endpoint-registry, NAT-mediator, and operations modules. Modules communicate +through application interfaces, not through transport DTOs or one another's +storage implementation. Contracts and the client SDK remain independently +packageable. + +The service is a connection control plane. A successful join authorization only +grants permission to attempt a direct connection. The game host remains the +final authority for player identity, capacity, bans, admission, and gameplay. +Rendezvous success is reported only after the host accepts a valid connection +ticket and LiteNetLib establishes the authenticated peer connection. + +## Domain glossary + +| Term | Definition | Lifetime and exposure | +| --- | --- | --- | +| `SessionListing` | Bounded public discovery data for one hosted game session. | Visible only while its lease and host presence are fresh. Never contains endpoints or credentials. | +| `Lease` | Renewable capability controlling the lifetime of a listing. | Secret, host-only, expires unless renewed. | +| `HostPresence` | Authenticated observation of the host's local and public UDP endpoints from its gameplay socket. | Internal, short-lived, never returned by browsing. | +| `JoinAttempt` | Authorization linking one client attempt to one compatible session. | Internal and short-lived; it is not authoritative game admission. | +| `PunchCapability` | Opaque, one-time credential scoped to attempt, role, tenant, and expiry. | Sent only to its intended peer; consumed at the UDP mediator. | +| `ConnectionTicket` | Compact signed credential presented to the host during the direct connection. | One-time, short-lived, and scoped to the attempt and host. | + +IDs are opaque and tenant-scoped. They are never canonical player, entity, or +world identities. + +## Trust boundaries + +```mermaid +flowchart LR + Browser["Untrusted browser/client"] -->|"HTTPS: browse/join"| Proxy["Reverse proxy"] + Host["Game host"] -->|"HTTPS: register/renew"| Proxy + Operator["Privileged operator"] -->|"separate authenticated route"| Proxy + Proxy -->|"normalized HTTP + trusted forwarding metadata"| Service["Rendezvous service"] + Host -->|"host gameplay UDP socket"| Mediator["UDP mediator module"] + Browser -->|"client gameplay UDP socket"| Mediator + Mediator <--> Service + Service -->|"read keys; never list or log values"| Secrets["Secret provider"] + Service -.->|"future authenticated state protocol"| Store["Future shared store"] + Service -->|"redacted events and aggregate metrics"| Ops["Observability systems"] +``` + +- Public HTTP input is hostile even after TLS termination. The proxy may be + trusted to terminate TLS and supply forwarding metadata only when its source + address is allowlisted; forwarded headers from other sources are discarded. +- Public UDP input is hostile even when structurally valid. HTTP-supplied + endpoints are claims, never proof. Public response targets come only from an + authenticated UDP packet's observed source. A private local candidate may be + carried inside that packet only under ADR 0002's bounded same-LAN rules. +- Operator routes use a separate authentication policy and network exposure. + Operator access does not bypass tenant scoping, audit, or secret redaction. +- The client SDK is convenience code in an untrusted process. Server decisions + never rely on client-side validation or secrecy. +- Game hosts are authoritative only for their own gameplay admission. A host + cannot enumerate or mutate another game/environment tenant. +- The secret provider is trusted with long-lived key material. The application + receives only the minimum named key version it needs. +- A future shared store is a distinct authenticated boundary. Moving state to it + does not make stored input trusted and requires a new availability ADR. + +## Connection data flow + +```mermaid +sequenceDiagram + participant H as Game host + participant R as Rendezvous HTTP + participant M as Rendezvous UDP mediator + participant C as Game client + H->>R: Register listing (publisher authorization) + R-->>H: Lease + host-presence capability + H->>M: Presence from gameplay UDP socket + M->>R: Store observed endpoint and freshness + H->>R: Renew lease + C->>R: Browse compatible visible listings + C->>R: Request join attempt + R-->>C: Client punch capability + R-->>H: Host attempt/capability via authenticated poll or stream + H->>M: Host capability from gameplay UDP socket + C->>M: Client capability from gameplay UDP socket + M->>M: Validate scope, freshness, expiry, and replay state + M-->>H: Introduce verified client endpoints + connection ticket + M-->>C: Introduce verified host endpoints + connection ticket + C->>H: Direct LiteNetLib connect + ticket + H->>H: Validate and consume ticket; apply game admission + H-->>C: Authenticated peer connection or rejection +``` + +The mediator does not forward normal gameplay packets. A connection attempt +that times out or is rejected returns a typed outcome to the caller. + +## Consequences + +- Directory and mediator can ship together without erasing their module boundary. +- Contracts cannot expose server storage or LiteNetLib implementation types. +- Tests must cover the three-party host/service/client flow; an HTTP-only test is + insufficient evidence of a successful connection. +- Splitting modules into processes requires an explicit protocol, shared-state + ownership, deterministic mediator routing, and a superseding ADR. diff --git a/docs/architecture/0002-publisher-trust-and-connection-policy.md b/docs/architecture/0002-publisher-trust-and-connection-policy.md new file mode 100644 index 0000000..23b2f47 --- /dev/null +++ b/docs/architecture/0002-publisher-trust-and-connection-policy.md @@ -0,0 +1,88 @@ +# ADR 0002: publisher trust, discovery, compatibility, and fallback + +- Status: Accepted +- Date: 2026-07-16 +- Tracking: #2 + +## Context + +Dedicated servers can protect provisioned credentials. Public game binaries +cannot. Discovery also needs rules that prevent accidental cross-game joins and +make the meaning of a successful authorization precise. + +## Decision + +### Publisher trust modes + +Each `GameId` and `EnvironmentId` is provisioned policy, never caller-created +free text. V1 supports two visibly distinct publisher modes: + +1. **Managed dedicated host.** A provisioned workload principal authenticates + with a rotatable credential held outside the game binary. It is scoped to + allowed games, environments, regions, and listing limits. Public or unlisted + discovery may be enabled by policy. +2. **Player-hosted session.** A short-lived publisher grant is minted by a + game-owned backend and is scoped to one game, environment, host, and expiry. + Rendezvous does not interpret it as player identity. If a game has no grant + issuer, it may opt into anonymous unlisted hosting with strict address and + concurrency limits; anonymous sessions can be joined only through an opaque + share code and never appear in public browsing. + +A reusable credential embedded in a downloadable client is not authentication +and is rejected as a provisioning design. Responses and metrics expose the +publisher trust mode so operators and games can apply different policy without +claiming anonymous hosts are authenticated identities. + +### Discovery and metadata + +- `Public` listings can appear only in tenant-scoped compatible browsing. +- `Unlisted` listings never appear in browse results and require a random, + unguessable share code. Unlisted does not mean private; join authorization and + host admission still apply. +- Browser responses contain display data only. They exclude raw endpoints, + internal IDs, lease credentials, punch capabilities, and connection tickets. +- Metadata is treated as hostile data. It is schema/budget validated, stored and + returned as data, and never rendered as markup by the SDK or TestClient. + +### Compatibility and address families + +- `NetworkProtocolVersion` must match exactly in v1. `BuildVersion` is bounded + display/diagnostic text and never overrides protocol compatibility. +- `GameId` and `EnvironmentId` must match exactly. Region is a browse filter and + preference, not a compatibility escape hatch. +- IPv4 direct connection and NAT punching are required for v1. +- Contracts carry an address-family discriminator. IPv6 direct connections may + use observed global IPv6 endpoints when both peers support them, but IPv6 NAT + traversal is not a v1 release requirement. +- Public candidates are derived only from the authenticated UDP packet's source. + For same-LAN attempts, that packet may additionally claim at most one private + unicast candidate per supported address family. A local claim is scoped to the + capability and is introduced only to the opposite role in the same authorized + attempt after both roles contribute. Loopback, link-local, multicast, + unspecified, documentation, and otherwise invalid destinations are rejected. + The SDK bounds probes per introduced candidate and lets callers disable local + candidates. HTTP-supplied endpoint claims are never introduced. + +### Authorization and fallback + +Join authorization means only that Rendezvous permits a scoped connection +attempt. It does not reserve a game slot and does not authenticate a player to +the game. The host validates and consumes the connection ticket, then applies +its own capacity, ban, identity, and gameplay rules. + +The SDK returns a typed outcome including success, cancellation, timeout, +incompatibility, stale host, service rejection, host rejection, and transport +failure. A game may provision an optional dedicated fallback endpoint. The SDK +reports it but never connects without an explicit caller decision. + +Gameplay relay is not part of v1. It remains a separate future service whose +need is evaluated from privacy-safe measured direct-connection failures. + +## Consequences + +- A player-hosted game needs a game-owned grant issuer for public discovery. +- Anonymous player hosting is useful for direct invitations but makes no user + identity claim and receives the strictest quotas. +- Games remain responsible for presenting unsafe user-authored text safely. +- The exact-match v1 rule favors predictable interoperation over flexible + version ranges; a later compatibility scheme must be versioned explicitly. diff --git a/docs/architecture/0003-state-privacy-availability-and-budgets.md b/docs/architecture/0003-state-privacy-availability-and-budgets.md new file mode 100644 index 0000000..7338692 --- /dev/null +++ b/docs/architecture/0003-state-privacy-availability-and-budgets.md @@ -0,0 +1,137 @@ +# ADR 0003: state, privacy, availability, and safety budgets + +- Status: Accepted +- Date: 2026-07-16 +- Tracking: #2 + +## Context + +V1 needs safe defaults before contracts and stores make them difficult to +change. The initial deployment is deliberately single-active and in-memory, so +its restart and availability behavior must be honest. + +## Decision + +### State and lifecycle + +All directory, lease, presence, attempt, capability, ticket-consumption, and +rate-limit state is ephemeral and held behind atomic store interfaces. V1 has +one active writer/service instance. A second instance may be a cold standby but +must not accept public traffic concurrently. + +```mermaid +stateDiagram-v2 + [*] --> Registered: authenticated register + Registered --> Visible: fresh lease and fresh UDP presence + Visible --> Registered: presence becomes stale + Visible --> Visible: lease renew + presence refresh + Registered --> Expired: lease expires + Visible --> Expired: lease expires + Registered --> Revoked: host or operator revokes + Visible --> Revoked: host or operator revokes + Expired --> [*] + Revoked --> [*] +``` + +Restart loses all ephemeral state, used capabilities, and listings. Readiness is +false until HTTP, UDP, policy, key material, and the state store are ready. SDK +publishers use jittered backoff and re-register after a restart; old credentials +remain invalid. The service drains by refusing new registrations/attempts, +allowing a bounded completion window, then cancelling remaining work. + +No horizontal scale is supported until shared atomic state and deterministic +mediator routing exist. A shared-state design is triggered when any of these is +true: + +- one measured supported node cannot sustain 150% of the 30-day peak load; +- the approved availability target exceeds what single-active operation can meet; +- planned maintenance without listing loss becomes a product requirement; or +- a region needs more than one active mediator endpoint. + +Relay remains independently triggered only when a representative real-network +canary shows direct-connect failure high enough to justify its privacy, abuse, +bandwidth, and operating cost. + +### Initial time and size budgets + +These are enforceable v1 ceilings, not suggestions. Contract issue #4 may lower +them but must not raise them without security review. + +| Budget | V1 ceiling | +| --- | --- | +| HTTP request body | 16 KiB after content decoding; compressed request bodies are rejected in v1 | +| Listing metadata | 4 KiB encoded JSON, at most 32 keys; key 64 UTF-8 bytes; scalar value 256 UTF-8 bytes; nesting depth 3 | +| Browser page | 100 listings and 256 KiB encoded response; opaque cursor; stable bounded sort | +| UDP datagram accepted | 1,200 bytes; oversized or fragmented application payloads are dropped without response | +| Opaque HTTP credential | 1,024 bytes encoded | +| UDP capability or ticket | 768 bytes encoded, with the complete datagram still at most 1,200 bytes | +| Clock skew | 30 seconds maximum when validating issued/not-before/expiry times | +| Lease lifetime | 60 seconds; renewal accepted from 30 seconds; no client-selected extension | +| Host presence freshness | 20 seconds | +| Join attempt lifetime | 30 seconds | +| Punch capability lifetime | 30 seconds and one successful use per role | +| Connection ticket lifetime | 20 seconds and one successful host consumption | +| Graceful drain | 30 seconds maximum | + +All work queues are bounded. Initial per-instance ceilings are 1,024 concurrent +HTTP requests, 4,096 queued UDP datagrams, and 10,000 active join attempts. +Overflow is rejected or dropped early with a metric; it never creates an +unbounded task, allocation, log entry, or retry loop. + +For an endpoint that has not proved possession of a valid capability, the UDP +mediator sends no response. Once both valid peer contributions exist, +authenticated mediation sends at most one introduction datagram to each peer. +The combined response bytes caused by the completing contribution must be no +more than twice that contribution's bytes, giving zero unverified amplification +and at most 2.0 verified byte amplification. Protocol padding or a smaller +response enforces the byte ratio. Responses are sent only to endpoints observed +from the corresponding authenticated gameplay socket, never to an arbitrary +HTTP-supplied address. + +### Supported and capacity profiles + +The development profile is functional, not a production capacity claim. The +initial production candidate is one Linux instance with 2 vCPU and 2 GiB RAM, +targeting 25,000 visible listings, 10,000 active attempts, 200 HTTP requests per +second, and 2,000 UDP datagrams per second while staying below 70% sustained CPU +and 75% memory. Issue #18 must measure and publish the actual supported profile; +production is blocked if the target is not met or the documented profile is not +reduced accordingly. + +The initial single-active service objective, after the real-network canary, is +99.5% monthly successful availability for valid in-profile requests, excluding +announced maintenance. In-profile latency objectives are p95 <= 200 ms for HTTP +and p95 <= 100 ms from the second valid UDP contribution to both introduction +datagrams. These are service objectives, not guarantees of NAT traversal. + +### Data classification and retention + +| Data | Classification | Retention and handling | +| --- | --- | --- | +| Raw public/local endpoints | Sensitive network data | In memory only while the lease/attempt requires it, then deleted within 10 minutes; never logged or exported as metric labels | +| Listing display metadata | Public-untrusted or unlisted-untrusted | In memory for the active lease; audit stores only schema/result and a listing ID, not metadata values | +| Lease/capability/ticket/key material | Secret | Opaque random credentials are retained only as keyed digests; signed credentials retain verification keys and consumption IDs, not issued plaintext; plaintext is returned only at creation and is never logged or traced | +| Principal and tenant IDs | Internal identifiers | Audit retention 30 days; access-controlled and never used as high-cardinality metric labels | +| Security/audit event | Confidential operations data | 30 days online, access-controlled; contains action, coarse result, tenant, principal, and correlation ID, but no raw endpoint or secret | +| Diagnostic attempt record | Sensitive diagnostic data | Disabled by default; when explicitly enabled, redacted record retained at most 24 hours; raw endpoints remain excluded | +| Aggregate outcome/capacity metrics | Operational aggregate | 13 months; only bounded dimensions such as game, environment, region, trust mode, and typed outcome | + +Logs use allowlisted fields rather than after-the-fact redaction. Correlation IDs +are random and are not credentials. Error responses are stable and do not reveal +whether a cross-tenant resource exists. + +## Owner decisions required before production + +Implementation can proceed with the baseline above. Production remains blocked +until the owner records: + +- the actual secret-provider and key-custody system for each environment; +- which games may enable anonymous unlisted player hosting; +- deployment regions, data-processing jurisdiction, and approval of the stated + 30-day audit/13-month aggregate retention periods; +- the per-game dedicated fallback endpoint policy; +- the measured supported profile and whether the 99.5% single-active objective + is sufficient or shared-state/high-availability work must be brought forward. + +These are configuration and launch decisions, not permission to weaken the +tenant, replay, endpoint-verification, or secret-handling controls. diff --git a/docs/architecture/README.md b/docs/architecture/README.md new file mode 100644 index 0000000..38c5825 --- /dev/null +++ b/docs/architecture/README.md @@ -0,0 +1,14 @@ +# Rendezvous architecture decisions + +These records define the v1 architecture baseline. A later change to a ratified +decision requires a superseding ADR and corresponding contract/test updates. + +- [ADR 0001: v1 control-plane boundaries and domain](0001-v1-control-plane-boundaries.md) +- [ADR 0002: publisher trust, discovery, compatibility, and fallback](0002-publisher-trust-and-connection-policy.md) +- [ADR 0003: state, privacy, availability, and safety budgets](0003-state-privacy-availability-and-budgets.md) +- [Threat model](../security/threat-model.md) +- [Security promise and test matrix](../security/control-matrix.md) + +These decisions intentionally leave gameplay authority, player identity, +simulation, persistence, social features, skill matchmaking, and gameplay +traffic with each game. Relay is future evidence-driven scope, not part of v1. diff --git a/docs/security/control-matrix.md b/docs/security/control-matrix.md new file mode 100644 index 0000000..b779ba0 --- /dev/null +++ b/docs/security/control-matrix.md @@ -0,0 +1,27 @@ +# README security promise and test matrix + +Tracking: #2 + +This matrix turns each security and lifecycle promise in the README into an +enforceable control and planned evidence. Issue numbers refer to the delivery +backlog where the control is implemented and verified. + +| README promise | Enforceable control | Planned evidence | +| --- | --- | --- | +| Per-game credentials and signing keys | Provisioned principals and versioned keys are scoped to game/environment; secrets come from a provider and never a public binary. (#5) | Cross-tenant authorization tests, rotation/overlap/revocation tests, and secret scans. | +| Short-lived, single-purpose tokens resistant to replay | Issuer fixes audience, tenant, attempt, role, issued/expiry times, nonce, and key ID; store atomically consumes nonce/ticket. (#4, #6, #10) | Golden vectors; expired, future, mutated, wrong-role, wrong-tenant, and concurrent replay tests. | +| Strict payload, metadata, and token size limits | ADR 0003 ceilings are checked before allocation/deserialization and again at domain construction. (#4, #15) | Boundary/property tests, malformed corpus, and allocation-aware fuzzing. | +| Registration, query, and introduction rate limits | Layered per-address, principal, tenant, and global token buckets with bounded queues and stable retry guidance. (#15) | Limit partition/isolation tests and overload/soak profiles. | +| Lease expiry removes abandoned servers | Visibility and join eligibility atomically require a fresh lease and fresh authenticated presence. (#6, #7) | Fake-clock expiry, renew/expire race, restart, and stale-host join tests. | +| Validate game, environment, room, and protocol boundaries | Every identifier is a validated type; store keys and authorization decisions include server-derived tenant scope; protocol is exact-match in v1. (#4-#10) | Contract, tenant-isolation, incompatible-version, and confused-deputy tests. | +| Structured audit events without secrets or reusable credentials | Allowlisted audit schema excludes metadata values, raw endpoints, tokens, and key material; event volume is bounded. (#16) | Captured-log/audit assertions and credential canary scans. | +| Public endpoint observation | Only authenticated UDP packets from the gameplay socket establish public endpoint ownership; bounded private local candidates follow ADR 0002 and HTTP claims are never introduced. (#11) | Spoofed-source, arbitrary-target, private-range, and same-LAN/external tests. | +| Authenticated join and punch tokens | Join issuance rechecks compatible visible listing; mediator validates scoped one-time capabilities; host consumes signed ticket. (#10-#12) | Deterministic three-party success, rejection, replay, mismatch, and timeout tests. | +| Clear timeouts and failure results | SDK owns explicit deadlines/cancellation and returns a closed typed outcome set; NAT introduction alone is not success. (#12, #13) | Fake-clock deadline/cancellation and host-rejection tests. | +| Isolation by game, environment, protocol, and region | Tenant and protocol are mandatory exact filters; region is bounded policy/filter data and cannot override tenant compatibility. (#5, #8, #15) | Cross-product browse/register/join isolation tests. | +| Operational health, metrics, logging, administration, and rate limiting | Separate liveness/readiness, bounded privacy-safe metrics/logs, authenticated operator controls, and overload signals. (#15, #16, #27) | Authorization matrix, redaction tests, dashboard queries, and failure-injection checks. | +| Service leaves gameplay path after direct connection | Mediator handles only presence/capability/introduction messages and has no gameplay forwarding API. (#4, #11) | Contract/API review, UDP unknown-message drop tests, and end-to-end traffic-path assertion. | +| Direct traversal is not guaranteed and requires fallback | SDK distinguishes traversal failure from service/host rejection and only returns configured fallback data for caller choice. Relay is absent from v1. (#13, #20, #24) | Typed-outcome tests and TestClient scripted fallback scenarios. | + +Release readiness requires the linked implementation tests to exist and pass; +the design documents alone do not satisfy the security promise. diff --git a/docs/security/threat-model.md b/docs/security/threat-model.md new file mode 100644 index 0000000..fcb6247 --- /dev/null +++ b/docs/security/threat-model.md @@ -0,0 +1,96 @@ +# Rendezvous v1 threat model + +Tracking: #2 + +## Scope and assets + +This model covers the public HTTP API, public LiteNetLib-compatible UDP mediator, +operator API, client SDK, game host integration, reverse proxy, secret provider, +observability pipeline, and the proposed future shared store. Gameplay traffic +after direct connection and game-owned identity/admission systems are outside +the service boundary, but their handoff is in scope. + +Assets include tenant isolation, service availability, signing and publisher +keys, lease and connection credentials, raw endpoints, unlisted share codes, +listing integrity, audit integrity, and the guarantee that Rendezvous does not +turn into a reflector or private-network probe. + +## Actors and assumptions + +- Anonymous Internet attackers can send arbitrary HTTP and UDP traffic, spoof + source addresses where their network permits it, scrape listings, and create + many identities or addresses. +- Malicious publishers possess credentials only for their assigned tenant and + may submit hostile metadata or attempt to target arbitrary endpoints. +- Malicious clients can obtain legitimate join credentials for sessions they can + see and may replay, race, mutate, or share those credentials. +- A compromised game client and its SDK are fully attacker-controlled. No + reusable secret in them is trustworthy. +- Operators are privileged but fallible. Their actions are authenticated, + constrained, and audited. +- The reverse proxy, secret provider, and build/release pipeline are trusted + dependencies. Their compromise is considered and mitigated but cannot be + completely contained by the application. + +## Abuse paths and controls + +```mermaid +flowchart TD + A["Attacker input"] --> H{"HTTP or UDP?"} + H -->|HTTP| V["Authenticate when required; validate tenant, schema, size, and rate"] + H -->|UDP| U["Parse bounded datagram; validate capability before response"] + V --> S{"Allowed and in quota?"} + U --> E{"Capability valid, fresh, scoped, unused, and endpoint observed?"} + S -->|No| R["Stable bounded rejection"] + E -->|No| D["Silent drop + bounded aggregate metric"] + S -->|Yes| State["Atomic ephemeral state transition"] + E -->|Yes| State + State --> O["Allowlisted audit event; no secrets/endpoints"] +``` + +| Threat | Example | Required prevention/detection | Planned evidence | +| --- | --- | --- | --- | +| Spoofing and reflection | Forged UDP source causes traffic to a victim | No response before valid capability proof; send responses only to observed authenticated sources; at most two responses and <=2.0 verified byte amplification | Packet-level spoof/reflection tests and amplification accounting | +| Private-network probing | Publisher supplies `127.0.0.1`, link-local, or another victim as a same-LAN candidate | Accept only bounded private-unicast claims inside a scoped authenticated UDP contribution; reject prohibited ranges; disclose only to the opposite role in that attempt; bound SDK probes | Endpoint classification matrix and three-party adverse tests | +| Capability/ticket replay | Reuse a captured token to repeat introductions or connect | Short expiry, role/tenant/attempt scope, atomic one-time consumption, bounded skew, key rotation | Concurrent replay and post-expiry tests with golden vectors | +| Cross-tenant access | Game A browses, renews, or joins Game B | Server-derived principal scope on every lookup and atomic mutation; indistinguishable not-found response | Tenant isolation tests across every endpoint/store operation | +| Listing spam and scraping | Flood registrations or enumerate public sessions | Trust-mode quotas, per-principal/address limits, bounded pages/cursors, rate limits, aggregate alerts | Rate-limit, cursor-tamper, and sustained-load tests | +| Metadata injection | Control characters or markup attack logs/UI | UTF-8/schema/size validation; store as data; exclude values from audit; SDK does not render markup | Malformed Unicode/JSON corpus and TestClient safe-display tests | +| Credential theft | Secret appears in log, URL, metric, crash, or package | Credentials in headers/bodies only; allowlisted logging; secret-provider indirection; no credential metric labels | Log-capture tests, repository/package scans, rotation exercise | +| Parser/resource exhaustion | Oversized, nested, fragmented, or high-rate input | Fixed ceilings, bounded parsers/queues/concurrency, early rejection/drop, no input-sized logging | Fuzz/property corpus, allocation limits, overload tests | +| Stale or crashed host | Dead listing remains joinable | Both lease and recent authenticated presence required; atomic expiry; join rechecks freshness | Fake-clock lifecycle and join-race tests | +| Clock manipulation | Token accepted outside intended lifetime | Server-issued timestamps, monotonic elapsed-time for local expiry, <=30 s wall-clock skew | Boundary and clock-jump tests | +| Operator misuse | Unauthorized enumeration/revocation or secret exposure | Separate strong auth/network policy, least privilege, tenant scope, immutable audit, secrets never readable through API | Authorization matrix and audit completeness tests | +| Reverse-proxy confusion | Forged forwarded address bypasses limits | Trust forwarding headers only from allowlisted proxies; direct traffic uses socket peer | Forwarded-header spoof tests | +| Store race | Renew/revoke/expire/replay operations interleave | Compare-and-swap/transactional interfaces and deterministic outcomes | Parallel race tests with a fake clock | +| Dependency/supply-chain compromise | Malicious or drifting package/build output | Central pinning, lock files, reproducible builds, vulnerability review, signed release provenance | Locked clean restore, dependency audit, artifact verification | +| Availability attack | Valid-looking traffic fills CPU, memory, queues, logs | Layered quotas, bounded queues/tasks, graceful overload, readiness/drain, capacity alerts | Load/soak/resilience gates and forced saturation tests | + +## Security invariants + +The implementation and its tests must preserve these invariants: + +1. No UDP response is sent to an endpoint that has not presented a valid scoped + capability from that observed endpoint. +2. No browse response contains an endpoint, secret, internal attempt ID, or + credential. +3. Every state lookup and mutation includes server-derived game/environment + scope; caller-supplied scope alone is never authoritative. +4. A listing is visible and joinable only while both lease and presence are + fresh at the atomic decision point. +5. A capability or ticket can cause at most one successful state transition for + its intended role and attempt. +6. Join authorization never bypasses host-owned final admission. +7. Input cannot create unbounded memory, work, response bytes, metric labels, or + log volume. +8. Raw endpoints and secrets never enter normal logs, traces, audit payloads, or + metric dimensions. + +## Residual risk + +Direct traversal cannot work through every NAT, firewall, carrier, or platform +policy. Rate limiting cannot eliminate distributed abuse. A compromised trusted +proxy, secret provider, operator identity, game grant issuer, or host credential +can act within its granted scope until detected and revoked. Unlisted share +codes can be disclosed by recipients. These risks are communicated as typed +outcomes and operational signals rather than hidden behind a success claim.