docs: ratify v1 architecture and threat model (#2)

Closes #2
This commit is contained in:
KyuubiYoru
2026-07-16 04:11:24 +02:00
parent 2959c7e845
commit 286fbfeb36
7 changed files with 478 additions and 0 deletions
@@ -0,0 +1,113 @@
# ADR 0001: v1 control-plane boundaries and domain
- Status: Accepted
- Date: 2026-07-16
- Tracking: #2
## Context
Rendezvous must help two game peers discover and attempt an authenticated direct
connection without becoming a game server, an identity provider, or a gameplay
traffic service. The HTTP API and UDP mediator share short-lived state and must
agree on authorization, endpoint freshness, and tenant scope.
## Decision
V1 is one ASP.NET Core deployable with separable directory, join-authorization,
endpoint-registry, NAT-mediator, and operations modules. Modules communicate
through application interfaces, not through transport DTOs or one another's
storage implementation. Contracts and the client SDK remain independently
packageable.
The service is a connection control plane. A successful join authorization only
grants permission to attempt a direct connection. The game host remains the
final authority for player identity, capacity, bans, admission, and gameplay.
Rendezvous success is reported only after the host accepts a valid connection
ticket and LiteNetLib establishes the authenticated peer connection.
## Domain glossary
| Term | Definition | Lifetime and exposure |
| --- | --- | --- |
| `SessionListing` | Bounded public discovery data for one hosted game session. | Visible only while its lease and host presence are fresh. Never contains endpoints or credentials. |
| `Lease` | Renewable capability controlling the lifetime of a listing. | Secret, host-only, expires unless renewed. |
| `HostPresence` | Authenticated observation of the host's local and public UDP endpoints from its gameplay socket. | Internal, short-lived, never returned by browsing. |
| `JoinAttempt` | Authorization linking one client attempt to one compatible session. | Internal and short-lived; it is not authoritative game admission. |
| `PunchCapability` | Opaque, one-time credential scoped to attempt, role, tenant, and expiry. | Sent only to its intended peer; consumed at the UDP mediator. |
| `ConnectionTicket` | Compact signed credential presented to the host during the direct connection. | One-time, short-lived, and scoped to the attempt and host. |
IDs are opaque and tenant-scoped. They are never canonical player, entity, or
world identities.
## Trust boundaries
```mermaid
flowchart LR
Browser["Untrusted browser/client"] -->|"HTTPS: browse/join"| Proxy["Reverse proxy"]
Host["Game host"] -->|"HTTPS: register/renew"| Proxy
Operator["Privileged operator"] -->|"separate authenticated route"| Proxy
Proxy -->|"normalized HTTP + trusted forwarding metadata"| Service["Rendezvous service"]
Host -->|"host gameplay UDP socket"| Mediator["UDP mediator module"]
Browser -->|"client gameplay UDP socket"| Mediator
Mediator <--> Service
Service -->|"read keys; never list or log values"| Secrets["Secret provider"]
Service -.->|"future authenticated state protocol"| Store["Future shared store"]
Service -->|"redacted events and aggregate metrics"| Ops["Observability systems"]
```
- Public HTTP input is hostile even after TLS termination. The proxy may be
trusted to terminate TLS and supply forwarding metadata only when its source
address is allowlisted; forwarded headers from other sources are discarded.
- Public UDP input is hostile even when structurally valid. HTTP-supplied
endpoints are claims, never proof. Public response targets come only from an
authenticated UDP packet's observed source. A private local candidate may be
carried inside that packet only under ADR 0002's bounded same-LAN rules.
- Operator routes use a separate authentication policy and network exposure.
Operator access does not bypass tenant scoping, audit, or secret redaction.
- The client SDK is convenience code in an untrusted process. Server decisions
never rely on client-side validation or secrecy.
- Game hosts are authoritative only for their own gameplay admission. A host
cannot enumerate or mutate another game/environment tenant.
- The secret provider is trusted with long-lived key material. The application
receives only the minimum named key version it needs.
- A future shared store is a distinct authenticated boundary. Moving state to it
does not make stored input trusted and requires a new availability ADR.
## Connection data flow
```mermaid
sequenceDiagram
participant H as Game host
participant R as Rendezvous HTTP
participant M as Rendezvous UDP mediator
participant C as Game client
H->>R: Register listing (publisher authorization)
R-->>H: Lease + host-presence capability
H->>M: Presence from gameplay UDP socket
M->>R: Store observed endpoint and freshness
H->>R: Renew lease
C->>R: Browse compatible visible listings
C->>R: Request join attempt
R-->>C: Client punch capability
R-->>H: Host attempt/capability via authenticated poll or stream
H->>M: Host capability from gameplay UDP socket
C->>M: Client capability from gameplay UDP socket
M->>M: Validate scope, freshness, expiry, and replay state
M-->>H: Introduce verified client endpoints + connection ticket
M-->>C: Introduce verified host endpoints + connection ticket
C->>H: Direct LiteNetLib connect + ticket
H->>H: Validate and consume ticket; apply game admission
H-->>C: Authenticated peer connection or rejection
```
The mediator does not forward normal gameplay packets. A connection attempt
that times out or is rejected returns a typed outcome to the caller.
## Consequences
- Directory and mediator can ship together without erasing their module boundary.
- Contracts cannot expose server storage or LiteNetLib implementation types.
- Tests must cover the three-party host/service/client flow; an HTTP-only test is
insufficient evidence of a successful connection.
- Splitting modules into processes requires an explicit protocol, shared-state
ownership, deterministic mediator routing, and a superseding ADR.
@@ -0,0 +1,88 @@
# ADR 0002: publisher trust, discovery, compatibility, and fallback
- Status: Accepted
- Date: 2026-07-16
- Tracking: #2
## Context
Dedicated servers can protect provisioned credentials. Public game binaries
cannot. Discovery also needs rules that prevent accidental cross-game joins and
make the meaning of a successful authorization precise.
## Decision
### Publisher trust modes
Each `GameId` and `EnvironmentId` is provisioned policy, never caller-created
free text. V1 supports two visibly distinct publisher modes:
1. **Managed dedicated host.** A provisioned workload principal authenticates
with a rotatable credential held outside the game binary. It is scoped to
allowed games, environments, regions, and listing limits. Public or unlisted
discovery may be enabled by policy.
2. **Player-hosted session.** A short-lived publisher grant is minted by a
game-owned backend and is scoped to one game, environment, host, and expiry.
Rendezvous does not interpret it as player identity. If a game has no grant
issuer, it may opt into anonymous unlisted hosting with strict address and
concurrency limits; anonymous sessions can be joined only through an opaque
share code and never appear in public browsing.
A reusable credential embedded in a downloadable client is not authentication
and is rejected as a provisioning design. Responses and metrics expose the
publisher trust mode so operators and games can apply different policy without
claiming anonymous hosts are authenticated identities.
### Discovery and metadata
- `Public` listings can appear only in tenant-scoped compatible browsing.
- `Unlisted` listings never appear in browse results and require a random,
unguessable share code. Unlisted does not mean private; join authorization and
host admission still apply.
- Browser responses contain display data only. They exclude raw endpoints,
internal IDs, lease credentials, punch capabilities, and connection tickets.
- Metadata is treated as hostile data. It is schema/budget validated, stored and
returned as data, and never rendered as markup by the SDK or TestClient.
### Compatibility and address families
- `NetworkProtocolVersion` must match exactly in v1. `BuildVersion` is bounded
display/diagnostic text and never overrides protocol compatibility.
- `GameId` and `EnvironmentId` must match exactly. Region is a browse filter and
preference, not a compatibility escape hatch.
- IPv4 direct connection and NAT punching are required for v1.
- Contracts carry an address-family discriminator. IPv6 direct connections may
use observed global IPv6 endpoints when both peers support them, but IPv6 NAT
traversal is not a v1 release requirement.
- Public candidates are derived only from the authenticated UDP packet's source.
For same-LAN attempts, that packet may additionally claim at most one private
unicast candidate per supported address family. A local claim is scoped to the
capability and is introduced only to the opposite role in the same authorized
attempt after both roles contribute. Loopback, link-local, multicast,
unspecified, documentation, and otherwise invalid destinations are rejected.
The SDK bounds probes per introduced candidate and lets callers disable local
candidates. HTTP-supplied endpoint claims are never introduced.
### Authorization and fallback
Join authorization means only that Rendezvous permits a scoped connection
attempt. It does not reserve a game slot and does not authenticate a player to
the game. The host validates and consumes the connection ticket, then applies
its own capacity, ban, identity, and gameplay rules.
The SDK returns a typed outcome including success, cancellation, timeout,
incompatibility, stale host, service rejection, host rejection, and transport
failure. A game may provision an optional dedicated fallback endpoint. The SDK
reports it but never connects without an explicit caller decision.
Gameplay relay is not part of v1. It remains a separate future service whose
need is evaluated from privacy-safe measured direct-connection failures.
## Consequences
- A player-hosted game needs a game-owned grant issuer for public discovery.
- Anonymous player hosting is useful for direct invitations but makes no user
identity claim and receives the strictest quotas.
- Games remain responsible for presenting unsafe user-authored text safely.
- The exact-match v1 rule favors predictable interoperation over flexible
version ranges; a later compatibility scheme must be versioned explicitly.
@@ -0,0 +1,137 @@
# ADR 0003: state, privacy, availability, and safety budgets
- Status: Accepted
- Date: 2026-07-16
- Tracking: #2
## Context
V1 needs safe defaults before contracts and stores make them difficult to
change. The initial deployment is deliberately single-active and in-memory, so
its restart and availability behavior must be honest.
## Decision
### State and lifecycle
All directory, lease, presence, attempt, capability, ticket-consumption, and
rate-limit state is ephemeral and held behind atomic store interfaces. V1 has
one active writer/service instance. A second instance may be a cold standby but
must not accept public traffic concurrently.
```mermaid
stateDiagram-v2
[*] --> Registered: authenticated register
Registered --> Visible: fresh lease and fresh UDP presence
Visible --> Registered: presence becomes stale
Visible --> Visible: lease renew + presence refresh
Registered --> Expired: lease expires
Visible --> Expired: lease expires
Registered --> Revoked: host or operator revokes
Visible --> Revoked: host or operator revokes
Expired --> [*]
Revoked --> [*]
```
Restart loses all ephemeral state, used capabilities, and listings. Readiness is
false until HTTP, UDP, policy, key material, and the state store are ready. SDK
publishers use jittered backoff and re-register after a restart; old credentials
remain invalid. The service drains by refusing new registrations/attempts,
allowing a bounded completion window, then cancelling remaining work.
No horizontal scale is supported until shared atomic state and deterministic
mediator routing exist. A shared-state design is triggered when any of these is
true:
- one measured supported node cannot sustain 150% of the 30-day peak load;
- the approved availability target exceeds what single-active operation can meet;
- planned maintenance without listing loss becomes a product requirement; or
- a region needs more than one active mediator endpoint.
Relay remains independently triggered only when a representative real-network
canary shows direct-connect failure high enough to justify its privacy, abuse,
bandwidth, and operating cost.
### Initial time and size budgets
These are enforceable v1 ceilings, not suggestions. Contract issue #4 may lower
them but must not raise them without security review.
| Budget | V1 ceiling |
| --- | --- |
| HTTP request body | 16 KiB after content decoding; compressed request bodies are rejected in v1 |
| Listing metadata | 4 KiB encoded JSON, at most 32 keys; key 64 UTF-8 bytes; scalar value 256 UTF-8 bytes; nesting depth 3 |
| Browser page | 100 listings and 256 KiB encoded response; opaque cursor; stable bounded sort |
| UDP datagram accepted | 1,200 bytes; oversized or fragmented application payloads are dropped without response |
| Opaque HTTP credential | 1,024 bytes encoded |
| UDP capability or ticket | 768 bytes encoded, with the complete datagram still at most 1,200 bytes |
| Clock skew | 30 seconds maximum when validating issued/not-before/expiry times |
| Lease lifetime | 60 seconds; renewal accepted from 30 seconds; no client-selected extension |
| Host presence freshness | 20 seconds |
| Join attempt lifetime | 30 seconds |
| Punch capability lifetime | 30 seconds and one successful use per role |
| Connection ticket lifetime | 20 seconds and one successful host consumption |
| Graceful drain | 30 seconds maximum |
All work queues are bounded. Initial per-instance ceilings are 1,024 concurrent
HTTP requests, 4,096 queued UDP datagrams, and 10,000 active join attempts.
Overflow is rejected or dropped early with a metric; it never creates an
unbounded task, allocation, log entry, or retry loop.
For an endpoint that has not proved possession of a valid capability, the UDP
mediator sends no response. Once both valid peer contributions exist,
authenticated mediation sends at most one introduction datagram to each peer.
The combined response bytes caused by the completing contribution must be no
more than twice that contribution's bytes, giving zero unverified amplification
and at most 2.0 verified byte amplification. Protocol padding or a smaller
response enforces the byte ratio. Responses are sent only to endpoints observed
from the corresponding authenticated gameplay socket, never to an arbitrary
HTTP-supplied address.
### Supported and capacity profiles
The development profile is functional, not a production capacity claim. The
initial production candidate is one Linux instance with 2 vCPU and 2 GiB RAM,
targeting 25,000 visible listings, 10,000 active attempts, 200 HTTP requests per
second, and 2,000 UDP datagrams per second while staying below 70% sustained CPU
and 75% memory. Issue #18 must measure and publish the actual supported profile;
production is blocked if the target is not met or the documented profile is not
reduced accordingly.
The initial single-active service objective, after the real-network canary, is
99.5% monthly successful availability for valid in-profile requests, excluding
announced maintenance. In-profile latency objectives are p95 <= 200 ms for HTTP
and p95 <= 100 ms from the second valid UDP contribution to both introduction
datagrams. These are service objectives, not guarantees of NAT traversal.
### Data classification and retention
| Data | Classification | Retention and handling |
| --- | --- | --- |
| Raw public/local endpoints | Sensitive network data | In memory only while the lease/attempt requires it, then deleted within 10 minutes; never logged or exported as metric labels |
| Listing display metadata | Public-untrusted or unlisted-untrusted | In memory for the active lease; audit stores only schema/result and a listing ID, not metadata values |
| Lease/capability/ticket/key material | Secret | Opaque random credentials are retained only as keyed digests; signed credentials retain verification keys and consumption IDs, not issued plaintext; plaintext is returned only at creation and is never logged or traced |
| Principal and tenant IDs | Internal identifiers | Audit retention 30 days; access-controlled and never used as high-cardinality metric labels |
| Security/audit event | Confidential operations data | 30 days online, access-controlled; contains action, coarse result, tenant, principal, and correlation ID, but no raw endpoint or secret |
| Diagnostic attempt record | Sensitive diagnostic data | Disabled by default; when explicitly enabled, redacted record retained at most 24 hours; raw endpoints remain excluded |
| Aggregate outcome/capacity metrics | Operational aggregate | 13 months; only bounded dimensions such as game, environment, region, trust mode, and typed outcome |
Logs use allowlisted fields rather than after-the-fact redaction. Correlation IDs
are random and are not credentials. Error responses are stable and do not reveal
whether a cross-tenant resource exists.
## Owner decisions required before production
Implementation can proceed with the baseline above. Production remains blocked
until the owner records:
- the actual secret-provider and key-custody system for each environment;
- which games may enable anonymous unlisted player hosting;
- deployment regions, data-processing jurisdiction, and approval of the stated
30-day audit/13-month aggregate retention periods;
- the per-game dedicated fallback endpoint policy;
- the measured supported profile and whether the 99.5% single-active objective
is sufficient or shared-state/high-availability work must be brought forward.
These are configuration and launch decisions, not permission to weaken the
tenant, replay, endpoint-verification, or secret-handling controls.
+14
View File
@@ -0,0 +1,14 @@
# Rendezvous architecture decisions
These records define the v1 architecture baseline. A later change to a ratified
decision requires a superseding ADR and corresponding contract/test updates.
- [ADR 0001: v1 control-plane boundaries and domain](0001-v1-control-plane-boundaries.md)
- [ADR 0002: publisher trust, discovery, compatibility, and fallback](0002-publisher-trust-and-connection-policy.md)
- [ADR 0003: state, privacy, availability, and safety budgets](0003-state-privacy-availability-and-budgets.md)
- [Threat model](../security/threat-model.md)
- [Security promise and test matrix](../security/control-matrix.md)
These decisions intentionally leave gameplay authority, player identity,
simulation, persistence, social features, skill matchmaking, and gameplay
traffic with each game. Relay is future evidence-driven scope, not part of v1.