* feat(jellycompat): install web assets at runtime * fix(jellycompat): recover stale web operation locks * fix(jellycompat): harden web component management * feat(admin): refine compat settings and restart status * chore(dev): add hot-reload docker compose stack * fix(dev): include npm in hot-reload backend * feat(admin): refine Jellyfin compatibility settings * feat(settings): improve jellyfin proxy summary * feat(settings): improve jellyfin web controls * fix(settings): update jellyfin web removal status * fix(settings): enable jellyfin web after install * feat(jellycompat): auto-select web ui version * test(api): update rate limit handler setup * feat(jellycompat): refine web ui install onboarding * fix(jellycompat): address web ui install review issues * fix(onboarding): mirror jellyfin api runtime status * fix(admin): remove global restart banner * fix(settings): gate restart required tracking * fix(jellyfin): ignore live settings for restart status * fix(jellyfin): avoid restart for live compat settings * fix(subtitles): normalize AI language codes * fix(catalog): support partial title search tokens * feat(branding): add white-label customization * Add push relay engineering plan - Document relay API contracts, APNs/FCM behavior, auth, storage, and ops - Capture implementation plan, provider references, decisions, and README --------- Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
69 KiB
Silo Push Relay — Decision Record
Provenance. This is the decision record for the Silo Push Relay. It resolves the implementation decisions that the engineering spec (
00-relay-spec.md) left open in §15 ("Open questions needing human / product input"), building on the phased build plan (01-implementation-plan.md) and the cited 2026 factual foundation (02-apns-fcm-2026-reference.md). Each decision below is firm and implementer-followable: pick it, build it. The external wire contract in../02-apns-relay.mdand../03-fcm-relay.mdis authoritative and unchanged — no decision here adds, renames, or repurposes a caller-visible field. Commands assume the repository root is the cwd.
These decisions reflect 2026-current APNs and FCM behavior and a deliberately narrow v1 cut line: a single-region, self-hosted-ecosystem relay whose two priorities are reliability and custody of the official Apple/Google push credentials. Where the spec already leaned a direction, this record ratifies and operationalizes it; where it left a real fork, this record picks one decisively and says exactly what that commits the team to.
Summary
| Decision | Recommendation (one line) | Confidence | Gates |
|---|---|---|---|
| APNs Go client (§15 #3) | Hand-roll on net/http + x/net/http2 + golang-jwt/jwt/v5 behind internal/apns; do not take a production dep on sideshow/apns2. |
High | Phase 3 |
| FCM credentials (§15 #4) | Default keyless via ADC (attached SA on GCP, WIF off-GCP); SA-JSON-in-secret-manager is the explicit last-resort fallback. | High | Phase 4 / Phase 6 |
apns-expiration policy (§15 #5) |
Relay-set finite TTL: 4h private_alert, 1h background_wake, parallel FCM android.ttl; relay-owned config, not a caller field. |
High | Phase 3 (APNs) / Phase 4 (FCM) |
| Per-account rate limits (§15 #1) | Ship hard-coded global defaults (config constants + env overrides); add per-account columns + relayctl quota only when logs prove need. |
High | Phase 2 |
| API-key hashing (§15 #6) | HMAC-SHA256(pepper, secret) — fast keyed hash, pepper in secret manager; not argon2id, not salted SHA-256. |
High | Phase 1 / Phase 2 |
| Op-log retention (§15 #8) | Native PostgreSQL declarative RANGE partitioning by day, retention by partition DROP via pg_partman; not bulk DELETE, not external offload. |
High | Phase 1 / Phase 5 |
Decision 1 — APNs Go client: hand-roll, do not depend on sideshow/apns2
Decision. Should the relay's APNs client adopt sideshow/apns2, or hand-roll the client on net/http + golang.org/x/net/http2 + golang-jwt/jwt/v5? (Spec §15 #3; plan §3 "Decision to make before Phase 3".)
Resolution. Hand-roll the APNs client on net/http with x/net/http2 and golang-jwt/jwt/v5, behind the existing internal/apns interface. Do not take a production dependency on sideshow/apns2. Mirror its verified internals — production host api.push.apple.com, sandbox host api.sandbox.push.apple.com, token TTL 3000 s with a 20-minute refresh floor, an HTTP/2 transport with ReadIdleTimeout/PingTimeout near 15 s, one client per process per (env, team) — and add the behavior apns2 lacks: single-flighted JWT regeneration keyed by team, an (env, topic) → .p8 key-selection map for Apple's Feb 2025 scoped keys with startup validation, a finite apns-expiration policy (Decision 3), and a connection-severed → 503 mapping with no transparent retry. Optionally keep apns2 v0.25.0 as a throwaway Phase 3 oracle, then drop it before Phase 6.
Alternatives considered.
| Option | Pros | Cons |
|---|---|---|
| Hand-roll (chosen) | No dependency on a stalled library; relay owns its whole upstream transport. Surface is small and fully specified. jwt/v5 and x/net/http2 are maintained and already needed (FCM, HTTP/2 tuning). Enables single-flight JWT regen, scoped-key selection, finite expiration, and connection-severed mapping that apns2 lacks. |
More code/tests up front for the JWT lifecycle, HTTP/2 pool, GOAWAY handling, and reason mapping. Must re-earn apns2's production hardening through tests. |
Adopt sideshow/apns2 v0.25.0 |
Fastest path to a working client; de-facto library with verified internals, battle-tested at throughput, MIT, small, vendorable. | Maintenance drift: no release since Oct 2024, one untagged commit in twelve months, none in 2026; 25 open issues / 8 PRs with no maintainer response in 2025–2026, including an unaddressed connection-leak issue (#238). Lacks single-flight regen, scoped-key selection, expiration policy, connection-severed semantics — so it would be wrapped anyway. Bus-factor risk. |
Adopt a fork (e.g. ad-mos/apns2) |
In principle routes around the inactive upstream; adds Live Activity. | Less current than upstream (untagged pseudo-version from Jul 2024, no stable tags) — bus factor is worse. Live Activity is out of scope. |
Rationale. Three points decide it for a reliability-first, multi-year, self-hosted relay. First, maintenance health is poor: v0.25.0 (Oct 2024) is still the newest tag (source: https://github.com/sideshow/apns2/releases, 2026-06-13), the only commit to master in twelve months is one untagged dep-bump PR #239 dated 2025-07-22 (source: https://github.com/sideshow/apns2/commits/master, 2026-06-13), and an open issue (#238, 2024-11-11) reports exactly the failure class that matters here — an HTTP/2 connection/goroutine leak in a process holding long-lived connections for days — with no maintainer response in 2025 or 2026 (source: https://github.com/sideshow/apns2/issues, 2026-06-13). Second, relay requirements already exceed apns2: single-flight JWT regen (to avoid 429 TooManyProviderTokenUpdates), the (env, topic) → .p8 scoped-key map with startup validation for Apple's Feb 2025 scoped keys, a finite apns-expiration, connection-severed → 503 with no transparent retry, and sandbox-only apns-unique-id capture — apns2 has none of these, so it would be wrapped regardless, leaving only small JWT/HTTP-2 glue. Third, that glue is low risk to own: the 2026 reference verified every constant, jwt/v5 and x/net/http2 are maintained and already needed, and Go's net/http2 handles GOAWAY and reconnect transparently on a reused client. Compatibility is a non-issue — apns2 master pins x/net v0.33.0 / go 1.18 but module MVS resolves up to the relay's x/net v0.56.0 and jwt v5.3.1 and builds under Go 1.25/1.26 (source: https://github.com/sideshow/apns2/blob/master/go.mod, 2026-06-13) — so the disqualifier is maintenance and missing behavior, not versions.
Consequences (what this commits).
- Deps.
go.modtakesgolang-jwt/jwt/v5 v5.3.1andgolang.org/x/net v0.56.0as production deps (both already planned in plan §3).apns2is not a production dep — at most a Phase 3 prototype dropped before Phase 6. - Code (
internal/apns).jwt.go— ES256.p8sign; in-memory cache; ~50-minute regen timer; 20-minute floor; eager regen on403 ExpiredProviderTokenwith a single retry; single-flight regen keyed by team.transport.go— per-(env, team) HTTP/2 transport with non-zeroReadIdleTimeout/PingTimeout, hardcoded host constants by environment enum, GOAWAY reconnect, USERTrust RSA via the system trust store.client.go— theinternal/apnsinterface implementation.errors.go— the spec §6.4 reason table, includingInvalidToken→BadDeviceToken, unrecognized reason →502, connection-severed →503.- Key-selection map and startup validation live in
internal/config. - Everything else depends only on the
APNsClientinterface, so the choice is reversible.
- Ops. A CVE or HTTP/2 bug is patched by bumping
x/netrather than waiting on an inactive repo. Test against a fake APNs upstream for GOAWAY and mid-stream reset, a JWT-expiry burst that must yield exactly one regen, plus the Apple sandbox integration test (Phase 3). - Schema. None — upstream-client-only; no Postgres/Redis or external-contract change.
Revisit if. Reverse toward apns2 only if, before Phase 3 ships, sideshow/apns2 cuts a fresh tagged release newer than v0.25.0 and shows sustained maintainer responsiveness — clearing the backlog and addressing the #238 leak class — over two to three months, resolving the bus-factor concern. Conversely, re-validate the hand-rolled JWT/HTTP-2 internals if Apple changes APNs provider auth (a new required header, a non-ES256 algorithm, a host change, or deprecating token auth), since the hand-rolled client now owns those mechanics.
Sources.
apns2latest release/tag isv0.25.0(Oct 24–25 2024); no newer tag exists (source: https://github.com/sideshow/apns2/releases, 2026-06-13).- Only commit to master in twelve months is untagged dep-bump PR #239 (2025-07-22); no 2026 commits (source: https://github.com/sideshow/apns2/commits/master, 2026-06-13).
- 25 open issues / 8 PRs; issue #238 (potential HTTP/2 connection/goroutine leak, 2024-11-11) has no maintainer response in 2025–2026 (source: https://github.com/sideshow/apns2/issues, 2026-06-13).
- pkg.go.dev flags
v0.25.0as not the latest module version; newest released version is ~20 months old (source: https://pkg.go.dev/github.com/sideshow/apns2, 2026-06-13). apns2mastergo.moddeclares go 1.18, pinsx/net v0.33.0andjwt v5.2.1; MVS resolves to the relay'sx/net v0.56.0/jwt v5.3.1, builds under Go 1.25/1.26 (source: https://github.com/sideshow/apns2/blob/master/go.mod, 2026-06-13).- Fork
ad-mos/apns2(Live Activity) is less current: untagged pseudo-version 2024-07-30, no stable tags (source: https://pkg.go.dev/github.com/ad-mos/apns2, 2026-06-13). x/net v0.56.0(HTTP/2ReadIdleTimeout/PingTimeouttuning) published 2026-06-09, actively maintained (source: https://pkg.go.dev/golang.org/x/net, 2026-06-13).- Verified
apns2 v0.25.0internals to mirror:HostProduction api.push.apple.com,HostDevelopment api.sandbox.push.apple.com,TokenTimeout 3000under a mutex,ReadIdleTimeout 15s, one client per process per env (source: https://github.com/sideshow/apns2/blob/v0.25.0/client.go, 2026-06-13).
Decision 2 — FCM credentials: default keyless via ADC, SA-JSON-in-secret-manager as last resort
Decision. Should the relay authenticate to FCM HTTP v1 via GCP keyless identity (Workload Identity Federation / attached service account) or via a service-account JSON key stored in a secret manager? (Spec §4.4, §15 #4; reference §4.5.)
Resolution. Default to keyless. The relay MUST authenticate via Application Default Credentials (ADC) using golang.org/x/oauth2/google with scope https://www.googleapis.com/auth/firebase.messaging, with the concrete credential source chosen by deployment:
- On GCP compute (GCE/GKE/Cloud Run): use an attached service account — zero key material; the metadata server mints short-lived tokens.
- Off-GCP (self-hosted LXC/VM — the documented Silo home-hosting reality): use Workload Identity Federation with a non-secret credential-configuration file plus an OIDC token from the host's own IdP — still no exported key.
- Only if WIF cannot be stood up (no usable OIDC issuer for the host): fall back to a service-account JSON key, which MUST live in a secret manager loaded once at startup into memory — never as a literal-key env var, never baked into the image, never on disk in the repo.
In all three cases the code path is identical: google.FindDefaultCredentials / DefaultTokenSource(ctx, "…/firebase.messaging") returns an auto-caching, auto-refreshing TokenSource; the relay does not branch on credential type. This makes keyless the default and JSON-in-secret-manager the explicit last resort, matching spec §4.4 / reference §4.5 with one refinement: even the off-GCP case is keyless via WIF first.
Alternatives considered.
| Option | Pros | Cons |
|---|---|---|
| Keyless via ADC — attached SA on GCP, WIF off-GCP; SA-JSON-in-secret-manager last resort (chosen) | No long-lived exportable secret in the common path — eliminates the highest-severity leak vector for a service whose whole reason to exist is custody of the official Google push credential (spec G3). Aligns with Google's current guidance that SA keys are "an exception rather than the norm" and that orgs created on/after 2024-05-03 block key creation by default. WIF's credential-config file carries no private key and need not be confidential. Tokens are short-lived (~1h) and auto-rotated — no manual rotation, no rotation-induced outage mode. Single code path through x/oauth2/google regardless of source. Org-policy iam.disableServiceAccountKeyCreation can be enforced relay-wide with no impact. |
WIF off-GCP requires the host to expose a usable OIDC token (or a supported IdP: AWS/Azure/GitHub/GitLab/K8s/AD/OIDC/SAML/X.509); a bare LXC with no IdP must stand one up or fall back to JSON. Slightly more one-time setup (pool + provider + attribute mapping). |
| Service-account JSON key in a secret manager (loaded at startup, never env/in-image) | Works on any host with no IdP or GCP compute — the universal fallback for a self-hosted ecosystem. Simple mental model. Same x/oauth2/google minting code via CredentialsFromJSONWithType. Secret-manager custody (vs env/in-image) is itself OWASP-aligned hardening, acceptable when keyless is genuinely unavailable. |
A long-lived, exportable private key is exactly what Google now treats as an exception and blocks by default for 2024-05-03+ orgs — if the official Silo Firebase project lives in such an org, the policy must be relaxed just to create it. Single high-value theft target (full FCM send authority), manual rotation burden, rotation-induced outage window. The untyped CredentialsFromJSON loader is deprecated; must use the typed variant. Fallback, not default. |
| Use the Firebase Admin SDK and let it resolve ADC | Batteries-included: auto token mint/cache/refresh, rich error helpers, also resolves ADC (inherits keyless transparently). | Orthogonal to the credential decision and explicitly not the relay's chosen client — plan §3 / reference §3.2 deliberately pick the lighter x/oauth2/google path for a fixed payload and direct control of connection reuse and error classification. Adopting it to "solve credentials" re-opens a settled client-library decision for no credential benefit (ADC resolution is identical either way). |
Rationale. The relay's single most sensitive asset is the official Google push credential (spec G3, §4.4); a fully compromised relay must not hand an attacker a portable, long-lived FCM send key. Keyless removes that artifact in the common path: on GCP the key never exists; off-GCP, WIF replaces it with a non-secret credential-config file plus a short-lived host-issued OIDC token that Google STS exchanges for a ~1h access token. Google's own current guidance is unambiguous — SA keys are "an exception rather than the norm," and organizations created on or after 2024-05-03 enforce iam.disableServiceAccountKeyCreation / disableServiceAccountKeyUpload by default (source: https://docs.cloud.google.com/iam/docs/best-practices-for-managing-service-account-keys, 2026-06-13) — so a brand-new project should treat keyless as the baseline. Reliability-first reinforces this: long-lived keys add a manual-rotation outage mode and a "key silently revoked/expired" failure class that keyless does not have. Crucially, the implementation cost of defaulting to keyless is near-zero, because x/oauth2/google ADC already resolves all three sources — GCE/GKE metadata, external_account WIF config files, and SA JSON — behind one TokenSource with the same scope, so the relay does not branch (source: https://pkg.go.dev/golang.org/x/oauth2/google, 2026-02-11). JSON-in-secret-manager remains an honest fallback for a host with no IdP, but is the exception: secret-manager-held and loaded once into memory, never an env-var key or image layer (spec §4.4 ranking secret-manager > mounted file > env var > in-image). This is a refinement, not a contradiction, of spec §4.4 ("On GCP, prefer keyless …; off-GCP, the JSON lives in the secret manager") — the verified addition is that off-GCP can and should try WIF (keyless) first, because WIF explicitly targets non-Google-Cloud workloads and its config file carries no secret.
Consequences (what this commits).
- Code.
internal/fcm/oauth.gobuilds theTokenSourceexclusively through ADC —google.DefaultTokenSource(ctx, "https://www.googleapis.com/auth/firebase.messaging")(orFindDefaultCredentials+creds.TokenSource) — and does not special-case credential type. The fallback SA-JSON path usesCredentialsFromJSONWithType(ctx, json, google.ServiceAccount, scope)(not the deprecatedCredentialsFromJSON) with the JSON sourced from the secret manager. - Config.
internal/configgains a singlefcm_credential_moderesolution: attached-SA (no config) > WIF credential-config file path > secret-manager SA JSON, mapped onto ADC (setGOOGLE_APPLICATION_CREDENTIALSto a mounted, non-secret WIF config, or inject the SA JSON into the secret-manager-backed loader). - Ops/deploy. On GCP: bind an attached SA with rights to send to the official Firebase project, and apply org policy
iam.disableServiceAccountKeyCreation. Off-GCP: stand up a workload-identity pool + provider, configure the host's OIDC issuer, deploy the WIFexternal_accountcredential-configuration file (it may be mounted/baked because this specific file contains no private key and need not be kept confidential, per Google's WIF docs — it is the only artifact exempt, and only because it carries no secret). This refines, not overrides, spec §11: the no-in-image rule still applies in full to every actual secret — the APNs.p8, the FCM service-account JSON, the HMAC pepper, and DB/Redis credentials remain strictly forbidden in the image under spec §11. The startup credential health-check (relayctl ping-upstream --provider fcm, spec §5.4, plan Phase 4/6) validates the resolvedTokenSourcemints a token regardless of source. Removes the manual key-rotation runbook in the common path; documents the WIF pool/provider setup and the JSON-fallback secret-manager procedure for the exception case. - Schema/contract. No DB schema change; no change to the external
02/03contract or topushwirepayloads.
Revisit if. (1) The relay must deploy on a host with neither GCP attached-SA nor any WIF-compatible identity source — then the SA-JSON-in-secret-manager fallback becomes the de facto default for that deployment and its rotation runbook must be activated. (2) The official Silo Firebase project is created under / moved into a Google org that enforces iam.disableServiceAccountKeyCreation by default (org created on/after 2024-05-03), making the JSON fallback unavailable without a policy exception — then keyless is mandatory. (3) x/oauth2/google changes its ADC resolution order or deprecates the typed-credentials API, requiring a client-library re-verification.
Sources.
- Google: "consider the use of service account keys as an exception rather than the norm"; recommended constraints Disable SA key creation/upload; "If your organization was created on or after May 3, 2024, these constraints are enforced by default." (source: https://docs.cloud.google.com/iam/docs/best-practices-for-managing-service-account-keys, 2026-06-13).
- WIF lets external workloads (AWS/Azure/GitHub/GitLab/K8s/on-prem AD/OIDC/SAML 2.0/X.509) access GCP without SA keys via OAuth2/STS exchange; "eliminates the maintenance and security burden associated with service account keys." (source: https://docs.cloud.google.com/iam/docs/workload-identity-federation, 2026-06-13).
- For off-GCP workloads, WIF uses a credential-config file that "doesn't contain a private key and doesn't need to be kept confidential"; ADC auto-discovers it via
GOOGLE_APPLICATION_CREDENTIALS; supported in Go (source: https://docs.cloud.google.com/iam/docs/workload-identity-federation-with-other-providers, 2026-06-11). x/oauth2/google v0.36.0FindDefaultCredentials/DefaultTokenSourceresolves, in order:GOOGLE_APPLICATION_CREDENTIALSJSON (incl. WIF/external_account), gcloud user creds, then GCE/GKE metadata (attached SA);CredentialsTypeincludesServiceAccountandExternalAccount;CredentialsFromJSONis deprecated — useCredentialsFromJSONWithType; sameTokenSourceregardless of source (source: https://pkg.go.dev/golang.org/x/oauth2/google, 2026-02-11).- FCM HTTP v1 auth docs recommend ADC on Compute Engine/GKE/App Engine/Functions, warn that referencing the SA JSON file "should be done with extreme care," recommend
GOOGLE_APPLICATION_CREDENTIALSas "more secure and … strongly recommended," cite Workload Identity as an alternative to downloaded keys, and specify scopehttps://www.googleapis.com/auth/firebase.messaging(source: https://firebase.google.com/docs/cloud-messaging/auth-server, 2026-06-13).
Decision 3 — apns-expiration policy: relay-set finite TTL for WAKE pushes
Decision. Should the relay set a finite apns-expiration (and parallel FCM android.ttl) on private_alert WAKE pushes so a stale wake doesn't deliver days later, or omit it and accept Apple's ~30-day store-and-retry? What concrete default, and does the policy live in the relay or stay caller-controlled per the 02/03 contract? (Spec §5.1, §15 #5; resolves spec open question 5.)
Resolution. Set a FINITE apns-expiration by default in the relay. Ship apns-expiration = now + 4 hours for the APNs private_alert mode and now + 1 hour for background_wake, with a parallel FCM android.ttl keyed by the FCM wire mode — "14400s" for private_data (the FCM counterpart of the APNs private_alert) / "3600s" for background_wake — for cross-provider parity. Both values are per-deployment config (relay-owned defaults, not caller-supplied) and do not change the external 02/03 wire contract, which has no expiration field — the relay sets the upstream header itself. Do not use apns-expiration=0 (single-attempt, no store) and do not omit it (Apple stores/retries up to ~30 days). This confirms and operationalizes spec open question 5; the firm refinement is to keep it relay-owned with config-tunable defaults rather than promoting it to a caller field.
Alternatives considered.
| Option | Pros | Cons |
|---|---|---|
Finite relay-set default (~4h private_alert / ~1h background_wake), config-tunable, parallel FCM ttl (chosen) |
A WAKE is only useful while it can still produce a near-real-time fetch; the inbox (../01-release-events-and-inbox.md) is the source of truth, so a wake missed for hours is re-driven by the caller outbox or the next foreground sync, not by a stale push. Avoids a device buzzing for an event already read/deleted server-side days ago. Bounds APNs/FCM store-and-retry use. Keeps the content-free posture intact (no new caller field, no contract change). A multi-day-offline device just gets a clean expiry. With collapse-id, a finite TTL compounds correctly: a newer wake supersedes the old one and the old one self-expires. Config-tunable so product can re-tune without a code/contract change. |
Bakes a partly product/UX call into infra: a device back online at hour 5 gets no wake and relies on its next foreground sync (acceptable — inbox is authoritative — but a real behavioral choice). One hardcoded default per mode cannot suit future urgency tiers. APNs (absolute epoch) vs FCM (duration) asymmetry means the relay must compute the APNs epoch from now()+TTL per request. |
Omit apns-expiration / omit android.ttl (Apple ~30-day store-and-retry; FCM 4-week default) |
Maximizes raw delivery probability; simplest code; matches provider default (no per-request epoch math). | Verified behavior: omitted apns-expiration → store/retry up to ~30 days; omitted FCM ttl → 4 weeks. For a content-free WAKE that exists only to trigger a live fetch, a wake fired days later is useless at best, confusing at worst. Spends store-and-retry budget on valueless pushes. Contradicts the "inbox is the source of truth" premise. Spec explicitly rejects the 30-day default. |
apns-expiration = 0 / android.ttl = 0s (deliver-once-now-or-discard) |
Strongest guarantee against stale wakes; trivial constant. | 0 means no storage at all: a device merely briefly offline (locked, poor signal, Doze) loses the wake entirely, defeating the point. Background pushes are already APNs-throttled (~2–3/hr) and may be delayed; with TTL=0 too many legitimate wakes drop. FCM ttl=0 also disables collapsible-message throttling — a behavior change with no upside. |
Promote expiration to a caller-controlled 02/03 field |
Maximum flexibility: the self-hosted server, which alone knows true urgency, picks the TTL per push; future-proofs urgency tiers without a relay redeploy. | Violates the locked 02/03 contract (additive-only, not a listed v1 capability) and widens the content-free surface with a new caller-controlled numeric channel visible to Apple/Google — exactly what §5.1 bounds (cf. the badge field). Adds validation/clamping burden. Premature: v1 has exactly one wake semantic. Can be added later additively if a real urgency-tier requirement emerges. |
Rationale. The decision hinges on what a private_alert IS in Silo: a content-free WAKE whose only job is to make the device fetch fresh notifications from the user server, where the inbox, not the push, is the source of truth (spec §5.1, ../01-release-events-and-inbox.md). That framing makes a wake's value decay with time — once hours old, the device's own next foreground sync (or a fresher wake in the same collapse series) has already reconciled state, so a late wake yields no benefit and a small confusion cost. Both extremes are verified-wrong for this use case: omitting the header invokes Apple's ~30-day store-and-retry (verbatim: stored "for 30 days or less") and FCM's 4-week default — far too long; expiration=0 / ttl=0s means deliver-once-or-discard with no storage, dropping wakes for a device merely briefly offline (especially bad given APNs already throttles background pushes). A finite few-hours TTL is the correct middle: long enough to ride out normal transient offline windows (Doze, locked phone, brief signal loss), short enough that a multi-day-offline device gets a clean expiry and relies on its authoritative next sync. background_wake gets the shorter ~1h TTL because it is even more time-sensitive and even more throttled (APNs requires apns-priority 5 for background pushes; priority 10 is an error). The collapse-id interaction reinforces this: collapse already means latest-wins on-device, and a finite expiration is the complementary mechanism for the cross-series / no-newer-wake case, so the two together ensure stale wakes neither pile up nor fire late. Keeping it relay-owned and config-tunable (not a caller field) preserves the locked 02/03 contract and the structural content-free guarantee while letting product re-tune durations — which is how the spec already resolved open question 5.
Consequences (what this commits).
- Code.
internal/pushwire/apns.gocomputesapns-expiration = now()+TTLas a UNIX epoch (UTC, seconds) per request from config, per APNs mode (private_alertvsbackground_wake).internal/pushwire/fcm.gosetsandroid.ttlas the parallel duration string keyed by the FCM wire mode —"14400s"forprivate_data(the FCM counterpart of the APNsprivate_alert; spec §5.2 lists FCM modes asprivate_data|background_wake) /"3600s"forbackground_wake. Note the mode-name asymmetry: an implementer buildinginternal/pushwire/fcm.gomust key onprivate_data, notprivate_alert, on the FCM path. Because the value isnow()-relative, the golden-file payload/header tests in Phase 3/4 must inject a fixed clock or assert the header is present and within tolerance rather than byte-exact; the mirroredsilo-serverpushwirecopy and its checksum/golden tests must use the same clock-injection approach so the two repos do not drift. - Config. Two new non-secret tuning values (
apns_expiration_private_alert,apns_expiration_background_wake) ininternal/config, with the FCMttlpaired/derived, defaulted in code, env/flag-overridable, loaded and validated at startup (reject values exceeding the APNs storage max or FCM 28-day max; reject<=0if omit/0 is disallowed by policy). - Contract. No change to the
02/03wire contract and no new caller field — preserving additive-only API rules and the content-free guarantee. - Ops. A stale-wake or missed-wake-rate concern becomes a config tune, not a logic redeploy. No schema change, no new dependency.
- Tests. Assert
background_wakenever emitsapns-priority 10alongside its shorter TTL, and that an out-of-range configured TTL fails startup validation.
Revisit if. (1) Silo introduces more than one wake semantic or an urgency/priority tier (today there is exactly one: wake = notifications.changed), at which point per-mode or per-request TTL — possibly an additive 02/03 caller field with a bounded max — becomes justified. (2) Production metrics show a material missed-wake rate attributable to the ~4h/~1h window being too short for the real offline distribution of Silo devices. (3) Apple or Google changes apns-expiration / android.ttl semantics or the storage-policy maximum (re-verify the cited pages near deployment). (4) The durable-inbox + foreground-sync assumption weakens (e.g., a future notification type whose value is NOT recoverable from the inbox), which would change whether a missed wake is truly harmless.
Sources.
- APNs verbatim: "If the value is 0, APNs attempts to deliver the notification only once and doesn't store it." and "If the value is nonzero, APNs stores the notification and tries to deliver it at least once, repeating the attempt as needed until the specified date." (source: https://developer.apple.com/documentation/usernotifications/sending-notification-requests-to-apns, 2026-06-13).
- Omitted
apns-expiration→ Apple stores per its storage policy: "it may store the notification for 30 days or less … APNs attempts to deliver … the next time the device activates and is available online." (source: https://developer.apple.com/documentation/usernotifications/sending-notification-requests-to-apns, 2026-06-13). apns-expirationcarries UNIX epoch seconds (UTC); an invalid value yields400 BadExpirationDate;apns-priorityfor background pushes MUST be 5 (priority 10 is an error) (source: https://developer.apple.com/documentation/usernotifications/sending-notification-requests-to-apns, 2026-06-13).apns-collapse-idmerges notifications on-device (latest of the same id wins), ≤ 64 bytes, independent of expiration — so finite expiration + collapse means a newer wake supersedes an older queued one AND an undelivered wake self-expires (source: https://developer.apple.com/documentation/usernotifications/sending-notification-requests-to-apns, 2026-06-13).- FCM
android.ttl: default when omitted is four weeks; max 2,419,200 s (28 days);ttl=0discards messages that can't be delivered immediately and disables collapsible-message throttling (source: https://firebase.google.com/docs/cloud-messaging/customize-messages/setting-message-lifespan, 2026-06-13). - FCM collapse: with a collapse key, an offline-stored message is replaced by the newer one; without one, both are stored; at most 4 distinct collapse keys per device while offline (source: https://firebase.google.com/docs/cloud-messaging/customize-messages/setting-message-lifespan, 2026-06-13).
- Spec already resolved this as open question 5: v1 ships finite, configurable defaults (~4h
private_alert, ~1hbackground_wake; parallel FCMandroid.ttl), relay-owned and tunable; the external02/03contract is unchanged (source: ./00-relay-spec.md, 2026-06-13).
Decision 4 — Per-account rate limits: hard-coded global defaults in v1
Decision. Ship hard-coded global rate-limit defaults (~10 req/s burst, 50k/day), or add per-account quota columns (rate_per_sec, burst, daily_quota) on relay_accounts settable via relayctl, seeded with the documented defaults? (Spec §9.2, §15 #1.)
Resolution. Ship hard-coded global defaults for v1: ~10 req/s burst (capacity-20 token bucket), 50,000 req/day (fixed-UTC), coarse per-(account, token) ~1 req/3s — enforced as named config constants with env-var overrides, NOT per-account DB columns. Keep all per-account enforcement plumbing exactly as specified (Redis bucket keyed rl:{account_id}, daily counter rl:day:{account_id}:{yyyymmdd}, the §9.5 shared-FCM-project circuit breaker). Add per-account override columns only when production logs prove at least one account legitimately needs a different cap. This matches the relay's own stated lean in spec §9.2 ("v1 ships a single global default and tunes from logs") and open Q#1, so it ratifies the plan rather than reopening it.
Make the later per-account path cheap by: (1) reading the three limit values from a single typed struct accountLimits{RatePerSec, Burst, DailyQuota} resolved per request, defaulting to the global constants, so introducing a per-account source later is a one-function change, not a call-site sweep; (2) keeping the Redis bucket/daily keys labeled by account_id already (done). When the trigger fires, add columns rate_per_sec int, burst int, daily_quota int (all NULLABLE, NULL = inherit global) to relay_accounts in one timestamped Goose migration, plus a relayctl quota set --account <id> [--rate-per-sec N] [--burst N] [--daily-quota N] / quota show / quota reset subcommand — a backward-compatible, additive change.
Alternatives considered.
| Option | Pros | Cons |
|---|---|---|
| Hard-coded global defaults (config constants + env overrides), no per-account columns in v1 (chosen) | Matches the relay's own documented lean (spec §9.2, open Q#1). Zero new schema/CLI surface — less code, fewer tests, smaller attack surface for a reliability-first first release. The abuse-containment requirement (OWASP API4:2023) is already met: the global default is enforced per-account (keyed by account_id in Redis) plus the §9.5 shared-project circuit breaker, so a stolen/misbehaving key is capped whether the number is a constant or a column. Operators still adjust globally via env var. Adding columns later is genuinely cheap (PG metadata-only ADD COLUMN). |
A single hot account legitimately exceeding 50k/day must be handled by raising the GLOBAL default (env var, affects all) or an emergency change — until override columns exist. Slightly less satisfying day-one operability. Mitigated by the §9.2 alert at ≥80% of the daily cap. |
Per-account columns on relay_accounts, settable via relayctl, seeded with defaults — in v1 |
Per-account tuning from day one; cleanest long-term shape; NULL-default inheritance is clean; avoids any future migration. | Builds an override surface against zero production evidence (the client already paces at 5 req/s, half the 10 req/s default, so real traffic won't press the limit). Adds schema + relayctl quota subcommands + tests + an inheritance/precedence rule and its edge cases to a reliability-first v1. More surface for a misconfiguration (daily_quota=0, absurd burst) to cause an outage or remove the abuse cap. The migration it "saves" is nearly free anyway. |
Per-account columns now but NOT exposed via relayctl (DB-only UPDATE) |
Future-proof schema without building/testing the CLI; emergency override via privileged DSN. | Worst of both: pays the schema + inheritance cost yet has no safe, audited workflow — raw UPDATEs bypass the relay_op_logs admin.* audit trail that relayctl provides (spec §5.6), contradicting the "all admin ops audited" posture. No advantage over adding columns + CLI together when the trigger fires. |
Rationale. Four things make global defaults the right v1 call. (1) The migration-cost argument for "add columns now to avoid pain later" does not hold: PostgreSQL ALTER TABLE ADD COLUMN with a constant default is a metadata-only operation (no row rewrite) since PG 11, and relay_accounts holds one row per opt-in installation — hundreds at most — so even a full rewrite is trivial; NULLABLE columns (NULL = inherit) avoid even needing a default, and adding columns + a relayctl quota subcommand later is backward-compatible (source: https://www.postgresql.org/docs/current/ddl-alter.html, 2026-06-13). (2) The security requirement (OWASP API4:2023, Unrestricted Resource Consumption) is satisfied by a per-account cap, which the global default already provides — every bucket is keyed by account_id, and the §9.5 shared-FCM-project circuit breaker bounds cross-tenant blast radius; OWASP recommends per-client throttling and business-tuned limits but does not require per-account configurability (source: https://owasp.org/API-Security/editions/2023/en/0xa4-unrestricted-resource-consumption/, 2026-06-13). (3) Real traffic won't press the default: the Silo server paces client-side at 5 req/s (02/03), half the 10 req/s relay backstop, so the relay limit exists to catch a misbehaving/compromised caller, not to be tuned per well-behaved tenant (source: ./00-relay-spec.md, 2026-06-13). (4) This is a reliability-first first release; the §15/§9 text already leans "single global default, tune from logs," so shipping that ratifies the plan and minimizes v1 surface. The correct sequencing is: ship one default, watch the ≥80%-of-daily-cap alert and per-account 429 metrics, and add targeted overrides only when a real account demands one — at which point the change is cheap and evidence-driven. The shared-project risk is unavoidable regardless of limit shape: FCM's default per-project quota is 600,000 messages/minute enforced per project (shared across all tenants using it), returning 429 RESOURCE_EXHAUSTED over quota — which is exactly why the §9.5 circuit breaker is required no matter whether the per-account number is a constant or a column (source: https://firebase.google.com/docs/cloud-messaging/throttling-and-quotas, 2026-06-12).
Consequences (what this commits).
- Schema. NO change to the spec §8.1 v1 schema —
relay_accountsships withoutrate_per_sec/burst/daily_quotacolumns. - Code. Limit values live as named config constants in
internal/config(per plan §2.4, "non-secret defaults … live as constants with env overrides") with env-var overrides; the rate-limit middleware (internal/ratelimit) reads them through a singleaccountLimitsstruct resolved per request, defaulting to the globals, so a future per-account source is a one-function swap, not a call-site sweep. - Redis. Keys remain account-scoped (
rl:{account_id},rl:day:{account_id}:{yyyymmdd}) — already specified, so enforcement granularity is unchanged. - CLI. NO new
relayctl quotasubcommand in v1; the §5.6 command table is unchanged. - Ops. Operators tune the SINGLE global default via env var + rolling restart; the §9.2 ≥80%-of-daily-cap alert and the
relay_rate_limited_total{kind=account|daily}/relay_daily_quota_used_ratiometrics are the feedback loop. The §9.5 shared-FCM-project circuit breaker is unaffected and still required. - Commits the team to a later additive migration +
relayctl quota set/show/resetCLI only if/when the revisit trigger fires (columns NULL = inherit global).
Revisit if. Add the per-account columns + relayctl quota set/show/reset CLI as soon as ANY of: (a) production metrics show one or more legitimate accounts repeatedly hitting the 50k/day cap or 10 req/s burst (relay_rate_limited_total{kind=daily|account} or relay_daily_quota_used_ratio >= 0.8 sustained for a real, non-abusive account) such that raising the GLOBAL default would over-provision everyone else and erode the abuse cap; (b) operators need to throttle or temporarily zero a SPECIFIC misbehaving account without affecting others (today only all-or-nothing revoke/disable exists); (c) the relay onboards tenant tiers with deliberately different quotas — at which point per-account quotas become a product requirement, not a tuning convenience.
Sources.
- PostgreSQL
ALTER TABLE ADD COLUMNwith a CONSTANT default is a fast metadata-only operation (no table rewrite; default materialized lazily); only a VOLATILE default forces a per-row update (source: https://www.postgresql.org/docs/current/ddl-alter.html, 2026-06-13). - OWASP API4:2023 recommends rate limiting per client/operation and business-tuned limits; it does NOT prescribe per-client configurable quota columns or a specific multi-tenant quota architecture — a per-tenant cap satisfies the control (source: https://owasp.org/API-Security/editions/2023/en/0xa4-unrestricted-resource-consumption/, 2026-06-13).
- FCM default per-project downstream quota is 600,000 messages/minute, enforced PER PROJECT (shared across tenants) via a per-minute token bucket; over quota returns
429 RESOURCE_EXHAUSTED— hence the shared-project circuit breaker is required regardless of per-account limit shape (source: https://firebase.google.com/docs/cloud-messaging/throttling-and-quotas, 2026-06-12). - Spec §9.2 states the v1 intent: "These are tunable per account (future column on
relay_accounts); v1 ships a single global default and tunes from logs," and lists per-account columns as open question #1 (source: ./00-relay-spec.md, 2026-06-13). - The Silo server already paces relay calls client-side at a default 5 req/s shared token bucket (
02/03"Relay pacing"), half the 10 req/s relay backstop (source: ./00-relay-spec.md, 2026-06-13).
Decision 5 — API-key hashing: HMAC-SHA256(pepper, secret)
Decision. For high-entropy random relay API keys (rk_<env>_<random>, ≥128 bits), what is the correct at-rest hashing in 2026: HMAC-SHA256(pepper) vs salted SHA-256 vs argon2id? (Spec §8.2, §15 #6; plan Phase 1; reference §4.1.)
Resolution. Store API keys as HMAC-SHA256(pepper, secret) — a fast keyed hash with a server-side pepper held in the secret manager, separate from PostgreSQL. The spec/plan already lean this way (spec §8.2, plan Phase 1, reference §4.1); keep it. Concretely: generate the secret as ≥32 bytes from crypto/rand, format rk_<env>_<base62/base64url>; store a non-secret, uniquely-indexed prefix (rk_<env>_<first 6–8 chars>) cleartext for O(1) lookup; store key_hash = HMAC-SHA256(pepper, presented_secret) (32 bytes, bytea); verify with crypto/subtle.ConstantTimeCompare against the stored hash (and against a fixed decoy hash on prefix-miss to keep the path constant-time). Do not use argon2id/bcrypt here — they are for low-entropy human passwords and only add per-request latency and (for bcrypt) a 72-byte input-truncation hazard with no security gain against a 256-bit keyspace. No per-key random salt is required because the full-entropy secret already makes the hash input unique and rainbow tables infeasible; the pepper is the deliberate, shared defense-in-depth lever instead.
Rotation note (load-bearing): a pepper change cannot be done in place — it requires re-hashing every live key (impossible without the plaintext, which is never stored) or, the workable path, a dual-pepper verification window (try the new pepper, fall back to the old) plus a pepper_version column, then re-issuing keys under the new pepper and retiring the old pepper once all live keys are re-issued. Because key secrets are unrecoverable, "re-hash all keys" is really "re-issue all keys," so v1 must (a) tag each key_hash row with a pepper_version, (b) support reading under N peppers during overlap, and (c) treat pepper rotation as a make-before-break, multi-key-reissue operation aligned with the existing credential-rotation runbook (spec §14.3). Pepper rotation should be rare (compromise-driven), not scheduled.
Alternatives considered.
| Option | Pros | Cons |
|---|---|---|
HMAC-SHA256(pepper, secret) — fast keyed hash with secret-manager pepper (chosen) |
Fast: sub-microsecond verify, no per-request CPU/RAM tax on the hot auth path (the only synchronous DB touch is the prefix lookup). Correct for high-entropy keys: a 256-bit random secret is ~2^256 keyspace, infeasible to brute-force regardless of hash speed. Defense-in-depth: if PostgreSQL leaks but the pepper does not, the stolen key_hash values are inert to an offline attacker. Matches OWASP's pepper guidance (shared secret, stored separately, applied via HMAC-SHA256). Plays naturally with prefix-indexed O(1) lookup + constant-time compare. Already the spec/plan choice — zero rework. |
Pepper rotation is heavier: cannot re-hash in place, so it needs a dual-pepper overlap window + pepper_version column + key re-issuance. The pepper is a single shared secret whose compromise affects all keys at once (mitigated: secret-manager-only, never in DB/image). Slightly more moving parts than plain SHA-256. |
| Per-key salted SHA-256 (fast hash, random salt per key, no pepper) | Also fast and correct for high-entropy keys; trivially simple; no pepper to provision/rotate; salt defeats cross-key precompute (near-moot at 256-bit). | No defense-in-depth on a DB-only leak: with salt+hash an attacker can verify guesses offline — but at 256-bit entropy that offline attack is itself infeasible, so the practical delta is small. The real loss is the cheap, high-value "DB leak but pepper intact ⇒ stolen hashes inert" property, which is exactly this design's headline threat (hostile/compromised relay operator with full DB, spec §3.5). For a credential-custody-focused service holding official Apple/Google push credentials, paying the modest rotation cost to keep that property is the right trade. |
| argon2id (or bcrypt) — slow memory-hard KDF | Industry-default "correct" answer for password storage; unimpeachable if these were human-chosen secrets; brute-force resistance for low-entropy inputs. | Wrong tool. OWASP's slow-KDF requirement is scoped to passwords ("Fast hashing algorithms such as SHA-256 are not suitable for password storage…") — the threat is low password entropy, which does not exist for a 256-bit crypto/rand key. Adds real per-request cost (argon2id at OWASP's m=19 MiB, t=2 is tens of ms and tens of MB per verification — a self-inflicted DoS/memory-amplification vector on an auth endpoint). bcrypt additionally truncates input at 72 bytes (needs an HMAC/SHA pre-hash anyway). No security benefit: cannot make an already-infeasible 2^256 search more infeasible. Spec and reference explicitly reject it. |
Rationale. The decision turns on one fact: the secret is high-entropy and machine-generated, not a human password. Slow KDFs (argon2id/bcrypt) exist to compensate for low password entropy by making each guess expensive — their entire value is throttling a feasible offline guessing attack. Against a ≥128-bit (here 256-bit) crypto/rand secret the guessing attack is not feasible to begin with (~2^256 ≈ 10^77 candidates), so the slow hash protects against nothing while imposing real cost; the GPU gap that justifies slow KDFs for an 8-char password (≈180 billion SHA-256 h/s vs ≈1000 argon2 h/s) is irrelevant when the keyspace itself is unreachable (source: https://www.thesslstore.com/blog/what-is-256-bit-encryption/, 2026-06-13). OWASP confines the fast-hash prohibition to "password storage" (source: https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html, 2026-06-13), and OWASP API2:2023 treats API keys as a distinct category requiring anti-brute-force at the endpoint — which the relay supplies via per-IP throttle + per-account rate limits, not a slow hash (source: https://owasp.org/API-Security/editions/2023/en/0xa2-broken-authentication/, 2026-06-13). So a fast hash is the correct family. Within fast hashes, the pepper-vs-salt choice is a defense-in-depth-vs-rotation-cost trade, and the pepper buys a property that maps directly onto this service's threat model: the headline adversary is a compromised relay operator / DB leak (spec §3.5), and the relay holds the official Apple .p8 and Google SA credentials for the whole ecosystem. With the pepper in the secret manager (never in PG, never in the image), a Postgres-only compromise yields key_hash values that are inert without the pepper key — an attacker cannot even mount the (already-infeasible) offline search, and cannot forge or recognize valid keys. OWASP defines a pepper as a secret "shared between stored passwords," stored separately, applied as an HMAC where "the pepper is acting as the HMAC key" — exactly this design (source: https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html, 2026-06-13). The rotation cost is real but bounded and rare (compromise-driven), and reuses the existing make-before-break credential-rotation runbook (spec §14.3). Salted SHA-256 is a defensible simpler fallback if the team judges pepper rotation not worth it, but it forfeits the DB-leak-resistance property that is this design's reason to care; argon2id is the only genuinely wrong choice and both spec and reference already reject it.
Consequences (what this commits).
- Schema/DB — amends spec §8.1 / plan Phase 1.
relay_api_keysstoreskey_hash bytea = HMAC-SHA256(pepper, secret)(32 bytes), a cleartext uniquely-indexedkey_prefix, andaccount_id/created_at/last_used_at/expires_at/revoked_at(spec §8.1). No salt column. This decision adds one column the reproduced §8.1relay_api_keystable does not have:pepper_version smallint NOT NULL DEFAULT 1(akey_idcolumn is an acceptable equivalent), included in the Phase 1 init migration from the start so dual-pepper rotation is possible later without a migration under pressure. This is an additive, backward-compatible deviation from the §8.1 schema reproduced in the spec and the plan's "exact schema from spec §8.1" Phase 1 statement — an implementer must add the column up front rather than discovering the gap during pepper rotation. - Code.
internal/accounts/keys.go:crypto/randsecret generation; HMAC-SHA256 hashing.internal/httpapi/middleware_auth.go:crypto/subtle.ConstantTimeComparefor verification; a fixed decoy-hash compare on prefix-miss to keep the auth path constant-time (no enumeration oracle); verification tries the active pepper and, during a rotation window, falls back to prior pepper(s) keyed bypepper_version. - Ops / secret management. The pepper is a first-class secret in the secret manager (KMS/Vault/cloud), never in PG, never in the image/env where avoidable, loaded once at startup (spec §11, reference §4.5); it must be backed up / escrowed because losing it invalidates every live key.
- Rotation runbook. Pepper rotation = make-before-break (provision new pepper alongside old, verify under both, re-issue all keys under the new pepper, retire old) — reuse the §14.3 credential-rotation pattern; document that "re-hash all keys" is really "re-issue all keys."
- Performance. Auth stays a sub-microsecond hash + one indexed lookup, preserving the stateless-fast hot path (spec §4.1). Commits to NOT pulling argon2/bcrypt onto the request path.
- Tests. Unit tests asserting the hash is
HMAC-SHA256(pepper, ·), constant-time compare, plaintext never stored, and unknown-prefix vs wrong-secret are timing-indistinguishable (plan Phase 1/2 acceptance criteria).
Revisit if. (1) The relay ever accepts a LOW-entropy or human-chosen secret (an operator passphrase, basic-auth admin password, shortened/customer-supplied key) — then a slow KDF (argon2id at OWASP's m=19 MiB, t=2, p=1 minimum) becomes mandatory for that credential class. (2) The pepper would have to live in the same trust boundary as PostgreSQL (same host/backup, in the DB) — its defense-in-depth value evaporates and plain salted SHA-256 becomes the honest, simpler choice. (3) A hard, citable, API-key-specific standard later prescribes a different at-rest scheme. (4) Pepper rotation proves operationally painful and pepper-compromise risk is judged low — drop to per-key salted SHA-256. (5) Key entropy is ever reduced below ~128 bits.
Sources.
- OWASP: "Fast hashing algorithms such as SHA-256 are not suitable for password storage because they allow attackers to perform large numbers of guesses quickly." (scoped to passwords) (source: https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html, 2026-06-13).
- OWASP argon2id recommended configs (e.g. m=19456 KiB, t=2, p=1) are a per-verification CPU/RAM tax that would hit every relay auth if wrongly applied to API keys (source: https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html, 2026-06-13).
- OWASP defines a pepper as a secret "shared between stored passwords," "stored separately from the password database," applied via HMAC where "the pepper is acting as the HMAC key" (source: https://cheatsheetseries.owasp.org/cheatsheets/Password_Storage_Cheat_Sheet.html, 2026-06-13).
- OWASP API2:2023 flags weakly-hashed credentials, requires anti-brute-force mechanisms on auth endpoints, and states "API keys should not be used for user authentication … only for API clients authentication." (source: https://owasp.org/API-Security/editions/2023/en/0xa2-broken-authentication/, 2026-06-13).
- A 256-bit random key has ~2^256 (~10^77) values — computationally infeasible to brute-force regardless of hash speed; security derives from key entropy, not hash slowness; the ~180-billion-vs-~1000 h/s gap is irrelevant when the keyspace is unreachable (source: https://www.thesslstore.com/blog/what-is-256-bit-encryption/, 2026-06-13).
- bcrypt truncates input at 72 bytes (a long
rk_key would be silently shortened, needing an HMAC/SHA pre-hash anyway), making a slow KDF not merely unnecessary but operationally hazardous for long high-entropy keys; fast HMAC-SHA256/SHA-256 is appropriate for quick auth checks on high-entropy tokens (source: https://mojoauth.com/compare-hashing-algorithms/hmac-sha256-vs-bcrypt, 2026-06-13).
Decision 6 — relay_op_logs retention: PostgreSQL daily RANGE partitioning, retention by partition DROP
Decision. For relay_op_logs at ~50k/day/account scale with 14-day-success / 30–90-day-failure retention, use native PostgreSQL declarative time partitioning (with partition drop), a simple indexed table with a periodic DELETE job, or offload to a log pipeline (Loki/ELK) — and what is the exact retention mechanism for v1? (Spec §8.1, §15 #8; plan Phase 5.)
Resolution. Ship v1 with native PostgreSQL declarative RANGE partitioning of relay_op_logs by occurred_at, using DAILY partitions, and enforce retention by DROPPING whole partitions on a schedule — not by row-level DELETE, and not by offloading to a log pipeline. Pre-create partitions ahead of time and drop those older than the retention horizon. Run the maintenance (create-ahead + drop-old) with pg_partman 5.4.3 invoked from a scheduled relayctl logs --gc / cron SELECT run_maintenance(), configured to DROP (not just DETACH) expired partitions. Keep retention simple in v1: a single uniform 90-day drop horizon for the whole table (the failure-retention ceiling), so no partition straddles two policies. The differential 14-day-success vs 90-day-failure trim is then a cheap row-level DELETE within the youngest partitions only (success rows older than 14 days), where the working set is tiny and bloat is bounded — the bulk reclamation that matters comes free from dropping 90-day-old partitions. This is the partition-drop mechanism the spec already calls "the recommended store strategy" (§8.1) and resolves open question §15 #8 in favor of partitioning over a plain DELETE job or external offload.
Alternatives considered.
| Option | Pros | Cons |
|---|---|---|
Native declarative RANGE partitioning by occurred_at (daily), retention by partition DROP via pg_partman/cron (chosen) |
Retention is metadata-only: DROP TABLE on a day partition "can very quickly delete millions of records" and "entirely avoids the VACUUM overhead caused by a bulk DELETE." No dead-tuple/index churn from the append workload — the exact failure mode a daily bulk-DELETE on a high-write table hits. Stays inside the existing Postgres dependency (no new store), honoring "operational simplicity that still scales." pg_partman (5.4.3, PG14+) automates create-ahead + drop and runs from plain cron via run_maintenance() with no background worker, fitting the Goose + scheduled-job model. Inserts still target one parent table; the redacted-by-construction write path is unchanged. Scales linearly; old partitions drop in O(1). |
Schema change: a partitioned table's PRIMARY KEY must include the partition key, so the current id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY becomes composite (occurred_at, id). Adds pg_partman (or hand-rolled DDL) plus a create-ahead job whose failure (no future partition) would block inserts — needs a monitor. Differential 14d/90d retention is no longer one query; a small DELETE runs inside recent partitions. DROP TABLE on a partition briefly needs an ACCESS EXCLUSIVE lock on the parent (mitigated by DETACH … CONCURRENTLY first, or accepting a sub-second lock on a low-contention audit table). |
Simple indexed table + periodic DELETE job (the spec's relayctl logs --gc baseline, as the only mechanism) |
Zero schema change — keeps the §8.1 table and surrogate bigint PK. Trivially expresses differential policy in one statement. Already a Phase 5 deliverable, so lowest-effort and a correct fallback. Fine at low account counts. | At 50k/day/account across accounts this is the anti-pattern the PG docs warn against: a large recurring DELETE leaves dead tuples and index bloat that standard VACUUM marks reusable but does not return to disk, so the table and its two occurred_at indexes bloat and slow the very append path. Reclaiming space needs VACUUM FULL (ACCESS EXCLUSIVE) or pg_repack — ongoing toil. The delete competes with hot-path inserts for I/O and locks. Works for months then the disk-growth surprise hits — exactly risk R12. |
| Offload to a log pipeline (Loki/ELK), Postgres holds only recent rows | Purpose-built for high-volume append-only logs with cheap time-based retention and rich ad-hoc querying; removes write pressure from Postgres; scales far past 50k/day/account. | Violates v1's operational-simplicity / self-hosted brief: a whole new stateful system to deploy, secure, scope per-account, and back up for a single-region relay. relay_op_logs is an audit/forensics table queried by relayctl and joined to relay_accounts (FK account_id); moving it out breaks the in-Postgres admin/audit model and the single-point redaction guarantees. Over-engineered for the stated scale; reintroduces a deliberately-avoided dependency. The right answer later, not for v1. |
Rationale. The workload is unambiguously a high-write, append-only, time-ordered table whose only deletes are time-based retention — the textbook case for declarative range partitioning. PostgreSQL's own current (v18) docs state the decision plainly: dropping a partition "is far faster than a bulk operation" and "entirely avoids the VACUUM overhead caused by a bulk DELETE," and recommend sizing partitions to match the retention cadence (source: https://www.postgresql.org/docs/current/ddl-partitioning.html, 2026-06-13). Reliability-first means the hot push path must not contend with a daily bulk DELETE or pay for its dead-tuple/index bloat; partition DROP makes retention a metadata operation that touches neither autovacuum nor the live indexes — whereas recurring large DELETEs accumulate bloat that VACUUM marks reusable but does not shrink on disk, requiring VACUUM FULL or pg_repack to reclaim (source: https://www.tigerdata.com/learn/how-to-reduce-bloat-in-large-postgresql-tables, 2026-06-13). Daily (not weekly/monthly) granularity is chosen because the success horizon is only 14 days and the failure horizon 90 days — daily children keep each partition small, make the 90-day drop precise, and keep the create-ahead window short. The plain-DELETE option is kept as the differential-trim tool inside young partitions and as a degraded fallback, not as bulk reclamation. A log pipeline is the correct answer at much larger multi-region scale, but for a single-region self-hosted v1 it adds a new stateful dependency that contradicts the "stays in Postgres" posture and severs the FK/audit/redaction model that lives in one place. pg_partman is the pragmatic automation: de-facto standard, current (5.4.3, 2026-03-05), PG14+, defaults to DETACH (so retention must be explicitly configured to DROP via retention_keep_table = false), and runs from ordinary cron via run_maintenance() so no background worker is required — matching the relay's Goose + scheduled-job model (source: https://github.com/pgpartman/pg_partman, 2026-06-13). The spec already names partitioning "the recommended store strategy"; this decision firms that into "partition + drop, daily, single 90-day horizon."
Consequences (what this commits).
- Schema (§8.1) — amends the plan.
relay_op_logsbecomesPARTITION BY RANGE (occurred_at)with daily child partitions; the surrogateid bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEYmust become a composite PRIMARY KEY(occurred_at, id)(id remains identity for ordering/correlation;occurred_atcarries the partition key). The two indexes —(account_id, occurred_at DESC)and(occurred_at DESC)— are created on the partitioned parent and propagate to children; theaccount_idFK withON DELETE SET NULLis unaffected. This decision explicitly supersedes spec §8.1's non-partitionedrelay_op_logsDDL (its single-columnid bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY) and the plan's "exact schema from spec §8.1" Phase 1 statement: the partitioned-from-birth design only works if the structural DDL lands in the Phase 1 init migration. Converting an already-populated non-partitioned table to partitioned later is exactly the painful full-table rewrite that partitioning exists to avoid, so it must not be deferred to Phase 5. - Migrations — Phase 1 (structural, amending the §8.1 reproduction). The Phase 1 init migration creates the partitioned parent + an initial window of daily partitions in place of the non-partitioned spec §8.1
relay_op_logsDDL and its single-column PK;pg_partmanis added viaCREATE EXTENSION+ apart_configrow (parent table, daily interval, premake N future days,retention = '90 days',retention_keep_table = falseto DROP). The composite(occurred_at, id)PK and theCREATE EXTENSION pg_partman+part_configrow are part of this Phase 1 init migration, not a later Phase 5 add-on. - Code / ops. The oplog writer is unchanged (still
INSERTinto the parent). A scheduled job —relayctl logs --gcor cronSELECT run_maintenance()— must run at least daily to (a) pre-create upcoming day partitions and (b) drop partitions older than 90 days; its failure (missing future partition) would reject inserts on the hot path, so it needs a dedicated alert in addition to the existing table-growth/disk alert (R12). The differential 14d-success trim runs as a boundedDELETEwithin recent partitions only. - Deps. Adds
pg_partman 5.4.3(PG14+) as a Postgres extension dependency; if the deployment cannot install extensions, fall back to hand-rolled create-ahead/drop DDL inrelayctl(same partition design, no extension). - Locking. Schedule drops at low-traffic times and/or
DETACH … CONCURRENTLYthenDROPto avoid the brief ACCESS EXCLUSIVE lock on the parent. - Phase split (amends the plan). The plan's Phase 5 "partitioning DDL (or documented partition/offload decision)" line item is resolved by this decision, but the structural DDL itself moves into Phase 1 (per the Schema/Migrations items above). Phase 5 owns only the operational pieces: the scheduled
run_maintenance()/relayctl logs --gccreate-ahead+drop job, the create-ahead-failure alert (a missing future partition rejects inserts on the hot path), and the bounded 14d-success differentialDELETEwithin recent partitions. The plan's Phase 1 "exact schema from spec §8.1" line and Phase 5 "partitioning DDL" deferral are both superseded.
Revisit if. (a) Sustained ingest materially exceeds the v1 envelope — e.g. tens of accounts pushing near 50k/day each so relay_op_logs exceeds ~100M live rows, or Postgres shows I/O/bloat pressure from op-log writes — at which point offloading to a dedicated log pipeline (Loki/ELK/ClickHouse) with Postgres keeping a short hot window becomes worth its added ops cost. (b) The deployment target cannot install pg_partman AND hand-rolled partition DDL proves error-prone. (c) The differential success/failure retention needs to grow into many distinct per-event horizons such that a single uniform partition-drop horizon no longer fits. (d) Operators need rich ad-hoc/full-text forensic search over op-logs beyond indexed Postgres queries.
Sources.
- PostgreSQL: "Dropping an individual partition using DROP TABLE, or doing ALTER TABLE DETACH PARTITION, is far faster than a bulk operation. These commands also entirely avoid the VACUUM overhead caused by a bulk DELETE." (PG 18) (source: https://www.postgresql.org/docs/current/ddl-partitioning.html, 2026-06-13).
- PostgreSQL: dropping the no-longer-needed partition "can very quickly delete millions of records because it doesn't have to individually delete every record";
DROP TABLEon a partition needs ACCESS EXCLUSIVE on the parent, whileDETACH PARTITION CONCURRENTLYneeds only SHARE UPDATE EXCLUSIVE (source: https://www.postgresql.org/docs/current/ddl-partitioning.html, 2026-06-13). - A unique/primary-key constraint on a declaratively partitioned table must include all partition-key columns — so partitioning by
occurred_atrequires the composite key (source: https://www.postgresql.org/docs/current/ddl-partitioning.html, 2026-06-13). pg_partmanlatest is 5.4.3 (2026-03-05); requires PostgreSQL ≥ 14; third-party extension; can run maintenance from external cron viarun_maintenance(); by default DETACHes rather than DROPs expired partitions, so drop-on-retention must be configured (retention_keep_table = false) (source: https://github.com/pgpartman/pg_partman, 2026-06-13).- For high-churn tables, recurring large DELETEs accumulate dead tuples and index bloat; VACUUM marks them reusable but does not shrink files, requiring
VACUUM FULLorpg_repack; partition DETACH+DROP needs no vacuum at all (source: https://www.tigerdata.com/learn/how-to-reduce-bloat-in-large-postgresql-tables, 2026-06-13). pg_partmanretention: "By default, partitions are not dropped (DROP), but are detached (DETACH)"; the configured interval drops partitions older than that interval; pg_partman never drops the last remaining child (source: https://www.keithf4.com/partman-retention/, 2026-06-13).
What this unblocks
Each decision retires a spec §15 open question and clears a specific implementation-plan phase to proceed without further product/architecture input:
| Decision | Retires | Unblocks |
|---|---|---|
| 1 — APNs client (hand-roll) | §15 #3, plan §3 "decision before Phase 3", risk R2 | Phase 3 — internal/apns (jwt.go/transport.go/client.go/errors.go) and POST /v1/apple/send can be built directly against net/http2 + jwt/v5 with no library evaluation gating the work; the Apple provisioning checklist item "decide apns2 vs hand-rolled" (plan §7.1) is closed. |
| 2 — FCM credentials (keyless-first) | §15 #4 | Phase 4 (internal/fcm/oauth.go builds the ADC TokenSource; POST /v1/fcm/send) and Phase 6 (secret-manager wiring, the GCP attached-SA / off-GCP WIF deploy paths, and the relayctl ping-upstream --provider fcm credential-health gate). |
3 — apns-expiration policy (finite, relay-owned) |
§15 #5 (spec open question 5) | Phase 3 (APNs apns-expiration epoch in internal/pushwire/apns.go) and Phase 4 (parallel FCM android.ttl in internal/pushwire/fcm.go); the two new internal/config TTL values and their clock-injected golden tests are now specified. |
| 4 — Rate limits (global defaults) | §15 #1 | Phase 2 — internal/ratelimit ships the global constants + env overrides and the accountLimits resolution struct with no new schema, no relayctl quota CLI, and no inheritance-precedence edge cases; spec §8.1 schema and §5.6 command table are unchanged. |
| 5 — API-key hashing (HMAC-pepper) | §15 #6 | Phase 1 (relay_api_keys schema with key_hash bytea + a pepper_version smallint NOT NULL DEFAULT 1 column added to the reproduced §8.1 table; the pepper as a first-class secret-manager secret) and Phase 2 (internal/httpapi/middleware_auth.go constant-time verify with decoy-hash and dual-pepper fallback). |
| 6 — Op-log retention (daily partition + DROP) | §15 #8 | Phase 1 (the structural DDL — partitioned relay_op_logs parent, composite (occurred_at, id) PK, CREATE EXTENSION pg_partman + part_config row — lands in the init migration, amending the plan's "exactly spec §8.1" Phase 1 statement and superseding the non-partitioned §8.1 DDL) and Phase 5 (only the relayctl logs --gc / run_maintenance() schedule, the create-ahead-failure alert, and the bounded differential 14d-success trim). |