docs(playback): monitoring & kill-switch plan + as-built coverage matrix
Capture the design intent and the shipped coverage for the stream monitoring + revocation kill-switch work, refreshed to match the final implementation and the monitoring → kill-switch → docs commit layout. - docs/superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md: the implementation plan of record (detection / enforcement / async brain split), with a Status banner and as-built deltas noting what shipped differently (durable Postgres mirror wired non-optional, shared Refuse/WatchAndCut helpers, monotonic expiry, ownership carry-forward, first-class monitoring fields). - docs/architecture/playback-paths-monitoring-kill-matrix.md: the as-built coverage matrix across server layout × playback type × route, GAP-1..GAP-9 resolution notes, restart-durability axis, and open verification items (VERIFY-3/4, operator config keys, compat download quota). Part of the stream monitoring & kill-switch epic.
This commit is contained in:
@@ -0,0 +1,433 @@
|
||||
# Playback Paths — Monitoring & Kill-Switch Coverage Matrix
|
||||
|
||||
> Reference checklist for stream monitoring + revocation kill-switch coverage.
|
||||
> Check every change to playback/monitoring/revocation against this. As-built
|
||||
> companion to the design plan
|
||||
> [`2026-07-04-stream-monitoring-and-kill-switch.md`](../superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md)
|
||||
> (that doc is the intent; this one is the shipped coverage). Status as of
|
||||
> 2026-07-05: monitoring, enforcement, and the async over-cap enforcer are all
|
||||
> shipped across two code commits — the *monitoring* commit (`feat(playback):
|
||||
> server-observed stream monitoring`) and the *kill-switch* commit
|
||||
> (`feat(playback): stream kill switch + async over-cap enforcer`). GAP-1..GAP-4
|
||||
> are all RESOLVED (see Findings): GAP-4 (kill list did not survive a restart) was
|
||||
> closed by wiring a durable Postgres mirror, and GAP-3 (in-flight cut for
|
||||
> integrated direct-play) was closed by adding `Unwrap()` to the native/compat
|
||||
> middleware writers so the cut reaches the socket. A post-review hardening pass
|
||||
> then closed GAP-5..GAP-9 found by adversarial re-review: user-kill cutoff
|
||||
> semantics, compat manifest/subtitle/mint-bypass coverage, terminate-by-id,
|
||||
> in-flight download cut, admin-list dedupe, and the metered-writer sendfile chain
|
||||
> (see Findings). That hardening is folded into the two code commits above — the
|
||||
> branch is organized as monitoring → kill-switch → this docs commit, not as a
|
||||
> running series of fix commits. (This doc references sibling commits by role, not
|
||||
> SHA — squash/rebase rewrites hashes, so a SHA citation goes stale the moment the
|
||||
> branch is amended.)
|
||||
|
||||
## Dimensions
|
||||
|
||||
- **Server layouts:** (A) integrated, no Redis · (B) integrated, with Redis ·
|
||||
(C) multi-node, with Redis (proxy/transcode edge nodes).
|
||||
- **Playback types:** direct play · remux · transcode (HLS).
|
||||
- **Routes:** native silo `/api/v1` · jellycompat (`:8096`).
|
||||
|
||||
## Key runtime facts (why cells collapse)
|
||||
|
||||
- **Integrated (A, B) serves bytes locally.** Node selection returns an empty
|
||||
plan when no edge nodes exist, and both native and jellycompat fall through to
|
||||
in-process serving (`nodepool/planner.go:248`, `jellycompat/handlers_playback.go:306`
|
||||
— the `proxyNode == nil` guard in `buildProxyRedirectURL`). A vs B differ only
|
||||
in *revocation propagation* (memory-only vs Redis) and *monitor source*
|
||||
(FuncSource vs MultiSource) — the **serve + enforcement points are identical**.
|
||||
So A ≡ B for this matrix; treated as one column "Integrated".
|
||||
- **Multi-node (C) serves bytes on edges.** Native playback hands the client a
|
||||
proxy-node URL at session start (`playback.go:1575`); jellycompat 302-redirects
|
||||
to a proxy node (`streams.go:120`). Edge serving = `internal/proxy` + `internal/transcodenode`.
|
||||
- **jellycompat shares the SessionManager.** A jellycompat play starts a real
|
||||
`playback.SessionManager` session (`streams.go:1127`), so the monitor/enforcer
|
||||
see it exactly like native. The compat `PlaybackSessionStore` is separate
|
||||
bookkeeping (PlaySessionId → recipe), not liveness/enforcement.
|
||||
- **Local transcode fallback:** in multi-node, if no transcode node is available,
|
||||
jellycompat falls back to LOCAL transcode (`streams.go:277`,
|
||||
`LocalTranscodeFallbackAllowed`) — i.e. the "Integrated jellycompat" serve path
|
||||
can also occur under layout C.
|
||||
|
||||
## The matrix
|
||||
|
||||
Legend: ✅ covered · ⚠️ covered with caveat · ❌ gap.
|
||||
|
||||
### Serving handler (where the bytes come from)
|
||||
|
||||
| Route | Type | Integrated (A/B) | Multi-node (C) |
|
||||
|---|---|---|---|
|
||||
| native | direct | `StreamHandler.HandleStream` → `ServeDirectPlay` (`stream.go:141`) | proxy `handleDirectPlay` (`proxy/server.go:171`) |
|
||||
| native | remux | `HandleStream` → `ServeRemux` (`stream.go:157`) | proxy `handleRemux` |
|
||||
| native | transcode | `HandleGetTranscodeSegment` local → `ServeFile` (`playback.go:2860`) | proxy → transcode node `handleSegment` |
|
||||
| jellycompat | direct | `HandleVideoStream` → `ServeDirectPlay` (`streams.go:153`) | 302 → proxy `handleDirectPlay` |
|
||||
| jellycompat | remux | `HandleVideoStream` → `ServeRemux` (`streams.go:151`) | 302 → proxy `handleRemux` |
|
||||
| jellycompat | transcode | `HandleHLSSegment` → `ServeFile` (`streams.go:543`) | 302 → proxy → transcode node |
|
||||
|
||||
### Monitoring (server-observed existence, reported to central)
|
||||
|
||||
The monitoring model separates **existence** (must be server-observed, never
|
||||
gated by a client report — the hidden-stream defense) from **timing** (client
|
||||
progress is fine, especially for native). A disguised re-streamer that pulls
|
||||
bytes but withholds/falsifies progress must still be counted.
|
||||
|
||||
| Route | Type | Integrated (A/B) — existence signal | Multi-node (C) — existence signal |
|
||||
|---|---|---|---|
|
||||
| native | direct | ✅ transport-count shield for the whole pour (`stream.go:136`) | ✅ edge tracker, byte-observed (`sessionByteWriter`) |
|
||||
| native | remux | ✅ transport-count shield (`stream.go:146`) | ✅ edge tracker, byte-observed |
|
||||
| native | transcode | ✅ per-segment transport marker around `ServeFile` (`playback.go:2860`) | ✅ edge tracker (Touch + AddBytes/segment) |
|
||||
| jellycompat | direct | ✅ transport-count shield (`streams.go:129`) | ✅ edge tracker, owner+route carried |
|
||||
| jellycompat | remux | ✅ transport-count shield | ✅ edge tracker, owner+route carried |
|
||||
| jellycompat | transcode | ✅ per-segment transport marker around `ServeFile` (`streams.go:543`) | ✅ edge tracker, owner+route carried |
|
||||
|
||||
- **Existence is server-observed on every cell.** Integrated: a session is
|
||||
unreapable while `activeTransportCount > 0`, and every byte-serving path now
|
||||
holds that marker — direct/remux for the whole pour, transcode for each segment
|
||||
serve (the segment markers were added to close a hidden-stream hole where a slow
|
||||
single-segment drain past the 45s grace could be reaped mid-serve; the compat
|
||||
HLS path previously refreshed liveness *only* from the client's progress POST,
|
||||
a direct violation, now also transport-marked). Edge: a Redis record exists for
|
||||
the whole connection (direct/remux) or while segments are pulled (transcode),
|
||||
advanced only by real bytes. **No path requires a client progress report to stay
|
||||
visible or counted.**
|
||||
- **Timing is secondary and may trust the client.** Native fully trusts client
|
||||
progress for position; the integrated `LastServedAt` is mapped from
|
||||
`SessionManager.LastActivityAt` (`cmd/silo/main.go` FuncSource → `handlers.LiveLocalSessions`),
|
||||
which client progress can advance. This never inflates the **count** (existence
|
||||
is the transport shield, not a heartbeat), so the cap can't be bypassed; it only
|
||||
affects which of a user's *own* over-cap sessions `selectVictims` trims first.
|
||||
The edge advances `LastServedAt` only on real bytes.
|
||||
- Report-to-central: **C** = async via Redis (edge writes `silo:sessions:*`,
|
||||
`streammonitor.RedisSource` reads). **A/B** = in-process SessionManager IS
|
||||
central; read via `streammonitor.FuncSource`. MultiSource unions both, deduped
|
||||
by session id.
|
||||
- Owner + attribution: the transcode node's start record now carries owner + route
|
||||
+ client (threaded via `TranscodeStartRequest`), and `streammonitor.mergeStreams`
|
||||
additionally backfills any missing owner/route/client from another record for the
|
||||
same session — so an ownerless-but-freshest node record can never bucket a stream
|
||||
under user 0 (which the enforcer skips) or drop route/client from the view.
|
||||
|
||||
### Monitoring data model (what central sees per stream)
|
||||
|
||||
First-class monitoring goal: see **all** active streams with enough to act on.
|
||||
Captured per live stream (`streammonitor.LiveStream` / `nodesessions.SessionInfo`):
|
||||
|
||||
| Field | Source | Notes |
|
||||
|---|---|---|
|
||||
| existence, session id | both | primary; server-observed |
|
||||
| user id, profile id | token claims / Session | primary; owner attribution |
|
||||
| media file id | claims / Session | primary; human title is a display-time lookup (follow-up) |
|
||||
| play method (direct/remux/transcode) | `Type` | primary |
|
||||
| **route (native/jellycompat)** | `Route` — `Session.Origin` (integrated) or token `Origin` claim (edge) | primary; native and jellycompat share the SessionManager and `Type`, so route is a distinct field seeded at the two `WithClientInfo` sites and carried to edges in the token |
|
||||
| node serving | NodeName/NodeURL | integrated stamps the local host |
|
||||
| client ip / client name | `Session.ClientIP/ClientName` (integrated) or edge request + `ClientName` claim | secondary; also the natural re-streaming fingerprint |
|
||||
| position | `Session.Position` | secondary timing |
|
||||
| bytes served | edge `AddBytes` | edge only; throughput signal |
|
||||
| hw/sw, resolution, codecs | Session / node record | secondary |
|
||||
|
||||
- **Admin visibility:** `HandleListSessions` unions the Redis edge records with the
|
||||
in-process integrated sessions (`LiveLocalSessions`), so a single-node integrated
|
||||
deployment is no longer blind (previously Redis-only). The union is deduped by
|
||||
session id (`streammonitor.DedupeSessionInfos`, mirroring `mergeStreams`) so a
|
||||
stream tracked by both the central manager and the edge serving it shows as ONE
|
||||
row — matching the enforcer's count. A `node_id` filter targets
|
||||
an edge, so integrated sessions appear only in the unfiltered listing.
|
||||
- **Downloads are an intentional exemption — with one asymmetry.** `/downloads/*`
|
||||
(native) and compat `/Items/{id}/Download` serve full media with no
|
||||
session/monitor record and do NOT count against the live-stream cap. The native
|
||||
route requires a quota-checked download row (`internal/downloads`
|
||||
concurrency/period limits); the compat route requires only compat auth — it is
|
||||
covered by NO quota today (bringing it under an Infuse-compatible download
|
||||
quota is an open follow-up). Both routes now arm the shared `WatchAndCut`, so a
|
||||
per-user stream revocation cuts an in-flight download pour, and the same hook
|
||||
deletes every compat login so reconnects need re-auth. Documented at the
|
||||
handlers.
|
||||
|
||||
### Kill switch (revocation enforced on the serve path)
|
||||
|
||||
| Route | Type | Integrated (A/B) | Multi-node (C) |
|
||||
|---|---|---|---|
|
||||
| native | direct | ✅ `guardRevocationCut` (Refuse + in-flight cut†) | ✅ `verifyToken` + `cutOnRevocation` (in-flight cut) |
|
||||
| native | remux | ✅ `guardRevocationCut` (Refuse + cut†) | ✅ verifyToken + cutOnRevocation |
|
||||
| native | transcode | ✅ `guardRevocation` (refused within one segment) | ✅ proxy verifyToken + node `refuseIfRevoked` + reconstruct guard‡ |
|
||||
| jellycompat | direct | ✅ `Refuse` + in-flight cut† (GAP-1/3 fixed) | ✅ via proxy; ✅ local fallback guarded |
|
||||
| jellycompat | remux | ✅ `Refuse` + in-flight cut† | ✅ via proxy; ✅ local fallback guarded |
|
||||
| jellycompat | transcode | ✅ `Refuse` per segment (GAP-1 fixed) | ✅ via proxy/node; ✅ local fallback guarded |
|
||||
|
||||
† In-flight cut uses `streamrevoke.Store.WatchAndCut` (SetWriteDeadline, checked on
|
||||
entry then every 5s). Works at the edge (the proxy's metered writer implements
|
||||
`Unwrap`) **and** integrated: the native (`statusWriter`, `requestStatusWriter`)
|
||||
and compat (`loggingResponseWriter`, `compatImageProxyTagResponseWriter`,
|
||||
`debugResponseWriter`) middleware writers now implement `Unwrap()`, so
|
||||
`http.NewResponseController` reaches the socket instead of no-oping (see GAP-3). A
|
||||
live-client smoke test is still worthwhile, but the previously-guaranteed
|
||||
integrated no-op is fixed; worst case it still degrades to a next-request `Refuse`.
|
||||
|
||||
‡ The transcode node's serve-path check is session-only (`refuseIfRevoked` passes
|
||||
`userID = 0`), so a per-*user* kill is enforced by the fronting proxy's `verifyToken`
|
||||
(which passes `claims.UserID`), not at the node's segment serve. The node's
|
||||
reconstruct guard *does* use the real `claims.UserID`, so a rebuild-after-restart of
|
||||
a user-killed session is blocked.
|
||||
|
||||
### Restart durability (kill survives a process restart / Redis loss)
|
||||
|
||||
This axis exists because the branch sits on top of PR #174 (restart-resilient
|
||||
playback): a session killed before a restart is **reconstructed** afterward from a
|
||||
durable recipe card, so if the kill list does not also survive the restart the
|
||||
reconstructed stream is silently re-served. The kill list must be at least as
|
||||
durable as the thing it kills.
|
||||
|
||||
| Deployment | Session survives restart (PR #174) | Kill survives restart | Kill survives Redis flush |
|
||||
|---|---|---|---|
|
||||
| A — integrated, no Redis | ✅ recipe card (PG) | ✅ durable PG mirror | ✅ (never used Redis) |
|
||||
| B — integrated, with Redis | ✅ | ✅ durable PG mirror (+ Redis warm) | ✅ durable PG mirror |
|
||||
| C — multi-node, with Redis | ✅ | ✅ central durable PG mirror; edges re-warm from Redis | ✅ central re-warms edges from PG→Redis on next write/poll |
|
||||
|
||||
\* Steady-state guarantee. Transient boot caveat: if the durable warm exhausts its
|
||||
bounded retry **and** Redis is empty, the kill list is empty until the first poll
|
||||
tick (≤60s). This fails *open* by design — see the "Warm is async-tolerant" bullet
|
||||
below.
|
||||
|
||||
- **Hot path unchanged.** `IsRevoked` is still a pure in-memory map read. Postgres
|
||||
is touched only on write (`Upsert`), on warm/reconcile (`ListActive` at
|
||||
`StartSync` and on the poll tick), and on trim (`Prune`) — never per request.
|
||||
- **Central-side only.** The durable mirror is wired into the integrated/api
|
||||
`streamrevoke.Store` (`cmd/silo/main.go`), not the edge/proxy/transcode nodes,
|
||||
which have no app DB and enforce via Redis pub/sub + poll.
|
||||
- **Bounded growth.** Rows are keyed by `(kind, id)`, so the async enforcer
|
||||
re-revoking every pass UPSERTs one row rather than accumulating; expired rows are
|
||||
physically reclaimed by `Prune` on the poll tick, and `ListActive` filters on
|
||||
`expires_at > now()`.
|
||||
- **Expiry is monotonic.** Both the in-memory `applyLocal` and the durable
|
||||
`Upsert` keep whichever copy expires LATER (`GREATEST`), never shortening an
|
||||
existing kill. This matters because the async over-cap enforcer re-revokes with a
|
||||
short 5m self-heal TTL on the same `KindSession` key an admin may have killed for
|
||||
24h; without monotonic expiry the enforcer would silently shrink the admin kill
|
||||
to 5m and reopen the restart-resurrection window this whole axis defends against.
|
||||
- **Warm is async-tolerant.** By product decision, a just-reconstructed revoked
|
||||
stream may serve a few seconds before the durable warm/next Refuse tick cuts it;
|
||||
no blocking startup ordering is required. Edge caveat: if the boot durable warm
|
||||
exhausts its bounded retry **and** Redis is empty, the kill list is empty until
|
||||
the first poll tick (≤60s). This fails *open* by design (a kill switch cannot
|
||||
fail closed without blocking all playback when the DB is briefly unavailable).
|
||||
- **Edge Redis-outage fails open.** Edge nodes (proxy/transcode) have no durable
|
||||
store and learn *new* kills only from Redis pub/sub + SCAN. While Redis is
|
||||
unreachable, already-cached kills persist but a kill issued *during* the outage
|
||||
does not reach the edge until Redis recovers — enforcement fails open there.
|
||||
|
||||
## Findings
|
||||
|
||||
> GAP-1..GAP-3 all shipped in the kill-switch commit (see the Status banner for
|
||||
> how commits are referenced by role rather than SHA).
|
||||
|
||||
**GAP-1 — RESOLVED.** jellycompat LOCAL serving now consults the
|
||||
shared `Store.Refuse` in `HandleVideoStream` (direct/remux) and `HandleHLSSegment`
|
||||
(per segment), closing the integrated / local-transcode-fallback hole.
|
||||
|
||||
**GAP-2 — RESOLVED.** `buildProxyRedirectURL` now stamps
|
||||
`uid`/`pid`/`mfid` onto the jellycompat proxy-redirect token (looked up from the
|
||||
upstream SessionManager session), so edge monitoring/kill attribution matches
|
||||
native. `AuthUserID = 0` no longer occurs for jellycompat.
|
||||
|
||||
**GAP-3 — RESOLVED.** The in-flight cut is shared as
|
||||
`streamrevoke.Store.WatchAndCut` and applied to native `/stream`
|
||||
(`guardRevocationCut`) and jellycompat `HandleVideoStream`. The original caveat —
|
||||
`SetWriteDeadline` might not reach the socket through the native/compat middleware
|
||||
chain — was **confirmed** as a guaranteed integrated no-op (four middleware writers
|
||||
wrapped `w` without `Unwrap()`) and then fixed by adding `Unwrap()` to
|
||||
`statusWriter`, `requestStatusWriter`, `loggingResponseWriter`,
|
||||
`compatImageProxyTagResponseWriter`, and `debugResponseWriter`. The cut now reaches
|
||||
the socket in integrated mode; if it ever no-ops again the stream still stops on
|
||||
its next request via `Refuse` (no regression). See VERIFY-1.
|
||||
|
||||
**GAP-4 — RESOLVED.** The kill list did not survive a process restart or a Redis
|
||||
flush, but PR #174's restart-resilient playback *does* reconstruct the session — so
|
||||
a stream revoked before a restart was silently re-served afterward. This was
|
||||
universal in deployment A (integrated, no Redis: the kill list was pure RAM) and
|
||||
occurred on any Redis loss in B/C. Fixed by wiring a concrete Postgres
|
||||
`DurableStore` (`internal/streamrevoke/durable_postgres.go`, table
|
||||
`stream_revocations`) into the central `streamrevoke.Store`: `Revoke` mirrors to
|
||||
Postgres, `StartSync` warms from it on boot, and the poll tick re-warms (heals a
|
||||
Redis flush) and `Prune`s expired rows. The `defaultTTL` (24h) is held `>=` the
|
||||
recipe-card `MaxTokenTTL` (24h) as an invariant so a kill cannot expire before its
|
||||
session can be reconstructed, and expiry is **monotonic** on every write (in-memory
|
||||
`applyLocal` + durable `Upsert` `GREATEST`) so the async enforcer's short 5m
|
||||
re-revoke can never shorten a longer admin kill on the same session key. Hot path
|
||||
(`IsRevoked`) stays an in-memory read.
|
||||
|
||||
**GAP-5 — RESOLVED (post-review).** User-kind revocations were a blanket ban:
|
||||
any admin user edit (via `OnUserSessionsRevoked`) 403'd that user's playback for
|
||||
24h even after re-login, unrevocably. Fixed with cutoff semantics — see
|
||||
"user-ID kills" above.
|
||||
|
||||
**GAP-6 — RESOLVED (post-review).** Compat coverage holes: `HandleMasterManifest`
|
||||
/ `HandleHLSManifest` had no revocation check and kept (or, post-restart,
|
||||
re-spawned) a killed session's ffmpeg via the ensure path; `HandleVideoStream`
|
||||
checked revocation only after `ensureUpstreamPlayback`, which can replace an
|
||||
unreconstructable killed session with a fresh id that passed the check (kill
|
||||
dodged by re-hitting the URL); `HandleSubtitleStream` ignored kills entirely.
|
||||
All four now `Refuse` up front.
|
||||
|
||||
**GAP-7 — RESOLVED (post-review).** Admin terminate 404'd before revoking when
|
||||
the session was absent from the central in-memory manager — exactly the
|
||||
edge-served / post-restart / progress-withholding streams the kill switch
|
||||
exists for (the admin list showed them via the Redis union; terminate couldn't
|
||||
touch them). Terminate now writes the revocation FIRST, keyed on the session id
|
||||
alone, and answers `202 {status: "revoked"}` when there is no local session for
|
||||
the cooperative command.
|
||||
|
||||
**GAP-8 — RESOLVED (post-review).** The admin session list double-counted every
|
||||
edge-served stream (central manager row + edge Redis row, no dedupe), unlike
|
||||
the enforcer's merged picture. Now deduped by session id with owner/attribution
|
||||
carry-forward (`streammonitor.DedupeSessionInfos`).
|
||||
|
||||
**GAP-9 — RESOLVED (post-review).** The "restore sendfile" commit was dead code
|
||||
end-to-end: every stream route runs inside `meterEgress`, and
|
||||
`meteredResponseWriter` hid `io.ReaderFrom`, so `sessionByteWriter.ReadFrom`'s
|
||||
fast path never fired and all direct-play/remux bytes went through a userspace
|
||||
copy. `meteredResponseWriter` now forwards `ReadFrom` (metering the returned
|
||||
total); a chain test locks the full production writer stack onto the sendfile
|
||||
path. Related tracker fixes in the same pass: `Touch`/`Track` preserve the
|
||||
first-seen `StartedAt` (was reset per segment/range request, corrupting the
|
||||
enforcer's tie-break and the admin start time), `AddBytes` no longer recreates
|
||||
entries for pruned sessions (slow permanent leak), the proxy→node segment pour
|
||||
attributes bytes incrementally (a slow drain can't go invisible mid-segment or
|
||||
post bytes to a dead record), and the transcode node marks serve activity so
|
||||
its record's `LastServedAt` is no longer frozen at start time.
|
||||
|
||||
**GAP-3 (original, superseded) — no in-flight cut for native/compat LOCAL direct-play.**
|
||||
The original finding: `guardRevocation` refused the next request/reconnect but could
|
||||
not hang up an in-flight `ServeFile` on the central process for a single long GET.
|
||||
Now closed — see GAP-3 above (shared `WatchAndCut` + middleware `Unwrap()`). Even
|
||||
if the cut degrades, admin terminate also fires the realtime command + session stop,
|
||||
so a killed direct-play stops on its next request at worst.
|
||||
|
||||
## Open verification items (check before prod)
|
||||
|
||||
- [x] **VERIFY-1 — in-flight cut reaches the socket on native + jellycompat — RESOLVED.**
|
||||
`streamrevoke.Store.WatchAndCut` cuts a revoked long-GET direct-play/remux by
|
||||
setting a zero write deadline via `http.NewResponseController(w)`, which walks the
|
||||
writer chain via `Unwrap()`. This was **confirmed broken** in integrated mode:
|
||||
`statusWriter`/`requestStatusWriter` (native) and `loggingResponseWriter`/
|
||||
`compatImageProxyTagResponseWriter`/`debugResponseWriter` (compat) wrapped `w`
|
||||
without `Unwrap()`, so `SetWriteDeadline` returned `ErrNotSupported` and the cut
|
||||
no-oped. Fixed by adding `Unwrap() http.ResponseWriter` to all five (pattern from
|
||||
`activitylog/middleware.go:186`). A live smoke test is still worthwhile: start a
|
||||
native and a jellycompat single-long-GET direct-play, admin-terminate it, confirm
|
||||
the connection drops within ~5s rather than only on the next request.
|
||||
- [x] **VERIFY-2 — download routes vs the stream cap — RESOLVED (documented exemption).**
|
||||
`/downloads/*` (native) and compat `/Items/{id}/Download` serve full media with no
|
||||
session/monitor record; they are intentionally exempt from the live-stream cap and
|
||||
governed by the separate download concurrency/period quota. Documented at both
|
||||
handlers and in the monitoring data model above. (Note: they also do not consult
|
||||
the kill switch — a per-user *stream* revocation does not stop an in-flight
|
||||
download. If a ban should also cut downloads, guard those handlers separately.)
|
||||
- [ ] **VERIFY-4 — transcode buffer-ahead evasion (HOLE 2).** A transcode client
|
||||
(native or compat) that buffers far ahead and pauses segment fetches for >45s (the
|
||||
active grace) is transiently reaped, then self-heals on the next fetch — so a
|
||||
determined abuser could time buffer-pauses to duck under the cap at the enforcer's
|
||||
tick. The per-segment transport marker fixes the mid-serve case but not the idle
|
||||
gap between segments. Fix needs a reaper change: treat a running transcode ffmpeg
|
||||
as liveness, or a transcode-specific grace ≥ the client's max buffer. Deferred
|
||||
because it touches the sensitive session-reaping path (cf. #279).
|
||||
- [ ] **VERIFY-3 — multi-replica central enforcement.** The async enforcer runs on
|
||||
every integrated/api process. Two *integrated* replicas behind a load balancer
|
||||
each see only their own in-process `FuncSource` (integrated mode writes no
|
||||
`silo:sessions:*` records), so a user split across both can exceed `max_streams`
|
||||
undetected; multiple *api* replicas each run an independent enforcer and revoke
|
||||
the same victims (idempotent, but redundant write load). A single-leader election
|
||||
for the enforcer, or having integrated nodes publish their local sessions to the
|
||||
shared Redis picture, would close this.
|
||||
|
||||
## Design assessment (shared vs duplicated code)
|
||||
|
||||
**Kill DECISION is well-centralized — keep it.** All revocation state and the
|
||||
`IsRevoked(sessionID, userID)` check live in one place (`internal/streamrevoke`).
|
||||
Session kills and user kills use the SAME store and the SAME check — `RevokeSession`
|
||||
and `RevokeUser` write into one `items` map; `IsRevoked` tests both keys. No
|
||||
duplication at the decision layer.
|
||||
|
||||
**Kill ENFORCEMENT — DONE.** The "extract sid/uid → refuse (403)"
|
||||
step is now a single shared method `streamrevoke.Store.Refuse(w, sessionID, userID)
|
||||
bool`. All four serve surfaces call it (proxy `verifyToken`, transcode node
|
||||
`refuseIfRevoked`, native `guardRevocation`, jellycompat serve handlers); each only
|
||||
owns its token extraction (path `{token}`, query `st`, or compat `PlaySessionId`).
|
||||
The in-flight connection-cut is the shared `WatchAndCut` (SetWriteDeadline), used by
|
||||
both the edge (`cutOnRevocation`) and the native/compat long-pour paths. ONE section
|
||||
per concern, as intended.
|
||||
|
||||
**Monitoring WRITE side is two implementations by necessity — leave it.** Edges
|
||||
have no SessionManager; the integrated process has no tracker. `nodesessions.Tracker`
|
||||
(edge) and `SessionManager` (integrated) are different runtimes, not duplicated
|
||||
logic, and they already converge at the single read layer (`streammonitor`,
|
||||
unioned by MultiSource). Forcing one implementation would be worse.
|
||||
|
||||
**user-ID kills — plumbing check (as asked).** Two distinct layers, correctly
|
||||
separate:
|
||||
- *Auth/login revocation* (pre-existing): `OnUserSessionsRevoked` revokes native
|
||||
auth sessions (`auth.SessionRepository.RevokeAllByUser`) and drops compat login
|
||||
tokens (`SessionStore.DeleteByUserID`). This governs "can this user authenticate",
|
||||
not "stop this byte stream".
|
||||
- *Stream revocation* (new): `streamRevocation.RevokeUser` writes a `KindUser`
|
||||
entry into the SAME `streamrevoke` store as session kills. Session-level and
|
||||
user-level STREAM kills already share all plumbing (one store, one `IsRevoked`).
|
||||
- **A user kill is a CUTOFF, not a ban (GAP-5 fix).** `IsRevoked` matches a
|
||||
`KindUser` entry only when the stream's credential predates the revocation:
|
||||
the stream token's `iat` on token-bearing surfaces (edge proxy, transcode
|
||||
node, native `?st=`), the request entry time on freshly-authenticated
|
||||
surfaces (native session auth, compat login — every pre-revocation login is
|
||||
reset by the same hook, so reaching a serve path afterward proves fresh
|
||||
auth). An in-flight pour always predates a later revocation, so mid-pour
|
||||
user kills still cut. An unknown (zero) credential time never matches (fail
|
||||
open). This is what makes it safe for `OnUserSessionsRevoked` — which fires
|
||||
on ANY admin edit of password/role/enabled/permissions/quality — to also
|
||||
write a stream kill: without the cutoff, a routine permission tweak 403'd
|
||||
the user's playback for the full 24h TTL after re-login, with no unrevoke
|
||||
path (expiry is deliberately monotonic).
|
||||
- These two layers should NOT be merged — conflating "may log in" with "this
|
||||
stream must die" would couple auth and playback. They are chained in the same
|
||||
hook, which is the right seam.
|
||||
|
||||
## Recommended follow-up order
|
||||
|
||||
GAP-1, GAP-2, GAP-3, and GAP-4 are all shipped (see Findings). Already-shipped
|
||||
hardening not to re-do: `Revoke` propagation uses `context.WithoutCancel`, so an
|
||||
aborted admin request can no longer strand a kill in central memory only.
|
||||
Remaining work:
|
||||
|
||||
1. **Operator config keys (small):** wire `auth.stream_revocation_poll` /
|
||||
`auth.stream_revocation_ttl` through the settings pipeline; defaults (60s / 24h)
|
||||
are currently hardcoded in `internal/streamrevoke/store.go`.
|
||||
2. **Transcode buffer-ahead liveness (VERIFY-4 / HOLE 2):** treat a running ffmpeg
|
||||
as liveness (or a transcode-specific grace) so a buffered-ahead transcode isn't
|
||||
transiently reaped and can't time buffer-pauses to duck the cap.
|
||||
3. **Media title enrichment:** resolve `MediaFileID → title` at admin display time
|
||||
so the view shows *what* is being watched, not just a numeric id (also enables a
|
||||
distinct-title re-streaming heuristic).
|
||||
4. **Multi-replica enforcement (VERIFY-3):** single-leader election for the async
|
||||
enforcer, or have integrated nodes publish local sessions to the shared Redis
|
||||
picture so the cap holds across replicas.
|
||||
5. **Self-heal a failed durable write:** `Revoke` mirrors to Postgres best-effort;
|
||||
a transient DB error at revoke time is logged but never repaired (the maintain
|
||||
tick only reads *from* durable, never re-mirrors live in-memory kills *to* it),
|
||||
so that one kill is lost on the next restart. Add a memory→durable re-mirror on
|
||||
the poll tick.
|
||||
6. **Invariant guard test:** `defaultTTL >= playback.MaxTokenTTL` is coupled only by
|
||||
a prose comment (streamrevoke can't import playback — import cycle). Add a test
|
||||
in a package that can import both, so raising `MaxTokenTTL` can't silently
|
||||
violate it.
|
||||
7. **Transcode-node ghost records:** the node's session-backed `Track` record is
|
||||
refreshed forever until an explicit stop/cleanup; if central dies or loses the
|
||||
session without the stop reaching the node, an owner-attributed ghost persists
|
||||
in Redis indefinitely — permanently +1 in the user's live count (the enforcer's
|
||||
freshness ordering trims the ghost first, so real streams survive, but the
|
||||
enforcer re-revokes it every pass and the admin view overcounts). Needs a
|
||||
node-side idle sweep tied to real serve activity (`MarkServed` now provides the
|
||||
signal).
|
||||
8. **Compat download quota:** compat `/Items/{id}/Download` is exempt from the
|
||||
stream cap AND covered by no download quota (native downloads require a
|
||||
quota-checked download row; the compat route requires only compat auth). An
|
||||
Infuse-compatible quota (plain GETs, Range requests, no download rows) is open
|
||||
product/design work; today's controls are compat auth, the library access
|
||||
filter, and the in-flight user-revocation cut.
|
||||
@@ -0,0 +1,179 @@
|
||||
# Stream Monitoring & Kill-Switch Implementation Plan
|
||||
|
||||
> **Status: IMPLEMENTED (with a minor operator-config follow-up open)** on branch `feat/sauron-async-enforcer`. All three phases shipped — Phase 1 (authoritative server-side monitoring), Phase 2 (revocation kill switch + edge/native/jellycompat enforcement), and Phase 3 (async over-cap enforcer). The one still-open item is operator ergonomics, not a phase: the `auth.stream_revocation_poll` / `auth.stream_revocation_ttl` keys are not yet wired through the settings pipeline (defaults 60s / 24h are hardcoded) — see "As-built deltas" and the coverage matrix's follow-up list. A durable Postgres mirror of the kill list was added **after** this plan was written, so kills now survive a server restart or a Redis flush (this plan's "optional Postgres durability" is no longer optional — it closes the gap where restart-resilient playback could reconstruct and re-serve an already-killed stream). For the **as-built** coverage across server layout × playback type × route, the resolved gap findings (GAP-1..GAP-4), and the open verification items, see the companion reference: [`playback-paths-monitoring-kill-matrix.md`](../../architecture/playback-paths-monitoring-kill-matrix.md). The task checkboxes below are retained as the original plan of record.
|
||||
|
||||
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Do not edit this plan file while implementing it.**
|
||||
|
||||
**Goal:** Gain control over stream abuse (concurrency-cap violations, admin termination, Stremio-style re-streaming) **without** putting the central database on the per-segment hot path and **without** changing the client-facing token protocol (the later-added `Origin` route claim is server-signed and opaque to clients — they only echo the token, so the wire protocol clients see is unchanged). Do it with two primitives plus an off-hot-path decision loop: (1) authoritative server-side monitoring that never trusts the client, and (2) a kill switch that stops any stream within ~120s and keeps it dead.
|
||||
|
||||
**Architecture:** Split the problem into *detection* and *enforcement*, and keep both off the hot path.
|
||||
|
||||
- **Detection = authoritative monitoring (no client trust).** Bytes can only reach a viewer by flowing through the edge (segment serving) or a held-open direct-play socket that the edge itself is writing. So the server *inherently* knows every live stream — it is the one sending the bytes. We make the existing `nodesessions` records the source of truth, refreshed by *real serve activity* (segment served / active transport / open connection), enriched with `last_served_at` and bytes, and mirror the same signal into integrated mode. No client progress report is trusted for liveness. A stream that "plays quietly" is still pulling segments (or holding a socket) and is therefore visible; one that stops pulling ages out on its own.
|
||||
- **Enforcement = a revocation kill switch.** A small revocation set lives in Redis (multi-node) or in-memory (integrated), mirrored to Postgres for durability (shipped as a wired part of the kill path, not an optional add-on — it is what lets a kill survive a restart/Redis flush). Edges cache it in memory, refreshed by pub/sub push (near-instant) plus a ≤60s poll (safety net). The **only** hot-path addition is a local in-memory set-membership check per segment — no per-segment Redis or Postgres. On a hit the edge refuses the request and cuts the connection; for direct-play the edge periodically re-checks during the long GET and aborts. A kill "stays dead" because the revocation entry outlives the token and every reconnect is refused.
|
||||
- **Decision = an async brain, fully off the hot path.** A central background evaluator reads the live monitoring picture and the per-user limits (`users.max_streams`), and decides what to kill: over-cap victims, admin-terminated sessions, and (pluggable) abuse heuristics. Every reason — exceeded limit, admin terminate, Stremio abuse — collapses to the *same* action: write a revocation. Enforcement follows within ~120s. The delay is acceptable for a media server.
|
||||
|
||||
The client-facing token **wire protocol is unchanged**. Its `sid`/`uid`/`pid` claims are the revocation keys; the server may add opaque server-signed claims (e.g. the later-added `Origin` route claim) that clients only echo. Nothing about mint sites, segment URLs, HLS manifests, or client behavior changes. This is why the approach is low-risk relative to a short-ticket redesign (see "Background — why not short tickets").
|
||||
|
||||
**Tech Stack:** Go, chi routers (proxy/transcode edge + native api), `redis/go-redis/v9`, Postgres via pgx + Goose migrations, the existing `nodeconfig.Watcher` settings pipeline, the `cache` package Redis client + `EventBus` pub/sub.
|
||||
|
||||
---
|
||||
|
||||
## Commands
|
||||
|
||||
Commands assume the repository root is the cwd.
|
||||
|
||||
- Build: `make build`
|
||||
- Backend only: `go build ./... && go vet ./...`
|
||||
- Lint: `make lint`
|
||||
- New migration: `make migrate-create NAME=revoked_stream_sessions`
|
||||
- Local services: `docker compose up -d postgres redis`
|
||||
|
||||
---
|
||||
|
||||
## Background — what we learned (why this shape)
|
||||
|
||||
Audited integration points this plan builds on (do not re-derive):
|
||||
|
||||
- **Stream token** is a stateless signed JWT, 24h TTL, verified at the edge by signature only, no store, no revocation today (`internal/streamtoken/token.go`; mints at `internal/api/handlers/playback.go:517,1601,2159` and `internal/jellycompat/handlers_playback.go:317`). Claims carry `sid` (`SessionID`), `uid`, `pid`, `mfid` — our revocation keys.
|
||||
- **Edge already sees every segment and already writes Redis session records.** Proxy `touchTranscodeSession` → `nodesessions.Tracker.Touch` on every manifest/segment (`internal/proxy/server.go:192`; `internal/nodesessions/tracker.go:137`), 60s TTL, refreshed every 30s. Direct-play/remux are tracked for the request lifetime (`proxy/server.go:143-145,156-158`). Records already carry `AuthUserID`/`ProfileID`/`MediaFileID` (`tracker.go:41-43`) but **no** `last_served_at` or bytes yet.
|
||||
- **Edge egress is already metered** but only as a node aggregate (`internal/proxy/egress.go`), not per session.
|
||||
- **Edge Redis is mandatory** and constructed once (`cmd/silo/main.go:576-580,605`), reused for the recipe store (`main.go:623`) — reuse the same client for a revocation store via a `SetRevocationStore(...)` wired exactly like `SetRecipeStore`.
|
||||
- **Central Redis** is `apiRedisClient` (`main.go:672`, may be nil on Redis-less single-node), injected as `deps.RedisClient` (`internal/api/router.go:135`). Central and edges point at the **same** Redis in multi-node (existing `silo:sessions:*` write/read contract proves it).
|
||||
- **Config auto-propagates to edges** via `nodeconfig.Watcher` reading `server_settings` (`internal/nodeconfig/watcher.go`); follow the `jellyfin_compat.session_ttl` precedent (`internal/config/config.go:217,233`; `internal/config/db_loader.go:381`) for any new key.
|
||||
- **Pub/sub already exists** on `cache.ChannelAdmin` and the watcher subscribes to it (`watcher.go:83-93`) — reuse this channel/pattern to push revocations to edges near-instantly.
|
||||
- **Revocation hook exists but is compat-only.** `OnUserSessionsRevoked` (`router.go:162`, `admin.go:101`, impl `cmd/silo/main.go:2047`) currently only deletes Jellyfin-compat sessions. Admin terminate (`internal/api/handlers/admin_playback_control.go:63` `HandleTerminateSession` → `handleSessionCommand`) currently only sends a realtime command + `stopPlaybackSessionByID` fallback; it does **not** revoke the stream credential at the edge. These are the exact hooks to extend.
|
||||
- **Concurrency counting already exists in-process** (`internal/playback/session.go` `SessionManager`, per-user `activeCountLocked`, `SessionLimitProvider` reading `users.max_streams`) but only for the local API process and only at admission. The async brain reuses the *limit provider* but computes the live count from the monitoring picture instead of the in-process map, so it works across nodes.
|
||||
|
||||
**Why not short tickets (the rejected alternative):** turning the 24h token into a short-lived play ticket would enforce the cap continuously, but (a) Silo's `master.m3u8` is a **VOD** playlist fetched once (`internal/playback/transcode.go:1231` `GenerateFullManifest`), so short tickets embedded in segment URLs would expire mid-movie and 401 the rest of the stream — the "renewal rides on playlist refresh" premise does not hold; (b) flipping the global token TTL is a break-all-playback change that cannot be behaviorally verified without real clients. The monitor-and-kill approach accepts a ~120s enforcement delay in exchange for zero client-protocol change and zero hot-path DB/Redis.
|
||||
|
||||
---
|
||||
|
||||
## File Structure
|
||||
|
||||
### Phase 1 — Authoritative monitoring (no client trust)
|
||||
|
||||
- Modify `internal/nodesessions/tracker.go`
|
||||
- Add `LastServedAt string` and `BytesServed int64` to `SessionInfo`.
|
||||
- Refresh `LastServedAt` (and accumulate `BytesServed`) on real serve activity, not just first touch. Keep the Redis write throttled (only re-`SET` when the record materially changes or on the 30s refresh) so this stays cheap.
|
||||
- Add a `Snapshot(ctx) ([]SessionInfo, error)` helper (edge-local view) if useful for `/status`.
|
||||
- Modify `internal/proxy/server.go`
|
||||
- Feed real bytes-served per session into the tracker (wrap the segment/direct/remux response writers, or reuse the `egressMeter` write hook keyed by `claims.SessionID`).
|
||||
- Ensure direct-play/remux held-open connections keep `LastServedAt` fresh periodically (a ticking update while `io.Copy`/`ServeFile` runs), so a long quiet pour is never invisible.
|
||||
- Modify `internal/proxy/egress.go`
|
||||
- Allow the metered writer to report per-session deltas to a callback (so the tracker can attribute bytes), in addition to the existing node-aggregate rate.
|
||||
- Modify `internal/transcodenode/server.go`
|
||||
- Mirror the same `last_served_at`/bytes tracking on the node's own segment serve path (`handleSegment`).
|
||||
- Modify `internal/api/handlers/stream.go` and `internal/api/handlers/playback.go`
|
||||
- Integrated mode: on the native segment/direct serve paths, record server-side serve activity (last-served/bytes) into the same monitoring surface (in-memory for single-node; the tracker if Redis present). Stop relying on client progress reports for liveness. Reuse `BeginTransport`/`EndTransport` (`internal/playback/session.go:577,593`) plus a serve-activity touch.
|
||||
- Modify `internal/api/handlers/nodes.go`
|
||||
- Extend `HandleListSessions` (`:303`) to surface `last_served_at`/bytes so the admin view and the brain see the authoritative picture.
|
||||
- Create `internal/streammonitor/monitor.go`
|
||||
- A thin central-side aggregator that returns a normalized "live streams" snapshot: read Redis `silo:sessions:*` (multi-node) or the in-process `SessionManager`/tracker (integrated), keyed by `uid`, with `sid`, method, `last_served_at`, bytes, node. This is the single input to the async brain and to admin monitoring.
|
||||
- Create `internal/streammonitor/monitor_test.go`
|
||||
- Table tests over synthetic session records: liveness aging, per-user grouping, multi-node de-dup by `sid`.
|
||||
|
||||
### Phase 2 — Kill switch (revocation) + enforcement
|
||||
|
||||
- Create `internal/streamrevoke/store.go`
|
||||
- Interface `Store` with `Revoke(ctx, key RevKey, reason string, until time.Time) error`, `IsRevoked(key RevKey) bool` (local, hot-path, in-memory), `List(ctx) ([]Revocation, error)`, `StartSync(ctx)`.
|
||||
- `RevKey` addresses a session (`sid`) or a user (`uid`); a user revocation matches all that user's sessions.
|
||||
- In-memory backend (integrated): a `map`/set behind an RWMutex with TTL entries; `Revoke` writes locally.
|
||||
- Redis backend (multi-node): `silo:revoked:sess:{sid}` / `silo:revoked:user:{uid}` keys with TTL = token remaining life (default 24h); `StartSync` warms an in-memory cache from a `SCAN` on start, subscribes to `cache.ChannelAdmin` for push updates, and polls every ≤60s as a safety net. `IsRevoked` reads only the in-memory cache (no per-call Redis).
|
||||
- Postgres durability mirror (see migration), wired central-side, so kills survive a Redis flush/restart; repopulate Redis + cache on startup.
|
||||
- Create `internal/streamrevoke/store_test.go`
|
||||
- Cover: local cache hit/miss, TTL expiry, user-key matches session, pub/sub push applies, poll reconciles, durability repopulation.
|
||||
- Create `migrations/sql/<timestamp>_revoked_stream_sessions.sql` (via `make migrate-create NAME=revoked_stream_sessions`)
|
||||
- Table `revoked_stream_sessions(kind text, key text, reason text, revoked_at timestamptz, expires_at timestamptz, primary key(kind,key))` for durable kills; index on `expires_at` for pruning.
|
||||
- Modify `cmd/silo/main.go`
|
||||
- Construct the revocation store on both sides: edge (reuse `redisClient` from `main.go:576`, wire `srv.SetRevocationStore(...)` next to `SetRecipeStore` at `:623` and the proxy equivalent); central (reuse `apiRedisClient` `:672`, fall back to in-memory when nil). Call `StartSync` for edges.
|
||||
- Extend the `OnUserSessionsRevoked` closure (`:2047`) to also `Revoke(user:{uid})` in the store, so account-level revocation kills streams too.
|
||||
- Modify `internal/proxy/server.go`
|
||||
- In `verifyToken` (`:126`) / each serve handler: after signature verify, `if revStore.IsRevoked(sid or uid) { close + 403 }`. For direct-play/remux, run a periodic re-check during the long serve (a deadline/ticker goroutine bound to the request context) that aborts `io.Copy`/`ServeFile` and closes the connection when revoked.
|
||||
- Modify `internal/transcodenode/server.go`
|
||||
- Same revocation check on `handleSegment`/`handleManifest`; refuse + stop feeding a revoked session (and tear down its ffmpeg via the existing session teardown, but ALSO keep the revocation so `reconstructFromToken` refuses to rebuild it — otherwise the node self-reconstructs the kill away, `server.go:620`).
|
||||
- Modify `internal/api/handlers/admin_playback_control.go`
|
||||
- In `handleSessionCommand` for `terminate` (`:63`): in addition to the realtime command + `stopPlaybackSessionByID`, call `revStore.Revoke(sess:{sid})` so termination sticks at the edge regardless of client cooperation. Add an explicit `stop`-vs-`terminate` semantic: `terminate` revokes; `stop` remains a cooperative request.
|
||||
- Modify `internal/api/router.go`
|
||||
- Inject the revocation store into the handlers that need it (`Dependencies` field, like `RedisClient`).
|
||||
- Modify `internal/config/config.go`, `internal/config/db_loader.go`, `internal/config/yaml_import.go`
|
||||
- Add `auth.stream_revocation_poll` (default `60s`) and `auth.stream_revocation_ttl` (default `24h`) following the `jellyfin_compat.session_ttl` precedent. Decide restart-required vs hot-reload (`internal/config/restart_keys.go`); prefer hot-reload.
|
||||
|
||||
### Phase 3 — Async decision engine (the brain)
|
||||
|
||||
- Create `internal/playback/enforcer.go` (or `internal/worker/stream_enforcer.go`, following the `internal/worker/cleanup.go` ticker precedent)
|
||||
- A background loop (central only) that every N seconds: pulls the live snapshot from `streammonitor`, groups by `uid`, reads each user's limit via the existing `SessionLimitProvider` / `users.max_streams`, and for any user over cap selects victims (default: most-recently-started beyond the cap) and calls `revStore.Revoke(sess:{sid}, "over_cap")`.
|
||||
- Pluggable rule interface `Rule(snapshot) []KillDecision` so admin/abuse rules compose. Ship two rules: `OverCapRule` and a stub `RestreamHeuristicRule` (documented, disabled by default) that flags sustained N-distinct-title concurrency / throughput anomalies per `uid`.
|
||||
- Structured logging + metrics for every kill: who, why, which rule.
|
||||
- Create `internal/playback/enforcer_test.go`
|
||||
- Deterministic tests: over-cap victim selection, no-op when under cap, idempotent re-revocation, limit-provider error → fail-open (do not kill).
|
||||
- Modify `cmd/silo/main.go`
|
||||
- Start the enforcer loop on central (integrated/api), wire it to `streammonitor` + the limit provider + `revStore`.
|
||||
- Modify `internal/api/handlers/admin_playback_control.go` / a new admin endpoint
|
||||
- Optional: expose the current kill decisions / recently revoked sessions for the admin UI ("why was this killed").
|
||||
|
||||
---
|
||||
|
||||
## Implementation Tasks
|
||||
|
||||
### Phase 1 — Authoritative monitoring
|
||||
|
||||
- [ ] Add `LastServedAt` + `BytesServed` to `nodesessions.SessionInfo`; refresh on real serve activity with a throttled Redis write.
|
||||
- [ ] Wrap edge serve writers (proxy segment/direct/remux) to attribute per-session bytes + keep `LastServedAt` fresh during long pours; extend `egressMeter` with a per-session delta callback.
|
||||
- [ ] Mirror serve-activity tracking on the transcode node (`handleSegment`).
|
||||
- [ ] Integrated mode: record server-side serve activity on native serve paths; stop trusting client progress for liveness.
|
||||
- [ ] Add `internal/streammonitor` aggregator returning the normalized per-`uid` live snapshot (Redis and in-process backends) + tests.
|
||||
- [ ] Surface `last_served_at`/bytes in `HandleListSessions` for the admin view.
|
||||
- [ ] `go build ./... && go vet ./...`; manual check that a playing stream appears with a moving `last_served_at` and disappears within the TTL after it stops — with the client sending no progress reports.
|
||||
|
||||
### Phase 2 — Kill switch
|
||||
|
||||
- [ ] Create `internal/streamrevoke` store: interface + in-memory backend + Redis backend (SCAN warm-up, pub/sub push, ≤60s poll, in-memory hot-path cache) + tests.
|
||||
- [ ] Add durable `revoked_stream_sessions` migration and the Postgres mirror + startup repopulation.
|
||||
- [ ] Wire the store into edges (`SetRevocationStore` on proxy + transcode Server, `StartSync`) and central (reuse `apiRedisClient`, in-memory fallback).
|
||||
- [ ] Enforce on the edge hot path: local `IsRevoked` check per segment (refuse + close); periodic re-check + abort for direct-play/remux long pours.
|
||||
- [ ] Ensure the transcode node does NOT self-reconstruct a revoked session (`reconstructFromToken` consults the store).
|
||||
- [ ] Extend `OnUserSessionsRevoked` to revoke `user:{uid}`; make admin `terminate` revoke `sess:{sid}` (keep `stop` cooperative).
|
||||
- [ ] Add config keys (`auth.stream_revocation_poll`, `auth.stream_revocation_ttl`) via the settings pipeline.
|
||||
- [ ] `go build ./... && go vet ./...`; manual check: revoke a live session → it stops within ~one poll interval and a reconnect with the same token is refused (stays dead); revoke a user → all their streams stop.
|
||||
|
||||
### Phase 3 — Async brain
|
||||
|
||||
- [ ] Create the enforcer loop + pluggable `Rule` interface; implement `OverCapRule` + stub `RestreamHeuristicRule`; fail-open on limit-provider errors.
|
||||
- [ ] Start the enforcer on central; wire snapshot + limit provider + revocation store.
|
||||
- [ ] Structured logging/metrics for every kill; optional admin "recent kills" endpoint.
|
||||
- [ ] `enforcer_test.go` deterministic tests.
|
||||
- [ ] `go build ./... && go vet ./...`; manual check: start N+1 concurrent streams for a capped user → within ~120s the over-cap stream(s) are killed and stay dead; under-cap users are untouched.
|
||||
|
||||
---
|
||||
|
||||
## Testing
|
||||
|
||||
- Unit: `internal/streammonitor` (aggregation/aging), `internal/streamrevoke` (cache/TTL/pub-sub/poll/durability), `internal/playback` enforcer (victim selection, fail-open).
|
||||
- Integration (manual, needs `docker compose up -d postgres redis` + a real client): verify authoritative monitoring with a client that sends no progress; verify kill-within-poll-interval + stays-dead for transcode, direct-play, and remux; verify multi-node (central issues kill, edge enforces) via the shared Redis.
|
||||
- Regression: existing playback is unaffected when nothing is revoked (the only hot-path addition is a local set lookup that returns false).
|
||||
|
||||
## Risks / follow-ups
|
||||
|
||||
- **Enforcement latency (~≤120s)** is intentional and accepted; document it so operators expect a short window before a kill takes effect.
|
||||
- **Direct-play long GET** requires a periodic in-serve revocation check + connection cut; a naive single-GET client with no resume may see an error rather than a clean stop — acceptable for a killed (abusive/over-cap) stream.
|
||||
- **Redis flush/restart** would drop in-flight revocations unless the Postgres mirror is in place; keep the mirror for durability (reliability-first).
|
||||
- **Integrated single-node without Redis** uses the in-memory revocation store and in-process monitoring; the durable mirror still applies. No cross-node concern there.
|
||||
- **Abuse heuristics** (`RestreamHeuristicRule`) ship disabled; tune against real traffic before enabling to avoid false kills.
|
||||
- **Bytes attribution** on the edge is best-effort for monitoring/heuristics; it must never gate the hot path.
|
||||
|
||||
## As-built deltas from this plan
|
||||
|
||||
The shipped implementation follows this plan; a few names/shapes differ (the coverage matrix is authoritative for the current state):
|
||||
|
||||
- The durable mirror landed as table `stream_revocations(kind, id, reason, revoked_at, expires_at)` (PK `(kind,id)`) via `internal/streamrevoke/durable_postgres.go`, and is **wired central-side and non-optional** — warmed on boot (bounded retry) and re-armed into Redis on the poll tick so multi-node edges reconverge after a Redis flush.
|
||||
- Enforcement is centralized in two shared helpers `streamrevoke.Store.Refuse` (403) and `WatchAndCut` (in-flight connection cut) used by every serve surface, rather than per-handler ad-hoc checks.
|
||||
- **Expiry is monotonic on every write.** `applyLocal` and the durable `Upsert` keep whichever copy expires later (`GREATEST`), so the async enforcer's short 5m self-heal TTL can never shorten a longer admin `RevokeSession` on the same session key (which would otherwise reopen the restart-resurrection window the durable mirror closes).
|
||||
- **In-flight cut required a middleware fix.** `WatchAndCut` uses `http.NewResponseController(w).SetWriteDeadline`, which walks `Unwrap()`. The native and jellycompat middleware `ResponseWriter` wrappers lacked `Unwrap()`, so the integrated cut was a guaranteed no-op until `Unwrap()` was added to `statusWriter`, `requestStatusWriter`, `loggingResponseWriter`, `compatImageProxyTagResponseWriter`, and `debugResponseWriter`.
|
||||
- **Ownership carry-forward in the monitor.** `streammonitor.mergeStreams` recovers a resolved `UserID`/`ProfileID`/`MediaFileID` (and backfills route/client) across the same-session dedupe, so the transcode node's ownerless start record can't bucket a stream under user 0 and exempt it from the cap.
|
||||
- **Existence vs timing, made explicit.** Monitoring separates *existence* (server-observed on every path — the transport-count shield integrated, byte-observed at the edge — so a hidden stream that withholds progress is still counted) from *timing* (client progress is trusted, native fully). The transcode serve paths (native `playback.go` and compat `HandleHLSSegment`) gained a per-segment transport marker so a slow single-segment drain can't be reaped mid-serve and the compat HLS path no longer depends on the client's progress POST for liveness. A buffer-ahead idle gap >45s is a remaining reaper follow-up (VERIFY-4).
|
||||
- **First-class monitoring fields.** Added a `Route` (native/jellycompat) dimension — seeded from `Session.Origin` and carried to edges via a token `Origin` claim — plus client IP/name and position, all surfaced through `SessionInfo`/`LiveStream`. The transcode node's start record is now owner+route+client attributed via `TranscodeStartRequest`. The admin session list unions the in-process integrated sessions (`LiveLocalSessions`) so single-node deployments aren't Redis-blind. Downloads remain an explicit cap exemption (separate quota).
|
||||
- `auth.stream_revocation_poll` / `auth.stream_revocation_ttl` operator config keys are not yet wired (defaults 60s / 24h are hardcoded) — remaining follow-up.
|
||||
|
||||
## AI-use disclosure
|
||||
|
||||
This plan was drafted with AI assistance (Claude), based on a read-only audit of the current codebase. No behavior was changed by writing it. The Status banner and As-built deltas section were added after implementation.
|
||||
Reference in New Issue
Block a user