docs(playback): monitoring & kill-switch plan + as-built coverage matrix

Capture the design intent and the shipped coverage for the stream monitoring +
revocation kill-switch work, refreshed to match the final implementation and the
monitoring → kill-switch → docs commit layout.

- docs/superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md: the
  implementation plan of record (detection / enforcement / async brain split),
  with a Status banner and as-built deltas noting what shipped differently
  (durable Postgres mirror wired non-optional, shared Refuse/WatchAndCut
  helpers, monotonic expiry, ownership carry-forward, first-class monitoring
  fields).
- docs/architecture/playback-paths-monitoring-kill-matrix.md: the as-built
  coverage matrix across server layout × playback type × route, GAP-1..GAP-9
  resolution notes, restart-durability axis, and open verification items
  (VERIFY-3/4, operator config keys, compat download quota).

Part of the stream monitoring & kill-switch epic.
This commit is contained in:
CoffeeKnyte
2026-07-29 12:26:03 +00:00
parent 0217cce2df
commit 53c9313b57
2 changed files with 612 additions and 0 deletions
@@ -0,0 +1,433 @@
# Playback Paths — Monitoring & Kill-Switch Coverage Matrix
> Reference checklist for stream monitoring + revocation kill-switch coverage.
> Check every change to playback/monitoring/revocation against this. As-built
> companion to the design plan
> [`2026-07-04-stream-monitoring-and-kill-switch.md`](../superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md)
> (that doc is the intent; this one is the shipped coverage). Status as of
> 2026-07-05: monitoring, enforcement, and the async over-cap enforcer are all
> shipped across two code commits — the *monitoring* commit (`feat(playback):
> server-observed stream monitoring`) and the *kill-switch* commit
> (`feat(playback): stream kill switch + async over-cap enforcer`). GAP-1..GAP-4
> are all RESOLVED (see Findings): GAP-4 (kill list did not survive a restart) was
> closed by wiring a durable Postgres mirror, and GAP-3 (in-flight cut for
> integrated direct-play) was closed by adding `Unwrap()` to the native/compat
> middleware writers so the cut reaches the socket. A post-review hardening pass
> then closed GAP-5..GAP-9 found by adversarial re-review: user-kill cutoff
> semantics, compat manifest/subtitle/mint-bypass coverage, terminate-by-id,
> in-flight download cut, admin-list dedupe, and the metered-writer sendfile chain
> (see Findings). That hardening is folded into the two code commits above — the
> branch is organized as monitoring → kill-switch → this docs commit, not as a
> running series of fix commits. (This doc references sibling commits by role, not
> SHA — squash/rebase rewrites hashes, so a SHA citation goes stale the moment the
> branch is amended.)
## Dimensions
- **Server layouts:** (A) integrated, no Redis · (B) integrated, with Redis ·
(C) multi-node, with Redis (proxy/transcode edge nodes).
- **Playback types:** direct play · remux · transcode (HLS).
- **Routes:** native silo `/api/v1` · jellycompat (`:8096`).
## Key runtime facts (why cells collapse)
- **Integrated (A, B) serves bytes locally.** Node selection returns an empty
plan when no edge nodes exist, and both native and jellycompat fall through to
in-process serving (`nodepool/planner.go:248`, `jellycompat/handlers_playback.go:306`
— the `proxyNode == nil` guard in `buildProxyRedirectURL`). A vs B differ only
in *revocation propagation* (memory-only vs Redis) and *monitor source*
(FuncSource vs MultiSource) — the **serve + enforcement points are identical**.
So A ≡ B for this matrix; treated as one column "Integrated".
- **Multi-node (C) serves bytes on edges.** Native playback hands the client a
proxy-node URL at session start (`playback.go:1575`); jellycompat 302-redirects
to a proxy node (`streams.go:120`). Edge serving = `internal/proxy` + `internal/transcodenode`.
- **jellycompat shares the SessionManager.** A jellycompat play starts a real
`playback.SessionManager` session (`streams.go:1127`), so the monitor/enforcer
see it exactly like native. The compat `PlaybackSessionStore` is separate
bookkeeping (PlaySessionId → recipe), not liveness/enforcement.
- **Local transcode fallback:** in multi-node, if no transcode node is available,
jellycompat falls back to LOCAL transcode (`streams.go:277`,
`LocalTranscodeFallbackAllowed`) — i.e. the "Integrated jellycompat" serve path
can also occur under layout C.
## The matrix
Legend: ✅ covered · ⚠️ covered with caveat · ❌ gap.
### Serving handler (where the bytes come from)
| Route | Type | Integrated (A/B) | Multi-node (C) |
|---|---|---|---|
| native | direct | `StreamHandler.HandleStream` → `ServeDirectPlay` (`stream.go:141`) | proxy `handleDirectPlay` (`proxy/server.go:171`) |
| native | remux | `HandleStream` → `ServeRemux` (`stream.go:157`) | proxy `handleRemux` |
| native | transcode | `HandleGetTranscodeSegment` local → `ServeFile` (`playback.go:2860`) | proxy → transcode node `handleSegment` |
| jellycompat | direct | `HandleVideoStream` → `ServeDirectPlay` (`streams.go:153`) | 302 → proxy `handleDirectPlay` |
| jellycompat | remux | `HandleVideoStream` → `ServeRemux` (`streams.go:151`) | 302 → proxy `handleRemux` |
| jellycompat | transcode | `HandleHLSSegment` → `ServeFile` (`streams.go:543`) | 302 → proxy → transcode node |
### Monitoring (server-observed existence, reported to central)
The monitoring model separates **existence** (must be server-observed, never
gated by a client report — the hidden-stream defense) from **timing** (client
progress is fine, especially for native). A disguised re-streamer that pulls
bytes but withholds/falsifies progress must still be counted.
| Route | Type | Integrated (A/B) — existence signal | Multi-node (C) — existence signal |
|---|---|---|---|
| native | direct | ✅ transport-count shield for the whole pour (`stream.go:136`) | ✅ edge tracker, byte-observed (`sessionByteWriter`) |
| native | remux | ✅ transport-count shield (`stream.go:146`) | ✅ edge tracker, byte-observed |
| native | transcode | ✅ per-segment transport marker around `ServeFile` (`playback.go:2860`) | ✅ edge tracker (Touch + AddBytes/segment) |
| jellycompat | direct | ✅ transport-count shield (`streams.go:129`) | ✅ edge tracker, owner+route carried |
| jellycompat | remux | ✅ transport-count shield | ✅ edge tracker, owner+route carried |
| jellycompat | transcode | ✅ per-segment transport marker around `ServeFile` (`streams.go:543`) | ✅ edge tracker, owner+route carried |
- **Existence is server-observed on every cell.** Integrated: a session is
unreapable while `activeTransportCount > 0`, and every byte-serving path now
holds that marker — direct/remux for the whole pour, transcode for each segment
serve (the segment markers were added to close a hidden-stream hole where a slow
single-segment drain past the 45s grace could be reaped mid-serve; the compat
HLS path previously refreshed liveness *only* from the client's progress POST,
a direct violation, now also transport-marked). Edge: a Redis record exists for
the whole connection (direct/remux) or while segments are pulled (transcode),
advanced only by real bytes. **No path requires a client progress report to stay
visible or counted.**
- **Timing is secondary and may trust the client.** Native fully trusts client
progress for position; the integrated `LastServedAt` is mapped from
`SessionManager.LastActivityAt` (`cmd/silo/main.go` FuncSource → `handlers.LiveLocalSessions`),
which client progress can advance. This never inflates the **count** (existence
is the transport shield, not a heartbeat), so the cap can't be bypassed; it only
affects which of a user's *own* over-cap sessions `selectVictims` trims first.
The edge advances `LastServedAt` only on real bytes.
- Report-to-central: **C** = async via Redis (edge writes `silo:sessions:*`,
`streammonitor.RedisSource` reads). **A/B** = in-process SessionManager IS
central; read via `streammonitor.FuncSource`. MultiSource unions both, deduped
by session id.
- Owner + attribution: the transcode node's start record now carries owner + route
+ client (threaded via `TranscodeStartRequest`), and `streammonitor.mergeStreams`
additionally backfills any missing owner/route/client from another record for the
same session — so an ownerless-but-freshest node record can never bucket a stream
under user 0 (which the enforcer skips) or drop route/client from the view.
### Monitoring data model (what central sees per stream)
First-class monitoring goal: see **all** active streams with enough to act on.
Captured per live stream (`streammonitor.LiveStream` / `nodesessions.SessionInfo`):
| Field | Source | Notes |
|---|---|---|
| existence, session id | both | primary; server-observed |
| user id, profile id | token claims / Session | primary; owner attribution |
| media file id | claims / Session | primary; human title is a display-time lookup (follow-up) |
| play method (direct/remux/transcode) | `Type` | primary |
| **route (native/jellycompat)** | `Route` — `Session.Origin` (integrated) or token `Origin` claim (edge) | primary; native and jellycompat share the SessionManager and `Type`, so route is a distinct field seeded at the two `WithClientInfo` sites and carried to edges in the token |
| node serving | NodeName/NodeURL | integrated stamps the local host |
| client ip / client name | `Session.ClientIP/ClientName` (integrated) or edge request + `ClientName` claim | secondary; also the natural re-streaming fingerprint |
| position | `Session.Position` | secondary timing |
| bytes served | edge `AddBytes` | edge only; throughput signal |
| hw/sw, resolution, codecs | Session / node record | secondary |
- **Admin visibility:** `HandleListSessions` unions the Redis edge records with the
in-process integrated sessions (`LiveLocalSessions`), so a single-node integrated
deployment is no longer blind (previously Redis-only). The union is deduped by
session id (`streammonitor.DedupeSessionInfos`, mirroring `mergeStreams`) so a
stream tracked by both the central manager and the edge serving it shows as ONE
row — matching the enforcer's count. A `node_id` filter targets
an edge, so integrated sessions appear only in the unfiltered listing.
- **Downloads are an intentional exemption — with one asymmetry.** `/downloads/*`
(native) and compat `/Items/{id}/Download` serve full media with no
session/monitor record and do NOT count against the live-stream cap. The native
route requires a quota-checked download row (`internal/downloads`
concurrency/period limits); the compat route requires only compat auth — it is
covered by NO quota today (bringing it under an Infuse-compatible download
quota is an open follow-up). Both routes now arm the shared `WatchAndCut`, so a
per-user stream revocation cuts an in-flight download pour, and the same hook
deletes every compat login so reconnects need re-auth. Documented at the
handlers.
### Kill switch (revocation enforced on the serve path)
| Route | Type | Integrated (A/B) | Multi-node (C) |
|---|---|---|---|
| native | direct | ✅ `guardRevocationCut` (Refuse + in-flight cut†) | ✅ `verifyToken` + `cutOnRevocation` (in-flight cut) |
| native | remux | ✅ `guardRevocationCut` (Refuse + cut†) | ✅ verifyToken + cutOnRevocation |
| native | transcode | ✅ `guardRevocation` (refused within one segment) | ✅ proxy verifyToken + node `refuseIfRevoked` + reconstruct guard‡ |
| jellycompat | direct | ✅ `Refuse` + in-flight cut† (GAP-1/3 fixed) | ✅ via proxy; ✅ local fallback guarded |
| jellycompat | remux | ✅ `Refuse` + in-flight cut† | ✅ via proxy; ✅ local fallback guarded |
| jellycompat | transcode | ✅ `Refuse` per segment (GAP-1 fixed) | ✅ via proxy/node; ✅ local fallback guarded |
† In-flight cut uses `streamrevoke.Store.WatchAndCut` (SetWriteDeadline, checked on
entry then every 5s). Works at the edge (the proxy's metered writer implements
`Unwrap`) **and** integrated: the native (`statusWriter`, `requestStatusWriter`)
and compat (`loggingResponseWriter`, `compatImageProxyTagResponseWriter`,
`debugResponseWriter`) middleware writers now implement `Unwrap()`, so
`http.NewResponseController` reaches the socket instead of no-oping (see GAP-3). A
live-client smoke test is still worthwhile, but the previously-guaranteed
integrated no-op is fixed; worst case it still degrades to a next-request `Refuse`.
‡ The transcode node's serve-path check is session-only (`refuseIfRevoked` passes
`userID = 0`), so a per-*user* kill is enforced by the fronting proxy's `verifyToken`
(which passes `claims.UserID`), not at the node's segment serve. The node's
reconstruct guard *does* use the real `claims.UserID`, so a rebuild-after-restart of
a user-killed session is blocked.
### Restart durability (kill survives a process restart / Redis loss)
This axis exists because the branch sits on top of PR #174 (restart-resilient
playback): a session killed before a restart is **reconstructed** afterward from a
durable recipe card, so if the kill list does not also survive the restart the
reconstructed stream is silently re-served. The kill list must be at least as
durable as the thing it kills.
| Deployment | Session survives restart (PR #174) | Kill survives restart | Kill survives Redis flush |
|---|---|---|---|
| A — integrated, no Redis | ✅ recipe card (PG) | ✅ durable PG mirror | ✅ (never used Redis) |
| B — integrated, with Redis | ✅ | ✅ durable PG mirror (+ Redis warm) | ✅ durable PG mirror |
| C — multi-node, with Redis | ✅ | ✅ central durable PG mirror; edges re-warm from Redis | ✅ central re-warms edges from PG→Redis on next write/poll |
\* Steady-state guarantee. Transient boot caveat: if the durable warm exhausts its
bounded retry **and** Redis is empty, the kill list is empty until the first poll
tick (≤60s). This fails *open* by design — see the "Warm is async-tolerant" bullet
below.
- **Hot path unchanged.** `IsRevoked` is still a pure in-memory map read. Postgres
is touched only on write (`Upsert`), on warm/reconcile (`ListActive` at
`StartSync` and on the poll tick), and on trim (`Prune`) — never per request.
- **Central-side only.** The durable mirror is wired into the integrated/api
`streamrevoke.Store` (`cmd/silo/main.go`), not the edge/proxy/transcode nodes,
which have no app DB and enforce via Redis pub/sub + poll.
- **Bounded growth.** Rows are keyed by `(kind, id)`, so the async enforcer
re-revoking every pass UPSERTs one row rather than accumulating; expired rows are
physically reclaimed by `Prune` on the poll tick, and `ListActive` filters on
`expires_at > now()`.
- **Expiry is monotonic.** Both the in-memory `applyLocal` and the durable
`Upsert` keep whichever copy expires LATER (`GREATEST`), never shortening an
existing kill. This matters because the async over-cap enforcer re-revokes with a
short 5m self-heal TTL on the same `KindSession` key an admin may have killed for
24h; without monotonic expiry the enforcer would silently shrink the admin kill
to 5m and reopen the restart-resurrection window this whole axis defends against.
- **Warm is async-tolerant.** By product decision, a just-reconstructed revoked
stream may serve a few seconds before the durable warm/next Refuse tick cuts it;
no blocking startup ordering is required. Edge caveat: if the boot durable warm
exhausts its bounded retry **and** Redis is empty, the kill list is empty until
the first poll tick (≤60s). This fails *open* by design (a kill switch cannot
fail closed without blocking all playback when the DB is briefly unavailable).
- **Edge Redis-outage fails open.** Edge nodes (proxy/transcode) have no durable
store and learn *new* kills only from Redis pub/sub + SCAN. While Redis is
unreachable, already-cached kills persist but a kill issued *during* the outage
does not reach the edge until Redis recovers — enforcement fails open there.
## Findings
> GAP-1..GAP-3 all shipped in the kill-switch commit (see the Status banner for
> how commits are referenced by role rather than SHA).
**GAP-1 — RESOLVED.** jellycompat LOCAL serving now consults the
shared `Store.Refuse` in `HandleVideoStream` (direct/remux) and `HandleHLSSegment`
(per segment), closing the integrated / local-transcode-fallback hole.
**GAP-2 — RESOLVED.** `buildProxyRedirectURL` now stamps
`uid`/`pid`/`mfid` onto the jellycompat proxy-redirect token (looked up from the
upstream SessionManager session), so edge monitoring/kill attribution matches
native. `AuthUserID = 0` no longer occurs for jellycompat.
**GAP-3 — RESOLVED.** The in-flight cut is shared as
`streamrevoke.Store.WatchAndCut` and applied to native `/stream`
(`guardRevocationCut`) and jellycompat `HandleVideoStream`. The original caveat —
`SetWriteDeadline` might not reach the socket through the native/compat middleware
chain — was **confirmed** as a guaranteed integrated no-op (four middleware writers
wrapped `w` without `Unwrap()`) and then fixed by adding `Unwrap()` to
`statusWriter`, `requestStatusWriter`, `loggingResponseWriter`,
`compatImageProxyTagResponseWriter`, and `debugResponseWriter`. The cut now reaches
the socket in integrated mode; if it ever no-ops again the stream still stops on
its next request via `Refuse` (no regression). See VERIFY-1.
**GAP-4 — RESOLVED.** The kill list did not survive a process restart or a Redis
flush, but PR #174's restart-resilient playback *does* reconstruct the session — so
a stream revoked before a restart was silently re-served afterward. This was
universal in deployment A (integrated, no Redis: the kill list was pure RAM) and
occurred on any Redis loss in B/C. Fixed by wiring a concrete Postgres
`DurableStore` (`internal/streamrevoke/durable_postgres.go`, table
`stream_revocations`) into the central `streamrevoke.Store`: `Revoke` mirrors to
Postgres, `StartSync` warms from it on boot, and the poll tick re-warms (heals a
Redis flush) and `Prune`s expired rows. The `defaultTTL` (24h) is held `>=` the
recipe-card `MaxTokenTTL` (24h) as an invariant so a kill cannot expire before its
session can be reconstructed, and expiry is **monotonic** on every write (in-memory
`applyLocal` + durable `Upsert` `GREATEST`) so the async enforcer's short 5m
re-revoke can never shorten a longer admin kill on the same session key. Hot path
(`IsRevoked`) stays an in-memory read.
**GAP-5 — RESOLVED (post-review).** User-kind revocations were a blanket ban:
any admin user edit (via `OnUserSessionsRevoked`) 403'd that user's playback for
24h even after re-login, unrevocably. Fixed with cutoff semantics — see
"user-ID kills" above.
**GAP-6 — RESOLVED (post-review).** Compat coverage holes: `HandleMasterManifest`
/ `HandleHLSManifest` had no revocation check and kept (or, post-restart,
re-spawned) a killed session's ffmpeg via the ensure path; `HandleVideoStream`
checked revocation only after `ensureUpstreamPlayback`, which can replace an
unreconstructable killed session with a fresh id that passed the check (kill
dodged by re-hitting the URL); `HandleSubtitleStream` ignored kills entirely.
All four now `Refuse` up front.
**GAP-7 — RESOLVED (post-review).** Admin terminate 404'd before revoking when
the session was absent from the central in-memory manager — exactly the
edge-served / post-restart / progress-withholding streams the kill switch
exists for (the admin list showed them via the Redis union; terminate couldn't
touch them). Terminate now writes the revocation FIRST, keyed on the session id
alone, and answers `202 {status: "revoked"}` when there is no local session for
the cooperative command.
**GAP-8 — RESOLVED (post-review).** The admin session list double-counted every
edge-served stream (central manager row + edge Redis row, no dedupe), unlike
the enforcer's merged picture. Now deduped by session id with owner/attribution
carry-forward (`streammonitor.DedupeSessionInfos`).
**GAP-9 — RESOLVED (post-review).** The "restore sendfile" commit was dead code
end-to-end: every stream route runs inside `meterEgress`, and
`meteredResponseWriter` hid `io.ReaderFrom`, so `sessionByteWriter.ReadFrom`'s
fast path never fired and all direct-play/remux bytes went through a userspace
copy. `meteredResponseWriter` now forwards `ReadFrom` (metering the returned
total); a chain test locks the full production writer stack onto the sendfile
path. Related tracker fixes in the same pass: `Touch`/`Track` preserve the
first-seen `StartedAt` (was reset per segment/range request, corrupting the
enforcer's tie-break and the admin start time), `AddBytes` no longer recreates
entries for pruned sessions (slow permanent leak), the proxy→node segment pour
attributes bytes incrementally (a slow drain can't go invisible mid-segment or
post bytes to a dead record), and the transcode node marks serve activity so
its record's `LastServedAt` is no longer frozen at start time.
**GAP-3 (original, superseded) — no in-flight cut for native/compat LOCAL direct-play.**
The original finding: `guardRevocation` refused the next request/reconnect but could
not hang up an in-flight `ServeFile` on the central process for a single long GET.
Now closed — see GAP-3 above (shared `WatchAndCut` + middleware `Unwrap()`). Even
if the cut degrades, admin terminate also fires the realtime command + session stop,
so a killed direct-play stops on its next request at worst.
## Open verification items (check before prod)
- [x] **VERIFY-1 — in-flight cut reaches the socket on native + jellycompat — RESOLVED.**
`streamrevoke.Store.WatchAndCut` cuts a revoked long-GET direct-play/remux by
setting a zero write deadline via `http.NewResponseController(w)`, which walks the
writer chain via `Unwrap()`. This was **confirmed broken** in integrated mode:
`statusWriter`/`requestStatusWriter` (native) and `loggingResponseWriter`/
`compatImageProxyTagResponseWriter`/`debugResponseWriter` (compat) wrapped `w`
without `Unwrap()`, so `SetWriteDeadline` returned `ErrNotSupported` and the cut
no-oped. Fixed by adding `Unwrap() http.ResponseWriter` to all five (pattern from
`activitylog/middleware.go:186`). A live smoke test is still worthwhile: start a
native and a jellycompat single-long-GET direct-play, admin-terminate it, confirm
the connection drops within ~5s rather than only on the next request.
- [x] **VERIFY-2 — download routes vs the stream cap — RESOLVED (documented exemption).**
`/downloads/*` (native) and compat `/Items/{id}/Download` serve full media with no
session/monitor record; they are intentionally exempt from the live-stream cap and
governed by the separate download concurrency/period quota. Documented at both
handlers and in the monitoring data model above. (Note: they also do not consult
the kill switch — a per-user *stream* revocation does not stop an in-flight
download. If a ban should also cut downloads, guard those handlers separately.)
- [ ] **VERIFY-4 — transcode buffer-ahead evasion (HOLE 2).** A transcode client
(native or compat) that buffers far ahead and pauses segment fetches for >45s (the
active grace) is transiently reaped, then self-heals on the next fetch — so a
determined abuser could time buffer-pauses to duck under the cap at the enforcer's
tick. The per-segment transport marker fixes the mid-serve case but not the idle
gap between segments. Fix needs a reaper change: treat a running transcode ffmpeg
as liveness, or a transcode-specific grace ≥ the client's max buffer. Deferred
because it touches the sensitive session-reaping path (cf. #279).
- [ ] **VERIFY-3 — multi-replica central enforcement.** The async enforcer runs on
every integrated/api process. Two *integrated* replicas behind a load balancer
each see only their own in-process `FuncSource` (integrated mode writes no
`silo:sessions:*` records), so a user split across both can exceed `max_streams`
undetected; multiple *api* replicas each run an independent enforcer and revoke
the same victims (idempotent, but redundant write load). A single-leader election
for the enforcer, or having integrated nodes publish their local sessions to the
shared Redis picture, would close this.
## Design assessment (shared vs duplicated code)
**Kill DECISION is well-centralized — keep it.** All revocation state and the
`IsRevoked(sessionID, userID)` check live in one place (`internal/streamrevoke`).
Session kills and user kills use the SAME store and the SAME check — `RevokeSession`
and `RevokeUser` write into one `items` map; `IsRevoked` tests both keys. No
duplication at the decision layer.
**Kill ENFORCEMENT — DONE.** The "extract sid/uid → refuse (403)"
step is now a single shared method `streamrevoke.Store.Refuse(w, sessionID, userID)
bool`. All four serve surfaces call it (proxy `verifyToken`, transcode node
`refuseIfRevoked`, native `guardRevocation`, jellycompat serve handlers); each only
owns its token extraction (path `{token}`, query `st`, or compat `PlaySessionId`).
The in-flight connection-cut is the shared `WatchAndCut` (SetWriteDeadline), used by
both the edge (`cutOnRevocation`) and the native/compat long-pour paths. ONE section
per concern, as intended.
**Monitoring WRITE side is two implementations by necessity — leave it.** Edges
have no SessionManager; the integrated process has no tracker. `nodesessions.Tracker`
(edge) and `SessionManager` (integrated) are different runtimes, not duplicated
logic, and they already converge at the single read layer (`streammonitor`,
unioned by MultiSource). Forcing one implementation would be worse.
**user-ID kills — plumbing check (as asked).** Two distinct layers, correctly
separate:
- *Auth/login revocation* (pre-existing): `OnUserSessionsRevoked` revokes native
auth sessions (`auth.SessionRepository.RevokeAllByUser`) and drops compat login
tokens (`SessionStore.DeleteByUserID`). This governs "can this user authenticate",
not "stop this byte stream".
- *Stream revocation* (new): `streamRevocation.RevokeUser` writes a `KindUser`
entry into the SAME `streamrevoke` store as session kills. Session-level and
user-level STREAM kills already share all plumbing (one store, one `IsRevoked`).
- **A user kill is a CUTOFF, not a ban (GAP-5 fix).** `IsRevoked` matches a
`KindUser` entry only when the stream's credential predates the revocation:
the stream token's `iat` on token-bearing surfaces (edge proxy, transcode
node, native `?st=`), the request entry time on freshly-authenticated
surfaces (native session auth, compat login — every pre-revocation login is
reset by the same hook, so reaching a serve path afterward proves fresh
auth). An in-flight pour always predates a later revocation, so mid-pour
user kills still cut. An unknown (zero) credential time never matches (fail
open). This is what makes it safe for `OnUserSessionsRevoked` — which fires
on ANY admin edit of password/role/enabled/permissions/quality — to also
write a stream kill: without the cutoff, a routine permission tweak 403'd
the user's playback for the full 24h TTL after re-login, with no unrevoke
path (expiry is deliberately monotonic).
- These two layers should NOT be merged — conflating "may log in" with "this
stream must die" would couple auth and playback. They are chained in the same
hook, which is the right seam.
## Recommended follow-up order
GAP-1, GAP-2, GAP-3, and GAP-4 are all shipped (see Findings). Already-shipped
hardening not to re-do: `Revoke` propagation uses `context.WithoutCancel`, so an
aborted admin request can no longer strand a kill in central memory only.
Remaining work:
1. **Operator config keys (small):** wire `auth.stream_revocation_poll` /
`auth.stream_revocation_ttl` through the settings pipeline; defaults (60s / 24h)
are currently hardcoded in `internal/streamrevoke/store.go`.
2. **Transcode buffer-ahead liveness (VERIFY-4 / HOLE 2):** treat a running ffmpeg
as liveness (or a transcode-specific grace) so a buffered-ahead transcode isn't
transiently reaped and can't time buffer-pauses to duck the cap.
3. **Media title enrichment:** resolve `MediaFileID → title` at admin display time
so the view shows *what* is being watched, not just a numeric id (also enables a
distinct-title re-streaming heuristic).
4. **Multi-replica enforcement (VERIFY-3):** single-leader election for the async
enforcer, or have integrated nodes publish local sessions to the shared Redis
picture so the cap holds across replicas.
5. **Self-heal a failed durable write:** `Revoke` mirrors to Postgres best-effort;
a transient DB error at revoke time is logged but never repaired (the maintain
tick only reads *from* durable, never re-mirrors live in-memory kills *to* it),
so that one kill is lost on the next restart. Add a memory→durable re-mirror on
the poll tick.
6. **Invariant guard test:** `defaultTTL >= playback.MaxTokenTTL` is coupled only by
a prose comment (streamrevoke can't import playback — import cycle). Add a test
in a package that can import both, so raising `MaxTokenTTL` can't silently
violate it.
7. **Transcode-node ghost records:** the node's session-backed `Track` record is
refreshed forever until an explicit stop/cleanup; if central dies or loses the
session without the stop reaching the node, an owner-attributed ghost persists
in Redis indefinitely — permanently +1 in the user's live count (the enforcer's
freshness ordering trims the ghost first, so real streams survive, but the
enforcer re-revokes it every pass and the admin view overcounts). Needs a
node-side idle sweep tied to real serve activity (`MarkServed` now provides the
signal).
8. **Compat download quota:** compat `/Items/{id}/Download` is exempt from the
stream cap AND covered by no download quota (native downloads require a
quota-checked download row; the compat route requires only compat auth). An
Infuse-compatible quota (plain GETs, Range requests, no download rows) is open
product/design work; today's controls are compat auth, the library access
filter, and the in-flight user-revocation cut.
@@ -0,0 +1,179 @@
# Stream Monitoring & Kill-Switch Implementation Plan
> **Status: IMPLEMENTED (with a minor operator-config follow-up open)** on branch `feat/sauron-async-enforcer`. All three phases shipped — Phase 1 (authoritative server-side monitoring), Phase 2 (revocation kill switch + edge/native/jellycompat enforcement), and Phase 3 (async over-cap enforcer). The one still-open item is operator ergonomics, not a phase: the `auth.stream_revocation_poll` / `auth.stream_revocation_ttl` keys are not yet wired through the settings pipeline (defaults 60s / 24h are hardcoded) — see "As-built deltas" and the coverage matrix's follow-up list. A durable Postgres mirror of the kill list was added **after** this plan was written, so kills now survive a server restart or a Redis flush (this plan's "optional Postgres durability" is no longer optional — it closes the gap where restart-resilient playback could reconstruct and re-serve an already-killed stream). For the **as-built** coverage across server layout × playback type × route, the resolved gap findings (GAP-1..GAP-4), and the open verification items, see the companion reference: [`playback-paths-monitoring-kill-matrix.md`](../../architecture/playback-paths-monitoring-kill-matrix.md). The task checkboxes below are retained as the original plan of record.
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Do not edit this plan file while implementing it.**
**Goal:** Gain control over stream abuse (concurrency-cap violations, admin termination, Stremio-style re-streaming) **without** putting the central database on the per-segment hot path and **without** changing the client-facing token protocol (the later-added `Origin` route claim is server-signed and opaque to clients — they only echo the token, so the wire protocol clients see is unchanged). Do it with two primitives plus an off-hot-path decision loop: (1) authoritative server-side monitoring that never trusts the client, and (2) a kill switch that stops any stream within ~120s and keeps it dead.
**Architecture:** Split the problem into *detection* and *enforcement*, and keep both off the hot path.
- **Detection = authoritative monitoring (no client trust).** Bytes can only reach a viewer by flowing through the edge (segment serving) or a held-open direct-play socket that the edge itself is writing. So the server *inherently* knows every live stream — it is the one sending the bytes. We make the existing `nodesessions` records the source of truth, refreshed by *real serve activity* (segment served / active transport / open connection), enriched with `last_served_at` and bytes, and mirror the same signal into integrated mode. No client progress report is trusted for liveness. A stream that "plays quietly" is still pulling segments (or holding a socket) and is therefore visible; one that stops pulling ages out on its own.
- **Enforcement = a revocation kill switch.** A small revocation set lives in Redis (multi-node) or in-memory (integrated), mirrored to Postgres for durability (shipped as a wired part of the kill path, not an optional add-on — it is what lets a kill survive a restart/Redis flush). Edges cache it in memory, refreshed by pub/sub push (near-instant) plus a ≤60s poll (safety net). The **only** hot-path addition is a local in-memory set-membership check per segment — no per-segment Redis or Postgres. On a hit the edge refuses the request and cuts the connection; for direct-play the edge periodically re-checks during the long GET and aborts. A kill "stays dead" because the revocation entry outlives the token and every reconnect is refused.
- **Decision = an async brain, fully off the hot path.** A central background evaluator reads the live monitoring picture and the per-user limits (`users.max_streams`), and decides what to kill: over-cap victims, admin-terminated sessions, and (pluggable) abuse heuristics. Every reason — exceeded limit, admin terminate, Stremio abuse — collapses to the *same* action: write a revocation. Enforcement follows within ~120s. The delay is acceptable for a media server.
The client-facing token **wire protocol is unchanged**. Its `sid`/`uid`/`pid` claims are the revocation keys; the server may add opaque server-signed claims (e.g. the later-added `Origin` route claim) that clients only echo. Nothing about mint sites, segment URLs, HLS manifests, or client behavior changes. This is why the approach is low-risk relative to a short-ticket redesign (see "Background — why not short tickets").
**Tech Stack:** Go, chi routers (proxy/transcode edge + native api), `redis/go-redis/v9`, Postgres via pgx + Goose migrations, the existing `nodeconfig.Watcher` settings pipeline, the `cache` package Redis client + `EventBus` pub/sub.
---
## Commands
Commands assume the repository root is the cwd.
- Build: `make build`
- Backend only: `go build ./... && go vet ./...`
- Lint: `make lint`
- New migration: `make migrate-create NAME=revoked_stream_sessions`
- Local services: `docker compose up -d postgres redis`
---
## Background — what we learned (why this shape)
Audited integration points this plan builds on (do not re-derive):
- **Stream token** is a stateless signed JWT, 24h TTL, verified at the edge by signature only, no store, no revocation today (`internal/streamtoken/token.go`; mints at `internal/api/handlers/playback.go:517,1601,2159` and `internal/jellycompat/handlers_playback.go:317`). Claims carry `sid` (`SessionID`), `uid`, `pid`, `mfid` — our revocation keys.
- **Edge already sees every segment and already writes Redis session records.** Proxy `touchTranscodeSession` → `nodesessions.Tracker.Touch` on every manifest/segment (`internal/proxy/server.go:192`; `internal/nodesessions/tracker.go:137`), 60s TTL, refreshed every 30s. Direct-play/remux are tracked for the request lifetime (`proxy/server.go:143-145,156-158`). Records already carry `AuthUserID`/`ProfileID`/`MediaFileID` (`tracker.go:41-43`) but **no** `last_served_at` or bytes yet.
- **Edge egress is already metered** but only as a node aggregate (`internal/proxy/egress.go`), not per session.
- **Edge Redis is mandatory** and constructed once (`cmd/silo/main.go:576-580,605`), reused for the recipe store (`main.go:623`) — reuse the same client for a revocation store via a `SetRevocationStore(...)` wired exactly like `SetRecipeStore`.
- **Central Redis** is `apiRedisClient` (`main.go:672`, may be nil on Redis-less single-node), injected as `deps.RedisClient` (`internal/api/router.go:135`). Central and edges point at the **same** Redis in multi-node (existing `silo:sessions:*` write/read contract proves it).
- **Config auto-propagates to edges** via `nodeconfig.Watcher` reading `server_settings` (`internal/nodeconfig/watcher.go`); follow the `jellyfin_compat.session_ttl` precedent (`internal/config/config.go:217,233`; `internal/config/db_loader.go:381`) for any new key.
- **Pub/sub already exists** on `cache.ChannelAdmin` and the watcher subscribes to it (`watcher.go:83-93`) — reuse this channel/pattern to push revocations to edges near-instantly.
- **Revocation hook exists but is compat-only.** `OnUserSessionsRevoked` (`router.go:162`, `admin.go:101`, impl `cmd/silo/main.go:2047`) currently only deletes Jellyfin-compat sessions. Admin terminate (`internal/api/handlers/admin_playback_control.go:63` `HandleTerminateSession` → `handleSessionCommand`) currently only sends a realtime command + `stopPlaybackSessionByID` fallback; it does **not** revoke the stream credential at the edge. These are the exact hooks to extend.
- **Concurrency counting already exists in-process** (`internal/playback/session.go` `SessionManager`, per-user `activeCountLocked`, `SessionLimitProvider` reading `users.max_streams`) but only for the local API process and only at admission. The async brain reuses the *limit provider* but computes the live count from the monitoring picture instead of the in-process map, so it works across nodes.
**Why not short tickets (the rejected alternative):** turning the 24h token into a short-lived play ticket would enforce the cap continuously, but (a) Silo's `master.m3u8` is a **VOD** playlist fetched once (`internal/playback/transcode.go:1231` `GenerateFullManifest`), so short tickets embedded in segment URLs would expire mid-movie and 401 the rest of the stream — the "renewal rides on playlist refresh" premise does not hold; (b) flipping the global token TTL is a break-all-playback change that cannot be behaviorally verified without real clients. The monitor-and-kill approach accepts a ~120s enforcement delay in exchange for zero client-protocol change and zero hot-path DB/Redis.
---
## File Structure
### Phase 1 — Authoritative monitoring (no client trust)
- Modify `internal/nodesessions/tracker.go`
- Add `LastServedAt string` and `BytesServed int64` to `SessionInfo`.
- Refresh `LastServedAt` (and accumulate `BytesServed`) on real serve activity, not just first touch. Keep the Redis write throttled (only re-`SET` when the record materially changes or on the 30s refresh) so this stays cheap.
- Add a `Snapshot(ctx) ([]SessionInfo, error)` helper (edge-local view) if useful for `/status`.
- Modify `internal/proxy/server.go`
- Feed real bytes-served per session into the tracker (wrap the segment/direct/remux response writers, or reuse the `egressMeter` write hook keyed by `claims.SessionID`).
- Ensure direct-play/remux held-open connections keep `LastServedAt` fresh periodically (a ticking update while `io.Copy`/`ServeFile` runs), so a long quiet pour is never invisible.
- Modify `internal/proxy/egress.go`
- Allow the metered writer to report per-session deltas to a callback (so the tracker can attribute bytes), in addition to the existing node-aggregate rate.
- Modify `internal/transcodenode/server.go`
- Mirror the same `last_served_at`/bytes tracking on the node's own segment serve path (`handleSegment`).
- Modify `internal/api/handlers/stream.go` and `internal/api/handlers/playback.go`
- Integrated mode: on the native segment/direct serve paths, record server-side serve activity (last-served/bytes) into the same monitoring surface (in-memory for single-node; the tracker if Redis present). Stop relying on client progress reports for liveness. Reuse `BeginTransport`/`EndTransport` (`internal/playback/session.go:577,593`) plus a serve-activity touch.
- Modify `internal/api/handlers/nodes.go`
- Extend `HandleListSessions` (`:303`) to surface `last_served_at`/bytes so the admin view and the brain see the authoritative picture.
- Create `internal/streammonitor/monitor.go`
- A thin central-side aggregator that returns a normalized "live streams" snapshot: read Redis `silo:sessions:*` (multi-node) or the in-process `SessionManager`/tracker (integrated), keyed by `uid`, with `sid`, method, `last_served_at`, bytes, node. This is the single input to the async brain and to admin monitoring.
- Create `internal/streammonitor/monitor_test.go`
- Table tests over synthetic session records: liveness aging, per-user grouping, multi-node de-dup by `sid`.
### Phase 2 — Kill switch (revocation) + enforcement
- Create `internal/streamrevoke/store.go`
- Interface `Store` with `Revoke(ctx, key RevKey, reason string, until time.Time) error`, `IsRevoked(key RevKey) bool` (local, hot-path, in-memory), `List(ctx) ([]Revocation, error)`, `StartSync(ctx)`.
- `RevKey` addresses a session (`sid`) or a user (`uid`); a user revocation matches all that user's sessions.
- In-memory backend (integrated): a `map`/set behind an RWMutex with TTL entries; `Revoke` writes locally.
- Redis backend (multi-node): `silo:revoked:sess:{sid}` / `silo:revoked:user:{uid}` keys with TTL = token remaining life (default 24h); `StartSync` warms an in-memory cache from a `SCAN` on start, subscribes to `cache.ChannelAdmin` for push updates, and polls every ≤60s as a safety net. `IsRevoked` reads only the in-memory cache (no per-call Redis).
- Postgres durability mirror (see migration), wired central-side, so kills survive a Redis flush/restart; repopulate Redis + cache on startup.
- Create `internal/streamrevoke/store_test.go`
- Cover: local cache hit/miss, TTL expiry, user-key matches session, pub/sub push applies, poll reconciles, durability repopulation.
- Create `migrations/sql/<timestamp>_revoked_stream_sessions.sql` (via `make migrate-create NAME=revoked_stream_sessions`)
- Table `revoked_stream_sessions(kind text, key text, reason text, revoked_at timestamptz, expires_at timestamptz, primary key(kind,key))` for durable kills; index on `expires_at` for pruning.
- Modify `cmd/silo/main.go`
- Construct the revocation store on both sides: edge (reuse `redisClient` from `main.go:576`, wire `srv.SetRevocationStore(...)` next to `SetRecipeStore` at `:623` and the proxy equivalent); central (reuse `apiRedisClient` `:672`, fall back to in-memory when nil). Call `StartSync` for edges.
- Extend the `OnUserSessionsRevoked` closure (`:2047`) to also `Revoke(user:{uid})` in the store, so account-level revocation kills streams too.
- Modify `internal/proxy/server.go`
- In `verifyToken` (`:126`) / each serve handler: after signature verify, `if revStore.IsRevoked(sid or uid) { close + 403 }`. For direct-play/remux, run a periodic re-check during the long serve (a deadline/ticker goroutine bound to the request context) that aborts `io.Copy`/`ServeFile` and closes the connection when revoked.
- Modify `internal/transcodenode/server.go`
- Same revocation check on `handleSegment`/`handleManifest`; refuse + stop feeding a revoked session (and tear down its ffmpeg via the existing session teardown, but ALSO keep the revocation so `reconstructFromToken` refuses to rebuild it — otherwise the node self-reconstructs the kill away, `server.go:620`).
- Modify `internal/api/handlers/admin_playback_control.go`
- In `handleSessionCommand` for `terminate` (`:63`): in addition to the realtime command + `stopPlaybackSessionByID`, call `revStore.Revoke(sess:{sid})` so termination sticks at the edge regardless of client cooperation. Add an explicit `stop`-vs-`terminate` semantic: `terminate` revokes; `stop` remains a cooperative request.
- Modify `internal/api/router.go`
- Inject the revocation store into the handlers that need it (`Dependencies` field, like `RedisClient`).
- Modify `internal/config/config.go`, `internal/config/db_loader.go`, `internal/config/yaml_import.go`
- Add `auth.stream_revocation_poll` (default `60s`) and `auth.stream_revocation_ttl` (default `24h`) following the `jellyfin_compat.session_ttl` precedent. Decide restart-required vs hot-reload (`internal/config/restart_keys.go`); prefer hot-reload.
### Phase 3 — Async decision engine (the brain)
- Create `internal/playback/enforcer.go` (or `internal/worker/stream_enforcer.go`, following the `internal/worker/cleanup.go` ticker precedent)
- A background loop (central only) that every N seconds: pulls the live snapshot from `streammonitor`, groups by `uid`, reads each user's limit via the existing `SessionLimitProvider` / `users.max_streams`, and for any user over cap selects victims (default: most-recently-started beyond the cap) and calls `revStore.Revoke(sess:{sid}, "over_cap")`.
- Pluggable rule interface `Rule(snapshot) []KillDecision` so admin/abuse rules compose. Ship two rules: `OverCapRule` and a stub `RestreamHeuristicRule` (documented, disabled by default) that flags sustained N-distinct-title concurrency / throughput anomalies per `uid`.
- Structured logging + metrics for every kill: who, why, which rule.
- Create `internal/playback/enforcer_test.go`
- Deterministic tests: over-cap victim selection, no-op when under cap, idempotent re-revocation, limit-provider error → fail-open (do not kill).
- Modify `cmd/silo/main.go`
- Start the enforcer loop on central (integrated/api), wire it to `streammonitor` + the limit provider + `revStore`.
- Modify `internal/api/handlers/admin_playback_control.go` / a new admin endpoint
- Optional: expose the current kill decisions / recently revoked sessions for the admin UI ("why was this killed").
---
## Implementation Tasks
### Phase 1 — Authoritative monitoring
- [ ] Add `LastServedAt` + `BytesServed` to `nodesessions.SessionInfo`; refresh on real serve activity with a throttled Redis write.
- [ ] Wrap edge serve writers (proxy segment/direct/remux) to attribute per-session bytes + keep `LastServedAt` fresh during long pours; extend `egressMeter` with a per-session delta callback.
- [ ] Mirror serve-activity tracking on the transcode node (`handleSegment`).
- [ ] Integrated mode: record server-side serve activity on native serve paths; stop trusting client progress for liveness.
- [ ] Add `internal/streammonitor` aggregator returning the normalized per-`uid` live snapshot (Redis and in-process backends) + tests.
- [ ] Surface `last_served_at`/bytes in `HandleListSessions` for the admin view.
- [ ] `go build ./... && go vet ./...`; manual check that a playing stream appears with a moving `last_served_at` and disappears within the TTL after it stops — with the client sending no progress reports.
### Phase 2 — Kill switch
- [ ] Create `internal/streamrevoke` store: interface + in-memory backend + Redis backend (SCAN warm-up, pub/sub push, ≤60s poll, in-memory hot-path cache) + tests.
- [ ] Add durable `revoked_stream_sessions` migration and the Postgres mirror + startup repopulation.
- [ ] Wire the store into edges (`SetRevocationStore` on proxy + transcode Server, `StartSync`) and central (reuse `apiRedisClient`, in-memory fallback).
- [ ] Enforce on the edge hot path: local `IsRevoked` check per segment (refuse + close); periodic re-check + abort for direct-play/remux long pours.
- [ ] Ensure the transcode node does NOT self-reconstruct a revoked session (`reconstructFromToken` consults the store).
- [ ] Extend `OnUserSessionsRevoked` to revoke `user:{uid}`; make admin `terminate` revoke `sess:{sid}` (keep `stop` cooperative).
- [ ] Add config keys (`auth.stream_revocation_poll`, `auth.stream_revocation_ttl`) via the settings pipeline.
- [ ] `go build ./... && go vet ./...`; manual check: revoke a live session → it stops within ~one poll interval and a reconnect with the same token is refused (stays dead); revoke a user → all their streams stop.
### Phase 3 — Async brain
- [ ] Create the enforcer loop + pluggable `Rule` interface; implement `OverCapRule` + stub `RestreamHeuristicRule`; fail-open on limit-provider errors.
- [ ] Start the enforcer on central; wire snapshot + limit provider + revocation store.
- [ ] Structured logging/metrics for every kill; optional admin "recent kills" endpoint.
- [ ] `enforcer_test.go` deterministic tests.
- [ ] `go build ./... && go vet ./...`; manual check: start N+1 concurrent streams for a capped user → within ~120s the over-cap stream(s) are killed and stay dead; under-cap users are untouched.
---
## Testing
- Unit: `internal/streammonitor` (aggregation/aging), `internal/streamrevoke` (cache/TTL/pub-sub/poll/durability), `internal/playback` enforcer (victim selection, fail-open).
- Integration (manual, needs `docker compose up -d postgres redis` + a real client): verify authoritative monitoring with a client that sends no progress; verify kill-within-poll-interval + stays-dead for transcode, direct-play, and remux; verify multi-node (central issues kill, edge enforces) via the shared Redis.
- Regression: existing playback is unaffected when nothing is revoked (the only hot-path addition is a local set lookup that returns false).
## Risks / follow-ups
- **Enforcement latency (~≤120s)** is intentional and accepted; document it so operators expect a short window before a kill takes effect.
- **Direct-play long GET** requires a periodic in-serve revocation check + connection cut; a naive single-GET client with no resume may see an error rather than a clean stop — acceptable for a killed (abusive/over-cap) stream.
- **Redis flush/restart** would drop in-flight revocations unless the Postgres mirror is in place; keep the mirror for durability (reliability-first).
- **Integrated single-node without Redis** uses the in-memory revocation store and in-process monitoring; the durable mirror still applies. No cross-node concern there.
- **Abuse heuristics** (`RestreamHeuristicRule`) ship disabled; tune against real traffic before enabling to avoid false kills.
- **Bytes attribution** on the edge is best-effort for monitoring/heuristics; it must never gate the hot path.
## As-built deltas from this plan
The shipped implementation follows this plan; a few names/shapes differ (the coverage matrix is authoritative for the current state):
- The durable mirror landed as table `stream_revocations(kind, id, reason, revoked_at, expires_at)` (PK `(kind,id)`) via `internal/streamrevoke/durable_postgres.go`, and is **wired central-side and non-optional** — warmed on boot (bounded retry) and re-armed into Redis on the poll tick so multi-node edges reconverge after a Redis flush.
- Enforcement is centralized in two shared helpers `streamrevoke.Store.Refuse` (403) and `WatchAndCut` (in-flight connection cut) used by every serve surface, rather than per-handler ad-hoc checks.
- **Expiry is monotonic on every write.** `applyLocal` and the durable `Upsert` keep whichever copy expires later (`GREATEST`), so the async enforcer's short 5m self-heal TTL can never shorten a longer admin `RevokeSession` on the same session key (which would otherwise reopen the restart-resurrection window the durable mirror closes).
- **In-flight cut required a middleware fix.** `WatchAndCut` uses `http.NewResponseController(w).SetWriteDeadline`, which walks `Unwrap()`. The native and jellycompat middleware `ResponseWriter` wrappers lacked `Unwrap()`, so the integrated cut was a guaranteed no-op until `Unwrap()` was added to `statusWriter`, `requestStatusWriter`, `loggingResponseWriter`, `compatImageProxyTagResponseWriter`, and `debugResponseWriter`.
- **Ownership carry-forward in the monitor.** `streammonitor.mergeStreams` recovers a resolved `UserID`/`ProfileID`/`MediaFileID` (and backfills route/client) across the same-session dedupe, so the transcode node's ownerless start record can't bucket a stream under user 0 and exempt it from the cap.
- **Existence vs timing, made explicit.** Monitoring separates *existence* (server-observed on every path — the transport-count shield integrated, byte-observed at the edge — so a hidden stream that withholds progress is still counted) from *timing* (client progress is trusted, native fully). The transcode serve paths (native `playback.go` and compat `HandleHLSSegment`) gained a per-segment transport marker so a slow single-segment drain can't be reaped mid-serve and the compat HLS path no longer depends on the client's progress POST for liveness. A buffer-ahead idle gap >45s is a remaining reaper follow-up (VERIFY-4).
- **First-class monitoring fields.** Added a `Route` (native/jellycompat) dimension — seeded from `Session.Origin` and carried to edges via a token `Origin` claim — plus client IP/name and position, all surfaced through `SessionInfo`/`LiveStream`. The transcode node's start record is now owner+route+client attributed via `TranscodeStartRequest`. The admin session list unions the in-process integrated sessions (`LiveLocalSessions`) so single-node deployments aren't Redis-blind. Downloads remain an explicit cap exemption (separate quota).
- `auth.stream_revocation_poll` / `auth.stream_revocation_ttl` operator config keys are not yet wired (defaults 60s / 24h are hardcoded) — remaining follow-up.
## AI-use disclosure
This plan was drafted with AI assistance (Claude), based on a read-only audit of the current codebase. No behavior was changed by writing it. The Status banner and As-built deltas section were added after implementation.