From 6ae1ea30c8eb4e748fe61ab4994ea26694b2e0ea Mon Sep 17 00:00:00 2001 From: CoffeeKnyte <67730400+CoffeeKnyte@users.noreply.github.com> Date: Tue, 7 Jul 2026 08:32:30 +0000 Subject: [PATCH] docs: adversarial abuse matrix + enforcement architecture decision - docs/architecture/stream-abuse-matrix.md: red-team of this branch against ~29 abuse stories, each scored on detection vs enforcement, including the three places the branch's own docs oversold the code. - docs/superpowers/plans/2026-07-07-abuse-cold-enforcer-architecture.md: the deny-list vs lease analysis and phased plan, with the final decision note: keep the single revocation pipeline with reason-scoped TTLs; the two-tier lease is dropped. Part of #306. --- docs/architecture/stream-abuse-matrix.md | 431 ++++++++++++++++++ ...-07-07-abuse-cold-enforcer-architecture.md | 303 ++++++++++++ 2 files changed, 734 insertions(+) create mode 100644 docs/architecture/stream-abuse-matrix.md create mode 100644 docs/superpowers/plans/2026-07-07-abuse-cold-enforcer-architecture.md diff --git a/docs/architecture/stream-abuse-matrix.md b/docs/architecture/stream-abuse-matrix.md new file mode 100644 index 00000000..18cfa223 --- /dev/null +++ b/docs/architecture/stream-abuse-matrix.md @@ -0,0 +1,431 @@ +# Stream & API Abuse — Coverage Matrix + +> **What this is.** A red-team assessment of how the `feat/sauron-async-enforcer` +> branch (server-observed stream monitoring + revocation kill switch + async +> over-cap enforcer) — layered on the current `main` — reacts to ~29 concrete +> abuse stories. Each story is scored on two axes the branch itself separates: +> +> - **Detection** — does the server *see* the abuse and *label it correctly* +> (right user, right stream, right kind)? +> - **Enforcement** — can the server *stop* it, and *how fast*? +> +> **Companion docs:** the intent lives in +> [`../superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md`](../superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md); +> the as-built coverage-by-path lives in +> [`playback-paths-monitoring-kill-matrix.md`](playback-paths-monitoring-kill-matrix.md). +> This doc is the *adversarial* companion: it assumes a motivated abuser and asks +> where the design ends. +> +> **Scope note.** The branch is scoped to **concurrent playback-stream abuse**. A +> large fraction of real-world "abuse" (library enumeration load, bulk ripping, +> API request floods, CPU exhaustion, re-broadcasting) lives *outside* that scope. +> Where that is the case this doc says so plainly and points at the subsystem that +> would actually own the defense (`internal/ratelimit`, `internal/downloads`, +> catalog browse, node scheduling) — several of which have real gaps of their own. + +## Legend + +| Symbol | Detection | Enforcement | +|---|---|---| +| ✅ | Seen and correctly attributed | Stopped (and how fast) | +| ⚠️ | Seen but partial / mislabeled / lossy | Partial, slow, best-effort, or config-dependent | +| ❌ | Invisible to the server | Cannot stop | + +Enforcement latency budget for the async path is **~120s** = enforcer tick +(`DefaultInterval = 30s`, `streamenforcer/enforcer.go`) + revocation propagation +(pub/sub push immediate; ≤60s poll safety net, `streamrevoke/store.go`) + in-flight +cut (5s `WatchAndCut` ticker for long pours; next segment for HLS). Synchronous +admission refusal (`ErrTooManyStreams`, `playback/session.go`) is **immediate**. + +--- + +## Three corrections the investigation turned up (read first) + +The branch docs oversell three things. The stories below are scored against the +**code**, not the plan. + +1. **The async enforcer does nothing for a *standard* user.** The enforcer's limit + comes from the **raw** `users.max_streams` column (`cmd/silo/main.go`, the + `LimitFunc`), and `limit <= 0` is treated as "unlimited → skip" + (`streamenforcer/enforcer.go`). Migrations `20260702180000` / `20260702190000` + set every standard user's `users.max_streams = 0` and delegate the real cap (5) + to the **Default Group**. So for a normal user the enforcer sees `limit 0` and + **never revokes**. The cap that actually holds is the *synchronous admission* + check, which uses the group-merged effective policy + (`access.EffectivePolicyForUser` → `strictestPositive`). The "async brain" only + bites users given an explicit per-user `max_streams > 0`. Treat this as a + **bug**, not a design choice — it means the branch's headline feature is + effectively dormant on a default install. + +2. **The re-streaming heuristic does not exist.** The plan describes shipping + `OverCapRule` + a disabled `RestreamHeuristicRule` behind a pluggable `Rule` + interface. In code, `streamenforcer/enforcer.go` is a single hardcoded over-cap + loop — no rule interface, no restream logic, not even a disabled stub. The + fingerprints a heuristic *would* need (`ClientIP`, `ClientName`, `MediaFileID`, + `BytesServed`) are captured in `streammonitor` but nothing consumes them for + detection. + +3. **There is no server-wide or per-node transcode/CPU cap.** Only a per-*user* + `max_transcodes` (default 2, `<=0` = unlimited) checked at admission. The one + ffmpeg semaphore (`reconstructSem`, sized to `NumCPU`) guards **only** the + restart-reconstruct path in `transcodenode/server.go`, **not** fresh starts. + +One thing the docs *under*-sell: `main` already ships a real general API rate +limiter (`internal/ratelimit/`), enabled by default — but it is mounted only on +the authenticated native surface + specific auth endpoints, and the entire +Jellyfin-compat surface is unthrottled (see Category E). + +--- + +## Summary matrix + +### Category A — Concurrency-cap abuse (the branch's core competency) + +| # | Abuse story | Detection | Enforcement | One-line verdict | +|---|---|---|---|---| +| 1 | User opens N+1 concurrent **transcodes** (over cap) | ✅ | ✅ immediate | Admission refuses the N+1 start synchronously | +| 2 | User opens N+1 concurrent **direct-play** streams | ✅ | ✅ immediate | Same admission gate; method-agnostic | +| 3 | **Shared account**, many devices streaming at once | ✅ | ✅ / ⚠️ | Capped per-user on one process; leaks across replicas (#7) | +| 4 | **Buffer-ahead then pause** fetches >45s to duck the cap | ❌ (window) | ❌ (window) | VERIFY-4 hole: transcode goes invisible 45s/60s, admits another | +| 5 | **Withhold/falsify progress** to hide a stream | ✅ | ✅ | Existence is byte-observed; progress only moves a display field | +| 6 | **Rapid session churn** between enforcer ticks | ⚠️ | ⚠️ | Admission blocks over-cap starts; no rate limit on `/start` | +| 7 | Split streams across **multiple integrated replicas** | ⚠️ | ⚠️ | Admission is per-process; enforcer is the only cross-node backstop (and see correction #1) | +| 8 | **Standard user** exceeds cap via the edge/multi-node path | ⚠️ | ⚠️ | Enforcer no-ops (`users.max_streams = 0`); only admission holds | + +### Category B — Admin control / kill switch + +| # | Abuse story | Detection | Enforcement | One-line verdict | +|---|---|---|---|---| +| 9 | Admin **terminates** a specific abusive live session | ✅ | ✅ | Terminate writes revocation first, keyed on sid; 202 even if no local session | +| 10 | Admin **bans a user** (revoke all their streams) | ✅ | ✅ | `RevokeUser` cutoff kills pre-ban streams incl. in-flight | +| 11 | Killed stream **reconnects** with the same token | ✅ | ✅ | Revocation outlives token; every reconnect `Refuse`d | +| 12 | Killed stream **survives a server restart** | ✅ | ✅ | Durable Postgres mirror re-warms the kill list; monotonic expiry | +| 13 | Kill issued **during an edge Redis outage** | ⚠️ | ⚠️ | Edge fails **open**: new kills don't reach edge until Redis recovers | + +### Category C — Re-streaming / redistribution + +| # | Abuse story | Detection | Enforcement | One-line verdict | +|---|---|---|---|---| +| 14 | Re-broadcast **one** Silo stream to many external viewers | ❌ | ❌ | Under-cap = untouched; no restream heuristic exists (correction #2) | +| 15 | **Many** concurrent Silo streams feeding a restream service | ⚠️ | ✅ ~120s | Caught only as raw over-count, not labeled re-streaming | +| 16 | **Mint & hoard** many 24h tokens, fan out later | ❌ | ❌ / ⚠️ | No mint cap, signature-only, stop doesn't revoke; caught only if fan-out becomes over-cap | + +### Category D — Ripping / bulk data exfiltration + +| # | Abuse story | Detection | Enforcement | One-line verdict | +|---|---|---|---|---| +| 17 | Rip whole library via **native download** endpoints | ⚠️ | ⚠️ | 3-concurrent gate only; **no volume cap**; loop create→complete defeats it | +| 18 | Rip via **compat `/Items/{id}/Download`** (Infuse) | ❌ | ❌ | **No quota at all** — auth + library filter only. Widest hole | +| 19 | Rip via **sequential direct-play GETs** (one at a time) | ⚠️ | ❌ | Cap counts concurrency, not volume; stays at 1 forever | +| 20 | Admin **stops an in-progress rip** mid-transfer | — | ⚠️ | Branch adds `WatchAndCut` on downloads (best-effort); admin must issue the revoke | + +### Category E — API / DB load abuse (non-stream; branch is orthogonal) + +| # | Abuse story | Detection | Enforcement | One-line verdict | +|---|---|---|---|---| +| 21 | **Infuse pulls the whole library list on every app start** | ❌ | ⚠️ | Per-page capped at 1000; full browse is uncached & unthrottled | +| 22 | **Thundering herd** — many clients cold-cache full enumeration | ❌ | ⚠️ | Latest/hub rails singleflight; general browse does not | +| 23 | Client **pages a 100k+ library as fast as possible** | ❌ | ❌ | Per-page COUNT + rising OFFSET; no page-rate limit, no cache | +| 24 | **Expensive sorted/filtered views** over the whole library | ❌ | ⚠️ | Some sorts index-optimized; `DateCreated`/filters = full top-N + COUNT, uncapped | +| 25 | Badly-coded client **retry storm** (thousands req/s) | ✅ | ⚠️ | Native authed surface throttled (per-IP 120rps); compat & `/auth/refresh` not | +| 26 | Tight-loop polling **`/auth/refresh`** | ✅ | ❌ | No rate-limit middleware on that route | +| 27 | **Credential stuffing** via compat `/Users/AuthenticateByName` | ✅ | ❌ | Compat router mounts no limiter; native login is capped 20/min | +| 28 | Many concurrent **HTTP connections / slowloris** | ❌ | ❌ | No connection cap / `LimitListener`; only Read/Write/Idle timeouts | +| 29 | Spawn many concurrent **transcodes to exhaust CPU** | ⚠️ | ❌ | Per-user `max_transcodes` only; no per-node/server-wide ffmpeg cap | + +--- + +## Detailed stories + +### Category A — Concurrency-cap abuse + +**A1. N+1 concurrent transcodes (over cap).** +The (N+1)th `StartSession` returns `ErrTooManyStreams` *synchronously* at admission +(`playback/session.go`, `inlineAdmissionErrorLocked`), using the group-merged +effective cap (5 for a standard user). The client never gets a session. The async +enforcer is a redundant backstop *only* for users with an explicit +`users.max_streams > 0` (correction #1). **Detection ✅ / Enforcement ✅ (immediate).** + +**A2. N+1 concurrent direct-play streams.** +Direct-play sessions are `Track`-backed and counted identically at admission — the +`MaxStreams` check is method-agnostic. Same immediate refusal. **✅ / ✅.** + +**A3. Shared account, many devices at once.** +All sessions carry the same `AuthUserID`; `streammonitor.ByUser` groups them and +`Client*` fields expose the fan-out. The per-user cap bounds simultaneous streams +to the effective `MaxStreams`. **Caveat:** the cap is per-process (see A7), so +spreading devices across integrated replicas multiplies the allowance. +**Detection ✅ / Enforcement ✅ on one process, ⚠️ across replicas.** + +**A4. Buffer-ahead then pause (VERIFY-4 hole). — REAL, UNFIXED.** +A transcode client that pre-buffers far ahead and stops fetching segments for +>45s (integrated active grace, `session.go`) / >60s (edge `sessionTTL`, +`nodesessions/tracker.go`) stops counting at admission +(`countsTowardLimitsLocked`) and is reaped from the snapshot by `CleanStale` / +idle-prune — while the client keeps playing from buffer. During that window the +server admits *another* stream. Staggered pre-buffer/pause rotation across devices +ducks the instantaneous cap. Direct-play/remux are immune (their `Track` records +are never idle-timed). **Detection ❌ (during window) / Enforcement ❌ (during +window).** Fix requires treating a running ffmpeg as liveness or a transcode-grace +≥ client buffer. + +**A5. Withhold/falsify progress reports.** +This *fails* as an evasion. Existence and liveness are **byte-observed** +(`BeginTransport`/`EndTransport` bracket every pour; `AddBytes`/`MarkServed` set +`LastServedAt`), never derived from client progress. Falsifying `Position` only +corrupts a secondary display field and, at worst, changes which of the user's +*own* over-cap sessions `selectVictims` trims first. **Detection ✅ / Enforcement ✅.** + +**A6. Rapid session churn between ticks.** +No rate limit exists on `/start` or `/transcode/start` (the `ratelimit` +middleware is wired only to auth endpoints). In integrated mode synchronous +admission still refuses over-cap starts instantly, so churn can't exceed the cap +locally. The gap is the multi-node/edge picture, where admission may have run on a +different node and the enforcer only reconciles every ~120s. +**Detection ⚠️ / Enforcement ⚠️ (integrated fast, edge lags).** + +**A7. Split across multiple integrated replicas.** +`activeCountLocked` counts only the local `SessionManager` map — there is no +distributed admission counter. Two integrated replicas each admit up to the cap +independently. The enforcer is the only cross-replica backstop, but (a) integrated +replicas write no `silo:sessions:*` records, so a remote enforcer can't see them, +and (b) per correction #1 it no-ops for standard users anyway. +**Detection ⚠️ / Enforcement ⚠️.** (VERIFY-3.) + +**A8. Standard user exceeds the cap on the edge path.** +Direct consequence of correction #1. The enforcer — the component *specifically* +built to enforce the cap across nodes — is dormant for `users.max_streams = 0`. +On a single integrated node admission still holds; in multi-node, an over-cap edge +situation that admission didn't catch locally will **not** be reconciled. **⚠️ / ⚠️.** + +### Category B — Admin control / kill switch + +**B9. Admin terminates a specific session.** +`terminate` writes the revocation **first**, keyed on the session id alone, then +issues the cooperative realtime command; it answers `202 {status:"revoked"}` even +when the session is absent from the central in-memory manager (GAP-7) — exactly the +edge-served / post-restart / progress-withholding streams the kill switch exists +for. **Detection ✅ / Enforcement ✅.** + +**B10. Admin bans a user.** +`RevokeUser` writes a `KindUser` entry; `IsRevoked` matches it for any of the +user's sessions whose credential predates the revocation (cutoff semantics, GAP-5), +so in-flight pours are cut (`WatchAndCut`) and reconnects `Refuse`d, while a +later legitimate re-login is not permanently banned. **✅ / ✅.** + +**B11. Killed stream reconnects with the same token.** +The revocation entry outlives the token TTL, so every reconnect with the same +credential is refused. "Stays dead." **✅ / ✅.** + +**B12. Killed stream survives a restart.** +The durable Postgres mirror (`streamrevoke/durable_postgres.go`, table +`stream_revocations`) is warmed on boot and re-armed on the poll tick, and expiry +is monotonic (`GREATEST`) so the enforcer's short 5m self-heal TTL can't shorten a +24h admin kill. This closes GAP-4, where PR #174's restart-resilient playback would +otherwise reconstruct and re-serve a killed stream. **✅ / ✅.** *Transient boot +caveat:* if the durable warm exhausts its bounded retry **and** Redis is empty, the +kill list is empty until the first poll tick (≤60s) — fails **open** by design. + +**B13. Kill issued during an edge Redis outage.** +Edge nodes have no durable store and learn *new* kills only from Redis pub/sub + +SCAN. While Redis is unreachable, already-cached kills persist but a kill issued +*during* the outage does not reach the edge until Redis recovers. **Detection ⚠️ / +Enforcement ⚠️ (fails open at the edge).** + +### Category C — Re-streaming / redistribution + +**C14. Re-broadcast one Silo stream to many external viewers. — NO DEFENSE.** +The classic Stremio/Plex-share abuse: one account pulls a single stream and a proxy +fans it out to hundreds. It stays at concurrency 1, so admission and the enforcer +never fire, and the re-streaming heuristic that would catch it *does not exist* +(correction #2). The fingerprints to build it (`ClientIP`, `ClientName`, +`BytesServed`, distinct `MediaFileID`) are captured but unused. **Detection ❌ / +Enforcement ❌.** + +**C15. Many concurrent streams feeding a restream service (over cap).** +This *is* caught — but only as a raw over-count, not labeled as re-streaming. +`ByUser` counts sessions per uid; the enforcer revokes victims beyond the limit +within ~120s (and admission refuses new over-cap starts immediately). Relies on +correct owner attribution (jellycompat/edge records must carry the owner or they +bucket under user 0 and are skipped — mitigated by `mergeStreams` owner-adoption). +**Detection ⚠️ (as over-cap, not as restream) / Enforcement ✅ ~120s** — *subject to +correction #1 if the user has no explicit per-user cap.* + +**C16. Mint & hoard 24h tokens, fan out later.** +Stream tokens are signature-only 24h JWTs (`streamtoken/token.go`) with no `jti`, +no one-time-use, no per-user mint counter. A normal Stop does **not** write a +revocation, so a token for a stopped session stays cryptographically valid for its +full 24h. Hoarding is free and invisible at mint time. Fan-out is caught *only* if +the aggregate live use trips the per-user over-cap enforcer — i.e. a low-concurrency +fan-out (few tokens, each widely re-streamed) collapses back into C14 and evades +everything. **Detection ❌ (at mint) / Enforcement ❌ (unless it becomes over-cap).** + +### Category D — Ripping / bulk data exfiltration + +**D17. Rip via native download endpoints.** +Bounded only by `download.max_concurrent_per_user` (default **3**, +`internal/downloads/limiter.go`). `download.max_per_period` and both bandwidth +knobs default to **0 = unlimited** — and the bandwidth managers are *rate* throttles +(token buckets), never *volume* caps. The concurrency check runs at record +creation, so a `create → complete → create` loop pulls the entire library 3 files +at a time forever. **No total-volume cap exists anywhere.** **Detection ⚠️ +(per-download rows, no cumulative signal) / Enforcement ⚠️ (weak concurrency gate).** + +**D18. Rip via compat `/Items/{id}/Download`. — WIDEST HOLE.** +`HandleDownload` (`jellycompat/streams.go`) requires only a compat session and the +per-item library-access filter, then serves the raw file with full Range support. +**No download row, no quota, no bandwidth throttle, no session/monitor record, no +byte cap.** An Infuse-style client can GET/Range every original file back-to-back, +unlimited, and the server keeps no record beyond generic HTTP logs. The branch's own +doc comment admits this and files it as an open follow-up. **Detection ❌ / +Enforcement ❌.** + +**D19. Rip via sequential direct-play GETs.** +Playing items one at a time keeps `activeCountLocked == 1`, always under the cap. +Because the cap counts concurrency and never cumulative bytes, sequential +direct-play of every raw file is a fully-permitted rip. Each play *is* an +attributable session, but nothing correlates sequential sessions into "this user is +ripping." **Detection ⚠️ / Enforcement ❌.** + +**D20. Admin stops an in-progress rip mid-transfer.** +On `main`, impossible — `Refuse` only 403s *new* requests; an open multi-GB GET +keeps pouring. The branch adds `streamrevoke.Store.WatchAndCut` (checks on entry + +every 5s, forces the socket via `SetWriteDeadline`) and arms it on native downloads +(`api/handlers/downloads.go`) and compat download/direct-play +(`jellycompat/streams.go`). So a **user revocation now cuts an in-flight download** +— best-effort (no-op if the writer chain lacks write-deadline support, then stops on +next request). The admin still has to *decide* to revoke; nothing auto-detects the +rip. **Enforcement ⚠️ (branch improves `main`'s ❌ to best-effort cut).** + +### Category E — API / DB load abuse (branch is orthogonal) + +**E21. Infuse pulls the whole library list on every app start.** +Per-page size is hard-capped (`compatBrowseMaxLimit = 1000`, +`jellycompat/content_direct.go`; native 100), so no single call dumps 100k. But the +general paged browse (`handleBrowseItems` → `BrowsePage`) is **uncached and +unthrottled** — every app-start re-runs fresh SQL, and full-page requests add a +per-page `COUNT` over the filtered set. The `Latest`/hub rails *are* cached (15-min +shared resolved-list cache + singleflight), but that's a rail, not the full list. +Nothing attributes listing load per client/device. **Detection ❌ / Enforcement ⚠️ +(first page cheap & capped; full enumeration pays full price each start).** + +**E22. Thundering herd on cold cache.** +For `Latest`/hub rails, `singleflight` collapses concurrent identical builds into +one DB hit. For the **general browse there is no singleflight and no cache** — N +concurrent full-library enumerations run N independent SQL workloads, bounded only +by the shared pool (`userdb.pool_max_open`, default 500), beyond which requests +queue/time out rather than shed gracefully. **Detection ❌ / Enforcement ⚠️ (only +the cached rails are protected).** + +**E23. Page a 100k+ library as fast as possible.** +Each page is capped at 1000 rows, but nothing caps the number of pages, the page +rate, or aggregate cost; deep pages incur a per-page filtered `COUNT` and rising +`OFFSET` cost. No rate limit on `/Items` (compat router mounts no throttle). +**Detection ❌ / Enforcement ❌ (beyond the per-page size cap).** + +**E24. Expensive sorted/filtered views over the whole library.** +Some sorts are index-optimized (`recently_added` walks a dedicated index; the +multi-library case fans into per-library index walks + in-memory merge). But +arbitrary client sorts map through `mapSortBy` — `SortBy=DateCreated` "forces a +full-library top-N heapsort" (code comment), and genre/name-prefix/person filters +add WHERE/JOIN predicates with no cost cap beyond the LIMIT, plus a full-filtered +`COUNT` per page when `include_total` is on (default). No query-cost estimator, no +cache for filtered browses, no rate limit. **Detection ❌ / Enforcement ⚠️ +(query-shape-dependent).** + +**E25. Badly-coded client retry storm.** +`main` *does* ship `internal/ratelimit/` (token-bucket, default-on): global 1000 +rps + per-IP 120 rps, mounted inside the **authenticated native group** +(`api/router.go`, after `RequireAuth`). So a storm against an authed native +endpoint is shed per-IP. But the limiter runs *after* auth, the entire +**jellycompat surface is unthrottled**, `/auth/refresh` has no limiter, the Redis +backend **fails open** on error, and the default memory backend is per-process (so +"global 1000 rps" is really per-instance). Attribution + telemetry exist +(Prometheus by path/status, access log with `client_ip`). **Detection ✅ / +Enforcement ⚠️ (native authed only).** + +**E26. Tight-loop polling `/auth/refresh`.** +Registered as a plain public route (`api/router.go`) with **no** `AuthEndpointHandler` +and **outside** the authenticated group — it hits neither the global, per-IP, nor +per-endpoint limiter. A tight refresh loop is unthrottled. **Detection ✅ (logged) / +Enforcement ❌.** + +**E27. Credential stuffing via compat `/Users/AuthenticateByName`.** +Native `/api/v1/auth/login` is capped at 20/min/IP via `AuthEndpointHandler`. The +Jellyfin-compat login has **zero throttling** — the compat router mounts no limiter +— and there is no account-level lockout/backoff anywhere. **Detection ✅ / +Enforcement ❌ (compat path wide open).** + +**E28. Many concurrent HTTP connections / slowloris.** +No connection cap: no `netutil.LimitListener`, no per-IP connection limit, no +handler semaphore. The `http.Server` sets only Read/Write/Idle timeouts (and the +compat + ABS servers even set `WriteTimeout: 0`). The rate limiter counts +*requests*, not *connections*. **Detection ❌ / Enforcement ❌.** + +**E29. Spawn many concurrent transcodes to exhaust CPU.** +Only the per-user `max_transcodes` (default 2, `<=0` = unlimited) is checked, at +admission. There is **no per-node or server-wide ffmpeg/CPU cap** on fresh starts +(the `reconstructSem` NumCPU semaphore guards only the restart-reconstruct path). +Node scheduling has an optional `MaxJobs`/`MaxBandwidthKbps` per node +(`nodepool`), but both default to unlimited. N users at a high per-user cap — or one +user with `max_transcodes <= 0` — can saturate the box. **Detection ⚠️ (sessions +visible, no CPU/process signal) / Enforcement ❌ (no aggregate cap).** + +--- + +## What the branch genuinely delivers vs. what it does not + +**Delivers (and it's solid):** +- Authoritative, **byte-observed** stream existence that a client cannot hide by + lying about or withholding progress (A5). +- A durable, restart-surviving, reconnect-proof **kill switch** for a *specific + session* or a *user*, enforced on every serve surface (native, compat, edge, + transcode node) via one shared `Refuse` + `WatchAndCut` (B9–B12, D20). +- Immediate **synchronous admission** refusal of over-cap starts, and an async + over-cap reconciler for the multi-node picture (A1, A2, C15) — *when a per-user + cap is set*. + +**Does not cover (by scope or by gap):** +- **Under-cap re-streaming** (C14) and **token hoarding/fan-out** (C16) — no + detection, no enforcement; the heuristic that was scoped for this is unimplemented. +- **Bulk ripping** by download (D17/D18) or sequential direct-play (D19) — the cap + is a *concurrency* gate, not a *volume* quota; compat download is entirely + unquota'd. +- **Library-enumeration DB load** (E21–E24) — orthogonal subsystem; per-page capped + but per-client-uncached, unthrottled, and unattributed. +- **API floods on the compat surface, `/auth/refresh`, and connection exhaustion** + (E25–E28) — rate limiter has real coverage gaps. +- **CPU/transcode exhaustion** (E29) — no aggregate cap. +- **The enforcer being dormant for standard users** (correction #1 / A8) — a bug + that should be fixed before the async path is trusted. + +## Recommended follow-ups (prioritized by exposure) + +1. **Fix the enforcer limit source** (correction #1): resolve the group-merged + effective cap in the `LimitFunc`, not the raw `users.max_streams` column, so the + async reconciler actually enforces the standard cap it was built for. +2. **Quota the compat download route** (D18): bring `/Items/{id}/Download` under an + Infuse-compatible download quota (plain GET/Range, no download rows) — today's + only control is auth + library filter. +3. **Add a per-user *volume* budget** (D17/D19): a rolling bytes-per-period cap that + spans downloads *and* direct-play, since concurrency caps can't bound a rip. +4. **Close VERIFY-4** (A4): treat a running transcode ffmpeg as liveness (or a + transcode-specific grace ≥ max client buffer) so buffer-and-pause can't duck the + cap. +5. **Extend rate limiting to the compat surface + `/auth/refresh`** and add a + connection cap / `LimitListener` (E25–E28). +6. **Add a per-node concurrent-transcode cap** and wire `nodepool.MaxJobs` defaults + (E29). +7. **Implement the restream heuristic** (C14/C16): the fingerprints already flow + through `streammonitor`; a distinct-viewer-IP / distinct-title / throughput rule + is the missing consumer. Ship disabled, tune against real traffic. +8. **Per-client listing-load accounting** (E21–E24): attribute browse cost per + device and add caching/singleflight to the general paged browse, not just the + Latest/hub rails. + +## AI-use disclosure + +This assessment was produced with AI assistance (Claude), from a read-only audit of +the `feat/sauron-async-enforcer` branch and current `main`. No code behavior was +changed. Findings cite the code as of the audit; the three corrections above are +where the branch's own plan/coverage docs overstate what the shipped code does. diff --git a/docs/superpowers/plans/2026-07-07-abuse-cold-enforcer-architecture.md b/docs/superpowers/plans/2026-07-07-abuse-cold-enforcer-architecture.md new file mode 100644 index 00000000..a7bee2f4 --- /dev/null +++ b/docs/superpowers/plans/2026-07-07-abuse-cold-enforcer-architecture.md @@ -0,0 +1,303 @@ +# Abuse Cold-Enforcer — Architecture Decision & Plan + +> **Branch:** `abuse-cold-enforcer` (based on `origin/main`). +> **Name meaning:** "cold" = every decision stays **off the hot path**; "abuse" = +> the scope widens from concurrency-only (the `feat/sauron-async-enforcer` branch) +> to the full set of abuse classes catalogued in +> [`../../architecture/stream-abuse-matrix.md`](../../architecture/stream-abuse-matrix.md). +> Commands assume the repository root is the cwd. +> +> **Inputs learned from:** the `feat/sauron-async-enforcer` branch (kill-switch +> implementation), its plan +> [`2026-07-04-stream-monitoring-and-kill-switch.md`](2026-07-04-stream-monitoring-and-kill-switch.md), +> its coverage matrix +> [`../../architecture/playback-paths-monitoring-kill-matrix.md`](../../architecture/playback-paths-monitoring-kill-matrix.md), +> and the adversarial assessment +> [`../../architecture/stream-abuse-matrix.md`](../../architecture/stream-abuse-matrix.md). + +--- + +> **DECISION UPDATE (2026-07-07, supersedes §2.3 and P2 below).** After adversarial +> review, the two-tier lease was **dropped**. Everything a lease provides is already +> latent in the deny-list pipeline: enforcer revocations carry a short self-healing +> TTL and are re-asserted every tick (reversibility + continuous re-evaluation), and +> the one genuine lease benefit — surviving a sustained outage — is captured by +> **reason-scoped revocation TTLs** (deterministic violations like a byte-budget +> breach write a durable revocation until the period rolls; over-cap keeps the short +> self-healing TTL; admin kills stay 24h + durable). Fail-closed enforcement was +> judged wrong for its only distinctive feeder (untuned heuristics — false positive + +> Redis blip would kill legitimate playback), and a second enforcement mechanism with +> a second failure posture was judged more fragile than the coverage it buys (one +> matrix row, B13, during a >5m Redis outage). **The architecture is the single +> pipeline: `streammonitor` → enforcer tick → `streamrevoke` → in-memory `IsRevoked`, +> with policy checks for *new* activity at synchronous admission.** Pillars 2–4 and +> phases P0/P1/P3/P4 stand unchanged; P1's enforcement action becomes a durable +> revocation instead of a Tier-1 flag. + +## 1. Non-negotiables (inherited, do not regress) + +1. **Off the hot path.** The per-segment / per-byte serve path may add at most one + local in-memory lookup. No per-request DB or Redis on the serve path. (This is + the "cold" in cold-enforcer.) +2. **No client-facing protocol change.** Clients (Jellyfin, Infuse, Silo native) + keep echoing the same stable stream token. Any short-lived state lives + **server-side**, keyed by the token's `sid`/`uid`. This constraint is what + killed the naïve short-ticket idea in the prior plan; §3 shows how to keep the + ticket *benefits* without violating it. +3. **Reliability first / predictable under failure.** A brain/Redis hiccup must not + silently break legitimate playback. Failure posture must be a tunable knob, not + an accident. +4. **Reuse, don't rewrite.** The sauron branch's *authoritative, byte-observed + monitoring* (`streammonitor`, `nodesessions`) and its *durable, reconnect-proof + revocation* (`streamrevoke`) are good primitives. Keep them; change how the + *decision* is made and *widen what is decided*. + +--- + +## 2. The core question: kill-list (deny) vs short-lived tickets (allow) + +The prior branch chose a **deny-list**: streams run by default; the async brain +observes abuse, writes a revocation, and the edge refuses on an in-memory +set-membership check. The rejected alternative was **short-lived tickets**: the +credential itself expires quickly and must be continually re-minted, so an +unauthorized stream dies when its ticket lapses. + +### 2.1 Why the branch rejected tickets — and why that reasoning is incomplete + +The plan rejected tickets because Silo's `master.m3u8` is **VOD, fetched once**, so +a short ticket embedded in segment URLs would expire mid-movie and 401 the rest of +the stream; and flipping the global token TTL is a break-all-playback change. + +Both objections are about putting the short TTL **in the client's credential**. +They do **not** apply if the short-lived state lives **server-side**. The client +keeps its stable 24h token; the *server* holds a short-lived **lease** keyed by that +token's `sid`. The edge hot path checks the local lease, not a client-presented +ticket. This is a "ticket" in every property that matters, minted server-to-server, +requiring zero client cooperation — which is exactly the gap the prior analysis +missed. + +### 2.2 Deny-list vs lease (server-side ticket) — honest trade + +| Property | Deny-list (current branch) | Lease / server-side ticket | +|---|---|---| +| Default state | **allow** (runs unless revoked) | **deny** (runs only while leased) | +| Hot-path cost | in-memory set lookup | in-memory `sid→expiry` lookup (identical class) | +| To stop a stream | write a revocation, propagate | **stop renewing** (or expire) the lease | +| Stop latency | detect → decide → propagate (~120s) | ≤ one lease TTL, **same path for every reason** | +| Maintenance load | write **only on kill** (rare, cheap) | **renew every active stream every TTL** (continuous) | +| Restart-resurrection | **must** bolt on a durable mirror (branch's GAP-4) | **free**: a reconstructed session has no lease until re-granted | +| Failure posture | **fail-open** — brain/Redis down ⇒ new kills don't land, abuse continues, but playback survives | **fail-closed** — brain/Redis down ⇒ leases lapse, abuse stops, but legit playback also stops | +| Policy checkpoint | scattered: each rule writes its own kill | **one** checkpoint: the renewal tick evaluates cap + ban + byte-budget + throughput at once | + +The two are duals. Deny-list optimizes **availability** (fail-open, cheap, +already-built) and is *reactive*. Lease optimizes **abuse control** (fail-closed, +self-expiring, restart-safe, one uniform action) and is *preventive* — at the cost +of continuous renewal load and an availability blast-radius if the brain stalls. + +### 2.3 Decision: a **two-tier** enforcer (deny by default, lease for suspects) + +Neither pure form is right for a media server that must not break legit playback +*and* must reliably stop abuse. Use both, graduated by suspicion: + +- **Tier 0 — normal streams: deny-list default (keep the branch as-is).** A stream + runs unless explicitly revoked. Hot path = the existing `IsRevoked` set lookup. + Fail-open. This preserves "playback never breaks because the brain hiccuped" for + the 99% case, honoring *reliability first*. +- **Tier 1 — flagged/suspect streams: promote to lease-required (fail-closed).** + When the cold brain flags a session — over a *soft* cap, anomalous throughput / + distinct-viewer fan-out (re-streaming), or over a rolling byte budget — it + **demotes** that session to lease-required: the edge now additionally requires a + fresh central lease for that `sid`, and the brain grants it only while policy + permits. Stop = simply stop renewing. This gives suspects the lease's good + properties (≤TTL cut, no restart-resurrection, and — crucially — **an outage can + no longer be used to escape an existing flag**, because a lapsed lease fails + closed) while the expensive renewal load is paid *only for the small suspect set*, + not every stream (honoring *performance first*). +- **Immediate admin kill stays deny-list + durable.** "Cut this now" must not wait + ≤TTL, so admin terminate/ban writes an immediate revocation and a **durable ban + ledger** (Postgres) so a banned *user* is never re-leased or re-admitted across a + restart. The durable store thus shrinks from "mirror every ephemeral kill" to + "record standing bans + quota ledgers" — a more natural durability boundary. + +This is the headline architecture. Everything below is the machinery to feed it and +the adjacent pillars the matrix demands. + +**What this reuses from `feat/sauron-async-enforcer`:** `streammonitor` (the +authoritative cross-node snapshot) is the input to the renewal tick unchanged; +`streamrevoke` becomes the Tier-0 + admin path unchanged; the pub/sub + Redis +`silo:sessions:*` transport is reused to carry a central-stamped `lease_until` field +(the lease is one extra field on records the brain already reads). The edge check +flips from `if revoked → refuse` to `if revoked → refuse; else if leaseRequired[sid] +&& now > leaseUntil[sid] → refuse`. Small, additive, off hot path. + +--- + +## 3. The four capability pillars (mapped to the matrix) + +The "kill-list vs ticket" decision only governs Pillar 1. The matrix demands three +more, independent of that choice. + +### Pillar 1 — Concurrency & policy decision engine (A1–A8, C15) + +- **Fix the dormant-enforcer bug first (matrix correction #1).** The sauron + enforcer reads the *raw* `users.max_streams` column, which migrations set to `0` + for standard users (real cap lives in the Default Group). `limit<=0` ⇒ skip, so + the async brain never fires for a normal user. Resolve the **group-merged + effective policy** (`access.EffectivePolicyForUser` → `strictestPositive`) in the + brain's `LimitFunc`, exactly as synchronous admission already does. +- **Keep synchronous admission** (`ErrTooManyStreams`) as the immediate preventive + gate (A1/A2). The cold brain is the cross-node backstop (A7/A8/C15) and now the + lease authority for suspects. +- **Multi-replica (A7, VERIFY-3):** have integrated replicas publish their local + sessions into the shared `silo:sessions:*` picture (or elect a single brain + leader) so the cross-node count is authoritative. The lease grant is the natural + serialization point. + +### Pillar 2 — Liveness & anti-evasion (A4, A5) + +- **A4 buffer-ahead-then-pause is the one open monitoring hole.** Fix it + model-independently: treat a **running transcode ffmpeg as liveness** (or a + transcode-specific grace ≥ the client's max buffer) so a session that pre-buffers + and pauses fetches >45s is neither reaped nor decounted, and cannot free a slot to + admit another. This touches the sensitive session-reaping path (cf. #279) — do it + behind a flag with tests. +- **A5 progress-withholding is already solved** (existence is byte-observed). Keep; + add a regression test so it stays that way. + +### Pillar 3 — Volume/quota accounting (C16, D17–D19) — new + +Concurrency caps count *streams*, never *bytes*, so every rip vector walks straight +through them. Add a **rolling per-user byte budget** that spans **both** downloads +and direct-play/stream egress: + +- Meter served bytes per `uid` (the edge already meters egress; attribute it per + session — the sauron branch added `BytesServed` to `SessionInfo`, reuse it) and + per download (already tracked). +- A rolling `bytes-per-period` ledger (durable, alongside the ban ledger). When a + user exceeds it, the brain flags them → Tier-1 lease-required → new streams/ + downloads are refused and in-flight ones cut (`WatchAndCut`) ≤TTL. This is the + only thing that stops **D19 sequential direct-play ripping** and **C16 token + fan-out that manifests as sustained egress**. +- **Close the compat-download hole (D18, widest gap):** route + `/Items/{id}/Download` through the same `QuantityLimiter` + byte budget as native + downloads (see §4 — today it has *no* quota at all). This is the single + highest-value fix in the whole plan. + +### Pillar 4 — Detection heuristics & perimeter (C14, E21–E29) — partly new, partly gap-closing + +- **Re-streaming heuristic (C14/C16) — build it (it does not exist; matrix + correction #2).** The fingerprints already flow through `streammonitor` + (`ClientIP`, `ClientName`, `MediaFileID`, `BytesServed`). Add a consumer rule: + distinct client-IPs per `sid`, sustained throughput vs. media bitrate, distinct + concurrent titles per `uid`. Output = a flag → Tier-1 demotion. Ship **disabled**; + tune against real traffic before enabling (false-kill risk). +- **CPU/transcode exhaustion (E29):** add a **per-node concurrent-transcode cap** + at the node's `handleStart` admission (today `reconstructSem` guards only the + restart path; fresh starts are unbounded). Wire `nodepool.MaxJobs` sensible + defaults. +- **API-flood perimeter (E25–E28):** extend the existing `internal/ratelimit` to + the **jellycompat surface** and `/auth/refresh` (both currently unthrottled), and + add a connection cap (`netutil.LimitListener` / per-IP). These live in existing + subsystems — gap-closing, not new architecture. +- **Library-list load (E21–E24):** out of scope for the stream enforcer, but the + cold philosophy applies — add per-client browse-cost accounting + a cache/ + singleflight on the *general* paged browse (today only the Latest/hub rails are + cached). Track separately; note here so it is not forgotten. + +--- + +## 4. Download-quota verification (answering the explicit question) + +**Claim under test:** the admin "max concurrent downloads per user" / "max downloads +per period" / "period duration" settings let an operator cap e.g. **20 downloads per +30 days**, and that cap actually holds. + +**Result: it HOLDS for the native download routes, and is BYPASSED on the compat +route.** + +- Config keys exist and are operator-settable: `download.max_concurrent_per_user` + (default 3), `download.max_per_period` (default **0 = unlimited**), + `download.period_duration` (default 24h) — `internal/config/db_loader.go:545–553`. + Setting `max_per_period=20`, `period_duration=720h` gives "20 per 30 days". +- The period check is enforced **at download-record creation** at all four creation + sites (single, direct, series batch, season batch) — + `internal/downloads/service.go:356,427,597,689` call `limiter.Check(...)` *before* + inserting rows. +- **Crucially, the period count is not defeatable by completing downloads.** + `CountByUserSince` counts every row with `created_at >= now-period AND status NOT + IN ('cancelled','failed')` (`internal/downloads/repo.go:190–200`) — i.e. + **completed downloads still consume the period quota**. Only the *concurrent* + gate (`CountActiveByUser`, `status IN ('queued','downloading','preparing')`, + `repo.go:176–186`) frees on completion. So a `create→complete→create` loop can + keep 3-in-flight forever but **cannot** exceed 20-in-30-days. Cancelled/failed are + excluded so transient failures don't burn quota — correct. + +**Caveats the operator must know:** + +1. **Compat `/Items/{id}/Download` bypasses all of it.** `HandleDownload` + (`internal/jellycompat/streams.go`) creates no download row and never calls + `QuantityLimiter` (grep confirms zero references). An Infuse-style client + downloads unlimited full files regardless of the 20/30d setting. → Pillar 2/§3 + fix: route it through the same limiter + byte budget. +2. **The period quota counts downloads, not bytes.** "20 downloads" of 20 different + 4K remuxes is ~1 TB; the quota bounds *count*, not *volume*. The Pillar-3 byte + budget is what bounds volume. +3. **Streaming-based ripping is unaffected.** Sequential direct-play (D19) creates + stream sessions, not download rows, so the download quota never sees it. +4. **Minor TOCTOU:** two concurrent `Check` calls can both pass before either + inserts; the batch path mitigates via `batchSize`, but a burst of single + creates could momentarily overshoot by the concurrency of the burst. Low + severity; note for a follow-up (advisory lock or insert-then-verify). + +**Bottom line for the operator:** yes, 20/30d works today for the native app's +downloads and can't be gamed by finishing downloads — but it does **not** cover +Infuse/compat downloads or streaming rips, and it counts files not bytes. The plan +closes those three gaps in Pillars 2–3. + +--- + +## 5. Matrix coverage after this plan + +| Matrix rows | Today (sauron branch) | After this plan | +|---|---|---| +| A1–A3 over-cap | ✅ admission (once bug #1 fixed) | ✅ + reliable cross-node backstop | +| A4 buffer-ahead | ❌ invisible window | ✅ ffmpeg-liveness (Pillar 2) | +| A5 progress-withhold | ✅ | ✅ (+ regression test) | +| A6–A8 churn / multi-replica / dormant enforcer | ⚠️ | ✅ bug-fix + replica publish + lease-for-suspects | +| B9–B12 admin kill/ban/restart | ✅ | ✅ (durable ban ledger, unchanged) | +| B13 Redis outage | ⚠️ fail-open (abuser escapes) | ✅ flagged suspects fail-closed; legit fail-open | +| C14 under-cap re-stream | ❌ | ⚠️→✅ heuristic (Pillar 4, ship disabled) | +| C15 over-cap re-stream | ✅ | ✅ + labeled | +| C16 token hoard/fan-out | ❌ | ⚠️ caught via byte budget when it manifests as egress | +| D17 native rip | ⚠️ count-only | ✅ byte budget (Pillar 3) | +| D18 compat rip | ❌ no quota | ✅ route through quota + budget | +| D19 sequential direct-play rip | ❌ | ✅ byte budget | +| D20 admin stop rip | ⚠️ WatchAndCut | ✅ (unchanged) | +| E21–E24 library load | ❌ | ⚠️ browse-cost accounting + cache (tracked separately) | +| E25–E28 API floods | ⚠️ gaps | ✅ extend ratelimit to compat/`/auth/refresh` + conn cap | +| E29 CPU exhaustion | ❌ | ✅ per-node transcode cap | + +--- + +## 6. Phased plan + +1. **P0 — correctness first (small, high value):** fix the dormant-enforcer limit + source (matrix #1); close the compat-download quota hole (D18); add the per-node + transcode cap (E29). No new architecture — all three are bounded fixes with + outsized payoff. +2. **P1 — byte budget (Pillar 3):** rolling per-user bytes-per-period ledger across + downloads + stream egress; wire into a flag. Bounds all rip vectors. +3. **P2 — two-tier lease (Pillar 1):** add the `lease_until` field + Tier-1 + demotion for flagged sessions; fail-closed for suspects, fail-open for normal. +4. **P3 — liveness (Pillar 2):** ffmpeg-as-liveness for A4, behind a flag + tests. +5. **P4 — detection (Pillar 4):** re-streaming heuristic (ship disabled) + perimeter + ratelimit gap-closing + library-browse cost accounting. + +Each phase is independently shippable and independently valuable; P0 alone closes +the three worst holes. + +## AI-use disclosure + +Drafted with AI assistance (Claude) from a read-only audit of both branches and the +four referenced docs, plus targeted verification of the download-quota enforcement +path. No code behavior changed by writing this plan.