Files
silo-server/internal/api
CoffeeKnyte e5bf0155ad fix(playback): make revocation state converge and cutoffs credential-accurate
Five defects in revocation state and credential semantics. Lands after the
tracker-lifecycle batch on purpose: raising the over-cap TTL is only safe once
the count feeding it is trustworthy.

#13 -- a longer old revocation suppressed a newer cutoff. applyLocal kept or
replaced the WHOLE record by expiry, so when the existing revocation expired
later the new one was dropped entirely, including its newer RevokedAt. The
durable upsert did the same, with a comment documenting it as intentional.
RevokedAt is the user-kill CUTOFF, so this left a credential issued between the
two cutoffs valid -- a second admin kill after a user re-authenticates silently
failed to cut them. The two fields now merge independently: ExpiresAt stays
monotonic, RevokedAt advances to the later value, and reason follows the newer
cutoff. Both superseded comments are replaced rather than left contradicting the
code. Session-kind revocation still ignores RevokedAt, so the enforcer's
re-revoke cannot weaken a session kill.

Also fixed while here: Redis received the merged record but pub/sub published the
raw input one, so under pub/sub-only delivery (Redis down) an edge got the newer
short record without the older long expiry and lost monotonicity. Both now carry
the merged record.

#7 + M1 -- Postgres could indefinitely block the urgent Redis kill.
RevokeWithWarnings held the global opMu across all propagation, stripped the
caller's deadline with WithoutCancel, and did the durable Postgres upsert BEFORE
Redis, on a pool with no statement timeout. The local kill still applied, so
playback on that process was fine -- but edge propagation, pub/sub, the admin
response and every later revoke/unrevoke stalled behind the lock. Redis and
pub/sub now go first, and the detached context is bounded. WithoutCancel is kept
deliberately: propagation must outlive an aborted admin request.

opMu scope is deliberately NOT narrowed. mirrorToRedis is an unconditional SET
with no atomic merge, so same-process serialization is what stops an older value
overwriting a newer one; narrowing the lock would also let an unrevoke interleave
with a revoke's propagation. Bounding the context caps how long the lock can be
held, which is the actual reported harm. The remaining cross-replica race -- two
central replicas racing the same SET -- is documented, not half-fixed; it needs
A6's shared picture.

A2 / #6 -- a missed unrevoke got resurrected. In-memory tombstones already
existed, but being process-local they did not survive a restart or reach a
replica that missed the pub/sub event, so maintain's durable self-heal
re-Upserted the surviving entry and the ban returned. Tombstones are now durable,
via two nullable columns on stream_revocations rather than a second table: a
tombstone is a state of the same key, and it needs its own expiry horizon
separate from the revocation's. The upsert rejects a stale replica's write while
a tombstone is live but lets a genuinely newer revocation clear it, and warm
paths apply tombstones BEFORE revocations so an un-banned key cannot be restored
as a live kill. Tombstones are pruned on the same sweep, so the table cannot grow
without bound.

A1 / #3 -- over-cap kills reopened after 5 minutes while the token stayed
reconstructable for 24h. The TTL now derives from playback.MaxTokenTTL rather
than duplicating 24h, behind a validated setting.

Critically, the enforcer uses a revoke-if-absent path rather than re-revoking.
Expiry is monotonic and the enforcer re-evaluates every 30s, so a plain long TTL
would slide expiry forward by another full lifetime on every pass -- making a
wrong kill effectively permanent for as long as any stale record persisted, with
only an explicit unrevoke to recover it. Admin Revoke keeps its monotonic
behaviour; only the enforcer's own repeat kill is non-extending. The setting is
documented as affecting future revocations only, since monotonic expiry means it
cannot shorten one already issued.

A3 / #5 -- the user cutoff compared against a fresh time.Now() taken at request
entry, so a request from a pre-cutoff login could look post-cutoff and escape the
kill. The credential time is now the access token's iat.

Two deliberate choices worth stating. API-key credentials carry no issue time, so
they pass the zero time and, per IsRevoked's documented contract, are never
matched by a user cutoff: a user kill provably cannot cut an API-key-owned pour.
That is an accepted, logged, documented hole -- and strictly better than
time.Now(), which actively defeats the cutoff. And jellycompat uses the compat
session's CreatedAt rather than the bridged Silo token's iat, because that token
refreshes without a new Jellyfin login, so its iat would advance on refresh and
let a refreshed credential slip past a cutoff.

Stream tokens are now bound to their route: a token whose SessionID does not match
the URL's session_id is rejected with 403 instead of being silently ignored,
matching the reconstruction helper that already refused a different session.

Per-login logout cuts remain out of scope -- they need per-login identity in the
stream credential. S5 (the (sessionID, userID, startedAt) clump) is rejected as
ceremony now that A3 is the iat option rather than the generation model.

#12 (closing an RSS feed does not cut its current pour) is deferred: it needs a
namespaced revocation id that cannot collide with real session ids, that id
threaded onto public feed requests, and protection against a new feed inheriting
an old tombstone.

Part of #305.
2026-07-30 12:27:05 +00:00
..