Files
silo-server/internal/streamenforcer/enforcer.go
T
CoffeeKnyte e5bf0155ad fix(playback): make revocation state converge and cutoffs credential-accurate
Five defects in revocation state and credential semantics. Lands after the
tracker-lifecycle batch on purpose: raising the over-cap TTL is only safe once
the count feeding it is trustworthy.

#13 -- a longer old revocation suppressed a newer cutoff. applyLocal kept or
replaced the WHOLE record by expiry, so when the existing revocation expired
later the new one was dropped entirely, including its newer RevokedAt. The
durable upsert did the same, with a comment documenting it as intentional.
RevokedAt is the user-kill CUTOFF, so this left a credential issued between the
two cutoffs valid -- a second admin kill after a user re-authenticates silently
failed to cut them. The two fields now merge independently: ExpiresAt stays
monotonic, RevokedAt advances to the later value, and reason follows the newer
cutoff. Both superseded comments are replaced rather than left contradicting the
code. Session-kind revocation still ignores RevokedAt, so the enforcer's
re-revoke cannot weaken a session kill.

Also fixed while here: Redis received the merged record but pub/sub published the
raw input one, so under pub/sub-only delivery (Redis down) an edge got the newer
short record without the older long expiry and lost monotonicity. Both now carry
the merged record.

#7 + M1 -- Postgres could indefinitely block the urgent Redis kill.
RevokeWithWarnings held the global opMu across all propagation, stripped the
caller's deadline with WithoutCancel, and did the durable Postgres upsert BEFORE
Redis, on a pool with no statement timeout. The local kill still applied, so
playback on that process was fine -- but edge propagation, pub/sub, the admin
response and every later revoke/unrevoke stalled behind the lock. Redis and
pub/sub now go first, and the detached context is bounded. WithoutCancel is kept
deliberately: propagation must outlive an aborted admin request.

opMu scope is deliberately NOT narrowed. mirrorToRedis is an unconditional SET
with no atomic merge, so same-process serialization is what stops an older value
overwriting a newer one; narrowing the lock would also let an unrevoke interleave
with a revoke's propagation. Bounding the context caps how long the lock can be
held, which is the actual reported harm. The remaining cross-replica race -- two
central replicas racing the same SET -- is documented, not half-fixed; it needs
A6's shared picture.

A2 / #6 -- a missed unrevoke got resurrected. In-memory tombstones already
existed, but being process-local they did not survive a restart or reach a
replica that missed the pub/sub event, so maintain's durable self-heal
re-Upserted the surviving entry and the ban returned. Tombstones are now durable,
via two nullable columns on stream_revocations rather than a second table: a
tombstone is a state of the same key, and it needs its own expiry horizon
separate from the revocation's. The upsert rejects a stale replica's write while
a tombstone is live but lets a genuinely newer revocation clear it, and warm
paths apply tombstones BEFORE revocations so an un-banned key cannot be restored
as a live kill. Tombstones are pruned on the same sweep, so the table cannot grow
without bound.

A1 / #3 -- over-cap kills reopened after 5 minutes while the token stayed
reconstructable for 24h. The TTL now derives from playback.MaxTokenTTL rather
than duplicating 24h, behind a validated setting.

Critically, the enforcer uses a revoke-if-absent path rather than re-revoking.
Expiry is monotonic and the enforcer re-evaluates every 30s, so a plain long TTL
would slide expiry forward by another full lifetime on every pass -- making a
wrong kill effectively permanent for as long as any stale record persisted, with
only an explicit unrevoke to recover it. Admin Revoke keeps its monotonic
behaviour; only the enforcer's own repeat kill is non-extending. The setting is
documented as affecting future revocations only, since monotonic expiry means it
cannot shorten one already issued.

A3 / #5 -- the user cutoff compared against a fresh time.Now() taken at request
entry, so a request from a pre-cutoff login could look post-cutoff and escape the
kill. The credential time is now the access token's iat.

Two deliberate choices worth stating. API-key credentials carry no issue time, so
they pass the zero time and, per IsRevoked's documented contract, are never
matched by a user cutoff: a user kill provably cannot cut an API-key-owned pour.
That is an accepted, logged, documented hole -- and strictly better than
time.Now(), which actively defeats the cutoff. And jellycompat uses the compat
session's CreatedAt rather than the bridged Silo token's iat, because that token
refreshes without a new Jellyfin login, so its iat would advance on refresh and
let a refreshed credential slip past a cutoff.

Stream tokens are now bound to their route: a token whose SessionID does not match
the URL's session_id is rejected with 403 instead of being silently ignored,
matching the reconstruction helper that already refused a different session.

Per-login logout cuts remain out of scope -- they need per-login identity in the
stream credential. S5 (the (sessionID, userID, startedAt) clump) is rejected as
ceremony now that A3 is the iat option rather than the generation model.

#12 (closing an RSS feed does not cut its current pour) is deferred: it needs a
namespaced revocation id that cannot collide with real session ids, that id
threaded onto public feed requests, and protection against a new feed inheriting
an old tombstone.

Part of #305.
2026-07-30 12:27:05 +00:00

207 lines
7.3 KiB
Go

// Package streamenforcer is the async decision loop ("the brain") of the
// monitor-and-kill design. It runs on central only, off the hot path: it reads
// the authoritative live-streams snapshot, compares each user's live count to
// that user's limit, and issues revocations for the over-cap sessions. The
// revocation kill switch (internal/streamrevoke) then stops them at the edge
// within one propagation/poll interval.
//
// Every enforcement reason (over-cap here, admin terminate and account
// revocation elsewhere) collapses to the same action: write a revocation. This
// package owns only the over-cap rule; other reasons call the revoker directly.
package streamenforcer
import (
"context"
"fmt"
"log/slog"
"sort"
"time"
"github.com/Silo-Server/silo-server/internal/playback"
"github.com/Silo-Server/silo-server/internal/streammonitor"
)
// DefaultInterval is how often the enforcer evaluates the live picture. The
// ~120s enforcement budget = this interval + the revocation propagation/poll.
const DefaultInterval = 30 * time.Second
const (
// RevocationTTLSetting controls the lifetime of future over-cap
// revocations. Changing it cannot shorten an existing monotonic kill; an
// explicit unrevoke is required to clear one.
RevocationTTLSetting = "playback.over_cap_revocation_ttl"
// MinRevocationTTL prevents a setting typo from making automated kills
// churn faster than the evaluation and propagation cadence.
MinRevocationTTL = 5 * time.Minute
// DefaultRevocationTTL matches the full reconstructable stream-token life.
// Derive it from the token contract instead of duplicating 24h here.
DefaultRevocationTTL = playback.MaxTokenTTL
MaxRevocationTTL = playback.MaxTokenTTL
)
// Revoker is the subset of *streamrevoke.Store the enforcer needs. It uses the
// if-absent variant so recurring evaluation never slides its own kill expiry.
type Revoker interface {
RevokeSessionForIfAbsent(ctx context.Context, sessionID, reason string, ttl time.Duration) (bool, error)
}
// LimitFunc returns the maximum concurrent streams allowed for a user. A return
// of <= 0 means "unlimited" (no enforcement for that user), matching the
// SessionManager's convention. An error means the limit is currently unknown;
// the enforcer fails OPEN (does not kill) so a limit-lookup blip never
// terminates legitimate playback.
type LimitFunc func(ctx context.Context, userID int) (maxStreams int, err error)
// Enforcer periodically trims over-cap streams.
type Enforcer struct {
source streammonitor.Source
limits LimitFunc
revoker Revoker
interval time.Duration
revocationTTL time.Duration
now func() time.Time
}
// New builds an enforcer. interval <= 0 uses DefaultInterval.
func New(source streammonitor.Source, limits LimitFunc, revoker Revoker, interval time.Duration) *Enforcer {
return NewWithRevocationTTL(source, limits, revoker, interval, DefaultRevocationTTL)
}
// NewWithRevocationTTL builds an enforcer with the validated lifetime used for
// future automated revocations.
func NewWithRevocationTTL(
source streammonitor.Source,
limits LimitFunc,
revoker Revoker,
interval, revocationTTL time.Duration,
) *Enforcer {
if interval <= 0 {
interval = DefaultInterval
}
if revocationTTL < MinRevocationTTL || revocationTTL > MaxRevocationTTL {
revocationTTL = DefaultRevocationTTL
}
return &Enforcer{
source: source,
limits: limits,
revoker: revoker,
interval: interval,
revocationTTL: revocationTTL,
now: time.Now,
}
}
// ParseRevocationTTL validates the server setting. Empty uses the long default.
func ParseRevocationTTL(raw string) (time.Duration, error) {
if raw == "" {
return DefaultRevocationTTL, nil
}
ttl, err := time.ParseDuration(raw)
if err != nil {
return 0, fmt.Errorf("%s must be a duration: %w", RevocationTTLSetting, err)
}
if ttl < MinRevocationTTL || ttl > MaxRevocationTTL {
return 0, fmt.Errorf(
"%s must be between %s and %s",
RevocationTTLSetting, MinRevocationTTL, MaxRevocationTTL,
)
}
return ttl, nil
}
// Start runs the evaluation loop until ctx is canceled. Non-blocking.
func (e *Enforcer) Start(ctx context.Context) {
if e == nil || e.source == nil || e.limits == nil || e.revoker == nil {
return
}
go func() {
ticker := time.NewTicker(e.interval)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case <-ticker.C:
e.evaluate(ctx)
}
}
}()
}
// evaluate runs one pass: snapshot → per-user over-cap check → revoke victims.
// Exported behavior is covered by EvaluateOnce for tests.
func (e *Enforcer) evaluate(ctx context.Context) {
if err := e.EvaluateOnce(ctx); err != nil {
slog.DebugContext(ctx, "stream enforcer evaluate failed", "error", err)
}
}
// EvaluateOnce performs a single enforcement pass and returns the number of
// sessions revoked. Deterministic and side-effect-scoped for testing.
func (e *Enforcer) EvaluateOnce(ctx context.Context) error {
snap, err := e.source.Snapshot(ctx)
if err != nil {
return err
}
for userID, streams := range snap.ByUser() {
if userID <= 0 || len(streams) == 0 {
// user 0 == records with no resolved owner; never enforce against it.
continue
}
limit, err := e.limits(ctx, userID)
if err != nil {
// Fail open: a limit-lookup error must never kill legitimate streams.
slog.DebugContext(ctx, "stream enforcer: limit lookup failed; skipping user",
"user_id", userID, "error", err)
continue
}
if limit <= 0 || len(streams) <= limit {
continue
}
for _, victim := range e.selectVictims(streams, limit) {
created, err := e.revoker.RevokeSessionForIfAbsent(
ctx, victim.SessionID, "over_concurrent_stream_limit", e.revocationTTL,
)
if err != nil {
slog.WarnContext(ctx, "stream enforcer: revoke failed",
"user_id", userID, "session_id", victim.SessionID, "error", err)
continue
}
if !created {
continue
}
slog.InfoContext(ctx, "stream enforcer: revoked over-cap session",
"user_id", userID, "session_id", victim.SessionID,
"limit", limit, "live", len(streams))
}
}
return nil
}
// selectVictims returns the streams beyond the limit, keeping the `limit`
// MOST-RECENTLY-SERVED sessions and trimming the rest. Ordering by real serve
// activity (LastServedAt), not StartedAt, matters: after a network blip a client
// reconnects with a new session while the old ghost lingers in the monitor for
// up to its TTL. Keeping the freshly-served sessions means the live reconnect
// survives and the stale ghost is the one trimmed (and it would have aged out
// anyway). Falls back to StartedAt, then session id, for deterministic ties.
func (e *Enforcer) selectVictims(streams []streammonitor.LiveStream, limit int) []streammonitor.LiveStream {
ordered := make([]streammonitor.LiveStream, len(streams))
copy(ordered, streams)
// Most-recently-served first.
sort.SliceStable(ordered, func(i, j int) bool {
if !ordered[i].LastServedAt.Equal(ordered[j].LastServedAt) {
return ordered[i].LastServedAt.After(ordered[j].LastServedAt)
}
if !ordered[i].StartedAt.Equal(ordered[j].StartedAt) {
return ordered[i].StartedAt.After(ordered[j].StartedAt)
}
return ordered[i].SessionID < ordered[j].SessionID
})
if limit >= len(ordered) {
return nil
}
// Keep the first `limit` (freshest); revoke the rest (stalest).
return ordered[limit:]
}