Files
silo-server/internal/playback/session.go
T
604bbf1a0f feat(playback): unified restart-resilient playback (native + jellycompat) (#174)
* feat(playback): unified restart-resilient playback via shared TranscodeManager

Make direct, remux, and native HLS transcode sessions survive a server
restart through one shared flow instead of per-method paths. A missing
in-memory session becomes a reconstruct trigger, not a 404: the server
rebuilds the session from a tiny durable recipe card plus the position the
client re-supplies on its next request.

- internal/playback/transcode_manager.go: shared TranscodeManager owning the
  transcodes map, recipe-card lifecycle, reconstruct single-flight +
  concurrency cap, LoadOrReconstructSession front door, ReconstructSession /
  ReconstructTranscode, and orphan cleanup. ~90% is logic moved out of the
  native handler (no behavior change), not new surface.
- internal/playback/recipecard.go + recipecard_postgres.go: RecipeCard with a
  PlayMethod discriminator (direct/remux/transcode; empty decodes as transcode
  for back-compat) behind a swappable, nil-safe RecipeStore interface backed by
  transcode_recipes.
- internal/playback/session.go: RegisterReconstructed inserts a rebuilt Session
  under its existing id (no UUID mint, no limit double-count, race-yielding).
- internal/playback/transcode.go: CloseProcess keeps the output dir so a
  reconstruct winner keeps serving; Close removes it.
- internal/api/handlers: drain the transcode lifecycle into the manager; wire
  reconstruct into the stream/segment serve paths; re-bind ownership to the live
  caller (refuse userID==0/mismatch); card-aware orphan cleanup.
- migrations: add transcode_recipes (expires_at TTL, filter-on-read, indexed).

Ownership stays two-factor: an authenticated caller AND a session.UserID that
matches; the card stores no secrets and identity is re-resolved per request.

Tests: recipe-card round-trip/legacy-decode/disabled-noop, RegisterReconstructed
insert/race/concurrency, close-vs-close-process dir semantics, the
LoadOrReconstructSession status matrix, and the reconstruct concurrency cap.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* feat(jellycompat): reconstruct transcodes across restart via shared manager

Bring Jellyfin (jellycompat) HLS playback onto the same restart-resilient flow
as the native path. Previously jellycompat owned a separate PlaybackHandler with
a private transcodes map and a duplicated transcode lifecycle that never grew the
reconstruct half, so an in-flight Jellyfin transcode died on restart and the next
segment request 404'd.

- Embed the shared playback.TranscodeManager and delete the duplicate lifecycle,
  so jellycompat gets reconstruct, the concurrency cap, the node-affinity rule,
  and the card lifecycle for free.
- internal/jellycompat/playback_sessions_postgres.go: DurableCompatPlaybackStore,
  a write-through cache over jellycompat_playback_sessions behind the new
  CompatPlaybackStore interface (nil pool degrades to cache-only). This persists
  the load-bearing PlaySessionId -> UpstreamSessionID mapping (plus media sources,
  route item id, seek) so it survives a restart instead of vanishing with the map.
- Write a recipe card on compat transcode start keyed by the upstream session id,
  using the native StreamAppUserID so the ownership re-bind matches; reconstruct
  the upstream session and the transcode seeked to the requested seg_NNNNN.
- migrations: add jellycompat_playback_sessions (expires_at TTL + compat_token
  index, full PlaybackSession in data JSONB).

Auth is mapped to the native user id before reconstruct so the same two-factor
ownership check and userID==0/mismatch refusal apply unchanged.

Tests: DB-gated (SILO_TEST_DATABASE_URL) durable-store round-trip proving a
session written by one instance reloads in a fresh one (the restart case), plus
a nil-pool cache-only path; existing handler tests updated to the manager.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* docs(playback): consolidate unified playback reconstruction design

Replace the three overlapping playback docs (the native Postgres
restart-resilience spec, the jellycompat plan, and the unification spec) with a
single self-contained design at
docs/superpowers/specs/unified-playback-reconstruct.md.

The doc leads with the unified design — the one-idea reconstruct model, a strong
visual flow of a restart mid-playback, the shared TranscodeManager + recipe card,
the two swappable durable stores, security, the concurrency cap and node-affinity
constraint, preconditions, and verification. The design history and rationale
(reconstruct-not-rehydrate, phased delivery, Redis-vs-Postgres, token-as-
descriptor, failure analysis) move to an appendix. It references no other md file.

AI-use: written with AI assistance.

* fix(playback): address review on restart-resilient playback

Four fixes from PR review of the unified reconstruction work:

- Rewrite the recipe card on audio-track change. HandleChangeAudioTrack only
  updated the in-memory session/transcode, so after a restart reconstruct
  resumed with the stale AudioTrackIndex/TranscodeAudio (and stale play method)
  from the start-time card. Re-save the card (direct/remux/transcode) with the
  switched state, mirroring the start-card pattern.
- Guard nil TranscodeManager in LoadOrReconstructSession and ReconstructSession.
  StreamHandler.TM is documented optional (tests/minimal setups); a missing
  session previously panicked in recipeEnabled instead of returning
  SessionMissing. ReconstructTranscode already guarded nil; make the two
  siblings consistent.
- Reject direct/remux cards in doReconstructTranscode before spawning ffmpeg, so
  a non-transcode card id can never enter the HLS reconstruction path.
- Log a non-success status from the remote transcode-node DELETE in
  CloseTranscodeSession; a 401/404/500 was previously silent.

AI-use: implemented with AI assistance.

* fix(playback): harden restart-resilient compat sessions

* feat(playback): token-carried reconstruction across restarts

Build on the shared TranscodeManager (introduced earlier in this branch) so a
playback session survives an API-server or transcode-node restart without the
client re-negotiating, and retire the Postgres transcode_recipes store in favor
of a recipe carried inside the signed stream token.

- RecipeCard encodes the byte-affecting encode parameters and rides inside the
  stream token; LoadOrReconstructSession rebuilds the in-memory Session (and,
  for integrated transcodes, the ffmpeg process) on a cold miss, single-flighted
  per session and paced by a spawn semaphore. Removes recipecard_postgres.go and
  the 20260617233705_add_transcode_recipes migration.
- transcodenode reconstructs a lost ffmpeg node-side from the forwarded token.
- TR-lease: proxy/streamauth enforce a revocation deny-marker on every served
  segment, with a 500ms Redis timeout, a bounded per-session "allowed" cache
  (3s TTL, expiry-first graceful eviction), and a degraded-fail-open counter.

Review hardening folded in:
- Manifest/segment handlers do the in-memory session lookup first and only
  verify the stream token on a reconstruct miss (token HMAC was per-segment).
- Copy-mode reconstruct never applies the encoded-only seg*dur seek, at spawn
  time or via the recovery path: RestartSeekTarget reports "unresolved" for a
  copy session whose manifest cannot yet map the segment, so the client retries
  instead of seeking to a fabricated source time.
- Crash teardown is a compare-and-delete (CloseTranscodeSessionIf returns
  whether it matched); the crash closure tears down the playback session only
  when it matched, so a session reconstructed under the same id is not killed.
- Reconstruct enforces the same per-user stream/transcode caps as a fresh start
  (RegisterReconstructedWithLimits), closing a token-replay slot bypass.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* feat(jellycompat): node-side transcode reconstruct via shared recipe store

Make Jellyfin-compat playback sessions survive a server or transcode-node
restart by reusing the shared TranscodeManager reconstruct path and a durable
recipe store, on top of the durable compat session store added earlier in this
branch.

- Node-side transcode reconstruct goes through the shared recipe store; the
  recipe is persisted to the control-plane store (Redis) when a dedicated
  transcode node is used so the node can rebuild ffmpeg after its own restart.
- Adopt the shared manager's API (3-arg OnFFmpegCrash carrying the dead session,
  guarded CloseTranscodeSessionIf, RegisterReconstructedWithLimits).

Review hardening folded in:
- Recipe lifecycle: noderecipe.Store gains Delete, called on deliberate
  teardown (stop, method-switch discard, node stop/force-reload) so a stopped
  session cannot be resurrected by a buffered request after a node restart;
  crash paths intentionally keep the recipe so a resume can reconstruct.
- Crash closure tears down the upstream session only when the guarded transcode
  close matched, so a reconstructed successor is never left orphaned.
- Copy-mode segment recovery surfaces a retryable not-found instead of a
  wrong-position restart, matching the native and node paths.
- Durable Update is now a SELECT ... FOR UPDATE transaction, removing the
  lost-update clobber that could silently drop a transcode recipe.
- Empty-token route resolution no longer falls back to an unbounded full-table
  scan; DB expiry filters bind the injected clock; the redundant re-Get is gone.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* docs(playback): consolidate restart-resilient playback design

Replace the superpowers spec with a single architecture record describing the
token-carried recipe card, the shared TranscodeManager reconstruct path for
direct/remux/transcode, the jellycompat durable session + node recipe store, and
the revocation-lease model with its fail-open tradeoff.

AI-use disclosure: written with AI assistance (Claude Code).

* docs(playback): correct jellycompat node-recipe rationale in comments

The noderecipe / transcode-node / jellycompat comments justified the Redis
recipe store with "a Jellyfin client cannot round-trip a token". The real
reason: the node-hop token is server-minted and could carry the recipe, but the
recipe is mutated in place under a stable session id (a /Sessions/Playing/Progress
audio switch restarts ffmpeg without re-minting the client's token) and a
third-party Jellyfin client cannot be driven to refresh a stale token, so the
node must reconstruct from a server-authoritative, node-reachable store.

Aligns the comments with docs/architecture/restart-resilient-playback.md §10.
Comment-only; no behavior change.

* refactor(playback): remove deny-lease revocation, defer to future PR

The deny-lease stream-revocation mechanism (the internal/streamauth
package, its silo:streamauth:<sid> Redis markers, the proxy Allowed()
enforcement, and the admin Stop/Terminate deny write) only ever enforced
on the offload-proxy topology and was a silent no-op on the integrated
single box and the dedicated transcode node. Rather than ship a partial
revocation feature that looks complete but isn't, remove it wholesale and
defer a uniform cross-topology revocation design to a dedicated follow-up.

Removed: internal/streamauth (package + tests); the LeaseDenier field,
StreamLeaseDenier interface, and denyStreamLease helper in playback.go;
the admin deny write; the router/main wiring; and the proxy verifyToken
Allowed() gate. The unified-reconstruct core (recipe-token,
LoadOrReconstructSession) is orthogonal and untouched.

Known limitation (now on every topology): admin Terminate and user Stop
tear down the live in-memory session and ffmpeg producer, but a still-valid
stream token can reconstruct the session until its 24h TTL expires. No
node-side byte-withholding ships in this PR.

docs/architecture/restart-resilient-playback.md is updated to mark the
revocation/deny-lease sections as deferred and to drop the overstated
"instant revocation on admin kill" claim.

* fix(playback): allow zero-caller bearer on transcode reconstruct

The authless HLS transcode delivery routes (master.m3u8 / segment) treat
the session UUID as the bearer credential, so a real request carries
requestUserID == 0. The live serve path already allows this, but
ReconstructSession hard-rejected a zero caller, so a request that worked
before a restart became SessionMissing -> 404 after the in-memory session
was gone, breaking the restart resilience these routes advertise.

Match the live-path contract in LoadOrReconstructSession: allow a zero
caller (UUID-as-bearer) and refuse only a non-zero caller that mismatches
the card owner. The reconstructed session is bound to card.UserID either
way. Adds TestReconstructSession_Ownership covering both cases.

* fix(jellycompat): re-persist recipe on local audio switch

A Jellyfin client switching audio on an integrated/local compat transcode
restarted live ffmpeg with the new track but did not re-persist
PlaybackSession.Recipe. The remote branch already re-persists via
startRemoteTranscode -> persistTranscodeRecipe. After a central restart,
reconstruct rebuilt ffmpeg from the stale Recipe.AudioTrackIndex, so the
integrated session resumed on the original audio track.

Persist the updated recipe (best-effort) after a successful Restart in the
local branch, mirroring the remote branch, so the durable
Recipe.AudioTrackIndex tracks live ffmpeg. Adds a regression test.

* fix(playback): strip stream token from proxied transcode-node URL

proxyToTranscodeNode appended the client's raw query string to the internal
transcode-node URL and logged that URL on transport failure. When a remote
transcode runs without a separate proxy node, that query carries
?st=<signed JWT> — a 24h bearer reconstruction descriptor exposing the
media path and recipe claims — placing the token into internal requests and
error logs.

Strip the "st" param before building targetURL, preserving any other query
params. The token is neither forwarded to the node nor present in the
logged URL. Header-forwarding of the token (so the node can reconstruct) is
a separate follow-up (#6).

* fix(playback): fail open on transient limit-provider error in reconstruct

During the reconstruct wave right after a restart (Postgres under peak
load), a transient limit-provider DB error was collapsed into a hard 404,
permanently stopping playback for a user within their limits. limitsForUser
wrapped any provider error, RegisterReconstructedWithLimits propagated it,
and ReconstructSession mapped every error to SessionMissing -> 404 -
indistinguishable from a genuine over-cap rejection.

Distinguish the two: tag provider errors with a new ErrLimitProviderUnavailable
sentinel and, during reconstruct, fail OPEN on a provider error (admit via
RegisterReconstructed + log a degraded warning) rather than refuse - mirroring
the reliability-first fail-open-on-dependency-error philosophy. A genuine
ErrTooManyStreams / ErrTooManyTranscodes over-cap still refuses. Adds tests
for both the fail-open and still-refused paths.

* fix(playback): forward stream token to transcode node as header

The dedicated transcode node's reconstruct path reads the stream token only
from the X-Silo-Stream-Token header, but proxyToTranscodeNode forwarded only
the node-API bearer token (and #5 now strips st from the URL). So when the
central API proxied to the node and the node self-restarted, it could not
reconstruct from the recipe-complete native token -> 404.

Capture st before stripping it from the URL, verify it at the API boundary
(streamtoken.Verify + SessionID match, mirroring the node's own check), and
forward it as X-Silo-Stream-Token. Best-effort: a missing/invalid token never
blocks the live proxy, and the token is still kept out of the forwarded URL
and logs.

* fix(playback): restart node ffmpeg on native remote audio switch

A native audio-track switch on an offloaded/remote transcode was a no-op at
the node yet returned 200 with a fresh URL: HandleChangeAudioTrack restarted
ffmpeg only when the API owned a LOCAL TranscodeSession, so for an offloaded
transcode the node kept serving the OLD audio (the node consults the token
only on a session miss). The replacement URL was also minted from identity-
only claims, so a later node restart 404'd.

For the offloaded transcode case (detected via session.TranscodeNodeURL),
POST a fresh /transcode/start to the node with the new AudioTrackIndex
(handleStart tears down and restarts ffmpeg) and mint the replacement proxy
URL from a full RecipeCard so reconstruct survives a node restart. The encode
recipe is derived from the durable session target fields plus the file,
mirroring HandleStartTranscode. A concrete SegmentDuration
(playback.DefaultSegmentDuration) is embedded rather than 0: the node's token
completeness gate treats SegmentDuration<=0 as incomplete and falls back to a
recipe store the native path never populates, which would 404 on a node
restart - the exact resilience this path provides. A failed node POST now
surfaces 502 rather than a false 200. Remux and non-offloaded (local)
transcode paths keep their prior identity-claim URLs unchanged.

Known limitation: Session does not persist the original SegmentDuration or
SubtitleTrackIndex/SubtitleBurnIn, so a remote audio switch resets subtitle
selection to none and assumes the default segment length; a client that
started with a non-default segment length will resegment on switch. Making
that state durable on the session is a follow-up.

* docs(playback): scrub stale deny-lease/revalidator comments

The deny-lease revocation mechanism and its "central revalidator" were removed
earlier in this branch, but four comments still described them as live
(transcode_manager.go, noderecipe/store.go, streamtoken/token.go,
proxy/server.go). Reword them to match the shipped behavior: ownership claims
are re-resolved at reconstruct, the noderecipe store shares Redis only with the
node-session tracker, and a sub-TTL hard cut depends on a node-side revocation
mechanism that is deferred to a future PR.

* fix(jellycompat): surface durable playback-session write failures

DurableCompatPlaybackStore.Update applied the in-memory mutation and then
swallowed every Postgres commit-failure path, returning nil. Callers that
promise restart resilience (persistTranscodeRecipe's recipe write, the
upstream-session binds in streams.go) were told the session was durably
persisted when only the cache held it, so a transient DB hiccup could leave
the next restart reloading a stale row (wrong audio track) or 404ing.

updateDB now returns the genuine DB round-trip error (begin/query/unmarshal/
marshal/exec/commit); Update propagates it while still applying the in-memory
mutation so live state stays correct. A nil pool and a genuinely absent/expired
row remain best-effort (return nil) — only real infrastructure failures
propagate, so existing rollback paths fire exactly when durability is lost.

Part of #174

* fix(playback): re-inject stream token into proxied transcode manifests

API-proxied remote transcode manifests dropped the reconstruct token from
their segment URLs, so playback died after a node or API restart. When a
remote transcode has no separate proxy node, the client loads its manifest via
the API-local path; proxyToTranscodeNode strips the signed token ("st") from
the forwarded URL (keeping it off node URLs and logs, forwarded only as the
X-Silo-Stream-Token header), and the node builds relative segment URIs from
that token-less query. The segment URLs the client received carried no token,
and the proxy only re-attached the header when an incoming segment request
already had "st" — which it never did — so a restart made those segments
non-reconstructable and they 404'd.

proxyToTranscodeNode now rewrites the manifest body at the boundary: every
segment and #EXT-X-MAP init URI gets the client-facing, API-verified token
re-appended (new playback.AppendManifestQueryParam helper), so the client's
later segment fetches carry "st" again and reconstruct after a restart. The
token still never reaches the node URL or its logs. Only 200 .m3u8 responses
are rewritten (Content-Length corrected); segments stream through untouched.

Part of #174

* fix(playback): preserve subtitle/cadence recipe across offloaded audio switch

Switching audio on a remote (offloaded) transcode with burned-in subtitles
silently dropped them, and reset a non-default segment cadence. The offloaded
audio-switch restart rebuilt the node start request from Session state, but
Session/SessionStreamState retained no subtitle or segment-duration state
(only the live local ts.Opts() and the RecipeCard did), so the branch
hard-coded SubtitleTrackIndex:-1, SubtitleBurnIn:false and
SegmentDuration:Default — signing that altered recipe into the replacement
stream token. An audio switch then changed bytes beyond audio selection, and
any later reconstruct kept the wrong no-subtitle/wrong-cadence recipe.

Persist the byte-affecting recipe on the session: SubtitleTrackIndex,
SubtitleBurnIn and SegmentDuration are added to Session/SessionStreamState,
populated at start (finalizeTranscodeStart) and on post-restart reconstruct
(ReconstructSession from the card), carried forward on every audio-switch
state update, and read back when rebuilding the offloaded node request and its
recipe card. The restart now reproduces the exact live stream. Also resolves
the M-4b non-default segment_duration reset.

Part of #174

* fix(playback): serialize transcode spawn paths with a per-session lock

Reconstruct was single-flighted only against other reconstructs, so a
restart-driven segment reconstruct racing a quality/seek/audio fresh start
could spawn two ffmpeg processes writing the same output directory at once —
segment corruption, partial-write closes, orphaned processes, and skewed
active-job accounting. The atomic register-after-spawn (GetOrRegister / the
reconstruct compare-on-register) prevented a map leak but not the concurrent
disk writers, because the losing path had already spawned. The dedicated
transcode node had the same split between handleStart and spawnReconstruct.

Add a refcounted per-session lifecycle lock to both TranscodeManager and the
node Server, held across "check existing -> spawn -> register":
- reconstruct (doReconstructTranscode / spawnReconstruct) re-checks under the
  lock and yields to any live session instead of spawning a duplicate;
- the native and jellycompat fresh-start paths take the lock around their
  spawn+register (the native path also closes any session a reconstruct rebuilt
  in the meantime so its fresh ffmpeg is the sole writer);
- the node handleStart holds it across teardown+spawn+register.
The refcount drops the map entry once no path holds/waits, keeping it bounded.
GetOrRegisterTranscodeSession is removed — the lock supersedes it and keeping a
register-after-spawn primitive would invite reintroducing the race.

Part of #174

* fix(playback): serialize restart re-spawn under the session lifecycle lock

TranscodeSession.Restart() releases s.mu across cancel -> wait-for-done ->
re-exec and spawns ffmpeg into opts.OutputDir without holding the per-session
lifecycle lock. LockSessionLifecycle's contract (fresh start, restart,
reconstruct) requires restart to hold it too, but all five callers invoked
Restart unlocked: native audio-switch and segment-recovery, compat
audio-switch and segment-recovery, and the transcode-node segment-recovery.

A restart racing another restart (audio-switch vs segment-recovery) or a
fresh-start/reconstruct could land two ffmpeg processes writing the same
segment directory -- mixed timelines, init.mp4/segment mismatch, and an
orphaned-but-still-writing ffmpeg -- the exact concurrent-writer corruption
the lifecycle lock exists to prevent.

Add RestartSessionLocked (TranscodeManager) and restartSessionLocked (node
Server) that hold LockSessionLifecycle only across the cancel->respawn
transition, re-check that the handle is still the live mapped session under
the lock, and return ErrSessionSuperseded rather than re-spawning a stale
handle. Route all five call sites through them. The lock is released before
callers wait on segments so recovery latency is unchanged.

Tests: gating (restart blocks until the lifecycle lock frees, then spawns),
concurrent-restart serialization, and superseded re-check on both the manager
(covers native + compat) and node lock owners.

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-02 14:23:14 -04:00

817 lines
24 KiB
Go

package playback
import (
"context"
"errors"
"fmt"
"strings"
"sync"
"time"
"github.com/google/uuid"
)
// Session represents an active playback session.
type Session struct {
ID string
UserID int
ProfileID string
MediaFileID int
RequestedMediaFileID int
PlayMethod PlayMethod
BasePlayMethod PlayMethod
TranscodeAudio bool // when true, remux should transcode audio to AAC
ClientIP string // resolved client IP for the playback session
ClientName string // reported playback client name, when available
ClientVersion string // reported playback client version, when available
ClientUserAgent string // trimmed request user agent for the playback session
TranscodeNodeURL string // URL of assigned transcode node (empty = local/integrated)
AudioTrackIndex int
StreamBitrateKbps int // currently delivered bitrate, when known
TargetResolution string // requested output resolution for transcodes
TargetVideoCodec string // requested output video codec for transcodes
TargetAudioCodec string // requested output audio codec when audio is transcoded
TargetBitrateKbps int // requested output bitrate cap for transcodes
TranscodeHWAccel string // effective hardware acceleration mode for transcodes
// Byte-affecting transcode recipe fields the offloaded restart path needs to
// rebuild the exact same stream after an audio switch. Local transcodes read
// these from the live ts.Opts(); offloaded transcodes own no local runtime, so
// the session is the only place to recover them (see HandleChangeAudioTrack).
SubtitleTrackIndex int // -1 = no subtitles
SubtitleBurnIn bool
SegmentDuration int // HLS segment length in seconds (cadence)
Position float64
IsPaused bool
HasWebSocket bool
HasRealtimeConnection bool
DisableProgressPersistence bool
StartedAt time.Time
UpdatedAt time.Time
LastActivityAt time.Time
activeTransportCount int
}
// SessionStreamState stores the mutable stream-specific details that can
// change after a session is created (audio track, client IP, transcode target,
// and reported bitrate).
type SessionStreamState struct {
PlayMethod PlayMethod
BasePlayMethod PlayMethod
AudioTrackIndex int
TranscodeAudio bool
ClientIP string
ClientName string
ClientVersion string
ClientUserAgent string
StreamBitrateKbps int
TargetResolution string
TargetVideoCodec string
TargetAudioCodec string
TargetBitrateKbps int
TranscodeHWAccel string
// Byte-affecting transcode recipe fields preserved so an offloaded restart
// (e.g. audio switch) can rebuild the exact same stream. SubtitleTrackIndex
// defaults to 0 on a zero-value state; callers that manage subtitles must set
// it explicitly (-1 for none) — burn-in is additionally gated by
// SubtitleBurnIn so a zero index never burns track 0 by accident.
SubtitleTrackIndex int
SubtitleBurnIn bool
SegmentDuration int
}
type clientInfoContextKey struct{}
// ClientInfo carries best-effort client metadata from request handling into
// the playback session manager.
type ClientInfo struct {
Name string
Version string
UserAgent string
}
// WithClientInfo stores playback client metadata on a context.
func WithClientInfo(ctx context.Context, info ClientInfo) context.Context {
if ctx == nil {
ctx = context.Background()
}
return context.WithValue(ctx, clientInfoContextKey{}, info)
}
// ClientInfoFromContext returns playback client metadata stored on a context.
func ClientInfoFromContext(ctx context.Context) ClientInfo {
if ctx == nil {
return ClientInfo{}
}
info, _ := ctx.Value(clientInfoContextKey{}).(ClientInfo)
return info
}
// SessionManager tracks active playback sessions and enforces stream limits.
type SessionManager struct {
sessions map[string]*Session
mu sync.RWMutex
maxStreams int
maxTranscodes int
limitProvider SessionLimitProvider
activeGrace time.Duration
pausedGrace time.Duration
expireHook func(*Session)
}
// SessionLimits stores per-user admission limits. Zero values mean unlimited.
type SessionLimits struct {
MaxStreams int
MaxTranscodes int
}
// SessionLimitProvider returns the current admission limits for a user.
type SessionLimitProvider func(ctx context.Context, userID int) (SessionLimits, error)
const (
// DefaultActiveSessionGrace is how long an unpaused session may go without
// observed playback activity before it stops counting toward limits.
DefaultActiveSessionGrace = 45 * time.Second
// DefaultPausedSessionGrace is the longer grace period for paused
// sessions. It must comfortably cover an intentional pause (dinner
// break, phone call): reaping a paused session kills its transcode
// and there is currently no revival path, so a too-short grace makes
// pressing Play after a long pause freeze the client (issue #243).
// Keep in sync with pausedSessionGrace in internal/worker/cleanup.go.
DefaultPausedSessionGrace = 30 * time.Minute
)
// NewSessionManager creates a SessionManager with the given concurrency limits.
// maxStreams limits total active streams per user.
// maxTranscodes limits concurrent transcode streams per user.
func NewSessionManager(maxStreams, maxTranscodes int) *SessionManager {
return &SessionManager{
sessions: make(map[string]*Session),
maxStreams: maxStreams,
maxTranscodes: maxTranscodes,
activeGrace: DefaultActiveSessionGrace,
pausedGrace: DefaultPausedSessionGrace,
}
}
// SetLimitProvider overrides the manager defaults with dynamic per-user
// limits. The constructor limits remain the fallback when no provider is set.
func (m *SessionManager) SetLimitProvider(provider SessionLimitProvider) {
m.mu.Lock()
defer m.mu.Unlock()
m.limitProvider = provider
}
// SetLivenessGracePeriods overrides the grace periods used by admission
// control and stale-session cleanup.
func (m *SessionManager) SetLivenessGracePeriods(active, paused time.Duration) {
m.mu.Lock()
defer m.mu.Unlock()
if active > 0 {
m.activeGrace = active
}
if paused > 0 {
m.pausedGrace = paused
}
}
// SetExpirationHook registers a callback that runs after a session is removed
// by stale cleanup. The hook executes outside the manager lock.
func (m *SessionManager) SetExpirationHook(fn func(*Session)) {
m.mu.Lock()
defer m.mu.Unlock()
m.expireHook = fn
}
func normalizeClientMetadataValue(value string, maxLen int) string {
value = strings.TrimSpace(value)
if maxLen > 0 && len(value) > maxLen {
value = value[:maxLen]
}
return value
}
// StartSession creates a new playback session using the same file as both the
// requested and effective source.
func (m *SessionManager) StartSession(userID int, profileID string, fileID int, method PlayMethod, transcodeAudio bool) (*Session, error) {
return m.StartSessionWithContext(context.Background(), userID, profileID, fileID, method, transcodeAudio)
}
// StartSessionWithContext creates a new playback session using the same file
// as both the requested and effective source.
func (m *SessionManager) StartSessionWithContext(
ctx context.Context,
userID int,
profileID string,
fileID int,
method PlayMethod,
transcodeAudio bool,
) (*Session, error) {
return m.StartSessionWithFilesContext(ctx, userID, profileID, fileID, fileID, method, transcodeAudio)
}
// StartSessionWithFiles creates a new playback session after checking
// concurrency limits. requestedFileID is the user's requested version while
// effectiveFileID is the file currently backing playback.
// Returns ErrTooManyStreams if the user has reached the max active stream count.
// Returns ErrTooManyTranscodes if the user has reached the max transcode count
// and the requested method is transcode.
func (m *SessionManager) StartSessionWithFiles(
userID int,
profileID string,
effectiveFileID int,
requestedFileID int,
method PlayMethod,
transcodeAudio bool,
) (*Session, error) {
return m.StartSessionWithFilesContext(context.Background(), userID, profileID, effectiveFileID, requestedFileID, method, transcodeAudio)
}
// StartSessionWithFilesContext creates a new playback session after checking
// concurrency limits with request-scoped limit lookup.
func (m *SessionManager) StartSessionWithFilesContext(
ctx context.Context,
userID int,
profileID string,
effectiveFileID int,
requestedFileID int,
method PlayMethod,
transcodeAudio bool,
) (*Session, error) {
if ctx == nil {
ctx = context.Background()
}
limits, err := m.limitsForUser(ctx, userID)
if err != nil {
return nil, err
}
m.mu.Lock()
defer m.mu.Unlock()
// Enforce stream limits.
if limits.MaxStreams > 0 && m.activeCountLocked(userID) >= limits.MaxStreams {
return nil, ErrTooManyStreams
}
// Enforce transcode limits.
if method == PlayTranscode && limits.MaxTranscodes > 0 && m.transcodeCountLocked(userID) >= limits.MaxTranscodes {
return nil, ErrTooManyTranscodes
}
now := time.Now()
clientInfo := ClientInfoFromContext(ctx)
s := &Session{
ID: uuid.New().String(),
UserID: userID,
ProfileID: profileID,
MediaFileID: effectiveFileID,
RequestedMediaFileID: requestedFileID,
PlayMethod: method,
BasePlayMethod: method,
TranscodeAudio: transcodeAudio,
Position: 0,
IsPaused: false,
ClientName: normalizeClientMetadataValue(clientInfo.Name, 128),
ClientVersion: normalizeClientMetadataValue(clientInfo.Version, 64),
ClientUserAgent: normalizeClientMetadataValue(clientInfo.UserAgent, 512),
StartedAt: now,
UpdatedAt: now,
LastActivityAt: now,
}
m.sessions[s.ID] = s
return s, nil
}
// RegisterReconstructed re-inserts a session under an existing ID after the
// in-memory state was lost (e.g. a server restart). Unlike StartSession* it
// does NOT mint a new UUID and does NOT run admission/limit accounting: the
// session already existed and was admitted before the restart, so counting it
// again would be wrong. If a live session with the same ID already exists
// (a concurrent reconstruct won the race), the existing one is returned and
// the caller's copy is discarded.
//
// The caller is responsible for having re-bound s.UserID to the live
// authenticated request before calling this — RegisterReconstructed performs
// no authorization itself.
func (m *SessionManager) RegisterReconstructed(s *Session) *Session {
if s == nil || s.ID == "" {
return s
}
m.mu.Lock()
defer m.mu.Unlock()
if existing, ok := m.sessions[s.ID]; ok {
return existing
}
now := time.Now()
if s.StartedAt.IsZero() {
s.StartedAt = now
}
s.UpdatedAt = now
s.LastActivityAt = now
m.sessions[s.ID] = s
return s
}
// RegisterReconstructedWithLimits is RegisterReconstructed plus the same per-user
// admission caps StartSession enforces. Token-carried reconstruct replays a
// signed recipe to rebuild a session lost to a restart; without a cap check a
// client could replay one token repeatedly (or after legitimately reaching its
// limit) and reconstruct past the per-user concurrent stream/transcode caps,
// since RegisterReconstructed skips admission accounting.
//
// Legitimately reconstructing a user's own surviving sessions still succeeds:
// the cap counts the user's *currently-live* sessions, and the one being rebuilt
// is not yet in the map, so the first MaxStreams reconstructs admit. Only the
// over-cap replay is refused (ErrTooManyStreams / ErrTooManyTranscodes). If an
// identical session id is already live (a concurrent reconstruct won), it is
// returned without re-counting. Caps are looked up via the same limit provider
// as StartSession.
func (m *SessionManager) RegisterReconstructedWithLimits(ctx context.Context, s *Session) (*Session, error) {
if s == nil || s.ID == "" {
return s, nil
}
if ctx == nil {
ctx = context.Background()
}
limits, err := m.limitsForUser(ctx, s.UserID)
if err != nil {
return nil, err
}
m.mu.Lock()
defer m.mu.Unlock()
if existing, ok := m.sessions[s.ID]; ok {
return existing, nil
}
// The session being reconstructed is not yet in the map, so the live counts
// reflect the user's *other* sessions; admitting one more must stay within cap.
if limits.MaxStreams > 0 && m.activeCountLocked(s.UserID) >= limits.MaxStreams {
return nil, ErrTooManyStreams
}
if s.PlayMethod == PlayTranscode && limits.MaxTranscodes > 0 &&
m.transcodeCountLocked(s.UserID) >= limits.MaxTranscodes {
return nil, ErrTooManyTranscodes
}
now := time.Now()
if s.StartedAt.IsZero() {
s.StartedAt = now
}
s.UpdatedAt = now
s.LastActivityAt = now
m.sessions[s.ID] = s
return s, nil
}
func (m *SessionManager) limitsForUser(ctx context.Context, userID int) (SessionLimits, error) {
m.mu.RLock()
provider := m.limitProvider
limits := SessionLimits{
MaxStreams: m.maxStreams,
MaxTranscodes: m.maxTranscodes,
}
m.mu.RUnlock()
if provider == nil {
return limits, nil
}
limits, err := provider(ctx, userID)
if err != nil {
// Tag provider failures with ErrLimitProviderUnavailable so the
// reconstruct admission path can distinguish a transient limit-lookup
// failure (which it may fail open on) from a genuine over-cap rejection.
return SessionLimits{}, fmt.Errorf("load session limits for user %d: %w",
userID, errors.Join(ErrLimitProviderUnavailable, err))
}
return limits, nil
}
// UpdateProgress updates the playback position and pause state for a session.
func (m *SessionManager) UpdateProgress(sessionID string, position float64, isPaused bool) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.Position = position
s.IsPaused = isPaused
m.touchSessionLocked(s)
return nil
}
// UpdateAudioTrack updates the audio track index and optionally the play
// method for a session. Used when switching audio tracks mid-playback.
func (m *SessionManager) UpdateAudioTrack(sessionID string, audioTrackIndex int, method PlayMethod) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.AudioTrackIndex = audioTrackIndex
s.BasePlayMethod = method
if s.PlayMethod != PlayTranscode || method == PlayTranscode {
s.PlayMethod = method
}
m.touchSessionLocked(s)
return nil
}
// UpdateStreamState updates the live stream details for a session. This keeps
// the session manager's authoritative copy in sync with user-driven changes
// like audio track switches and quality changes.
func (m *SessionManager) UpdateStreamState(sessionID string, state SessionStreamState) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
if state.PlayMethod != "" {
s.PlayMethod = state.PlayMethod
}
if state.BasePlayMethod != "" {
s.BasePlayMethod = state.BasePlayMethod
}
s.AudioTrackIndex = state.AudioTrackIndex
s.TranscodeAudio = state.TranscodeAudio
s.ClientIP = state.ClientIP
if value := normalizeClientMetadataValue(state.ClientName, 128); value != "" {
s.ClientName = value
}
if value := normalizeClientMetadataValue(state.ClientVersion, 64); value != "" {
s.ClientVersion = value
}
if value := normalizeClientMetadataValue(state.ClientUserAgent, 512); value != "" {
s.ClientUserAgent = value
}
s.StreamBitrateKbps = state.StreamBitrateKbps
s.TargetResolution = state.TargetResolution
s.TargetVideoCodec = state.TargetVideoCodec
s.TargetAudioCodec = state.TargetAudioCodec
s.TargetBitrateKbps = state.TargetBitrateKbps
s.TranscodeHWAccel = state.TranscodeHWAccel
s.SubtitleTrackIndex = state.SubtitleTrackIndex
s.SubtitleBurnIn = state.SubtitleBurnIn
s.SegmentDuration = state.SegmentDuration
m.touchSessionLocked(s)
return nil
}
// SetTranscodeNodeURL assigns a transcode node URL to an existing session.
func (m *SessionManager) SetTranscodeNodeURL(sessionID, url string) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.TranscodeNodeURL = url
m.touchSessionLocked(s)
return nil
}
// SetEffectiveMediaFileID updates the currently delivered source file while
// preserving the originally requested file selection.
func (m *SessionManager) SetEffectiveMediaFileID(sessionID string, fileID int) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
if fileID > 0 {
s.MediaFileID = fileID
}
m.touchSessionLocked(s)
return nil
}
// SetWebSocket marks whether a WebSocket liveness connection is active for a session.
func (m *SessionManager) SetWebSocket(sessionID string, connected bool) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.HasWebSocket = connected
m.touchSessionLocked(s)
return nil
}
// SetRealtimeConnection marks whether a realtime control connection is active for a session.
func (m *SessionManager) SetRealtimeConnection(sessionID string, connected bool) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.HasRealtimeConnection = connected
// The admin/session sync layer still exposes a generic websocket flag.
s.HasWebSocket = connected
m.touchSessionLocked(s)
return nil
}
// SetProgressPersistenceDisabled controls whether session progress updates and
// stop events should write resume/history state. This is useful for players
// whose resume timeline is not the same as the session's file-local timeline.
func (m *SessionManager) SetProgressPersistenceDisabled(sessionID string, disabled bool) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.DisableProgressPersistence = disabled
m.touchSessionLocked(s)
return nil
}
// TouchActivity refreshes the session's activity timestamp without changing
// any other playback state.
func (m *SessionManager) TouchActivity(sessionID string) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
m.touchSessionLocked(s)
return nil
}
// BeginTransport increments the count of in-flight media transport requests
// for the session and refreshes its activity timestamp.
func (m *SessionManager) BeginTransport(sessionID string) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
s.activeTransportCount++
m.touchSessionLocked(s)
return nil
}
// EndTransport decrements the count of in-flight media transport requests for
// the session and refreshes its activity timestamp.
func (m *SessionManager) EndTransport(sessionID string) error {
m.mu.Lock()
defer m.mu.Unlock()
s, ok := m.sessions[sessionID]
if !ok {
return ErrSessionNotFound
}
if s.activeTransportCount > 0 {
s.activeTransportCount--
}
m.touchSessionLocked(s)
return nil
}
// StopSession removes a session from the manager.
func (m *SessionManager) StopSession(sessionID string) error {
m.mu.Lock()
defer m.mu.Unlock()
if _, ok := m.sessions[sessionID]; !ok {
return ErrSessionNotFound
}
delete(m.sessions, sessionID)
return nil
}
// GetSession returns the session with the given ID, or ErrSessionNotFound.
func (m *SessionManager) GetSession(sessionID string) (*Session, error) {
m.mu.RLock()
defer m.mu.RUnlock()
s, ok := m.sessions[sessionID]
if !ok {
return nil, ErrSessionNotFound
}
// Return a copy to avoid races.
cp := *s
return &cp, nil
}
// GetUserSessions returns all active sessions for a user.
func (m *SessionManager) GetUserSessions(userID int) []*Session {
m.mu.RLock()
defer m.mu.RUnlock()
var result []*Session
for _, s := range m.sessions {
if s.UserID == userID {
cp := *s
result = append(result, &cp)
}
}
return result
}
// GetSessionsByMediaFileID returns active sessions associated with the given file.
func (m *SessionManager) GetSessionsByMediaFileID(fileID int) []*Session {
m.mu.RLock()
defer m.mu.RUnlock()
if fileID <= 0 {
return nil
}
var result []*Session
for _, s := range m.sessions {
if s.MediaFileID != fileID && s.RequestedMediaFileID != fileID {
continue
}
cp := *s
result = append(result, &cp)
}
return result
}
// ActiveCount returns the number of active sessions for a user.
func (m *SessionManager) ActiveCount(userID int) int {
m.mu.RLock()
defer m.mu.RUnlock()
return m.activeCountLocked(userID)
}
// TranscodeCount returns the number of active transcode sessions for a user.
func (m *SessionManager) TranscodeCount(userID int) int {
m.mu.RLock()
defer m.mu.RUnlock()
return m.transcodeCountLocked(userID)
}
// activeCountLocked counts active sessions for a user. Caller must hold the lock.
func (m *SessionManager) activeCountLocked(userID int) int {
now := time.Now()
count := 0
for _, s := range m.sessions {
if s.UserID == userID && m.countsTowardLimitsLocked(s, now) {
count++
}
}
return count
}
// transcodeCountLocked counts transcode sessions for a user. Caller must hold the lock.
func (m *SessionManager) transcodeCountLocked(userID int) int {
now := time.Now()
count := 0
for _, s := range m.sessions {
if s.UserID == userID && s.PlayMethod == PlayTranscode && m.countsTowardLimitsLocked(s, now) {
count++
}
}
return count
}
// AllSessions returns a snapshot of all active sessions. Each session is
// copied to avoid data races with concurrent updates.
func (m *SessionManager) AllSessions() []*Session {
m.mu.RLock()
defer m.mu.RUnlock()
result := make([]*Session, 0, len(m.sessions))
for _, s := range m.sessions {
cp := *s
result = append(result, &cp)
}
return result
}
// CleanExpired removes sessions whose last playback activity exceeds maxIdle.
// Paused sessions receive a 3x grace period for backwards compatibility.
func (m *SessionManager) CleanExpired(maxIdle time.Duration) []*Session {
return m.CleanInactive(maxIdle, maxIdle*3)
}
// CleanStale removes sessions that have exceeded the manager's configured
// liveness grace windows.
func (m *SessionManager) CleanStale() []*Session {
m.mu.RLock()
active := m.activeGrace
paused := m.pausedGrace
m.mu.RUnlock()
return m.CleanInactive(active, paused)
}
// CleanInactive removes sessions whose last playback activity exceeds the
// provided grace period. Sessions with an active media transport request are
// preserved even if they have not emitted a recent heartbeat yet.
func (m *SessionManager) CleanInactive(activeIdle, pausedIdle time.Duration) []*Session {
m.mu.Lock()
now := time.Now()
var expired []*Session
for id, s := range m.sessions {
if s.activeTransportCount > 0 {
continue
}
if m.sessionIsInactiveLocked(s, now, activeIdle, pausedIdle) {
cp := *s
expired = append(expired, &cp)
delete(m.sessions, id)
}
}
hook := m.expireHook
m.mu.Unlock()
if hook != nil {
for _, s := range expired {
hook(s)
}
}
return expired
}
func (m *SessionManager) touchSessionLocked(s *Session) {
now := time.Now()
s.LastActivityAt = now
s.UpdatedAt = now
}
func (m *SessionManager) countsTowardLimitsLocked(s *Session, now time.Time) bool {
if s == nil {
return false
}
if s.activeTransportCount > 0 {
return true
}
return !m.sessionIsInactiveLocked(s, now, m.activeGrace, m.pausedGrace)
}
func (m *SessionManager) sessionIsInactiveLocked(s *Session, now time.Time, activeIdle, pausedIdle time.Duration) bool {
if s == nil {
return true
}
lastActivity := s.LastActivityAt
if lastActivity.IsZero() {
lastActivity = s.UpdatedAt
}
if lastActivity.IsZero() {
lastActivity = s.StartedAt
}
if lastActivity.IsZero() {
return false
}
grace := activeIdle
if s.IsPaused {
grace = pausedIdle
}
if grace <= 0 {
return !lastActivity.After(now)
}
return !lastActivity.Add(grace).After(now)
}
// String returns a human-readable summary of a session.
func (s *Session) String() string {
return fmt.Sprintf("Session{id=%s user=%d file=%d method=%s pos=%.1f paused=%v}",
s.ID, s.UserID, s.MediaFileID, s.PlayMethod, s.Position, s.IsPaused)
}