Commit Graph
9 Commits
Author SHA1 Message Date
c52ca7dd7a feat(admin): identify compat sessions and Android devices in the live session view (#495)
* feat(admin): identify Android devices by model in live session view

Android clients that send a bare default User-Agent (e.g.
"Dalvik/2.1.0 (Linux; U; Android 11; AFTKRT Build/RS8180.3729N)")
showed up as "Dalvik" in the admin live-session view, which tells an
operator nothing about the device.

Parse the model code out of the UA (the token between the last ';' and
"Build/") and map the Amazon Fire TV family and NVIDIA Shield to product
names. Unknown but parseable models fall back to "Android · <MODEL>"
instead of "Dalvik"; multi-word models like "Pixel 7" are preserved
whole. This is display-only: the session still stores the raw model code
in its user agent, and no response field or contract changes.

* feat(admin): mark Jellyfin-compat sessions with the JF pill by origin

The admin "JF" pill was derived at read time by substring-matching a
token list against the client name / user agent. A real Jellyfin
client that authenticates through the compat surface but sends a bare
User-Agent and no MediaBrowser client name (e.g. a Fire TV app) got no
pill, even though it plainly came through the Jellyfin API.

Stamp compat origin as immutable identity at session creation and carry
it through to the admin view:

- ClientInfo.IsCompat is set true in the jellycompat auth path; newSession
  copies it onto Session.IsJellyfinCompat.
- The flag rides the durable RecipeCard (next to the client metadata that
  already exists so the pill survives reconstruction) and is restored in
  ReconstructSession, so a server restart keeps the pill.
- buildLiveSessionSync -> worker.SessionSync -> a new compat_origin column
  on playback_sessions_sync (added migration); the reconciler upserts,
  reloads, and compares it so origin changes still publish and unchanged
  rows do not churn.
- The handler ORs the stored origin with the existing name/UA heuristic,
  which stays as a fallback for rows written before this column existed.

is_jellyfin_client keeps the same name and type on the wire; it is only
sourced more accurately.

* fix(admin): correct Android device labels

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-28 21:27:58 -04:00
203a18ae83 feat(observability): OpenTelemetry logs+traces with secret redaction and slog standardization (#290)
* feat(observability): OpenTelemetry logs+traces with secret redaction

Part of #265. Adds opt-in OpenTelemetry (logs + traces) alongside the existing
stderr + opslog pipeline, plus secret redaction on all sinks. Default-off: with
no OTEL_* / SILO_OTEL_ENABLED config, behavior is unchanged.

Bootstrap (internal/telemetry):
- Setup() builds one shared resource, a TracerProvider (parent-based trace-id
  ratio sampler), a LoggerProvider, and the W3C TraceContext+Baggage propagator
  from env. It installs NO MeterProvider — metrics stay on Prometheus, and the
  built-in no-op global MeterProvider keeps the trace instrumentation libs from
  double-emitting. Shutdown is deferred with a flush timeout.
- Logs are bridged via otelslog fan-out (slog.MultiHandler), level-gated by the
  shared LevelVar and best-effort so a failing collector can't break the console
  or DB branches. stderr + opslog stay untouched.

Secret redaction (internal/logredact):
- A slog.Handler masks secret-keyed attributes (password, token, api_key,
  authorization, cookie, ...) — including .With-bound attrs, nested groups,
  secret-keyed group subtrees, and values behind a LogValuer — on the console
  and OTLP sinks, with a no-op fast path when a record has no secret keys.
  opslog.shouldRedact delegates to logredact.SecretKey so all sinks share one
  marker list.

Rotation is infra-managed (no custom file sink): container runtime for stderr,
collector/backend for OTLP, opslog partition-pruning for the DB. Documented in
docs/architecture/observability.md.

Verification: go build ./..., go vet, gofmt -l — clean; go test
./internal/telemetry/ ./internal/logredact/ -race pass.

AI-use disclosure: implemented with AI assistance (Claude Code), including
adversarial reviews that hardened the bootstrap and fixed two redaction leak
paths; reviewed by the author.

* refactor(observability): slog context+component sweep, sloglint gate (phase 3)

Part of #265. Builds on the OTel bootstrap + redaction commit.

Standardizes every log call site onto the context-carrying slog variants so
records correlate with the active OpenTelemetry trace, and locks the standard
in with a machine gate so future code (human- or AI-authored) can't drift back.

- Call-site sweep: converted the remaining slog.<Level>(...) calls to the
  slog.<Level>Context(ctx, ...) form wherever a context.Context is in scope
  (background/init calls with no ctx are left as-is), across 183 files. Applied
  via a type-aware AST codemod. Log levels and message strings are preserved
  verbatim; a component attr (canonical per-package name) is added to direct
  package-level slog calls. Bound-logger calls keep their existing .With
  bindings. The main.go and telemetry package conversions rode with their file
  in the previous commit to keep each file within a single commit.
- Enforcement (.golangci.yml): enable sloglint with context=scope, static-msg,
  key-naming-case=snake, no-mixed-args. After the sweep all four report zero
  violations repo-wide (tests included), so make lint / CI now blocks any
  regression to the non-context form. The gate ships with the sweep because it
  cannot be green until the legacy sites are converted.

Metrics remain on Prometheus; no behavior change to /metrics or Grafana.

Verification: go build ./..., go vet ./..., gofmt -l — clean; sloglint (all 4
rules) 0 violations repo-wide; log levels verified unchanged.

AI-use disclosure: implemented with AI assistance (Claude Code), including the
codemod; reviewed by the author.

* fix(observability): honor per-signal OTLP protocol and secret WithGroup names

Two Codex review findings on PR #290:

- telemetry: OTEL_EXPORTER_OTLP_{TRACES,LOGS}_PROTOCOL now override the
  generic OTEL_EXPORTER_OTLP_PROTOCOL per signal, so mixed collector
  setups (e.g. HTTP logs + gRPC traces) build the right exporter.
- logredact: entering a group whose name is secret-bearing (e.g.
  WithGroup("authorization")) now masks every leaf in that subtree,
  matching how slog.Group("authorization", ...) is masked as a whole.

* fix(observability): address review feedback on telemetry bootstrap

- Telemetry setup failure no longer kills boot: Setup returns usable
  no-op providers alongside the error and main logs and continues with
  telemetry disabled, honoring the best-effort contract.
- Honor OTEL_TRACES_SAMPLER (always_on/off, traceidratio, parentbased_*
  variants); unsupported values fall back to parentbased_traceidratio.
- Attach node identity as semconv service.instance.id instead of the
  non-semconv node.name.
- Rename opslog retention-scope log attrs to target_component/target_level
  so they no longer collide with the canonical component routing key, and
  tag those lines with component=opslog.
- Fix stale levelGated comment casing; use WarnContext in the telemetry
  shutdown defer; document the LogValuer double-resolve on the redaction
  slow path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 08:53:52 -04:00
604bbf1a0f feat(playback): unified restart-resilient playback (native + jellycompat) (#174)
* feat(playback): unified restart-resilient playback via shared TranscodeManager

Make direct, remux, and native HLS transcode sessions survive a server
restart through one shared flow instead of per-method paths. A missing
in-memory session becomes a reconstruct trigger, not a 404: the server
rebuilds the session from a tiny durable recipe card plus the position the
client re-supplies on its next request.

- internal/playback/transcode_manager.go: shared TranscodeManager owning the
  transcodes map, recipe-card lifecycle, reconstruct single-flight +
  concurrency cap, LoadOrReconstructSession front door, ReconstructSession /
  ReconstructTranscode, and orphan cleanup. ~90% is logic moved out of the
  native handler (no behavior change), not new surface.
- internal/playback/recipecard.go + recipecard_postgres.go: RecipeCard with a
  PlayMethod discriminator (direct/remux/transcode; empty decodes as transcode
  for back-compat) behind a swappable, nil-safe RecipeStore interface backed by
  transcode_recipes.
- internal/playback/session.go: RegisterReconstructed inserts a rebuilt Session
  under its existing id (no UUID mint, no limit double-count, race-yielding).
- internal/playback/transcode.go: CloseProcess keeps the output dir so a
  reconstruct winner keeps serving; Close removes it.
- internal/api/handlers: drain the transcode lifecycle into the manager; wire
  reconstruct into the stream/segment serve paths; re-bind ownership to the live
  caller (refuse userID==0/mismatch); card-aware orphan cleanup.
- migrations: add transcode_recipes (expires_at TTL, filter-on-read, indexed).

Ownership stays two-factor: an authenticated caller AND a session.UserID that
matches; the card stores no secrets and identity is re-resolved per request.

Tests: recipe-card round-trip/legacy-decode/disabled-noop, RegisterReconstructed
insert/race/concurrency, close-vs-close-process dir semantics, the
LoadOrReconstructSession status matrix, and the reconstruct concurrency cap.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* feat(jellycompat): reconstruct transcodes across restart via shared manager

Bring Jellyfin (jellycompat) HLS playback onto the same restart-resilient flow
as the native path. Previously jellycompat owned a separate PlaybackHandler with
a private transcodes map and a duplicated transcode lifecycle that never grew the
reconstruct half, so an in-flight Jellyfin transcode died on restart and the next
segment request 404'd.

- Embed the shared playback.TranscodeManager and delete the duplicate lifecycle,
  so jellycompat gets reconstruct, the concurrency cap, the node-affinity rule,
  and the card lifecycle for free.
- internal/jellycompat/playback_sessions_postgres.go: DurableCompatPlaybackStore,
  a write-through cache over jellycompat_playback_sessions behind the new
  CompatPlaybackStore interface (nil pool degrades to cache-only). This persists
  the load-bearing PlaySessionId -> UpstreamSessionID mapping (plus media sources,
  route item id, seek) so it survives a restart instead of vanishing with the map.
- Write a recipe card on compat transcode start keyed by the upstream session id,
  using the native StreamAppUserID so the ownership re-bind matches; reconstruct
  the upstream session and the transcode seeked to the requested seg_NNNNN.
- migrations: add jellycompat_playback_sessions (expires_at TTL + compat_token
  index, full PlaybackSession in data JSONB).

Auth is mapped to the native user id before reconstruct so the same two-factor
ownership check and userID==0/mismatch refusal apply unchanged.

Tests: DB-gated (SILO_TEST_DATABASE_URL) durable-store round-trip proving a
session written by one instance reloads in a fresh one (the restart case), plus
a nil-pool cache-only path; existing handler tests updated to the manager.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* docs(playback): consolidate unified playback reconstruction design

Replace the three overlapping playback docs (the native Postgres
restart-resilience spec, the jellycompat plan, and the unification spec) with a
single self-contained design at
docs/superpowers/specs/unified-playback-reconstruct.md.

The doc leads with the unified design — the one-idea reconstruct model, a strong
visual flow of a restart mid-playback, the shared TranscodeManager + recipe card,
the two swappable durable stores, security, the concurrency cap and node-affinity
constraint, preconditions, and verification. The design history and rationale
(reconstruct-not-rehydrate, phased delivery, Redis-vs-Postgres, token-as-
descriptor, failure analysis) move to an appendix. It references no other md file.

AI-use: written with AI assistance.

* fix(playback): address review on restart-resilient playback

Four fixes from PR review of the unified reconstruction work:

- Rewrite the recipe card on audio-track change. HandleChangeAudioTrack only
  updated the in-memory session/transcode, so after a restart reconstruct
  resumed with the stale AudioTrackIndex/TranscodeAudio (and stale play method)
  from the start-time card. Re-save the card (direct/remux/transcode) with the
  switched state, mirroring the start-card pattern.
- Guard nil TranscodeManager in LoadOrReconstructSession and ReconstructSession.
  StreamHandler.TM is documented optional (tests/minimal setups); a missing
  session previously panicked in recipeEnabled instead of returning
  SessionMissing. ReconstructTranscode already guarded nil; make the two
  siblings consistent.
- Reject direct/remux cards in doReconstructTranscode before spawning ffmpeg, so
  a non-transcode card id can never enter the HLS reconstruction path.
- Log a non-success status from the remote transcode-node DELETE in
  CloseTranscodeSession; a 401/404/500 was previously silent.

AI-use: implemented with AI assistance.

* fix(playback): harden restart-resilient compat sessions

* feat(playback): token-carried reconstruction across restarts

Build on the shared TranscodeManager (introduced earlier in this branch) so a
playback session survives an API-server or transcode-node restart without the
client re-negotiating, and retire the Postgres transcode_recipes store in favor
of a recipe carried inside the signed stream token.

- RecipeCard encodes the byte-affecting encode parameters and rides inside the
  stream token; LoadOrReconstructSession rebuilds the in-memory Session (and,
  for integrated transcodes, the ffmpeg process) on a cold miss, single-flighted
  per session and paced by a spawn semaphore. Removes recipecard_postgres.go and
  the 20260617233705_add_transcode_recipes migration.
- transcodenode reconstructs a lost ffmpeg node-side from the forwarded token.
- TR-lease: proxy/streamauth enforce a revocation deny-marker on every served
  segment, with a 500ms Redis timeout, a bounded per-session "allowed" cache
  (3s TTL, expiry-first graceful eviction), and a degraded-fail-open counter.

Review hardening folded in:
- Manifest/segment handlers do the in-memory session lookup first and only
  verify the stream token on a reconstruct miss (token HMAC was per-segment).
- Copy-mode reconstruct never applies the encoded-only seg*dur seek, at spawn
  time or via the recovery path: RestartSeekTarget reports "unresolved" for a
  copy session whose manifest cannot yet map the segment, so the client retries
  instead of seeking to a fabricated source time.
- Crash teardown is a compare-and-delete (CloseTranscodeSessionIf returns
  whether it matched); the crash closure tears down the playback session only
  when it matched, so a session reconstructed under the same id is not killed.
- Reconstruct enforces the same per-user stream/transcode caps as a fresh start
  (RegisterReconstructedWithLimits), closing a token-replay slot bypass.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* feat(jellycompat): node-side transcode reconstruct via shared recipe store

Make Jellyfin-compat playback sessions survive a server or transcode-node
restart by reusing the shared TranscodeManager reconstruct path and a durable
recipe store, on top of the durable compat session store added earlier in this
branch.

- Node-side transcode reconstruct goes through the shared recipe store; the
  recipe is persisted to the control-plane store (Redis) when a dedicated
  transcode node is used so the node can rebuild ffmpeg after its own restart.
- Adopt the shared manager's API (3-arg OnFFmpegCrash carrying the dead session,
  guarded CloseTranscodeSessionIf, RegisterReconstructedWithLimits).

Review hardening folded in:
- Recipe lifecycle: noderecipe.Store gains Delete, called on deliberate
  teardown (stop, method-switch discard, node stop/force-reload) so a stopped
  session cannot be resurrected by a buffered request after a node restart;
  crash paths intentionally keep the recipe so a resume can reconstruct.
- Crash closure tears down the upstream session only when the guarded transcode
  close matched, so a reconstructed successor is never left orphaned.
- Copy-mode segment recovery surfaces a retryable not-found instead of a
  wrong-position restart, matching the native and node paths.
- Durable Update is now a SELECT ... FOR UPDATE transaction, removing the
  lost-update clobber that could silently drop a transcode recipe.
- Empty-token route resolution no longer falls back to an unbounded full-table
  scan; DB expiry filters bind the injected clock; the redundant re-Get is gone.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* docs(playback): consolidate restart-resilient playback design

Replace the superpowers spec with a single architecture record describing the
token-carried recipe card, the shared TranscodeManager reconstruct path for
direct/remux/transcode, the jellycompat durable session + node recipe store, and
the revocation-lease model with its fail-open tradeoff.

AI-use disclosure: written with AI assistance (Claude Code).

* docs(playback): correct jellycompat node-recipe rationale in comments

The noderecipe / transcode-node / jellycompat comments justified the Redis
recipe store with "a Jellyfin client cannot round-trip a token". The real
reason: the node-hop token is server-minted and could carry the recipe, but the
recipe is mutated in place under a stable session id (a /Sessions/Playing/Progress
audio switch restarts ffmpeg without re-minting the client's token) and a
third-party Jellyfin client cannot be driven to refresh a stale token, so the
node must reconstruct from a server-authoritative, node-reachable store.

Aligns the comments with docs/architecture/restart-resilient-playback.md §10.
Comment-only; no behavior change.

* refactor(playback): remove deny-lease revocation, defer to future PR

The deny-lease stream-revocation mechanism (the internal/streamauth
package, its silo:streamauth:<sid> Redis markers, the proxy Allowed()
enforcement, and the admin Stop/Terminate deny write) only ever enforced
on the offload-proxy topology and was a silent no-op on the integrated
single box and the dedicated transcode node. Rather than ship a partial
revocation feature that looks complete but isn't, remove it wholesale and
defer a uniform cross-topology revocation design to a dedicated follow-up.

Removed: internal/streamauth (package + tests); the LeaseDenier field,
StreamLeaseDenier interface, and denyStreamLease helper in playback.go;
the admin deny write; the router/main wiring; and the proxy verifyToken
Allowed() gate. The unified-reconstruct core (recipe-token,
LoadOrReconstructSession) is orthogonal and untouched.

Known limitation (now on every topology): admin Terminate and user Stop
tear down the live in-memory session and ffmpeg producer, but a still-valid
stream token can reconstruct the session until its 24h TTL expires. No
node-side byte-withholding ships in this PR.

docs/architecture/restart-resilient-playback.md is updated to mark the
revocation/deny-lease sections as deferred and to drop the overstated
"instant revocation on admin kill" claim.

* fix(playback): allow zero-caller bearer on transcode reconstruct

The authless HLS transcode delivery routes (master.m3u8 / segment) treat
the session UUID as the bearer credential, so a real request carries
requestUserID == 0. The live serve path already allows this, but
ReconstructSession hard-rejected a zero caller, so a request that worked
before a restart became SessionMissing -> 404 after the in-memory session
was gone, breaking the restart resilience these routes advertise.

Match the live-path contract in LoadOrReconstructSession: allow a zero
caller (UUID-as-bearer) and refuse only a non-zero caller that mismatches
the card owner. The reconstructed session is bound to card.UserID either
way. Adds TestReconstructSession_Ownership covering both cases.

* fix(jellycompat): re-persist recipe on local audio switch

A Jellyfin client switching audio on an integrated/local compat transcode
restarted live ffmpeg with the new track but did not re-persist
PlaybackSession.Recipe. The remote branch already re-persists via
startRemoteTranscode -> persistTranscodeRecipe. After a central restart,
reconstruct rebuilt ffmpeg from the stale Recipe.AudioTrackIndex, so the
integrated session resumed on the original audio track.

Persist the updated recipe (best-effort) after a successful Restart in the
local branch, mirroring the remote branch, so the durable
Recipe.AudioTrackIndex tracks live ffmpeg. Adds a regression test.

* fix(playback): strip stream token from proxied transcode-node URL

proxyToTranscodeNode appended the client's raw query string to the internal
transcode-node URL and logged that URL on transport failure. When a remote
transcode runs without a separate proxy node, that query carries
?st=<signed JWT> — a 24h bearer reconstruction descriptor exposing the
media path and recipe claims — placing the token into internal requests and
error logs.

Strip the "st" param before building targetURL, preserving any other query
params. The token is neither forwarded to the node nor present in the
logged URL. Header-forwarding of the token (so the node can reconstruct) is
a separate follow-up (#6).

* fix(playback): fail open on transient limit-provider error in reconstruct

During the reconstruct wave right after a restart (Postgres under peak
load), a transient limit-provider DB error was collapsed into a hard 404,
permanently stopping playback for a user within their limits. limitsForUser
wrapped any provider error, RegisterReconstructedWithLimits propagated it,
and ReconstructSession mapped every error to SessionMissing -> 404 -
indistinguishable from a genuine over-cap rejection.

Distinguish the two: tag provider errors with a new ErrLimitProviderUnavailable
sentinel and, during reconstruct, fail OPEN on a provider error (admit via
RegisterReconstructed + log a degraded warning) rather than refuse - mirroring
the reliability-first fail-open-on-dependency-error philosophy. A genuine
ErrTooManyStreams / ErrTooManyTranscodes over-cap still refuses. Adds tests
for both the fail-open and still-refused paths.

* fix(playback): forward stream token to transcode node as header

The dedicated transcode node's reconstruct path reads the stream token only
from the X-Silo-Stream-Token header, but proxyToTranscodeNode forwarded only
the node-API bearer token (and #5 now strips st from the URL). So when the
central API proxied to the node and the node self-restarted, it could not
reconstruct from the recipe-complete native token -> 404.

Capture st before stripping it from the URL, verify it at the API boundary
(streamtoken.Verify + SessionID match, mirroring the node's own check), and
forward it as X-Silo-Stream-Token. Best-effort: a missing/invalid token never
blocks the live proxy, and the token is still kept out of the forwarded URL
and logs.

* fix(playback): restart node ffmpeg on native remote audio switch

A native audio-track switch on an offloaded/remote transcode was a no-op at
the node yet returned 200 with a fresh URL: HandleChangeAudioTrack restarted
ffmpeg only when the API owned a LOCAL TranscodeSession, so for an offloaded
transcode the node kept serving the OLD audio (the node consults the token
only on a session miss). The replacement URL was also minted from identity-
only claims, so a later node restart 404'd.

For the offloaded transcode case (detected via session.TranscodeNodeURL),
POST a fresh /transcode/start to the node with the new AudioTrackIndex
(handleStart tears down and restarts ffmpeg) and mint the replacement proxy
URL from a full RecipeCard so reconstruct survives a node restart. The encode
recipe is derived from the durable session target fields plus the file,
mirroring HandleStartTranscode. A concrete SegmentDuration
(playback.DefaultSegmentDuration) is embedded rather than 0: the node's token
completeness gate treats SegmentDuration<=0 as incomplete and falls back to a
recipe store the native path never populates, which would 404 on a node
restart - the exact resilience this path provides. A failed node POST now
surfaces 502 rather than a false 200. Remux and non-offloaded (local)
transcode paths keep their prior identity-claim URLs unchanged.

Known limitation: Session does not persist the original SegmentDuration or
SubtitleTrackIndex/SubtitleBurnIn, so a remote audio switch resets subtitle
selection to none and assumes the default segment length; a client that
started with a non-default segment length will resegment on switch. Making
that state durable on the session is a follow-up.

* docs(playback): scrub stale deny-lease/revalidator comments

The deny-lease revocation mechanism and its "central revalidator" were removed
earlier in this branch, but four comments still described them as live
(transcode_manager.go, noderecipe/store.go, streamtoken/token.go,
proxy/server.go). Reword them to match the shipped behavior: ownership claims
are re-resolved at reconstruct, the noderecipe store shares Redis only with the
node-session tracker, and a sub-TTL hard cut depends on a node-side revocation
mechanism that is deferred to a future PR.

* fix(jellycompat): surface durable playback-session write failures

DurableCompatPlaybackStore.Update applied the in-memory mutation and then
swallowed every Postgres commit-failure path, returning nil. Callers that
promise restart resilience (persistTranscodeRecipe's recipe write, the
upstream-session binds in streams.go) were told the session was durably
persisted when only the cache held it, so a transient DB hiccup could leave
the next restart reloading a stale row (wrong audio track) or 404ing.

updateDB now returns the genuine DB round-trip error (begin/query/unmarshal/
marshal/exec/commit); Update propagates it while still applying the in-memory
mutation so live state stays correct. A nil pool and a genuinely absent/expired
row remain best-effort (return nil) — only real infrastructure failures
propagate, so existing rollback paths fire exactly when durability is lost.

Part of #174

* fix(playback): re-inject stream token into proxied transcode manifests

API-proxied remote transcode manifests dropped the reconstruct token from
their segment URLs, so playback died after a node or API restart. When a
remote transcode has no separate proxy node, the client loads its manifest via
the API-local path; proxyToTranscodeNode strips the signed token ("st") from
the forwarded URL (keeping it off node URLs and logs, forwarded only as the
X-Silo-Stream-Token header), and the node builds relative segment URIs from
that token-less query. The segment URLs the client received carried no token,
and the proxy only re-attached the header when an incoming segment request
already had "st" — which it never did — so a restart made those segments
non-reconstructable and they 404'd.

proxyToTranscodeNode now rewrites the manifest body at the boundary: every
segment and #EXT-X-MAP init URI gets the client-facing, API-verified token
re-appended (new playback.AppendManifestQueryParam helper), so the client's
later segment fetches carry "st" again and reconstruct after a restart. The
token still never reaches the node URL or its logs. Only 200 .m3u8 responses
are rewritten (Content-Length corrected); segments stream through untouched.

Part of #174

* fix(playback): preserve subtitle/cadence recipe across offloaded audio switch

Switching audio on a remote (offloaded) transcode with burned-in subtitles
silently dropped them, and reset a non-default segment cadence. The offloaded
audio-switch restart rebuilt the node start request from Session state, but
Session/SessionStreamState retained no subtitle or segment-duration state
(only the live local ts.Opts() and the RecipeCard did), so the branch
hard-coded SubtitleTrackIndex:-1, SubtitleBurnIn:false and
SegmentDuration:Default — signing that altered recipe into the replacement
stream token. An audio switch then changed bytes beyond audio selection, and
any later reconstruct kept the wrong no-subtitle/wrong-cadence recipe.

Persist the byte-affecting recipe on the session: SubtitleTrackIndex,
SubtitleBurnIn and SegmentDuration are added to Session/SessionStreamState,
populated at start (finalizeTranscodeStart) and on post-restart reconstruct
(ReconstructSession from the card), carried forward on every audio-switch
state update, and read back when rebuilding the offloaded node request and its
recipe card. The restart now reproduces the exact live stream. Also resolves
the M-4b non-default segment_duration reset.

Part of #174

* fix(playback): serialize transcode spawn paths with a per-session lock

Reconstruct was single-flighted only against other reconstructs, so a
restart-driven segment reconstruct racing a quality/seek/audio fresh start
could spawn two ffmpeg processes writing the same output directory at once —
segment corruption, partial-write closes, orphaned processes, and skewed
active-job accounting. The atomic register-after-spawn (GetOrRegister / the
reconstruct compare-on-register) prevented a map leak but not the concurrent
disk writers, because the losing path had already spawned. The dedicated
transcode node had the same split between handleStart and spawnReconstruct.

Add a refcounted per-session lifecycle lock to both TranscodeManager and the
node Server, held across "check existing -> spawn -> register":
- reconstruct (doReconstructTranscode / spawnReconstruct) re-checks under the
  lock and yields to any live session instead of spawning a duplicate;
- the native and jellycompat fresh-start paths take the lock around their
  spawn+register (the native path also closes any session a reconstruct rebuilt
  in the meantime so its fresh ffmpeg is the sole writer);
- the node handleStart holds it across teardown+spawn+register.
The refcount drops the map entry once no path holds/waits, keeping it bounded.
GetOrRegisterTranscodeSession is removed — the lock supersedes it and keeping a
register-after-spawn primitive would invite reintroducing the race.

Part of #174

* fix(playback): serialize restart re-spawn under the session lifecycle lock

TranscodeSession.Restart() releases s.mu across cancel -> wait-for-done ->
re-exec and spawns ffmpeg into opts.OutputDir without holding the per-session
lifecycle lock. LockSessionLifecycle's contract (fresh start, restart,
reconstruct) requires restart to hold it too, but all five callers invoked
Restart unlocked: native audio-switch and segment-recovery, compat
audio-switch and segment-recovery, and the transcode-node segment-recovery.

A restart racing another restart (audio-switch vs segment-recovery) or a
fresh-start/reconstruct could land two ffmpeg processes writing the same
segment directory -- mixed timelines, init.mp4/segment mismatch, and an
orphaned-but-still-writing ffmpeg -- the exact concurrent-writer corruption
the lifecycle lock exists to prevent.

Add RestartSessionLocked (TranscodeManager) and restartSessionLocked (node
Server) that hold LockSessionLifecycle only across the cancel->respawn
transition, re-check that the handle is still the live mapped session under
the lock, and return ErrSessionSuperseded rather than re-spawning a stale
handle. Route all five call sites through them. The lock is released before
callers wait on segments so recovery latency is unchanged.

Tests: gating (restart blocks until the lifecycle lock frees, then spawns),
concurrent-restart serialization, and superseded re-check on both the manager
(covers native + compat) and node lock owners.

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-02 14:23:14 -04:00
a3aa93534d fix(jellycompat): authenticate Android TV token-less direct play (#200)
* fix(jellycompat): authenticate Android TV token-less direct play

Stock Jellyfin Android TV ignores the api_key-bearing DirectStreamUrl
returned from PlaybackInfo and builds its own direct-play URL with no
auth header, no api_key/ApiKey query param, and no PlaySessionId. The
media request arrives via the player's HTTP stack (okhttp) with
auth_kind=none, so PlaybackSessionAuth 401s it — the client retries,
falls back to a transcode that stalls, and surfaces "player error".

Add a third fallback in PlaybackSessionAuth, scoped strictly to the
direct-play video stream routes (/Videos/{id}/stream[.{container}]) via
the chi route pattern so /Items/{id}/Download stays protected: anchor
auth on the PlaybackSession negotiated during the preceding (already
authenticated) PlaybackInfo, looked up by mediaSourceId when present
(else the route item id), and resolve its CompatToken.

Covered by tests: token-less direct play with a matching session
succeeds, no matching session 401s, and Download is not loosened.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(jellycompat): require session item match + expand direct-play tests

Address CodeRabbit review on PR #200:
- Require the matched PlaybackSession's RouteItemID to equal the requested
  route item before authorizing, so a mediaSourceId cannot authorize a
  stream for a different item.
- Seed the compat session in the 401 tests so they fail on route/session
  scoping rather than a missing session.
- Table-drive the positive test across /Videos/{id}/stream and
  /Videos/{id}/stream.{container}, plus the route-item lookup branch when
  mediaSourceId is absent; add a cross-item denial test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 12:27:37 -04:00
Quick 227986094b Propagate playback client metadata through session sync 2026-06-23 14:26:14 -04:00
544a262924 fix(jellycompat): accept Jellyfin current ApiKey query param for stream auth (#186)
Silo's ExtractToken only honored the legacy 'api_key' query parameter and
the X-Emby-Token/X-Mediabrowser-Token headers. Real Jellyfin's
AuthorizationContext treats 'ApiKey' (PascalCase) as the current,
always-enabled query token and 'api_key' as legacy (gated behind
EnableLegacyAuthorization). Native Jellyfin clients that build their own
direct-play /Videos/{id}/stream URLs (incl. Jellyfin Android TV) send
'ApiKey', which Silo rejected — the request arrived with auth_kind=none and
the route returned 401, surfacing as a client 'playbackerror'.

Match both 'ApiKey' and 'api_key' case-insensitively in ExtractToken and in
the authKind log classifier. Strictly additive: existing header and
api_key paths are unchanged.

Verified against the live server: /Videos/{id}/stream?...&ApiKey=<tok>
returned 401 before and is accepted after; &api_key=<tok> and the
X-Emby-Token header continue to return 206.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-21 16:20:04 -04:00
cf4e080bf4 fix(jellycompat): Wholphin (jellyfin-sdk-kotlin) playback & genre compatibility (#100)
* fix(jellycompat): match MediaSourceId across UUID formats (compact vs dashed)

* fix(jellycompat): honor ImageTypes=Backdrop as a filter on /Items

Wholphin genre cards request /Items?imageTypes=Backdrop&limit=1&sortBy=Random
and assume every returned item has a backdrop. Silo ignored ImageTypes, so a
random pick could lack a backdrop (BackdropImageTags: null), crashing Wholphin.

Push the filter down to the catalog browse SQL
(NULLIF(BTRIM(backdrop_path),'') IS NOT NULL) so random/limited selections only
ever consider backdrop-having items; empty genres correctly return [].

* fix(jellycompat): case-insensitive PlaySessionId + api_key in stream auth

Wholphin's jellyfin-sdk-kotlin builds its own direct-play URL
(/Videos/{id}/stream?static=true&playSessionId=...&mediaSourceId=...) with a
lowercase 'playSessionId', no api_key, and no auth header (ExoPlayer's data
source drops it). PlaybackSessionAuth read 'PlaySessionId'/'PlaySessionID'
case-sensitively, so the fallback never matched -> 401 on every direct-play
stream -> forced (often failing) transcode fallback. Resolve PlaySessionId via
newCaseInsensitiveQuery, and likewise accept case-variant api_key in
ExtractToken.

* fix(jellycompat): support Wholphin season item queries

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-06-08 21:05:33 -04:00
d3626fca00 feat(jellycompat): add /Users endpoint and unified sa_ admin-key auth (#91)
* fix(api): allow HEAD on /api/v1/direct-download

Firefox (and some download managers) issue a HEAD request before
starting a download. The route only registered GET, so HEAD returned
405 Method Not Allowed and the browser aborted the download.

Mirrors the pattern already used by /stream/{session_id}, which
registers both GET and HEAD on the same handler. ServeDirect is built
on http.ServeContent / ServeFile, which natively handle HEAD by
writing headers without a body, so no handler changes are needed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: design jellyfin autoscan scan compatibility

* refactor(scan): extract scan trigger resolver

* refactor(api): share scan target resolution

* feat(jellycompat): accept admin api keys for autoscan

* feat(jellycompat): add autoscan media update route

* docs: document jellyfin autoscan setup

* fix(jellycompat): harden autoscan auth and batch scan enqueue

- Reject nil API keys and bound last-used update with a 5s timeout
- Stop leaking internal queue errors in autoscan responses
- Batch scan enqueues via new CreateBatch and reuse folder list across path resolves

* chore: add planning docs and requests updates

- Add plans for date-named episodes and Jellyfin autoscan compat
- Update requests handlers, service, and UI hooks
- Remove Makefile.local.example

* refactor(scantrigger): drop redundant Target.LibraryID field

- Read library ID from Target.Folder.ID everywhere
- Guard scan queue enqueue against nil Folder
- Simplify admin API key auth error plumbing

* fix(catalog): gate search overview-only matches behind title FTS

- Always apply stats CTE + CROSS JOIN so single-word queries no longer flood results with description-only hits
- Require overview_rank >= 0.15 for overview-only fallback rows
- Switch title gate from contiguous LIKE to title_rank > 0 so reordered-token title matches aren't demoted

* ci(docker): build image on push to main via self-hosted runner

- Trigger Docker image builds on pushes to main instead of nightly cron
- Run on self-hosted Linux runner
- Drop the `nightly` tag

* docs(specs): add design for TMDB-backed request section in search

Adds the design for surfacing requestable TMDB results inside the main
catalog search (Cmd+K dialog and full results page) as a clearly
delimited "Request to Add" section that never blocks or displaces
library results.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(specs): address Codex adversarial review for search request section

Splits discovery eligibility from submission eligibility so blocked
and quota-exhausted viewers still see the requestable section with
disabled per-row CTAs, matching the documented behavior. Documents
the required extensions to useRequestSearch — signal forwarding,
viewer-identity-keyed cache, and invalidation on auth/profile/
settings/limit changes — so the planned 5-minute staleTime is
safe and cancellation works as described.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(plans): add implementation plan for search request section

Twelve TDD tasks covering: api() signal contract test, useCanRequest
hook, viewer-keyed requestKeys.search, useRequestSearch extension
(signal + viewer key + 5min staleTime + enabled override), invalidation
cascade tests, RequestPosterCard optional onRequest, RequestToAddSection
component (dialog + grid variants), GlobalSearch and Catalog wiring
with empty-state suppression for the library-0/TMDB-pending edge case,
final lint/test pass, and manual smoke. Notes a single deviation from
the spec: submitDisabledReason is null in the initial implementation,
with per-row request data driving disabled UI.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(plans): address Codex adversarial review for search request section

Fixes the high-severity finding that RequestToAddSection's internal
useRequestSearch call was not gated on discoveryEnabled, allowing
/api/v1/requests/search and TMDB lookups to fire for users without
request access. The plan now (1) passes { enabled: discoveryEnabled }
to the section's hook, (2) gates the parent mount in GlobalSearch and
Catalog on canRequest.discoveryEnabled as defense in depth, and (3)
adds tests asserting both the enabled forwarding and the no-mount
behavior when discovery is disabled.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(api): pin AbortSignal forwarding contract on api()

* feat(hooks): add useCanRequest gating hook for discovery eligibility

* refactor(keys): add viewerKey to requestKeys.search

* feat(requests): key useRequestSearch by viewer, forward signal, raise staleTime

* test(requests): document viewer-keyed cache isolation and invalidation cascade

* feat(request-card): make onRequest optional on discover variant

* feat(search): add RequestToAddSection dialog variant

* feat(search): add RequestToAddSection grid variant for Catalog page

* feat(search): render RequestToAddSection in the Cmd+K dialog with empty-state suppression

* feat(catalog): render RequestToAddSection grid with empty-state suppression

* chore(web): format request search section

* test(web): avoid unsupported Array.at in search request tests

* commit message

{"subject":"fix(search): prevent empty-state flash before TMDB fallback renders","body":"- Add isResolving to useCanRequest and gate empty states on it across GlobalSearch and Catalog\n- Debounce TMDB query in Catalog and hide ItemGrid when the request section may rescue an empty library\n- Track per-card submit state in RequestToAddSection grid so concurrent requests don't trample each other\n- Suppress anonymous TMDB request-search fetches to avoid cross-viewer cache leakage"}

* fix(jellycompat): tolerate autoscan sidecar updates

* fix(webhooksync): skip events for unmapped external users

- Require explicit profile mapping instead of falling back to the default profile
- Update settings UI copy to reflect that unmapped users are ignored

* refactor(admin): show per-section loading and error states

- Replace page-level loading gate with skeletons per section on dashboard and stats
- Surface query errors inline instead of blocking the whole page
- Disable "Scan All Libraries" when no libraries are configured

* fix(search): address request search review feedback

* feat(subtitles): restore upload management

* feat(auth): add assignable user permissions

* feat(auth): expose user permissions

* feat(api): authorize item metadata curation

* feat(api): route metadata curation by permission

* feat(web): add permission helpers

* feat(web): assign metadata curation permission

* fix(web): keep device profile hooks unconditional

* feat(web): show metadata tools to curators

* fix(auth): address metadata curation review issues

* fix(auth): tighten curator job response review fixes

* docs: add metadata curation permission plan

* test(auth): expand session revocation coverage

* docs: add PageBack component design spec

* fix(auth): gate media file paths on metadata curation permission

- Allow curators (not just admins) to view media file paths and locations
- Apply library access filter to file-level access checks

* fix(metadata): break duplicate provider candidate ties

- Score candidate metadata completeness and auto-match the richer duplicate when title/year/type tie
- Enrich near-duplicate candidates via the provider chain before initial match selection
- Seed both movie and series match queues for mixed-type libraries and wait for TV queue settle
- Add taskmanager worker test coverage and a plan doc for the tie-breaker work

* feat(ui): add shared PageBack component for consistent back navigation

Replace the eight inconsistent back affordances across user-facing pages
with a single absolute-positioned chevron pill, so the control lives in
the same screen position regardless of title length or hero content.

DetailBreadcrumb keeps its textual hierarchy path but no longer owns the
back chevron; PageBack does. DetailHero gains a topNav slot consumed by
Movie/Series/Season/Episode/Request detail pages. Non-hero pages drop
their bespoke back buttons and add PageBack inside a relative wrapper.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(ui): add floating variant to PageBack for sticky nav

- Add `floating` prop to pin PageBack to viewport on lg+ screens
- Switch styling from glass-subtle to glass with shadow for better contrast
- Use floating variant on SettingsLayout

* fix(ui): make PageBack destinations deterministic

* feat(episode-carousel): highlight currently viewed episode

- Add "Now Viewing" badge with pulsing indicator on current episode
- Replace border with primary-color ring for current episode card
- Set aria-current="page" on links to the current episode

* feat(jellycompat): sign image tags and accept them without session

- HMAC-sign image tags using the configured JWT secret
- Serve item/season/episode images via signed tag without requiring a session or cache hit

* fix(jellycompat): harden signed image tags

* fix(jellycompat): stabilize signed image tags across restarts

{"subject":"fix(jellycompat): stabilize signed image tags across restarts","body":"- Sign library poster and episode parent series image tags from canonical paths/thumbhashes instead of presigned URLs so tags survive restarts\n- Accept signed canonical tags in the image handler without a session and fall back to legacy URL-derived cache tags\n- Always fetch series detail for episodes to build stable parent image tags"}

* fix(metadata): accept exact cross-provider match ties

* Optimize episode added_at sorting

* fix(libraryingest): treat drainer shutdown cancel as clean stop

TV/series full scans (libraries with new or updated items) were recorded as
"cancelled" with an empty error message and never completed matching.

When the file-walk finishes, the ingest executor waits out a settle window and
then calls stopDrainers() to shut down the concurrent match goroutines. That
cancels the drainer context while a ProcessBatchByFolderAndPathPrefix call may
still be in flight. The drainer treated the resulting context.Canceled as a
fatal error: it pushed the error to drainerErrCh and called cancel() on the
whole scan context, so scanqueue.process() mapped it to cancelRun().

Large/slow libraries (many series, slow provider lookups) keep a batch in
flight continuously, so stopDrainers() almost always landed mid-call and the
scan was cancelled; small/fast libraries were usually idle at that instant and
completed normally.

Treat a cancelled drainer context as a deliberate shutdown: return cleanly
without escalating. Genuine external cancellation still reaches the run via the
main goroutine's scanCtx checks, so real cancels are not swallowed.

Adds a regression test (settle window made injectable) that fails against the
old handler with 'concurrent match scope ...: context canceled' and passes
with the fix.

* perf(catalog): add episode browse index fast path

* fix(catalog): address episode catalog review feedback

* Fix settle-window drainer cancellation in library ingest

* fix(catalog): support relative date filters

* fix(collections): cap smart collection results

* fix(library): preserve episode browse url

* fix(catalog): use season posters for episode cards

* feat(sections): show episode context in cards

* feat(calendar): show local episode airtimes

* Enable DRI passthrough in docker compose

Co-authored-by: Codex <noreply@openai.com>

* feat(ui): paginate the Ambiguous Roots table

Match the sibling tables on the Admin Libraries page (Troubleshooting/unmatched):
use the existing usePagination hook + PaginationBar (10/page, auto-hidden when
<=10 rows), render pag.rows, and reset to page 0 when the library selector or
search filter changes.

* feat(admin): enlarge match-candidate posters + hover-to-enlarge

Unmatched-item match dialog rendered candidate posters at 44x64px, too small
to identify a film. Bump to 64x96 (2:3) and add a portaled Tooltip hover preview
(192x288) using the existing poster image, so operators can tell candidates apart.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(admin): search unmatched items across the whole table, not just the page

The unmatched-items search filtered only the current page's rows client-side.
Push the query server-side: HandleListUnmatchedItems takes an optional 'q' param
and filters title/library/type/status with parameterized ILIKE across all rows,
paginating the filtered set. Frontend hook takes a debounced search, resets to
page 1 on change, keeps the section mounted while searching. Also fixes stale
test mocks that returned the pre-pagination array shape instead of {items,total}.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(admin): prettier-format match dialog; reset unmatched page in onChange

Run prettier over the Tooltip-wrapped poster JSX, and reset the unmatched-items
page in the search input's onChange rather than a useEffect (avoids the
react-hooks/set-state-in-effect warning / cascading renders).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(admin): wire QueryClientProvider + missing hook mocks so the suite runs

The AdminLibraries test file failed all 6 tests with 'No QueryClient set' on
this branch and at the parent commit -- pre-existing infrastructure gap. With
that fixed, several hooks that the page imports (useCancelLibraryScans,
useLibraryRoots, useUpsertLibraryRootOverride, useDeleteLibraryRootOverride,
useActiveScans) and the UNMATCHED_PAGE_SIZE constant also needed mocking. One
stale assertion on the renamed 'Root path' header is updated; the deeper
troubleshooting test, which mocked useSkippedLibraryRoots but the section was
refactored to useLibraryRoots(_, 'ambiguous'), is skipped with a TODO -- a real
rewrite is needed and is out of scope for this MR.

5 of 6 tests now run and pass; the 6th is properly flagged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(admin): search all unmatched item library memberships

* fix(naming): strip unsubstituted Sonarr tokens ({TvdbId}/{imdb-}) from titles

These tokens survived the provider-tag regex ([\w]+ doesn't match braces),
polluting parsed titles (e.g. 'A Girl & Her Guard Dog [tvdb-{TvdbId}]') so
they could not score-match. Broaden the regex to drop {...} and empty tokens.

* fix(naming): numeric-only titles are not bare provider IDs

'86' / '22 7' were parsed as trailing tvdb ids, tripping the trusted-ID gate
so the correct title match was rejected. Require a letter in the name before
treating a trailing number as a bare id.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(naming): document bare-id trade-off; cover CJK title + movies numeric

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(matcher): auto-accept a year-corroborated single distinct show

A search that resolves to one distinct show (one candidate, or the same
title+year returned once per source as unmerged TVDB/TMDB rows) whose year
matches the parsed year is now auto-accepted via the existing top-ranked
candidate, even when the fuzzy title score is in the 55-69 band. The 55/70/15
thresholds are unchanged; this only adds a year-gated acceptance for
effectively-unique results (recovers lone-correct-result items like 1201 (1993)).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(matcher): exercise the single-distinct-show guard properly + conflicting-ID case; doc notes

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(matcher): tolerate concurrent-merge ErrItemNotFound in series episode-link ensure

A scan drainer and the background MatchWorker can process the same folder
concurrently. When a provider-ID merge moves a series' episodes to the survivor
and deletes the source, an in-flight ensureSeriesEpisodeLinks(sourceID) hits
catalog.ErrItemNotFound and was failing the whole scan. The episodes are already
reattached, so this is benign: log and continue (matching the lenient call sites)
instead of failing. Genuine errors still abort.

* diag(matcher): debug-log per-candidate match scores

Adds a DEBUG-gated log in selectInitialMatchCandidate printing each scored
candidate (title/year/type/sources/provider_ids/score) against the hint, so
operators can see why an item did or didn't auto-match. Zero-cost when debug
logging is off; no change to matching behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(matcher): resolve cross-source ties by library provider priority

Accept a year-corroborated single distinct show when the TOP tie-group (within
15 pts of best) is one show across providers, ignoring low-score noise below it,
and pick the winner by the library's metadata-provider chain order (providerPriority,
highest-first; falls back to top-scored). Recovers items like '100 Days Wild'
that are returned identically by TVDB and TMDB. Thresholds (55/70/15) unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(matcher): accept cross-source-corroborated ties without a hint year

When the top tie-group is one distinct show returned by 2+ distinct providers
(candidatesAreSingleDistinctShow already verifies matching title+year), accept it
even if the hint has no parsed year (year-less folders like '100 Deeds for Eddie
McDowd'). Multi-source agreement substitutes for the year guard; lone single-source
no-year results stay subject to the single-candidate >=70 gate. Thresholds unchanged.

* diag(matcher): debug-log provider search query + per-provider result counts

Adds DEBUG logs in the ModeInitialMatch search path: each provider's result
count for the query, and the assembled raw/candidate totals. Lets us see when a
provider search returns zero ('no metadata found') vs a scoring/tie issue.
Zero behavior change.

* fix(naming): parse bare bracketed IMDb IDs ([tt10011226]/{tt...})

Folders tagged with a bare IMDb id in brackets (Plex/Kodi style, e.g.
'17 Blocks (2021) [tt10011226]') had the id silently dropped — folderIDPattern
needs an 'imdb-' prefix and trailingImdbIDPattern needs an un-bracketed trailing
tt-id. Recognize bracketed bare tt-ids so these items get the trusted-ID match
path instead of falling to title+year scoring.

* feat(metadata): match sole exact-title candidate despite year off by <=2

Folder years routinely differ from provider release years by a year or two
(festival vs wide release, regional dates), zeroing the year bonus and leaving
a lone exact-title candidate at 63-68 — just under the single-candidate >=70
gate (e.g. Dead Reckoning 1947 vs 1946, 17 Blocks 2021 vs 2019, Stasi FC). Add
title corroboration to the existing lone-result rule: a sole distinct show whose
normalized title exactly matches and whose year is within +/-2 is accepted. The
55 floor still rejects low-similarity titles (e.g. Hotel Transylvania Puppy! vs
Puppy!). No 55/70/15 threshold change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(lint): gofmt single-space alignment in root_inference.go var block

When inferProviderTagRe was broadened to handle unsubstituted Sonarr token
placeholders ({TvdbId}/{imdb-}), the regex grew long enough that gofmt prefers
single-space rather than column-aligned spacing across the var block.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(metadata): tighten matcher and bare IMDb parsing

* docs: add resilient library deletion design spec

Batched, deadlock-retrying rewrite of delete_library to replace the
single multi-minute transaction that deadlocks on large libraries.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs: add implementation plan for resilient library deletion

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(catalog): add deadlock-retry helper for batched deletes

* test(catalog): clarify cancel-path expectation in retry test

* feat(catalog): add deleteInBatches loop helper

* refactor(catalog): make image-dir helpers querier-agnostic

Add rowQuerier interface satisfied by both *pgxpool.Pool and pgx.Tx.
Split collectImageDirs into collectRawImageDirs (raw collection) and a
thin wrapper that filters via filterUnreferencedImageDirs. Both helpers
now accept rowQuerier so a later task can call them from pool-level
batch deletes without an open transaction.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(catalog): delete libraries in deadlock-retrying batches

Replaces the single multi-minute delete transaction with phased, batched
autocommit deletes (orphan items, media files, memberships, folder row),
each retried on deadlock. Holds only short locks, survives concurrent
writers, and is resumable on failure.

* refactor(catalog): wrap orphan-batch iteration error

* fix(catalog): clamp still/poster/logo backdrops to largest cached variant

Episode stills used as backdrops only exist at w500/w300 in the cache, so
requesting a w1280/w1920 backdrop width 404s. Add catalog.BackdropVariantPath
+ imageTypeFromCachedPath and route featured (w1920) and Continue Watching /
Next Up (w1280) backdrops through it; still/poster/logo paths clamp to their
type's largest cached variant while real backdrops keep the requested width.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* perf(startup): defer non-critical init off the HTTP listener path

Collect catalog-size-dependent seeding (metadata match queues, legacy
series-group cleanup) and the watch-provider scrobble sweep into a
backgroundInit slice that runs sequentially in a background goroutine after
the server is ready, instead of blocking startup before the listener accepts
connections. Steps log failures and stop early on shutdown.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(migrations): make air_timezone column add idempotent

Use ADD COLUMN IF NOT EXISTS so re-running 162 on a database that already has
the column is a no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* build: stamp git revision via Makefile ldflags

`make build` did not inject buildinfo's `revisionOverride`/`dirtyOverride`
ldflags (the Dockerfile already does), so binaries built via make report
their version as "unavailable" in the admin Build panel whenever Go's VCS
metadata isn't embedded. Mirror the Dockerfile by computing the git
revision + dirty state and passing them through `-ldflags -X`.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(web): accessibility & UX fixes from visual QA pass

Accessibility (WCAG AA):
- Lighten the Standard-theme `--muted-foreground` (#6e6e78 -> #9696a0,
  ~3.4:1 -> >=5.3:1) and darken the light-theme equivalent so secondary
  text meets 1.4.3 contrast app-wide; the opt-in High Contrast mode is no
  longer the only conformant path.
- Give icon-only controls accessible names (4.1.2): the password show/hide
  toggle (also drop tabIndex={-1} so it's keyboard reachable), and the
  Edit/Delete/health/copy/refresh actions across the Users, Libraries,
  Nodes, API Keys, Catalog Maintenance and Job History admin tables.
- Fix the Switch off-state (invisible track -> visible border + fill) and
  the PlaybackSettings SettingRow label association (the <label htmlFor>
  pointed at a wrapping <div>; the id now lands on the Switch/SelectTrigger).
- Login: wrap the card in <main> and add an <h1>; Profiles: add an
  accessible PIN-protected label and a corner lock badge.
- Player + catalog: role="status" on the initial loading overlay; scope the
  catalog count ("0 in library" for search) and announce it via aria-live;
  trim the verbose poster-link name to the title.

UX / consistency:
- Emphasize overdue scheduled tasks (warning colour + icon + word, not
  colour alone).
- Per-source catalog subtitles instead of one shared string.
- Add a Reconnect affordance when the admin log stream drops (it does not
  auto-retry).
- Page titles for Watch Party + all admin sub-pages (incl. plugins); admin
  heading capitalisation normalised to Title Case.
- Show "dev build" instead of "unavailable" when no build revision is
  stamped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(catalog): re-check orphan status when deleting library items

Orphan detection moved outside the media_items delete in the batched
library-delete rewrite, opening a TOCTOU race: a concurrent scan/import
could attach one of the collected content IDs to another library between
collectOrphanBatch and the delete, after which the unconditional
`DELETE FROM media_items WHERE content_id = ANY($1)` would still remove the
shared row and cascade away the newly-added membership — dropping the item
from the other library. Re-check the orphan invariant inside the delete
(NOT EXISTS a membership in another folder) and count rows actually deleted.

Addresses PR #21 review (P1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(metadata): persist a cleared air_timezone instead of skipping it

Clearing a previously-set air timezone sent JSON null, which decodes to a
nil *string that UpdateMetadata treats as "skip this column", so the old
value remained. The dialog now sends "" (accepted by ValidateAirTimezone),
and UpdateMetadata maps air_timezone through NULLIF so an empty value
persists as SQL NULL (matching the nullable column) rather than "".

Addresses PR #21 review (P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(startup): sweep open scrobbles before accepting playback

The open-scrobble sweep was queued in the deferred background-init list,
which runs concurrently with the HTTP listener; a resume immediately after
restart could start new scrobbles before the previous process's open
sessions were stopped, leaving overlapping/stale scrobbles on remote
providers. Run the sweep synchronously before the listener starts, bounded
by a 30s timeout so an unreachable provider can't hang startup (the heavier
non-critical init stays deferred).

Addresses PR #21 review (P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(catalog): scan air_timezone in paginated item queries

scanItemsWithTotal was not updated for the new air_timezone column, yet the
shared column lists it reads (itemColumns, qualifiedListItemColumns) include
it. Search and BrowseFavorites build their SELECTs from those lists with
COUNT(*) OVER (), so each row carried one more column than the scan had
destinations and every call failed at scan time with a pgx mismatch. Add the
missing &item.AirTimezone target between AirTime and ShowStatus.

Found during PR #21 review (critical: Search/Favorites runtime regression).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* revert(startup): keep open-scrobble sweep deferred for fast startup

Reverts 62563f5c. Running the sweep synchronously before the listener could
add up to 30s to restart-before-playback when a watch provider is
unreachable, which regresses the deliberate startup-deferral from a1a6c6cf.
Prefer the fast-startup behavior and accept the small window where a resume
immediately after restart may create a duplicate scrobble; the sweep returns
to the deferred background-init list. (Panic-safety for that list is added in
a follow-up commit.)

Per PR #21 review decision.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(startup): recover from panics in deferred background init

The deferred background-init steps run in a detached goroutine after the
HTTP listener is already accepting connections. An unrecovered panic in any
step (queue seeding, legacy cleanup, scrobble sweep) would crash the entire
live server. Wrap each step in a recover that logs the panic with a stack
and continues to the next step.

Found during PR #21 review.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(auth): make usernames and emails case-insensitive

Login identifiers were compared case-sensitively, so "John" and "john"
were distinct accounts and a user could not log in unless they matched the
exact casing used at registration.

Convert users.username and users.email to the citext type (migration 165).
citext compares case-insensitively while preserving the originally stored
casing for display, so the existing unique constraints become
case-insensitive and `WHERE username = $1` / `email = $1` lookups match
regardless of case with no change to the query code itself.

Also add auth.NormalizeUsername/NormalizeEmail (trim-only; case preserved),
applied at the repository chokepoints (Create, Update, GetByUsername,
GetByEmail) and before validation in the create paths, so surrounding
whitespace no longer defeats matching or creates lookalike accounts.

Verified non-destructively against the dev DB: mixed-case lookups resolve
to the same row, case-variant inserts are rejected by the unique
constraint, and the down migration cleanly reverts to text.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(sections): add trending_discover home section

A library-agnostic home section that surfaces external global trending
(TMDB or Trakt, admin-selectable) mixing movies + series, matched to
titles in the viewer's enabled libraries. TMDB uses /trending/all/{window}
(natively mixed); Trakt merges trending movies + shows. Fetched live with
a 1h in-process cache, so no background job or stored collection — and no
per-library duplication.

Appears in the admin section gallery via its recipe presets (TMDB Trending
Today/This Week, Trakt Trending); featured -> hero via the existing flag.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: add trending_discover persistent snapshot design

Replace the in-process 1h trending cache with a background-refreshed,
persisted snapshot for reliability under upstream failure and sync-run
observability.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: add trending_discover persistent snapshot implementation plan

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(sections): tidy trending_discover fetch and cache helpers

Extract newTrendingEntry, reuse orderMediaItems, and collapse concurrent
cache-miss loads with singleflight. Baseline for the persistent snapshot work.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sections): add trending_discover_snapshots table

* feat(sections): add trending snapshot model and repository

* feat(sections): list enabled trending_discover section configs

* feat(sections): add trending refresher with persisted snapshots

* feat(tasks): add refresh_trending_discover task

* refactor(sections): read trending_discover from persisted snapshot

* feat: wire trending refresh task and snapshot reader

* chore(sections): satisfy lint (wrap trakt errors, lift source/window constants)

* fix(migrations): renumber trending_discover_snapshots 166 -> 167

The shared dev DB already recorded version 166 (166_trending_blend_collection_type
from another branch), so the integer-version migration runner silently skipped our
166 and the table was never created — the trending section errored out empty.
167 is the next free version.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sections): harden trending refresher per PR review

- Interleave Trakt movies/shows by rank so the mixed row shows both types
  instead of burying all series past the display limit.
- Treat any Trakt sub-fetch failure as fatal (errors.Join) so a partial
  result never overwrites the last-good snapshot with a media type missing.
- Skip non-title entries (TMDB trending/all returns media_type "person") in
  both ID batching and ordering so they can't match an unrelated library title.
- Guard the refresh task against a nil refresher.
- Tests: person skip, Trakt interleave, Trakt partial-failure preserves
  last-good, snapshot read error propagation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(catalog): project air_timezone in episode catalog subquery

episodeCatalogSelectBody is the derived "mi" relation that episode catalog
hydration and preview read qualifiedListItemColumns("mi") from. The local
episode airtimes feature (808f265f) added air_timezone to the shared column
lists but not to this hand-written subquery, so the outer projection
referenced mi.air_timezone, which the subquery never exposed.

Postgres returns SQLSTATE 42703 (undefined_column), which is not one of the
codes episodeCatalogEntriesUnavailable treats as "fast path unavailable" (it
only catches 42P01/42883), so episode catalog requests failed with HTTP 500
instead of degrading. movie and series scopes query media_items directly, so
the column is present there and only episode scope broke.

Add si.air_timezone to the subquery, and add a regression test asserting that
episodeCatalogSelectBody exposes every column qualifiedListItemColumns reads
off mi, so future additions to the shared column lists cannot silently drift
from the episode read model again.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(calendar): order events by viewer-local wall-clock time

The local-airtime change re-sorted calendar events in Go using air_at,
the absolute UTC instant, which is nil whenever air_timezone is unset.
Since air_timezone is only inferred for a few networks/countries, most
events fell through to the alphabetical title tiebreak while still
displaying their raw air_time, so each day appeared scrambled.

Sort each local day by the wall-clock time the viewer actually sees,
mirroring the client: zoned events convert air_at into the viewer
timezone, unzoned events use the raw air_time, and date-only entries
(no air_time) sort last. The timezone reasoning lives in the new
catalog.CalendarEventLocalTime helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(calendar): add presets design spec and implementation plan

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): generalize personal filter to an id-set restriction

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): add per-profile followed/favorites/watchlist/watched resolvers

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): resolve presets to id-sets and overlay watched status

Also drops the now-unused Filter/UserID/ProfileID fields from the
blendUpcomingIntoDiscoverRows CalendarFilter literal in recommendations.go,
which only wants an unrestricted windowed query.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): wire popular and trending sources into calendar handler

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): add watched field to CalendarEvent type

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): preset selector with responsive pills, persistence, empty-state nudge

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): dim and check-mark already-watched event cards

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(calendar): hide server-wide Popular preset in web UI for now

Popular reflects server-wide watch counts, which are sparse on a
low-traffic server. Hidden from the selector, URL allowlist, and
empty-state nudge; backend filter and the CalendarFilter type are
left intact so re-enabling is a one-line change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(calendar): simplify preset handler and reuse storage util

- Extract hardcoded trending snapshot source/window to named constants.
- Collapse the three identical personal-preset nil-checks into one case.
- Persist the selected preset through the shared storage util (try/catch
  wrapped) instead of raw localStorage with manual SSR guards.
- Derive KNOWN_FILTERS from PRESET_OPTIONS so the lists can't drift.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(jellycompat): add /Users endpoint and unified sa_ admin-key auth

A user could not connect Tunarr to the Jellyfin-compat API. Tunarr's media-source
health check probes both /Users/Me and /Users; for API-key connections (its
recommended method) there is no "me", so it relies on GET /Users, which we did
not implement. Adding it alone was insufficient: a Silo sa_ admin key only
authorized two autoscan routes, so an API-key source would pass the ping but 401
on every browse/stream call.

This adds GET /Users and unifies auth so an sa_ admin key authorizes the same
browse/stream routes a session token does — matching Jellyfin, where an API key
authorizes every endpoint:

- GET /Users returns the caller's own user as a single-element list
  (current-profile-only; Silo is multi-account, so listing all users would leak
  across households). Behavior verified against real Jellyfin 10.11.8.
- An sa_ admin key synthesizes a compat session bound to the account's primary
  profile, injected into request context so existing handlers work unchanged.
  Applied to the browse group (RequireSessionOrAPIKeySession) and the stream
  group (PlaybackSessionAuth); /Library/VirtualFolders keeps its admin-bool path.
- The key + owning user are re-validated on every request (revocation is
  immediate); only the primary-profile lookup is cached. HLS follow-ups that
  carry only PlaySessionId resolve the negotiated session's sa_ CompatToken.

Validated end-to-end on dev: connect -> list user -> libraries -> browse ->
PlaybackInfo -> HLS stream, including PlaySessionId-only follow-ups.

Security notes: an admin key acts as the primary (parent) profile, so it bypasses
child-profile parental/PIN restrictions and can mutate the primary profile's
watch state — acceptable for an admin-trust key.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Code <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Silo Server Migration <noreply@silo-server.invalid>
Co-authored-by: zZebrahz <zzebrahz@gmail.com>
Co-authored-by: CoffeeKnyte <67730400+CoffeeKnyte@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Silo Server Developer <warmasterx555@gmail.com>
2026-06-08 14:25:10 -04:00
Silo Server Migration c085b12fd1 Initial Silo migration 2026-05-22 23:26:56 -04:00