Files
silo-server/cmd/silo/main.go
T
dc4b9a0909 feat(settings): add the cross-platform settings contract and its manifest (#479)
* docs(settings): define the cross-platform settings contract

Turns the audit in #376 into a decision-complete design for how user settings
work across the server, bundled web client, Apple clients, and Android clients.

Today there are three partial contracts - the server registry, the web client's
own manifest, and independently owned key constants in each native client - and
they have measurably drifted. The root enabler is that keyUsesUserScope returns
true for any unregistered key, so a client can invent a production setting
unilaterally and the server stores it as an unvalidated string.

The design decides:

Ownership. Every production user-facing setting needs a server-owned manifest
entry, even when the value is stored only on one client. The single exception is
private local.<client>.* diagnostics, bounded by five conditions.

Types and scopes. Native JSON values instead of strings. Five remote scopes plus
client_local, and each definition declares its own resolution order rather than
inheriting a global precedence.

Preferences versus restrictions. internal/policy already resolves
max_playback_quality and metadata-language limits over the same controls this
contract resolves preferences for. Definitions declare constrained_by, the
effective response reports the permitted value alongside the user's stored one,
and a mutation exceeding a restriction is stored rather than rejected - a capped
4K preference should take effect the day the cap lifts, not be destroyed by it.

Compatibility. Widening a scope, adding an enum member, or widening a range is
additive and revision-tagged; narrowing anything needs a new key. introduced_in
is a manifest revision attached to individual enum members and scopes, not just
whole definitions, so a newer client never offers a choice an older server will
reject.

Rollout. One coordinated breaking release, with no compatibility shim,
projection, or client fallback. After the cutover no future setting requires
coordination. No settings version check goes in the authenticated middleware and
nothing returns 426: deleting the old routes already produces the break, and a
gate would be more code in four repos for the same outcome while permanently
coupling every endpoint to one subsystem's versioning.

Scope placement. Appearance and date/time move from account to profile scope.
Account scope was an artifact of pre-profile storage; leaving it there means a
household shares one theme and text size, and any non-child profile can restyle
everyone else.

Read path. Batched context resolution, index requirements, a session-snapshot
rule, and a no-regression benchmark gating storage consolidation - profile_series
resolution is per-item, so a season view would otherwise issue one request per
episode.

Verified against the current server, Apple, and Android implementations. Two
findings shape it: the unknown-key extension bag is real, and v1 scope reads NOT
LOCKED, so removing the legacy surface needs no amendment if it lands before
lock.

Related to #376.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the canonical settings contract manifest

First implementation step for the cross-platform settings contract (#376).
Adds the artifact everything else depends on: the manifest, its JSON Schema,
the object value schemas, and a Go loader that validates the whole thing at
load time. No routes, no storage, no behavior change — nothing reads this yet.

contracts/settings/v1/ holds the artifact at a stable path because clients
vendor it and generate bindings from it. The embed directive has to sit beside
it (go:embed cannot reach outside its own directory), so that directory is a
tiny Go package containing nothing else; loading and validation live in
internal/settingscontract.

38 definitions: 35 remote, 3 contract-known client_local. That covers every key
the legacy registry accepts, every unregistered key the extension bag was
silently accepting from the web client, every unregistered device key Android
writes, and the profile preference columns that become settings.

Registering the previously-unregistered keys is where the drift shows up, and
the manifest records each case in a notes field:

- ui_theme, ui_text_scale, ui_text_weight, ui_high_contrast,
  ui_custom_theme_vars, and ui_custom_css reached the server only because
  keyUsesUserScope returns true for any unregistered key. They are now typed,
  renamed to the dotted convention every other key uses, and moved to profile
  scope per the design.
- player.match_frame_rate and player.sleep_timer_default_minutes are written by
  Android against a server that does not register them, so every write and reset
  is currently rejected. Registered.
- player.next_up_prompt_seconds is Android's alias for
  playback.next_up_prompt_seconds and does not become a definition; the test
  matrix pins it as a migration alias.
- player.playback_speed is capped at 3.0, matching the server rather than
  Android's 4.0.
- subtitle_appearance becomes playback.subtitle_appearance. Every other
  canonical key carries a domain prefix, and preserving accidental key names is
  an explicit non-goal of the design.

Validation is deliberately stricter than the schema can express. Beyond shape,
it enforces that a resolution order ends in "default", that it only resolves
scopes the definition allows, and — the one most likely to bite — that every
writable scope is actually read, so a setting cannot accept writes at a scope it
will never honor. Defaults are validated against their own value schema, so a
default that violates its own range or enum fails at load. Revision tags are
checked to never run ahead of the manifest revision, which is what makes
revision-aware client filtering trustworthy. Ceiling and floor policy
constraints are rejected on unordered types, where capping would silently do
nothing; playback.preferred_quality's enum is therefore ordered ascending.

ValidateValue is the single validation path, so the mutation endpoint, the
migration, and the manifest's own default checks cannot diverge later. Numbers
decode through json.Number so an integer setting rejects 30.5 rather than
truncating, and object values validate against their referenced JSON Schema
instead of accepting arbitrary JSON the way validateJSONSetting does today.

Canonicalization implements RFC 8785 over the value domain the contract uses:
sorted keys, no insignificant whitespace, ECMAScript number formatting. The
digest is the ETag, and PublicBytes strips maintainer notes so the served
manifest never carries internal commentary.

Promotes santhosh-tekuri/jsonschema/v6 from indirect to direct.

Verification: 124 tests pass across 16 cases; golangci-lint clean;
make verify-local-paths passes. Two failures in internal/api/handlers
(TestRemoveJellyfinCompatWebDisablesWebSetting, the playback v3 seek recovery
test) reproduce unchanged on main and are unrelated.

Part of #376.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): give ui.theme a device override

Theme joins text scale, text weight, and high contrast as a profile default
with an optional per-device override, resolving profile_device -> profile ->
default. The right theme is partly a function of the screen and the room — a
light theme on a phone in daylight, a dark one on a TV at night — which is the
same reasoning the other three appearance keys already used.

All four appearance settings now cascade consistently, which also means one
rule to explain in the UI rather than "these three follow the device, that one
does not".

ui.custom_theme_vars and ui.custom_css stay profile-wide. They are authored
styling rather than a contextual preference, so a profile's custom tokens still
apply on top of whichever theme a device resolves to. Recorded in the
definition notes because it is a visible consequence: vars tuned against a dark
theme will sit on top of a light one if a device overrides the theme. Widening
those to profile_device later is an additive revision bump if it turns out to
matter.

Part of #376.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web): tag local appearance caches with their owning account

The theme, text scale, text weight, high contrast, custom theme variable
and custom CSS caches in localStorage were untagged, so on a shared
browser a second account inherited the first account's appearance: with
no server value of its own, every fallback resolved to whatever the
previous account had stored, and the leftover `silo-theme` key also
suppressed the admin-configured default theme for the new account.

DateTimeFormatProvider already solved this by stamping its cache with the
authenticated user id and refusing another account's values. Extract that
mechanism into `createOwnedCache` in utils/storage.ts (where key
namespacing lives) and put all three groups behind it, so appearance and
custom theme get the same protection instead of a third copy of the rule.

- Each group carries its own owner stamp. A shared stamp would be unsafe:
  the groups are written by hooks nested inside each other, and effects
  run inner-first, so whichever hook stamped first would vouch for the
  other's still-stale values.
- A null owner (auth bootstrapping, or signed out) still trusts the
  cache, which keeps the warm start and the login screen's last look.
- An unstamped cache is not trusted once an account is known, so existing
  users take a one-time appearance reset on first load rather than a
  chance of seeing someone else's settings.
- When a foreign cache is detected the values are dropped and the empty
  cache is handed to the new account, so a later single save cannot
  re-trust the rest of the previous account's state.

Owner is the user id because /settings is user-scoped server side; it
lives in one helper (`appearanceCacheOwner`) so it can be widened if
appearance moves to profile scope. `shouldLoadApiTheme` is gone: it had
become a synonym for `appearanceCacheOwner(...) !== null` with no callers
left.

Part of #376

AI-use disclosure: implemented with Claude Code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): make the settings contract enforceable and fix the appearance cache

The contract manifest landed as a document nothing checked. This makes it a
mechanism, and fixes the one defect in the change set that hurt users on merge
rather than at cutover.

Web appearance cache. useTheme cleared the cache for any account whose stamp
did not match and never repopulated it — the only writers were the four
user-action setters — so every upgrading user lost their warm start on every
load, not once, and x-large-text and high-contrast users lost theirs too. The
owner-stamp protocol is replaced with per-account key namespacing
(`silo-theme:7`): a foreign value is absent rather than present-and-distrusted,
so nothing has to be deleted, the first account keeps its warm start, and there
is no shared stamp for a second tab, a stale debounce timer, or an out-of-order
effect to race on. Widening ownership to profile scope, which this manifest
requires, is now a change to appearanceCacheOwner alone. Adds the API-to-cache
mirror useTheme was missing, cancels pending debounced writes across an account
change, and re-seeds provider state during render so no frame paints the
previous account's look.

Canonicalization. writeCanonical used json.Marshal, which HTML-escapes < > and
&, and canonicalNumber used Go's 'g' format — both diverge from RFC 8785, so
the first label containing an ampersand or bound below 1e-4 would have forked
the server's ETag from every conforming client. Output is now byte-identical to
ECMAScript String() across the edge cases, verified against node. The ETag also
covers the value schemas, which decide what the server accepts and previously
could change while the tag stood still. All four derived representations are
memoized; a conditional GET no longer costs a full parse and re-serialize.

Validation. strictUnmarshal's decoder.More() answered false for a stray ] or },
so `true]` validated as a boolean. Enum matching compared fmt.Sprintf tokens, so
the string "3" satisfied an integer member. Declared steps were never enforced.
The language pattern rejected tags both mobile platforms emit unprompted
(en_US, ca-ES-valencia, ar-EG-u-nu-latn) and never normalized case, so en-US and
en-us were two rows for one preference; NormalizeValue now canonicalizes on the
shared path.

Manifest. show_forced_subtitles defaulted false where the server column is NOT
NULL DEFAULT true, which would have turned forced subtitles off for every
profile that never touched it. preferred_quality declared 13 members where the
planner speaks 6 and collapses the rest to auto. metadata_language's allowlist
was bound to the very column it migrates from. subtitle-appearance pinned
fontFamily to three families while Apple stores any installed system font.
Registers five user-facing settings the clients already ship, and corrects three
notes that described Android behaviour that was not true.

Enforcement. The package had no non-test callers, so MustLoad never ran; it now
loads and logs at startup. The inventory test compared the manifest against a
hand-copied map and could not see the drift it named; it now iterates
settingsRegistry and checks defaults too — both verified to fail on injected
drift. Adds .github/workflows/ci.yml, the repo's first CI that runs go test,
go vet, gofmt, and the frontend suite. Known pre-existing failures are named
individually in the Makefile so everything else stays gated and the list can
only shrink.

Part of #135

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): align the sleep timer default and range with the shipped client

Android is the only client that implements this setting. It clamps to 0..240
and defaults to 30. The manifest said 0..480 with a default of 0, so a
manifest-driven UI would have offered durations no client can store, and every
user who never opened the picker would have had the preset silently turned off
at cutover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: give the new workflow the deps it actually needs

The first run exposed two gaps in the workflow itself. go build ./... fails
without libvips headers, because h2non/bimg binds libvips through cgo and
pkg-config; the Dockerfile installs the same package. And pnpm/action-setup
resolves its version from package.json, but there is no package.json at the
repo root — the packageManager field lives in web/package.json, and a job's
defaults.run.working-directory does not apply to an action's inputs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(web): stop the diagnostics download test depending on the Node version

new Response(blob) reads the body through blob.stream(), which jsdom's Blob
does not implement on Node 22 — the version the Dockerfile builds with. The
test passed locally on Node 24 and threw "object.stream is not a function" in
CI. Nothing in it asserts on the body, only that the object URL and filename
reach the anchor, so a string body is equivalent and works on both.

Surfaced by the CI workflow added in this branch, which is the first thing in
this repo to run the frontend suite anywhere but a developer's machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(build): copy the settings contract into the container build context

Both Dockerfiles copy cmd/, internal/, migrations/ and web/embed.go, but the
manifest lives in contracts/settings/v1 — an embedded Go package that sits
outside internal/ because clients vendor those files. The image build therefore
fails with "no required module provides package .../contracts/settings/v1".

Caught deploying to the dev box. Nothing had built an image since the manifest
landed: the Docker workflow only runs on pushes to main and workflow_dispatch,
and CI's go build runs against a full checkout, so neither gate covers the
container context. This would have broken the published image on merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): enforce the language-tag and step constraints the manifest declares

A sweep of all 43 manifest definitions against the running server (160 checks:
declared default, both boundaries, and deliberate violations for each remote
key) found two places where the live registry accepts what the contract
forbids. Both are fixed by calling the contract's own validators rather than
adding a second implementation.

playback.audio_language was checked as "32 characters or fewer", so the server
stored "!!!" for a field the manifest declares as language_tag — a value track
matching would then silently never match. It now requires a well-formed tag via
settingscontract.NormalizeLanguageTag. The empty string is still accepted: the
string-only endpoint has no way to send null, and both Android and web send ""
to clear the choice, so rejecting it would break clearing the preference.

player.playback_speed declared step 0.05 and nothing enforced it, so 0.26 was
stored — a value no client's stepper can represent and that every client would
silently snap on the next write. settingscontract.StepAligned is now exported
and used by both the contract validator and the registry, so there is one
definition of "on step" rather than two that can drift.

This gives the contract its first production consumer beyond the startup load,
which is the direction Phase 2 continues in.

Also fixes a genuinely flaky test that the new CI gate would have hit
intermittently: TestRemoveJellyfinCompatWebDisablesWebSetting used t.TempDir as
the install root, but the endpoint returns 202 and its goroutine keeps writing
there after the test body returns, so cleanup tripped "directory not empty"
roughly one run in four. Confirmed pre-existing and unrelated to settings; the
suite now passes six consecutive full-package runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): keep widened numeric bounds resolvable at older revisions

A bound was one scalar plus the revision that introduced it, which discards
the value it replaced. Widening a maximum from 240 to 480 at revision 3 left
a revision-3 client with no correct answer against a revision-1 server:
honoring 480 offers values that server rejects, and filtering the tagged
bound out leaves the setting unbounded. Since clients are specified to filter
their pinned contract against the server's advertised revision, the bound has
to carry what it used to be.

Bounds now hold their full history, oldest first, and AtRevision hands back
the limit a given peer actually enforces. A bound nobody has widened still
serializes as a bare number, so the manifest reads the same and untouched
entries do not churn the ETag.

Validation gains the rules the representation makes checkable: a maximum may
only grow and a minimum may only shrink, history is strictly ordered, later
entries must say when they arrived, and the first entry cannot predate the
definition. That last rule is the lower bound allowed_scopes already
enforced; the same gap is closed for enum members, which could previously
claim to predate the definition containing them.

Reported by Codex review on #479.

* fix(settings): accept the partial subtitle appearance objects already stored

The schema required all nine properties, but the current API accepts and
round-trips sparse objects — settings_device_test.go stores
{"fontSize":"xxlarge"} and reads it back — and the web client has always
merged whatever it gets over DEFAULT_SUBTITLE_APPEARANCE. Requiring the full
object would have made the cutover migration quarantine preferences users
really set, or block on them.

Every property is now optional and a stored value is documented as a sparse
override merged over the definition's complete default. An empty object is
still rejected: an override that overrides nothing is the same state as no
override, which the contract represents as unset.

Cross-scope resolution is deliberately unchanged. A device override still
replaces the profile's object rather than merging into it, because a device
override means "draw subtitles this way on this screen", not "amend the
profile" — and that is what the server does today.

Reported by Codex review on #479.

* fix(jellycompat): scan the parent directory when a sidecar changes

Autoscan matched scantrigger rejections by comparing RequestError.Message
against literal strings. One of those messages became "Unsupported media file
extension for library type" and the copy in handlers_autoscan.go did not, so
the comparison silently stopped matching.

The effect is user-visible: a Jellyfin client posting a change for Movie.nfo
or poster.jpg gets a 400 and the batch is abandoned, when the sidecar should
have resolved to a scan of the directory containing it. Three tests covered
exactly this and had been excluded rather than read.

RequestError now carries a Reason the caller can switch on. Message stays
prose for the client reading the response — it is meant to be reworded, and
nothing should break when it is.

Also makes two tests honest about asynchronous work. The Jellyfin Web
teardown deleted its install root while the operation goroutine was still
writing to it, where a late write recreates a path RemoveAll already walked
past; it now waits for the operation's terminal state, which required
exporting CurrentWebOperation. And the direct-play If-Range test pinned size
and mtime so ctime was the only remaining validator, then read it back inside
a single coarse-clock tick — it failed about 85% of the time on main for a
reason unrelated to what it tests, and now rewrites until the stamp moves.

With those fixed, GOTEST_KNOWN_FAILURES is empty and gone: make test-go runs
the whole Go suite. The one test that cannot pass yet —
TestHandleReplanPlaybackV3SeekFailureRecoveryNeverChangesMediaVersion, which
has failed since the commit that introduced it and describes unimplemented v3
planner behavior — carries a t.Skip explaining that where the test is, rather
than a regex in the Makefile.

Reported by CodeRabbit review on #479.

* fix(settings): reject JSON the decoder would otherwise rewrite

Two cases where encoding/json accepts input by quietly changing it, which is
the one thing a contract promising byte-identical agreement between peers
cannot tolerate.

Duplicate object properties. jsonschema.UnmarshalJSON keeps the last
occurrence, so {"fontSize":"small","fontSize":"large"} validated and stored
"large". Which one wins is a property of the parser, not of the contract: a
client generated against a different JSON library can disagree about what it
just sent, and the canonical form cannot represent the duplicate at all.

Lone surrogates. An unpaired \ud800 became U+FFFD and canonicalization
reported success, so the server would issue canonical bytes and an ETag for
an artifact a conforming implementation must refuse — RFC 8785 requires
terminating here. Substitution also means the value read back is not the
value written.

Both checks run before the decode that would hide them, on the shared
decodeJSON path that the manifest, its public projection and every value
schema go through, and again on the object branch of ValidateValue, which
uses a different decoder.

Reported by Codex review on #479.

* ci: gate Go lint on the lines a branch changes

AGENTS.md told contributors CI ran the same checks as `make lint`, and the Go
job ran only gofmt and vet. A change failing the documented Go lint gate
passed all three jobs.

Running the linter as-is is not an option: the tree has ~296 findings today,
which is why this half of `make lint` was never enforced. Blocking every PR
on a cleanup nobody has scheduled gets the gate deleted again, so CI runs
with --new-from-merge-base and only the lines a branch touches have to be
clean. The count can then only fall.

golangci-lint is built from source at a pinned version rather than
downloaded. A released binary refuses to run against a Go newer than the one
it was built with, and go.mod here tracks Go closely enough that the current
release already fails that way on 1.26.4.

.golangci.yml declared version 2 while still using v1's issues.exclude-rules
key. Current golangci-lint ignores it, so the "allow repeated strings and
unchecked cleanup errors in tests" exclusions silently did not apply — 16
findings in test files that the config says to skip. Moved to
linters.exclusions, which `golangci-lint config verify` accepts.

The four lines this surfaced in scantrigger are fixed rather than excluded:
its repeated status codes and messages are now named constants, so one
condition cannot end up worded two ways.

Also drops the workflow token to contents:read and stops persisting
credentials in the three checkouts, neither of which any job needs.

Reported by CodeRabbit and Codex review on #479.

* docs(v1): record the settings removal as a pre-lock exception

The design removes the legacy /api/v1/settings routes and the profile DTO
preference fields, while AGENTS.md states /api/v1 is additive-only and
removals go through Deprecation/Sunset. Read together those contradict.

They do not actually conflict: v1-scope.md scopes the additive-only rule to
"when the scope locks", and the scope is still open, so a removal taken now
is in scope and there is no amendment process to invoke yet. But that
reasoning lived only in the settings design, where nobody checking the API
policy would find it.

v1-scope.md now carries a pre-lock removals table naming what goes and why
waiting is worse, and states the deadline the argument depends on: a removal
listed there must ship before lock or fall back to Deprecation/Sunset.
AGENTS.md points at the table and says to treat an unlisted removal as a
mistake.

Reported by CodeRabbit review on #479.

* fix(settings): clear the remaining review findings

Small, unrelated except that each was raised on #479.

compileObjectSchemas parsed every non-directory file under schemas/ as a JSON
Schema, so a stray editor backup or .DS_Store would panic the server at
startup through MustLoad. schema_ref can only name a .json file; anything
else is skipped.

cmd/silo used MustLoad while the ETag check beside it and every other startup
failure use log.Fatalf. It now fails the same way, so a bad contract prints
an error instead of a stack trace.

TestRegistryDefaultsMatchTheContract called scalarDefault before handling
null, and scalarDefault rejects null as non-scalar — so the subtest skipped
and the comparison after it was unreachable. A nullable contract default
could disagree with a non-empty registry default and nothing failed.
Confirmed by injecting that drift, which now reports it.

The three appearance providers each adapted the auth context to
AppearanceAuth with identical code, putting the shape of auth back in three
places that widening cache ownership would have to find. useAppearanceCacheOwner
now does it once.

useTheme.test.ts cleared storage.KEYS between cases, but appearanceCache
writes namespaced keys and an owner pointer that are not in that list, so
both survived and the suite was order-dependent. It clears the store, as
storage.test.ts already did.

The abs_smart_collection_store comment is reworded rather than given back its
SQL quotes: gofmt folds a pair of apostrophes in a doc comment into a
typographic quote, which is how it became one in the first place.

Reported by CodeRabbit review on #479.

* feat(settings): add canonical typed storage for the settings contract

The cross-platform settings contract needs one typed store behind it before a
resolver, routes or a migration can exist. This adds that storage to both
user-store backends and holds them to identical behavior.

PostgreSQL gets user_setting_values with the scope CHECK constraints, the five
partial unique indexes that enforce one explicit value per identity, and the
covering indexes the one-query read path needs, plus user_setting_mutations for
mutation_id idempotency and the inert user_setting_migration_rejects audit
table. The per-user SQLite store gets the same shape minus user_id, since that
database is already user-scoped.

The UserStore interface grows the typed operations: read one explicit value at
one scope, collect every candidate row for a resolution request in a single
query, upsert with a revision increment, unset, and the idempotency receipt
operations. The resolution read deliberately returns unranked candidates so the
resolver can rank in Go — one query per request, never one per scope, which the
pgx query-count test pins.

Delete behavior is application-enforced. Neither backend can inherit it from
constraints: the SQLite store declares no foreign keys, and library, series and
device columns are not FK targets in Postgres either. Profile deletion cascades
to profile-anchored values while account scope survives, forgetting a device
clears its profile_device values alongside the legacy overrides, and the
library/series purges remove only what is scoped to that entity.

The shared conformance suite covers all of it, including the set-versus-unset
distinction for false, 0, "" and null, so a divergence between the two backends
fails a test rather than reaching a client.

Part of #376

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(settings): pin the settings-value schema constraints in both backends

Completes the storage track. The conformance suite exercises the store API,
which validates identities in Go before any SQL runs — so nothing noticed
whether the CHECK constraints and partial unique indexes actually existed.
The one-time migration writes these rows in bulk without going through the
per-request path, so the schema is the only thing guarding it.

Adds constraint tests to both backends covering every scope's column
requirements, rejection of an unknown scope, a profile that does not exist,
non-JSON values, and each of the five partial unique indexes.

Also clears the lint the storage commit did not get to: sql.ErrNoRows and
pgx.ErrNoRows compared with == rather than errors.Is (which fails on a
wrapped error), an unchecked rows.Close, and repeated fixture literals in the
shared suite now named so a backend that confuses two scope columns fails on
the assertion rather than on a typo.

* fix(settings): close the review findings in the validator and the theme cache

Four defects the existing tests did not reach.

The web theme resolver compared the server's value against the appearance
cache and fell back when they agreed, but the mirroring effect writes the
server's value into that same cache — so the comparison held on the first
render and stopped holding on the second, reverting an explicitly chosen
theme to the default. The server's value is this account's own stored
choice, so it now simply wins. The regression test re-renders rather than
asserting on the first paint, which is why the original one passed.

golangci-lint's exclusions.paths is a path regex, not a directory list, so
a bare `web` also excluded internal/jellycompat/web_component.go,
internal/webhooksync/, internal/notifications/webhook*.go and eleven other
non-test files that were being linted before. Anchored.

json.Number is a string kind, so `"1.5"` unmarshalled into it happily and
Float64 parsed the quoted digits: a numeric setting validated as a JSON
string and NormalizeValue stored the quoted form into jsonb. Rejected.

The lone-surrogate check ran only on the object branch, so a lone surrogate
in ui.custom_css decoded to U+FFFD on SQLite and was refused outright by
Postgres jsonb — the two backends disagreeing about whether the same value
could be stored. Hoisted to cover every type.

The strict language-tag validation this branch added is correct, but it
rejects what the shipped Android client sends; the companion fix is
silo-android 4aeb78b4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(auth): stop TestJWT_TamperedToken passing a valid signature

The test overwrote the last character of the signature with "X". An
HMAC-SHA256 signature is 32 bytes, so its base64url encoding is 43
characters and the final one carries only four significant bits — U, V, W
and X all decode to the same trailing byte. Roughly one token in sixteen
was therefore left byte-identical and validly signed, and the test failed
because ValidateToken correctly accepted it.

Measured at 3098/50000 (6.2%) over distinct signatures; it just failed the
Go job on this branch for reasons unrelated to the branch. Flipping a
character in the middle of the signature is 0/50000.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): reject raw invalid UTF-8, not just escaped surrogates

The previous commit hoisted the lone-surrogate check to cover every value
type, but that only closes the escaped path. A raw 0xff byte inside a
quoted string — what an HTTP body carries when a client encodes text in the
wrong charset — is not an escape, so the surrogate scan never sees it, while
encoding/json still substitutes U+FFFD and reports success. NormalizeValue
then stores the original bytes, which SQLite's json_valid accepts and
Postgres jsonb refuses: the same backend divergence, reached the other way.

Found by the Codex review bot on the previous commit's own diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): size the library page state bound to what the web client writes

ui.library_page_state's `search` was bounded at 256 characters. The web
client serializes an advanced library view as URLSearchParams, encoding each
filter rule as three groups[i][rules][j][field|op|value] keys — measured at
216 characters for one rule, 518 for three, 820 for five.

The current endpoint validates this key by checking only that it parses, so
those oversized values are already stored in production. Typing them at the
declared bound would have failed the migration for anyone who had saved a
view with more than one filter rule, and rejected the equivalent write
afterwards.

Raised to 4096, which clears ten rules with room to spare while staying a
real bound. The test pins it against the key shapes
libraryPageSearchParams.ts actually emits rather than a round number.

Reported by the Codex review bot; the lengths above were measured by calling
serializeLibraryPageSearchParams, not estimated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): split quality into two axes and register the orphan keys

Two manifest changes the cutover needs.

**Quality becomes resolution + bitrate.** The legacy ladder values
(1080p-high, 720p-medium, 1080p-8, 420p, 328p) were never a third dimension
— they are a bitrate spelled into the resolution string. The web player
already decomposes them: useTranscodeQuality.ts defines 1080p-high as
{resolution: 1080p, bitrate: 10000} and sends the two separately, so the
compound form never reached the wire. Downloads went further and kept only
a bitrate ladder.

So playback.preferred_quality keeps the six clean resolutions and
playback.max_bitrate_kbps becomes the second axis, nullable because
"uncapped" is a real answer and a numeric sentinel would need widening
every time hardware improves. Clients compose their own presets from the
pair, which means retuning what "High" means is a client release rather
than a contract break. Migration decomposes each legacy value losslessly,
so none of them lands in the rejects table.

**The five extension-bag keys are now definitions.** card_overlays,
next_up_mode, sidebar_pins, disabled_library_ids and library_order reached
the server only through the unknown-key path, stored as unvalidated
strings. Two of them the server reads back — next_up_mode decides home
section assembly and card_overlays falls back to an admin default — so
they cannot be demoted to client-local. Registering them is what lets the
extension bag close.

Adds three schemas for their shapes and a test that exercises every
schema_ref against a real value: each of these is nullable with a null
default, so the existing default-validation test returns at the null branch
without ever compiling the reference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the canonical resolution engine

One answer to "what is this setting, for this profile, on this device, for
this content". Before this, each caller carried its own ladder:
catalog/detail.go resolved subtitles across four levels by hand and audio
across three, handlers/settings.go had a two-level device/user resolution
with a lazy write-back inside a GET, and jellycompat read profile columns
directly. Those disagreed about precedence, which is the drift the contract
exists to remove.

Resolution is one batched read regardless of how many keys, libraries, or
series are in play — ranking happens in Go against each definition's
declared resolution_order. Five sequential index lookups per key per item
is the implementation the design rejects, and a season view is exactly
where it would have shown up.

An absent identity drops its scope rather than erroring, so one code path
serves an identified client, an anonymous jellycompat seed, and a batch
spanning many series. Rows for a foreign profile, device, library or series
are ignored even though the batched read returns them.

Constraints narrow without destroying: a capped 4K preference resolves to
the cap, reports itself constrained, and keeps the authored value so it
takes effect the day the cap lifts. Two cases needed care — null on a
nullable numeric means unbounded, so a ceiling must cap it rather than rank
it equal and let the value that most needs capping slip past; and an
allowlist falls back to a permitted member rather than the definition's
default, which may itself be outside the list.

Adds ValueSchema.CompareValues to the contract package, since ordering
values is what makes a ceiling or floor mean anything and value semantics
belong with the schema that declares them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the one-time migration planner

The conversion rules from legacy settings storage to canonical values, as
ordinary Go rather than twice in two SQL dialects. Both backends read their
own rows, hand them to Plan, and write what comes back — so the decisions
are testable without a database and SQLite and Postgres cannot drift apart
in what they decide.

The rules that needed care, each pinned by a test:

Column defaults are not choices. quality_preference is NOT NULL DEFAULT
'1080p' while the contract defaults to auto, so migrating the column
unconditionally would pin every profile in the install to 1080p having
never chosen it — and that stored value would then outrank the contract
default forever. Same for language 'en', subtitle_mode 'auto', and
show_forced_subtitles true.

The empty string is unset, not a value. The legacy string API had no way to
send null, so both Android and web spell "clear my choice" as "". Storing
that would make a cleared setting outrank the default.

Legacy quality decomposes rather than rejects. Every compound value maps to
a resolution and a bitrate from the ladder in useTranscodeQuality.ts, so
nothing lands in the rejects table.

Account rows fan out to every profile, which is the account-to-profile move
the contract makes for appearance and search scope: a household that shared
one theme each end up owning theirs.

Legacy strings become typed JSON — "true" to true, "30" to 30 — or every
generated binding would fail to decode what the migration wrote.

Nullability differs per backend, so profile columns arrive as pointers and
the caller resolves "chose the default" versus "never written" when it
reads. jellycompat's DisplayPreferences blobs ride the same table under
synthetic keys and are left alone; they are that subsystem's storage.

Everything that cannot convert is recorded with a reason rather than
dropped, and a final test asserts every planned row would be accepted by
the mutation endpoint's own validation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): run the one-time migration on the SQLite backend

Wires the planner to real storage as userdb migration V15. V14 created the
tables; this fills them.

It runs inside runMigrations' existing transaction, so a database either
comes out fully migrated or untouched — a partial migration is the one
state neither the operator's backup nor a rollback covers. Pinned by a test
that rolls back and asserts nothing was left behind.

Two things the wiring had to get right that the planner could not see:

Reject identities are JSON. Postgres declares that column jsonb NOT NULL
and SQLite guards it with a json_valid CHECK, so the free-form
"profile=p1 device=d1" the planner emitted would have failed to insert — on
exactly the rows the table exists to record. They are structured documents
now, which is also queryable.

Subtitle and audio preferences are two tables keyed the same way, so they
merge into one per-series record before planning. Converting them
independently would have produced two rows racing for the same identity.

Every legacy read tolerates a missing table, since this runs against
databases created at any schema version, and preferred_metadata_language is
deliberately absent: that column exists only in the Postgres schema.

Tested end to end against a real database rather than only through the
planner — the rows land, satisfy the scope CHECK and the partial unique
indexes, and hold valid JSON. Also covers the empty-install case and
asserts a second run fails rather than silently doubling every value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): run the one-time migration on the Postgres backend

The mirror of userdb V15, registered with goose as a Go migration rather
than SQL: the conversion validates every value against its own definition
and re-encodes it as typed JSON, and one legacy quality string becomes two
rows — neither is expressible in SQL without duplicating the manifest. The
rules stay in internal/settingsmigrate, so the two backends cannot disagree.

RunTx, so the whole backfill lands in goose's transaction. The down
migration empties the canonical tables; the legacy ones are never touched
by the up, which is what keeps the cutover reversible until the follow-up
migration drops the superseded columns.

preferred_metadata_language is read here and only here — the column exists
in this schema and not in SQLite's, so this is the sole source for
catalog.metadata_language.

Verified against a real Postgres: the full goose chain runs, 1080p-high
decomposes to ("1080p", 10000), values land as typed jsonb rather than
strings (jsonb_typeof reports number), rejects carry a queryable jsonb
identity, and the composite profile foreign key refuses a row naming a
profile that does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the canonical settings API

The routes that make the typed storage reachable. Until now the manifest,
the resolver and the migration all existed with nothing able to call them.

GET /settings/contract serves the public manifest behind an ETag — clients
vendor a pinned copy and generate bindings from it, so the common request
asks "still the same contract?" rather than transferring it. Its
capabilities sibling reports revision and supported scopes for feature
detection instead of version sniffing.

/settings/values/{key} reads, writes and clears an explicit value at one
named scope, which is what a reset affordance needs: "did I set this here"
is a different question from "what applies", and the old endpoint could
only answer a blurred version of both. Scope comes from the query while
profile and device come from session headers, so one profile cannot address
another's settings by naming it.

/settings/values/effective resolves any number of keys in one request, with
the resolution ladder and the source of each answer reported so a client can
offer "reset this device's override" against the exact row holding it.
Asking for no keys returns every remote setting, which is what a settings
screen wants.

Writes are idempotent when a client sends X-Silo-Mutation-Id: a retry after
a dropped response replays the receipt, and reusing an id with different
content is a conflict rather than a silent overwrite of the wrong thing.

Three things the string-only endpoint could not do, each pinned by a test:
an unknown key is refused rather than stored in the extension bag, values
are checked against their declared type and range, and a write to a scope
the definition does not allow is rejected.

Registered before the catch-all /{key} routes, which would otherwise
swallow "contract" and "values" as setting names. The legacy endpoints stay
live for now; deleting them is the next commit, once their consumers move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): generate typed bindings for all four languages

One generator rather than one per repo. The point of the contract is that
four codebases agree on keys, types, scopes and defaults, and four
independently written generators would be four chances to disagree.

Go and TypeScript land in this repo; Kotlin and Swift are written into the
sibling client checkouts, skipped with a note when they are not present so
a server-only developer can still run it. Output is sorted by key so an
unrelated manifest edit does not produce spurious diffs.

The Kotlin output is the interesting one: it generates the DeviceSettings
allowlist Android maintained by hand, plus the BOOLEAN_KEYS/INT_KEYS/
DOUBLE_KEYS classification it kept as a *second* hand-maintained table that
had to agree with the first. Both are manifest questions now, so the whole
class of "wrote a local key to the server" and "flushed a value the store
could not parse" bugs stops being possible by construction.

The TypeScript output carries the full definition table — labels, controls,
enum members, bounds — so web/src/lib/settingsManifest.ts can be deleted
rather than kept in sync: it declared 17 definitions against the contract's
49, with its own two-scope model that does not match the contract's five.

make verify-settings-bindings fails when the committed output disagrees
with the manifest, wired into CI, so a manifest change cannot merge leaving
every client reading stale keys.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(web): add the two-axis quality picker and typed settings hooks

Quality becomes one picker over two stored values.

The server holds a resolution cap and a bandwidth cap independently, which
is what the player has always sent on the wire — useTranscodeQuality.ts has
decomposed 1080p-high into {resolution, bitrate} for as long as it has
existed. Presets live in the client rather than the contract so retuning
what "High" means is a one-line edit here instead of a contract change four
codebases have to agree on, and an older server keeps working because it
only ever sees the two axes it already understands.

A combination no preset covers still gets a truthful label rather than a
picker showing the wrong entry: reachable by setting the axes separately
through the API, or from a legacy value whose bitrate is off this ladder.
Choosing an uncapped preset clears the bitrate rather than storing a
sentinel, so "no cap" stays the absence of a value at every layer.

Adds hooks over the canonical API alongside the legacy ones rather than
replacing them wholesale — a key that is not in the manifest cannot be
expressed, because SettingKey is generated from it, and the default for an
unset value comes from the generated table rather than a literal at the
call site. That last part is what stops the flip-off bug the Apple client
carries a hand-written guard for.

A test asserts every preset composes values the contract actually accepts,
so a preset naming a resolution outside the enum or a bitrate outside the
declared bounds fails here rather than 400ing when a user picks it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): resolve catalog playback preferences through the contract

catalog/detail.go held the two hardest ladders in the codebase: subtitles
resolved across four levels by hand, audio across three, each partially
overriding the last through Has* flags. Both now call the canonical
resolver, so the precedence lives in the manifest and this file cannot
disagree with the contract about which override wins. Adding a scope is a
manifest change rather than another branch here.

The subtitle track signature stays on its specialized table — it identifies
a concrete track rather than expressing a preference, so it is not a
setting.

Resolution keeps the memoization the old lookups had: the audio resolver
still reads once per profile and once per library rather than once per
file, which is what kept a many-track audiobook detail page fast. The test
that guards it now counts resolver reads instead of GetProfile calls, since
the guarantee is about scaling with file count rather than about which
method does the reading.

Four tests seeded the profile column directly. That column is a migration
source now, not a read path, so they seed the canonical value instead —
they were passing against storage nothing reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): close the unknown-key extension bag

keyUsesUserScope returned true for any key the registry did not know, so a
client could invent a production setting unilaterally and the server stored
it as an unvalidated string. That is how six ui.* settings and five orphan
keys reached production untyped, and it is the root enabler the design
names.

An unknown key is no longer a user setting, so the legacy write path
rejects it and the canonical API — which validates every value against its
own definition — is the only way to store something new.

jellycompat's DisplayPreferences blobs ride the same table under synthetic
keys and keep working: they are that subsystem's storage rather than user
settings, and they move to dedicated storage in the follow-up rather than
being dropped here.

Also repoints the DisplayPreferences seed at the canonical resolver.
Resolved at profile scope with no device on purpose — Jellyfin clients do
not carry Silo's device identity, so a device override leaking into the
seed would hand one device's settings to every Jellyfin client on the
account.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): enforce viewer quality caps through resolver constraints

constraintsFor was the unwired half of the preferences-versus-restrictions
seam: it returned nil, so a profile capped at 1080p by policy still resolved
its stored 2160p preference at face value through the effective endpoint.

The settings routes are mounted inside RequireViewerAccess, so the resolved
access scope is already on the request context. Scope.MaxPlaybackQuality
holds a literal member of the contract's quality enum ("1080p"/"2160p"),
which is exactly what the manifest binds playback.preferred_quality's
ceiling to under policy_input "max_playback_quality" — so the wiring is a
direct map with no translation table. An empty value means the policy sets
no cap, expressed by returning nil so the resolver leaves the preference
alone.

catalog.metadata_language deliberately stays unconstrained: the manifest
notes record that the allowlist draft was circular (the policy input it
would bind to is populated from the very preference it would narrow).

The handler test covers both halves of the seam: a 2160p preference under a
1080p cap resolves to the cap with constrained:true/ceiling and the authored
value reported in stored_value, the stored row itself is not rewritten, and
an uncapped viewer gets the preference unchanged with no constraint noise.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): publish user_settings change events

Add a user_settings realtime channel so clients learn when a setting
changed on another device without polling. The channel is modeled on
user_state: non-admin subscribable, per-user addressed envelopes, null
snapshot.

SettingValuesHandler gains an EventsHub and publishes
user_settings.changed after every successful PUT and DELETE on
/settings/values/{key}. The payload carries only key, scope and
profile_id — never the value. Admins receive every user's user-scoped
events, so a value in the payload would leak private settings to
admins; interested clients re-fetch over the scoped REST API instead.
The payload is always non-empty because an empty Data falls back to a
null snapshot in the hub. A nil hub (tests) skips publishing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): sweep expired mutation receipts daily

Setting-mutation idempotency receipts were written with an expires_at that
nothing enforced, so the table grew forever. Add a hidden daily system task
(05:00) that walks every login account, opens its user store, and calls
DeleteExpiredSettingMutations. A user whose store fails to open or sweep is
logged and skipped so one broken store cannot stall retention for everyone
else; the delete is idempotent, so the next run repairs anything missed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(settings): resolve metadata language canonically in access and policy

Repoint the last legacy column readers onto canonical contract resolution
(settings cutover task A4a):

- access.Resolver and policy.ViewerResolver now resolve
  catalog.metadata_language through settingsresolve (profile scope ->
  contract default) via a shared access.PreferredMetadataLanguage helper,
  instead of reading user_profiles.preferred_metadata_language. Resolution
  is deliberately unconstrained: the policy input this preference feeds is
  the one a constraint would have to reference, which is circular — see the
  key's manifest notes.
- playback start now resolves playback.audio_language canonically for the
  profile default instead of reading user_profiles.language, matching the
  catalog detail path. Series and library override handling is unchanged.
- items.go needed no change: it already consumes the resolver-produced
  scope.PreferredMetadataLanguage.

The legacy columns keep their values but are no longer read on these
paths; a profile with only a column value now resolves to the contract
default, and a stored canonical value wins. Tests pin both directions in
access, policy (including scope parity, where the column is now a decoy),
and the playback handler. Read cost is one batched store read per
resolution, same as the profile-row read it replaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(jellycompat): give DisplayPreferences its own table

The Jellyfin DisplayPreferences blobs rode the legacy user_settings
key/value table under synthetic jellycompat:* keys, which forced the
legacy settings API to carry a prefix carve-out in its otherwise-closed
unknown-key gate. They are the compat subsystem's storage, not user
settings: the contract neither validates nor resolves them.

Move them to a dedicated jellycompat_displayprefs table in both
backends, keyed by (prefs id, client) per user, with the blob stored as
opaque text served back byte-for-byte (deliberately not jsonb, which
would re-serialize it). The data-copy migrations — per-user SQLite V16
and a paired SQL + Go goose migration for Postgres — are transactional
and harmless to re-run, and both drive their key parsing and row
classification from the new internal/jellycompat/displayprefs package
so the backends cannot diverge, following the internal/settingsmigrate
precedent. A jellycompat:* row that does not parse as a DisplayPrefs
key (only ever writable through the removed carve-out) is recorded in
user_setting_migration_rejects rather than silently deleted.

With the last non-settings tenant gone, the jellycompatSettingPrefix
carve-out is deleted: the legacy settings endpoints now refuse
jellycompat:* keys like any other unknown key and never surface them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): serve admin user-settings through the canonical API

Replace the ten string-registry /admin/users/{id}/settings* and
device-settings* routes with the canonical contract surface: one list of
every explicit value the target user has stored across all scopes, and
set/delete at an explicit scope named in the query string.

The admin handlers live on SettingValuesHandler and share the session
routes' implementation rather than duplicating it — the same key/scope
parsing, identity validation, contract scope allowance, value
normalization and mutation-receipt idempotency, factored into
keyedScopeFromRequest/completeIdentity and setValueAt/deleteValueAt.
The only admin-specific parts are the target user coming from the path,
profile and device ids coming from the query (an admin holds no session
claim to the user being inspected, so its named profile is checked to
exist), and change events attributed to the target user so their
clients refresh.

The list is a new UserStore read, ListAllSettingValues, implemented in
both backends and pinned by the shared storetest conformance suite:
the admin surface wants the stored truth (which overrides exist, for a
per-row reset affordance), which no resolution-shaped read answers.

The ten removed routes are recorded in the pre-lock removals table in
docs/architecture/v1-scope.md per the v1 API rules; the web admin
device-overrides page moves onto the new surface in the Phase B
rewrite inside this same unmerged PR.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(settings): add the cross-platform conformance fixture and its Go and web runners

contracts/settings/v1/conformance.json is the spec's named drift gate: 21
hand-authored cases of {keys, stored rows, context, constraints, expected
effective value + source}, every one executable against the shipped manifest.
They pin the semantics most likely to drift across four resolver
implementations: the full resolution ladder (series > library > device >
profile > default), an absent identity dropping its scopes, foreign-identity
rows never resolving, ceiling caps that report the authored value with
constrained:true, the ordered-enum sentinels (auto below every cap, original
above), null-on-a-nullable-numeric meaning unbounded and being brought down by
a ceiling but ignored by a floor, allowlist falling back to the first allowed
member rather than the (possibly forbidden) default, and
playback.subtitle_appearance resolving device > profile only with the sparse
device object replacing, not merging. Cases may inject a constraint binding
onto a copy of a real definition so constraint kinds no shipped definition
carries stay testable.

The Go runner (internal/settingsresolve/conformance_test.go) resolves each
case through the real resolver against the embedded manifest. The web runner
(web/src/lib/settingsConformance.test.ts) runs the same cases through a new
client-side resolver, web/src/lib/settingsResolve.ts, which mirrors the
server's semantics; the TypeScript bindings now carry each definition's
ordered flag and constrained_by binding so that resolver derives constraint
behavior from the contract instead of hardcoding it. Both runners reject
unknown fixture fields — schema drift in the fixture itself is drift — and
both refuse a fixture authored against a different manifest revision.

The fixture travels with the bindings: make settings-bindings vendors the copy
the web runner reads, and make verify-settings-bindings fails CI when that
copy goes stale. The Kotlin and Swift copies land together with their runners
in the client repos, which will pick their own test-resource paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): review pass over the phase A stack

Fixes the eight adversarially-confirmed defects the review of the
unpushed phase A stack (40e0f77a..1f2c7fe4) found, each with a test
that fails without its fix.

Writers left behind by the language cutover (high). 22e9d7f1 made
access, policy and playback start resolve catalog.metadata_language and
playback.audio_language exclusively from user_setting_values, but
POST/PUT /profiles — the write path the shipped web UI uses — still
wrote only the legacy columns, so a language change after the one-time
backfill never took effect (a stale backfilled row, or the contract
default, won forever). Profile mutations now mirror their preference
fields into the canonical profile-scope rows through the same contract
validation /settings/values applies (audio, subtitle and metadata
language, subtitle mode, forced subtitles; the empty string clears the
row, matching the migration's unset spelling), publish
user_settings.changed for each row moved, and 400 on a value the
canonical endpoint would refuse. quality_preference is deliberately not
mirrored: the server never resolves the legacy column and the two-axis
picker already writes canonically.

Web admin settings 404s (high + medium). facad78d removed the ten
/admin/users/{id}/settings* and device-settings* routes but shipped no
web changes, so the user-detail settings and device-overrides tabs and
the devices-page override editor were dead. The seven admin hooks now
speak the canonical values API: one list across all scopes feeds both
tabs, mutations address an explicit scope identity, values re-type
through the generated contract (display stringifies for the
registry-era controls), device rows are enriched with device and
profile names client-side, and the removed bulk device reset becomes
per-key deletes that treat 404 as already-reset.

Silent metadata-language degrade (medium). PreferredMetadataLanguage
now logs a warning with the profile and error when contract load or
store resolution fails, so pool exhaustion is distinguishable from "no
preference"; the healthy paths stay quiet.

Displayprefs move data loss (medium). Under READ COMMITTED the blanket
pattern DELETEs in moveDisplayPrefs/unmoveDisplayPrefs could destroy a
row an old-binary instance committed between the SELECT and the DELETE
during a rolling deploy — reproduced against real Postgres. Both
directions now delete only the exact rows they read (rejects restore by
primary key), leaving a late row stranded for a re-run to pick up.

Coverage the review proved missing (medium x3): admin mutations are now
tested to attribute change events to the target user, not the acting
admin (the exact regression passed the whole suite before); the
user_settings websocket channel is subscribed through the real events
websocket, failing if the channel is dropped from either
allowedChannelsForRole or AllChannels; and the conformance fixture
gains three locked-constraint cases (replace, equal-value pass-through,
locked default) so the Go and TypeScript locked branches — previously
executable by no test on either platform — are pinned by the shared
drift gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): read and write appearance and format preferences through the settings contract

Move the four identity-sensitive preference hooks — useTheme,
useCustomTheme, useDateTimeFormat, useSearchMediaScope — off the legacy
string-only /settings endpoints and onto the canonical settings API.
Each surface now reads through one batched useEffectiveSettings call and
writes via useSetSettingValue at scope "profile", matching what the
generated manifest declares: ui.theme / ui.text_scale / ui.text_weight /
ui.high_contrast are profile-scoped with a profile_device override the
effective read already resolves (no device-override UI exists, so writes
stay profile-wide), and ui.custom_theme_vars / ui.custom_css /
ui.date_format / ui.time_format / search.media_scope are profile-wide.
Keys come from the generated SETTING_KEYS table, so a typo'd or
unmanifested key can no longer be expressed.

Because the canonical effective endpoint always answers — resolving
unset keys to the contract default with source "default" — the hooks now
use the source to distinguish "the profile chose this" from "nobody
stored anything". That preserves the admin-default theme layering and
keeps resolved-but-unchosen values out of the warm-start mirror.

ui.theme moving account→profile scope means the appearance warm-start
cache must not be shared by sibling profiles on one account, so
appearanceCacheOwner widens its token from the user id to user id plus
active profile id. Every cache read/write already resolves through that
one function, so no call site could be left behind; the API→cache
mirror, the render-time re-seed on identity change, and the debounced
write cancellation all follow automatically. The ownership tests now
cover profile switches within one account: no theme/text-scale/CSS leaks
between profiles, each profile's warm start survives the switch, and a
debounce armed by one profile never persists under its sibling.

Part of the Phase B settings-contract cutover; the legacy hooks in
queries/settings.ts keep their remaining callers until B4 deletes them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): store library, sidebar, and overlay preferences through the settings contract

Phase B2 of the settings-contract cutover: the query-layer preference
stores — sidebar pins, library page state, disabled libraries, library
order, and card overlay prefs — move off the legacy string-valued
/settings endpoints onto the canonical values API, using generated
SETTING_KEYS and each definition's declared scope (profile for pins,
visibility, order, and overlays; profile_device for page state and the
remember toggle).

Values are now written as typed JSON matching the contract schemas
(sidebar-pins.json, library-page-state.json, library-id-list.json,
card-overlays.json) instead of JSON-encoded strings, so the encoding the
migration produced keeps validating. Every parser accepts both the
canonical object value and the legacy string encoding, so nothing breaks
while caches or older rows still hold strings.

Semantics preserved deliberately:
- Sidebar pin toggles keep their optimistic update with the
  revision-guarded rollback, now layered on the effective-settings cache
  entry (effectiveSettingsQueryKey is exported for exactly this).
- The remember-library-pages toggle clears the device override to
  inherit again rather than storing the default, via
  useClearSettingValue; the canonical DELETE's 404 for "nothing stored"
  is treated as already-done, matching the legacy delete's idempotency.
- Overlay prefs keep the admin default / kill-switch layering: the
  contract default null means "no preference expressed", which is what
  lets /settings/overlay-config defaults apply, and only a stored value
  overrides them.
- Library visibility/order keep their optimistic local state with
  rollback on error; ids are normalized client-side with the same rules
  library-id-list.json enforces.

parseDisabledLibraryIDs/parseLibraryOrder collapse into one
parseLibraryIDList (they were byte-identical), and the serialize helpers
disappear with the string encoding. Legacy hooks in queries/settings.ts
stay for the remaining consumers until B4.

Part of the settings-contract cutover (see
docs/superpowers/specs/2026-07-10-cross-platform-user-settings-contract-design.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): write playback, subtitle, and library preferences at canonical scopes

The settings screens and the player panels were the last web surfaces still
speaking the legacy string API, and each carried its own idea of where a
preference lives. Playback and subtitle behavior wrote profile columns through
PUT /profiles; auto-play and next-up wrote untyped strings; subtitle appearance
went through three bespoke routes that existed only because the string API had
no way to express an object-valued setting per device. All of them now read one
batched effective resolution and write typed JSON at an explicit scope.

Where each preference lands follows the manifest rather than the endpoint that
happened to hold it:

  - Playback and subtitle defaults, and next-up mode, write at profile.
  - Subtitle appearance writes playback.subtitle_appearance at profile_device,
    replacing /settings/subtitle_appearance/effective and the PUT/DELETE pair on
    /settings/device/subtitle_appearance. One hook now owns that value for the
    settings screen, the in-player panel, and the cue renderer, which before
    each parsed the effective response separately.
  - Per-library edits write at profile_library with the library identity, one
    key at a time. The legacy endpoint replaced a composite row, so clearing one
    field meant re-sending the other three and losing any concurrent change to
    them; independent per-key writes have no such coupling, and "inherit" is a
    delete rather than a sentinel.
  - The in-player series choice splits along the line the contract draws:
    language and mode are preferences and move to profile_series, while the
    track index and signature stay on /subtitle-prefs because they identify a
    concrete track rather than expressing a preference.

Controls render from the generated SETTING_DEFINITIONS. The hand-written
registry beside it had drifted — it declared several profile-only keys as device
overrides, and disagreed with the manifest about the bounds of two sliders — so
the display helpers now derive control shape, options, bounds, and the
device-overridable key list from the contract. (Deleting settingsManifest.ts
itself is B4; nothing outside its own test imports it any more.)

Two follow-on fixes fell out of reading the contract rather than the registry.
playback.auto_skip_recap and playback.auto_play_next_preview are declared at
profile_device but only the intro override was ever consulted, so a device
override on either silently did nothing; the player resolves all three now.
And per-library "Original Language" is gone: the contract types these as BCP 47
tags, and the phase-A migration already rejects "original" at profile_library,
so offering it would have written a value the server refuses.

Risk worth naming: LibrarySettings decides "overrides" from the resolved source
rather than by comparing values, which is what keeps three distinct cases apart
— a library row holding the same value as the profile is still an override, and
a library row holding null is an explicit "no subtitles" rather than an absent
choice. A screen that compared values would collapse the first into "inherits"
and the second into "unset".

Part of #135

AI-assisted: authored with Claude Code; reviewed and verified by the committer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): refresh settings from the user_settings channel

A canonical settings write reached only the tab that made it. The server
already publishes user_settings.changed on every write and delete, but no
web client subscribed, so a preference changed on a phone or by an admin
sat stale here until a manual reload or the 5-minute staleTime expired.

Subscribe the channel and treat the frame purely as an invalidation
signal. The payload carries the key, the scope and the profile — never a
value, because admins receive other accounts' user-scoped events and a
value there would leak private settings. Marking the value queries stale
lets react-query refetch only what a mounted screen is reading, and a
burst of writes coalesces into one fetch per key rather than one per
event.

A profile-addressed change to a profile other than the signed-in one is
dropped: it cannot alter what this tab resolves. Account-scoped changes
carry no profile and always invalidate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): render settings from the generated contract

web/src/lib/settingsManifest.ts was a hand-written table of labels,
controls, defaults and bounds sitting beside the generated contract, and
it had already drifted: it declared profile-scoped keys as device
overrides, disagreed with the server on the type and range of several
keys, and enumerated a language subset narrower than the one the player
speaks. lib/settingsDisplay.ts has derived all of that from
SETTING_DEFINITIONS since the contract landed, and nothing but the
manifest's own test still imported it.

Delete the manifest and its test. The one piece it owned that the
contract cannot express is the language list — language settings are
typed as BCP 47 rather than as an enum, so there is no member list to
render — which moves to lib/languageOptions.ts and is now derived from
the shared player language list. Two shapes ship: NAMED_LANGUAGE_OPTIONS
for a control that spells its own unset entry, and LANGUAGE_OPTIONS with
the leading "no preference" row for a nullable setting.

The per-library editor's LANGUAGE_OPTIONS re-export goes with it, so
every language dropdown in settings now iterates one list in one shape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): show canonical device overrides in admin devices

The device detail panel read its override rows from
GET /admin/devices/{user}/{device}, whose `settings` array still comes
out of the legacy user_device_settings table. The settings-contract
migration folded that table into user_setting_values and nothing writes
to it any more, so an override created since the cutover — including one
the admin had just saved through this very panel — was invisible here,
while the migrated rows stayed visible. The panel's own writes go to the
canonical route, which made the list look like it silently dropped
edits.

Read the overrides from the canonical values API instead, filtered to
device scope and to this device. Both storage generations show, because
the migration moved the legacy rows into the same table. The detail
endpoint is still the source for registration metadata — device name,
owner, which profiles have used it — which is not a setting and has no
canonical equivalent.

The override count and last-updated readouts move to the canonical rows
for the same reason: override_count is computed over the legacy table and
would disagree with the rows rendered underneath it. "Reset all for
device" has no bulk canonical route, so it keeps issuing one delete per
key, now over the keys that actually exist. The reset button also takes
the profile id from the tab rather than from its first row, which a
profile registered on the device with no override yet does not have.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): delete the legacy settings hooks

hooks/queries/settings.ts spoke the string-only registry API: every
value a string, scope implied by which function you called, and an
unknown key silently accepted. Phase B moved every consumer onto the
canonical value hooks, and the last importer left was the file's own
test — so both go together, along with the client functions they were
the only callers of.

hooks/queries/libraryPlaybackPreferences.ts goes with them. It wrapped
GET/PUT/DELETE /library-playback-prefs, which LibrarySettings replaced
with profile_library-scoped canonical writes; nothing in web has called
it since. The server route stays for now — the Android and Apple clients
may still use it — but the web type and query keys have no reason to
linger.

settingsKeys keeps only `all` (the prefix the canonical invalidation
targets) and the plugin entries, which are a different system. The
list/detail/deviceDetail/effective builders described the registry's
cache layout and had no remaining callers; effectiveSettingsQueryKey in
settingValues.ts owns the canonical shape.

hooks/useSettingsForm.ts is deliberately untouched: it edits admin
server_settings through /admin/settings, which is a separate surface
from the per-user contract, and has more than twenty live consumers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): review pass over the phase B adoption

Phase B moved the web client onto the canonical settings surface. Three
scope mistakes slipped in, all of the same shape: a value written at a
scope no UI can reach, shadowing the one the user can edit.

Auto-play next. The post-roll toggle wrote profile_device while Settings →
Playback wrote profile, and the contract resolves the device row above the
profile row. Turning auto-play off in the player therefore made the
settings switch permanently inert — it saved a profile value the device row
kept shadowing and snapped straight back, with no web affordance able to
clear the device row. Both surfaces now share useAutoPlayNextSetting, which
writes the profile and clears any device row (also the only way a migrated
per-device override becomes reachable). Before Phase B both writers used
useSetDeviceSetting, so they could not disagree; this restores that
invariant at the scope the rest of the Playback screen edits.

In-player subtitle picks. handleSubtitleChanged wrote three canonical keys
at profile_series, the top of the resolution ladder, while "Auto" on the
item page still deleted only the legacy /subtitle-prefs row — so the reset
silently stopped working and the abandoned language kept resolving for
every episode of the series, forever. One of the three,
show_forced_subtitles, was worse: the player has no forced-subtitle
control, so the value it wrote back was the *resolved* one, which for a
viewer who never expressed a preference is the contract default. That
pinned the default above the profile-scope toggle on the Subtitles screen.
The written set now comes from SERIES_SUBTITLE_SETTING_KEYS — language and
mode only, both derived from the user's actual choice — and
useDeleteSubtitlePreference clears exactly that list, so the writer and the
reset cannot drift. show_forced_subtitles still rides the legacy composite
row, which is keyed to a concrete track selection and is not part of the
canonical ladder.

Admin user settings. The tab now lists every non-device canonical row,
which includes the object-valued profile settings (sidebar pins, card
overlays, disabled libraries, library order, custom theme vars). It gated
only on `definition`, and controlKindFor has no `object` branch, so those
fell through to RegistrySettingControl's select — rendering a user's pins
as a one-entry "Unset" dropdown whose only option nulls them. It now uses
the same isStructuredSetting guard the device tab got, routing them to a
raw JSON editor.

Tests: each fix has a test that fails without it, verified by reverting the
fix in place. The auto-play and subtitle tests resolve through
lib/settingsResolve rather than a canned answer, so the scope-precedence
assertions exercise the real ladder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): serve profile preference fields from canonical resolution

PUT /settings/values?scope=profile writes only user_setting_values, but
GET /profiles still served the legacy user_profiles columns. A preference
saved through the canonical API was therefore invisible in every profile
DTO reader on every platform — Apple's shipped build reads exactly those
fields — while profiles_settings_sync.go mirrored one way only, legacy
column write to canonical row.

Serve those five fields (language, preferred_metadata_language,
subtitle_language, subtitle_mode, show_forced_subtitles) by resolving
their canonical keys through the settingsresolve seam at profile scope,
falling back to the contract default rather than to the stale column.
This matches the cutover direction taken everywhere else: the legacy
columns stay written but stop being read, so "clear this preference"
cannot resurface a pre-cutover value the one-time backfill already
converted. The write paths that accept these fields and mirror them are
unchanged; this is read-side only, and the DTO's field names and types
are untouched.

Resolution is batched. A profile list serves the whole household, so
SettingResolutionQuery.ProfileID becomes ProfileIDs and the new
Resolver.ResolveProfiles ranks every profile against one candidate set —
one store read per list request instead of one per profile. Both backends
carry the widened predicate and the shared storetest conformance suite
gains a household case, so they cannot drift on it.

quality_preference stays column-backed: the legacy column is one compound
value while the contract splits it across playback.preferred_quality and
playback.max_bitrate_kbps, so there is no lossless read. The auto_skip_*
and auto_play_next_preview fields stay column-backed too — the sync path
never mirrored them, so their canonical rows can lag the columns.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): repair CI findings after the main merge

CI runs checks the local loop does not: golangci-lint (not installed
here) flagged two unchecked Close errors in the new websocket test, and
tsc -b (the tests were only vitest-run locally) rejected strict
indexed-access in four test files touched by the review passes. The
merge also brought main's onboarding tour, whose SettingControl wrote
through the legacy useSetSetting hook this branch deletes — it now
writes the canonical scoped mutation, re-typing the tour's string values
through the generated contract like the admin surface does.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(settings): satisfy the incremental lint pass

golangci-lint reports findings incrementally, so these three surfaced
only after the previous fix: errors.Is for the pgx.ErrNoRows compare
(wrapped errors), and named constants for the repeated "values"
response key and the "usersettings" prefs id goconst flagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(database): pin the read value when deleting moved displayprefs rows

Under READ COMMITTED the move's DELETE takes its own snapshot, so during
a rolling deploy an old-binary instance could update a jellycompat row
between the migration's SELECT and its delete — and the (user_id, key)
predicate would destroy the newer value after copying only the older
one. Naming the value the transaction actually read makes such a row
survive as a stranded legacy row instead, the same disposition a
late-inserted row already had.

Extends the concurrent-write migration test to commit an update to an
already-read row during the stall and assert the newer value survives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(api): make canonical mutation writes honest under failure

Three review findings on the canonical settings endpoints:

- Idempotency receipts were recorded via defer, so a failed upsert
  still left a receipt and the client's retry replayed a success for a
  write that never happened. The receipt is now written only after the
  upsert lands, and it stores the actual response — revision and
  updated_at included — so a replay is byte-identical instead of a
  reconstruction of the input with revision 0.

- The mutation envelope accepted trailing JSON after the first
  document, leaving the interpreted mutation parser-dependent. The
  decoder now requires EOF after the envelope.

- Resolving a device-aware key without X-Silo-Device-Id silently
  skipped every stored device override and passed the profile fallback
  off as the effective value. The effective endpoint now fails closed
  with 400, matching the write path's existing requirement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): stop the contract rejecting values shipped clients store

Four bounds in the contract were narrower than what a shipped client
already produces, so real stored preferences would fail validation or
be quarantined at migration:

- The BCP 47 grammar rejected extlang tags (zh-cmn) and private-use-only
  tags (x-private) the legacy length-only validator accepted, turning an
  existing 204 into a 400. The pattern now covers both, and
  NormalizeLanguageTag cases a script correctly after an extlang and
  leaves private-use content lowercase.

- subtitle_appearance.fontFamily allowlisted ASCII, contradicting its
  own description: Apple clients store CTFontManager family names
  verbatim and those are routinely CJK. The pattern now excludes unsafe
  characters instead of allowlisting ASCII.

- theme-var-overrides capped CSS values at 128 characters, which real
  multi-stop gradients exceed; the web importer stores them unchecked.
  Raised to 1024.

Plus one tightening the review asked for: card-overlays.order now
declares uniqueItems, matching library-id-list, so an overlay cannot be
rendered twice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(userstore): reject non-canonical identities and bound resolution batches

Three review findings on the canonical settings storage layer:

- SettingIdentity.Validate trimmed ids only to check emptiness, so a
  padded id like " p1 " validated, persisted verbatim, and was then
  invisible to resolution queries, which bind trimmed forms — a
  silently orphaned row. Validation now rejects any id that is not in
  canonical trimmed form, pinned in the shared conformance suite so
  both backends hold the line.

- The effective-values endpoint accepted unbounded library_ids and
  series_ids lists; the SQLite backend expands each id into a bound
  parameter, so a crafted batch could exhaust the host-parameter budget
  and fail the whole resolution. The request boundary now caps the
  combined content ids at 200.

- pickForScope's doc comment promised ties broken "by the most
  specific id in the request order" while the implementation sorts by
  ascending library then series id; the comment now describes the
  actual (deliberately deterministic-only) behavior.

Plus: the pgstore conformance cleanups now assert the ON DELETE CASCADE
they rely on instead of discarding the delete error, so a dropped FK
can no longer leak seeded rows into the shared test database silently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): close the discovery gaps around the canonical API

Three review findings:

- Canonical profile_device writes never touched the device registry, so
  a device that only ever wrote through /settings/values was invisible
  to ListDevices and the admin device surfaces — undiscoverable and
  unforgettable. Device-scope writes now refresh the registry from the
  request's device headers, throttled the same way the legacy route is.

- The contract spec tells clients to probe GET /settings/manifest (and
  /settings/capability), and to read a 404 as "pre-contract server";
  the router only exposed /settings/contract*. The documented paths now
  alias the same handlers.

- The plugin proxy's X-Silo-Theme header came from the legacy
  account-level user_settings.ui_theme row, so a profile's theme change
  through the canonical API never reached plugins and profiles sharing
  an account were indistinguishable. The lookup now resolves the
  canonical profile-scoped ui.theme row (falling back to the legacy row
  for stores the backfill has not covered) using the request's active
  profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): emit revision metadata in generated bindings and verify the TS one

Two review findings on the generator surface:

- The bindings dropped every introduced_in tag, so a client generated
  from revision N could not filter its pinned contract down to an older
  server's advertised revision — the promised negotiation had no data.
  The TypeScript definitions now carry introducedIn per definition,
  per scope, per enum member, and the full history of any widened
  numeric bound. (Go/Kotlin/Swift emit keys, not definition tables, so
  they only need the Revision constant they already have.)

- make verify-settings-bindings compared only the generated Go file and
  the conformance fixture, so a manifest change could merge with a
  stale web/src/lib/settingsContract.ts. The target now regenerates and
  diffs the TypeScript binding too, through the same prettier config
  the bindings target applies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: drop the accidentally committed settingsgen binary

24ee9952 checked in a 5.5 MB compiled settingsgen alongside its source.
The binary is a local build artifact — cmd/settingsgen is the source of
truth and make settings-bindings runs it with go run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): close the migration planner's data-loss and crash findings

Four review findings on the one-time legacy-to-canonical migration:

- A device holding both player.next_up_prompt_seconds and its
  playback.* rename canonicalized to one identity, and both backends
  insert bare — a unique violation that failed NewUserDB (SQLite) or
  aborted the goose migration (Postgres). Plan now ends with a
  deterministic dedup keyed on the canonical identity; a canonically
  keyed row beats a renamed alias, since the runtime writes the
  canonical spelling first and only best-effort-deletes the alias.

- The four auto-skip profile columns (auto_skip_intro/credits/recap,
  auto_play_next_preview) were never read, so an explicit true silently
  became the contract default false. They now migrate — explicit true
  only, so an untouched false column does not become a choice.

- Profiles with language 'en' emitted no playback.audio_language row
  because the column default was suppressed, but that default WAS the
  effective behavior: the old playback path preferred English, while
  the canonical null default skips language matching entirely. English
  now migrates as an explicit row. The other suppressed defaults stay
  suppressed — their empty-string defaults already meant unset.

- Stored v1 card_overlays documents were quarantined because the
  planner validated them against the v2-only schema; the web parser has
  upgraded v1 at read time all along. The planner now applies the same
  v1-to-v2 upgrade before validation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): delete canonical library-scoped values with the library

The canonical settings schema deliberately has no FK on library_id or
series_id, and the migration comment promised the owning delete paths
would clean these rows up — but nothing called
DeleteSettingValuesForLibrary/-Series outside stores and tests, so a
deleted library left orphaned profile_library preferences in every
user's store forever.

Adds userstore.SettingValuesCleaner, a per-user best-effort sweep in the
mutation-sweeper's mold, and wires it into the library delete job. The
series-side cleanup is exposed on the same cleaner for the scanner's
orphan pruning to adopt; series have no single delete executor today.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): keep off-step playback speeds working until cutover

The legacy device endpoint gained step enforcement mid-branch, turning
an existing in-range PUT of 0.26 from 204 into 400 — a behavior change
on a live /api/v1 endpoint before the coordinated break, which the v1
rules forbid. The legacy validator is back to range-only; the typed
mutation endpoint keeps enforcing the manifest's step.

The migration planner now snaps stored off-step numbers onto their
definition's step grid instead of quarantining them: a stored 0.26 is a
real preference, and every client's stepper was going to snap it on the
next write anyway.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): serve canonical values to the readers the cutover stranded

Four P1 review findings where the web writes canonical rows the server
never reads — and the legacy keys those readers use are now unwritable,
so the values are frozen and user edits silently do nothing:

- access.DisabledLibraryIDs and the policy viewer resolver read the
  legacy account key while the library screen writes profile-scoped
  ui.disabled_library_ids. Both now resolve the canonical profile row,
  falling back to the legacy key only when no canonical row exists.

- The sections fetcher and handler read the legacy next_up_mode account
  key while the playback screen writes ui.next_up_mode. Same ladder,
  behind one shared sections.NextUpMode helper.

- Profile creation committed the profile and then synced settings
  non-atomically, so a mid-sync failure left a profile the retry could
  not recreate (name conflict) with preferences that read as contract
  defaults forever. The create path now compensates by deleting the
  profile it created.

- The mounted legacy PUT /subtitle-prefs/{series_id} wrote only
  user_subtitle_preferences, but item detail resolves those three keys
  canonically, so a post-upgrade client's "subtitles off" returned 204
  and was ignored. The legacy handler now dual-writes the canonical
  profile_series rows, and its delete clears them.

Plus the migration's disposition for stranded Apple device-scope audio
language rows: nothing read them before the contract, so promoting them
to real overrides would change track selection at upgrade. They are
recorded in the rejects table instead of copied.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): make canonical settings writes take effect

Three P1 review findings on the web half of the cutover:

- Every profile-default editor reads the resolved value but writes the
  profile row, so a device override — left by the migration converting
  legacy user_device_settings, or written by another client — kept
  shadowing the save and snapped the control back with no affordance to
  remove it. useAutoPlayNextSetting already solved this for one key;
  that logic is now a shared useProfileDefaultWriter used by the
  playback screen, the quality picker, subtitle behavior, and the four
  appearance setters. It only clears when the key is device-scopable
  and the resolved value actually came from a device row.

- The appearance cache only ever grew: when the effective response
  resolved a key to "default" — because another client deleted it —
  the namespaced entry and local state survived and kept winning the
  fallback, so a removal never reached this browser. The mirror now
  runs both ways, clearing only on an explicit default answer (silence
  is not a deletion) and only within the current identity's namespace.
  Custom theme vars and CSS do the same, except while a local draft is
  unsaved.

- The quality picker wrote the canonical two-axis keys while playback
  still derived its cap from currentProfile.quality_preference, a
  legacy compound column the canonical write deliberately does not
  mirror — so choosing a quality changed nothing about what played. The
  watch route and both item-detail pages now read
  playback.preferred_quality, falling back to the profile column until
  the settings read resolves so playback never blocks on it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): repair the type error and the bindings gate's job placement

Two breaks from the previous commits:

- useTheme referenced storage.StorageKey, but storage is a value, not a
  namespace — the Web job's tsc caught what the local incremental
  typecheck had already cached past. Imported the type properly.

- verify-settings-bindings gained a prettier step, and the Go job that
  runs it has no pnpm, so the check failed on its own tooling rather
  than on a stale binding. Split the web half into
  verify-settings-bindings-web and moved it to the Web job, which has
  pnpm; verify-settings-bindings-all runs both locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): close the second-round review findings

Three from the review of the pushed work:

- The live profile sync omitted auto_skip_intro/credits/recap and
  auto_play_next_preview, which my own change made load-bearing: the
  player now resolves those keys canonically, so a legacy PUT /profiles
  moved the columns, returned 200, and changed nothing about playback.
  All four now mirror on write. The DTO read block keeps its shape —
  clients pin it — and its columns are what the sync keeps current.

- The effective endpoint dropped unknown keys silently, letting a
  client fill the gap with its own vendored default and present a value
  this server would refuse to store. Unknown keys now 404 by name.

- Two sidebar-pin toggles in flight at once could commit in either
  order, and the server upsert is last-write-wins, so the first request
  landing second restored the pre-toggle document. The writes are now
  chained, and each link reads the document when it runs, so a queued
  toggle sends the newest state rather than the one it was queued with.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(database): make the settings-contract deploy reversible

Rolling back this release meant restoring a backup, for a reason that
was not obvious: the DisplayPreferences move deletes the jellycompat
rows from user_settings once it has copied them, and the previous
binary reads exactly those rows. An older server therefore starts
cleanly and silently serves defaults, so every Jellyfin client's saved
view preferences look reset.

The down functions were already written and correct — nothing could
invoke them. The backfill and the DisplayPreferences move are Go
migrations registered in-process, so the standalone goose CLI in the
Makefile cannot see them, and the server exposed only --migrate-only
and --migrate-status.

Adds MigrateDownTo, the --migrate-down-to flag, and a make target, plus
a rehearsal test that seeds a legacy row the way the old binary wrote
it, applies the move, rolls back, and asserts the row returns
byte-for-byte.

Documents the ordering in the spec's cutover section, including the two
caveats an operator needs beforehand: take a backup, and the per-user
SQLite backend cannot be rolled back at all — its migrations have no
down path and an older binary refuses to open a newer database, so
those installs restore from backup rather than degrade.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): skip legacy rows whose profile was deleted

The dev-server migration aborted on real data:

  writing playback.subtitle_appearance at profile_device for user 1:
  violates foreign key constraint user_setting_values_profile_fkey

user_device_settings carries an ON DELETE CASCADE on (user_id,
profile_id) today, but rows written before that constraint outlived the
profiles they belonged to — that install had 46 such rows across 14
deleted profiles. The planner copied them faithfully and the canonical
table, which declares the same foreign key, refused them; because the
backfill runs in one transaction, the whole migration failed and the
server could not start.

An override belonging to a profile nobody can select is not a preference
anyone can be shown or reset, so Plan now drops those rows rather than
repairing them, recording each in user_setting_migration_rejects so an
operator can see what was left behind. Account-scope rows carry no
profile and pass through untouched.

Verified by replaying that install's 514 device rows through the
planner: 9 rows would have hit the constraint before, 0 after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): address canonical cutover review findings

* fix(settings): address latest review findings

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:52:41 -04:00

3346 lines
125 KiB
Go

package main
import (
"context"
"crypto/rand"
"encoding/base64"
"encoding/json"
"flag"
"fmt"
"io"
"io/fs"
"log"
"log/slog"
"net/http"
"os"
"os/signal"
"path/filepath"
"runtime/debug"
"sort"
"strconv"
"strings"
"sync/atomic"
"syscall"
"time"
"github.com/go-chi/chi/v5"
chimiddleware "github.com/go-chi/chi/v5/middleware"
"github.com/google/uuid"
"github.com/hashicorp/go-hclog"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/prometheus/client_golang/prometheus/promhttp"
pluginv1 "github.com/Silo-Server/silo-plugin-sdk/pkg/pluginproto/silo/plugin/v1"
sdkcapability "github.com/Silo-Server/silo-plugin-sdk/pkg/pluginsdk/capability"
"github.com/Silo-Server/silo-server/internal/access"
"github.com/Silo-Server/silo-server/internal/activitylog"
"github.com/Silo-Server/silo-server/internal/adminjob"
"github.com/Silo-Server/silo-server/internal/api"
"github.com/Silo-Server/silo-server/internal/api/handlers"
"github.com/Silo-Server/silo-server/internal/audiobooks"
"github.com/Silo-Server/silo-server/internal/audiobooks/podcastfeed"
"github.com/Silo-Server/silo-server/internal/auth"
"github.com/Silo-Server/silo-server/internal/autoscan"
"github.com/Silo-Server/silo-server/internal/branding"
"github.com/Silo-Server/silo-server/internal/cache"
"github.com/Silo-Server/silo-server/internal/catalog"
"github.com/Silo-Server/silo-server/internal/catalogseed"
"github.com/Silo-Server/silo-server/internal/chapterthumbs"
"github.com/Silo-Server/silo-server/internal/clientip"
"github.com/Silo-Server/silo-server/internal/config"
"github.com/Silo-Server/silo-server/internal/database"
"github.com/Silo-Server/silo-server/internal/diagnostics"
"github.com/Silo-Server/silo-server/internal/downloads"
"github.com/Silo-Server/silo-server/internal/ebooks"
evt "github.com/Silo-Server/silo-server/internal/events"
"github.com/Silo-Server/silo-server/internal/historyimport"
"github.com/Silo-Server/silo-server/internal/imagecache"
"github.com/Silo-Server/silo-server/internal/intromarkers"
"github.com/Silo-Server/silo-server/internal/jellycompat"
"github.com/Silo-Server/silo-server/internal/libraryingest"
"github.com/Silo-Server/silo-server/internal/literaryworks"
"github.com/Silo-Server/silo-server/internal/logfilter"
"github.com/Silo-Server/silo-server/internal/logredact"
"github.com/Silo-Server/silo-server/internal/logstream"
"github.com/Silo-Server/silo-server/internal/mail"
"github.com/Silo-Server/silo-server/internal/manga"
"github.com/Silo-Server/silo-server/internal/markers"
"github.com/Silo-Server/silo-server/internal/mdblist"
"github.com/Silo-Server/silo-server/internal/metadata"
// Built-in metadata providers self-register into the metadata package's
// builtin registry on import; buildProviders resolves their seeded chain
// entries in-process (no gRPC).
_ "github.com/Silo-Server/silo-server/internal/metadata/nfo"
"github.com/Silo-Server/silo-server/internal/models"
"github.com/Silo-Server/silo-server/internal/nodeconfig"
"github.com/Silo-Server/silo-server/internal/nodepool"
"github.com/Silo-Server/silo-server/internal/noderecipe"
"github.com/Silo-Server/silo-server/internal/nodesessions"
"github.com/Silo-Server/silo-server/internal/notifications"
"github.com/Silo-Server/silo-server/internal/opslog"
"github.com/Silo-Server/silo-server/internal/partman"
"github.com/Silo-Server/silo-server/internal/playback"
"github.com/Silo-Server/silo-server/internal/pluginhost"
"github.com/Silo-Server/silo-server/internal/plugins"
"github.com/Silo-Server/silo-server/internal/policy"
"github.com/Silo-Server/silo-server/internal/proxy"
"github.com/Silo-Server/silo-server/internal/ratelimit"
"github.com/Silo-Server/silo-server/internal/recommendations"
mediarequests "github.com/Silo-Server/silo-server/internal/requests"
"github.com/Silo-Server/silo-server/internal/s3client"
"github.com/Silo-Server/silo-server/internal/scanner"
"github.com/Silo-Server/silo-server/internal/scanqueue"
"github.com/Silo-Server/silo-server/internal/secret"
"github.com/Silo-Server/silo-server/internal/sections"
"github.com/Silo-Server/silo-server/internal/server"
"github.com/Silo-Server/silo-server/internal/settingscontract"
"github.com/Silo-Server/silo-server/internal/subtitles"
"github.com/Silo-Server/silo-server/internal/taskmanager"
taskrepository "github.com/Silo-Server/silo-server/internal/taskmanager/repository"
"github.com/Silo-Server/silo-server/internal/taskmanager/tasks"
"github.com/Silo-Server/silo-server/internal/taskmanager/triggers"
"github.com/Silo-Server/silo-server/internal/telemetry"
"github.com/Silo-Server/silo-server/internal/transcodenode"
"github.com/Silo-Server/silo-server/internal/usercollections"
"github.com/Silo-Server/silo-server/internal/userdb"
"github.com/Silo-Server/silo-server/internal/userstore"
"github.com/Silo-Server/silo-server/internal/userstore/pgstore"
"github.com/Silo-Server/silo-server/internal/watchlist"
"github.com/Silo-Server/silo-server/internal/watchstate"
"github.com/Silo-Server/silo-server/internal/watchsync"
watchmdblist "github.com/Silo-Server/silo-server/internal/watchsync/providers/mdblist"
"github.com/Silo-Server/silo-server/internal/watchsync/providers/simkl"
"github.com/Silo-Server/silo-server/internal/watchsync/providers/trakt"
"github.com/Silo-Server/silo-server/internal/worker"
"github.com/Silo-Server/silo-server/migrations"
siloweb "github.com/Silo-Server/silo-server/web"
)
// resolveNodeIdentity returns a stable node identifier used by the
// heartbeat writer, reconciler, and shutdown cleanup. Resolution order:
// SILO_NODE_NAME > NODE_NAME > os.Hostname().
func resolveNodeIdentity() string {
if v := os.Getenv("SILO_NODE_NAME"); v != "" {
return v
}
if v := os.Getenv("NODE_NAME"); v != "" {
return v
}
h, _ := os.Hostname()
return h
}
func resolvePluginCacheDir() string {
if v := strings.TrimSpace(os.Getenv("SILO_PLUGIN_CACHE_DIR")); v != "" {
return v
}
return filepath.Join(os.TempDir(), "silo-plugins")
}
func buildBaseHandler(format string, level slog.Leveler, otelHandler slog.Handler) slog.Handler {
opts := &slog.HandlerOptions{Level: level}
var console slog.Handler
if strings.EqualFold(format, "json") {
console = slog.NewJSONHandler(os.Stderr, opts)
} else {
console = slog.NewTextHandler(os.Stderr, opts)
}
if otelHandler == nil {
// Redact secrets before they reach stderr (the opslog DB path redacts
// separately when flattening rows).
return logredact.New(console)
}
// Fan out to the console and the OTel bridge. The OTel branch is level-gated
// by the shared level var so console and OTLP share one verbosity knob (see
// telemetry.LevelGated) — otherwise slog.MultiHandler.Enabled would OR the
// branches and export Debug records while stderr stays silent. The whole
// fan-out is wrapped in secret redaction so console and OTLP both emit
// masked output (the opslog DB path redacts separately).
return logredact.New(telemetry.FanOut(console, telemetry.LevelGated(otelHandler, level)))
}
func parseLogLevel(level string) slog.Level {
switch strings.ToLower(level) {
case "debug":
return slog.LevelDebug
case "warn", "warning":
return slog.LevelWarn
case "error":
return slog.LevelError
default:
return slog.LevelInfo
}
}
func mustGetSetting(store interface {
Get(context.Context, string) (string, error)
}, ctx context.Context, key, fallback string) string {
value, err := store.Get(ctx, key)
if err != nil || strings.TrimSpace(value) == "" {
return fallback
}
return value
}
func configureOperationalLogging(
ctx context.Context,
pool *pgxpool.Pool,
settingsRepo catalog.SettingsStore,
redisCfg config.RedisConfig,
logStreamHub *logstream.Hub,
filteredHandler slog.Handler,
nodeID string,
) (opslog.Writer, *opslog.Repo, *partman.Manager) {
if err := opslog.SeedDefaults(ctx, settingsRepo); err != nil {
log.Fatalf("seed opslog defaults: %v", err)
}
if err := diagnostics.SeedDefaults(ctx, settingsRepo); err != nil {
log.Fatalf("seed diagnostics defaults: %v", err)
}
opsPM := partman.NewManager(pool, "operational_logs", partman.Daily, 3)
if err := opsPM.EnsureFuturePartitions(ctx); err != nil {
// Non-fatal: a partition hiccup must not crash-loop the server (see the
// operational_logs partition incident). Writes fall back to the default
// partition and the periodic cleanup retries EnsureFuturePartitions.
slog.WarnContext(ctx, "ensure operational log partitions; continuing in degraded mode", "component", "app", "error", err)
}
var operationalWriter opslog.Writer
operationalConsumer := opslog.NewConsumer(pool, nil, logStreamHub)
if redisCfg.URL != "" {
redisClient, redisErr := cache.NewRedisClient(redisCfg)
if redisErr == nil && redisClient != nil {
operationalWriter = opslog.NewRedisWriter(redisClient)
operationalConsumer = opslog.NewConsumer(pool, redisClient, logStreamHub)
go operationalConsumer.RunRedis(ctx)
}
}
if operationalWriter == nil {
memWriter := opslog.NewMemoryWriter(10000)
operationalWriter = memWriter
go operationalConsumer.RunMemory(ctx, memWriter.Chan())
}
opsCaptureLevel := slog.LevelInfo
switch strings.ToLower(strings.TrimSpace(mustGetSetting(settingsRepo, ctx, "opslog.capture_level", "info"))) {
case "debug":
opsCaptureLevel = slog.LevelDebug
case "warn", "warning":
opsCaptureLevel = slog.LevelWarn
case "error":
opsCaptureLevel = slog.LevelError
}
slog.SetDefault(slog.New(opslog.NewHandler(filteredHandler, operationalWriter, opsCaptureLevel, nodeID)))
return operationalWriter, opslog.NewRepo(pool), opsPM
}
func maybeApplyPostgresTuning(ctx context.Context, pool *pgxpool.Pool, appMaxConnections int, mode string) {
switch strings.ToLower(strings.TrimSpace(mode)) {
case "", "integrated", "api":
default:
return
}
opts, err := database.LoadPostgresTuneOptionsFromEnv(appMaxConnections)
if err != nil {
slog.WarnContext(ctx, "postgres auto-tuning disabled", "component", "app", "error", err)
return
}
if !opts.Enabled {
return
}
tuneCtx, cancel := context.WithTimeout(ctx, 30*time.Second)
defer cancel()
result, err := database.ApplyPostgresTuning(tuneCtx, pool, opts)
for _, failure := range result.Failures {
slog.WarnContext(ctx, "postgres auto-tuning setting failed", "component", "app",
"name", failure.Name,
"value", failure.Value,
"error", failure.Err,
)
}
if err != nil {
slog.WarnContext(ctx, "postgres auto-tuning failed", "component", "app",
"error", err,
"applied", result.Applied,
"failures", len(result.Failures),
)
return
}
slog.InfoContext(ctx, "postgres auto-tuning applied", "component", "app",
"profile", opts.Profile,
"postgres_major", result.PostgresMajorVersion,
"settings", result.Applied,
"resets", len(result.Reset),
"failures", len(result.Failures),
"memory_budget_bytes", opts.MemoryBudgetBytes,
"detected_memory_bytes", opts.DetectedMemoryBytes,
"memory_source", opts.MemorySource,
"memory_budget_percent", opts.MemoryBudgetPercent,
"cpus", opts.CPUs,
"connections", opts.Connections,
"storage", opts.Storage,
"db_size", result.DBSize,
"database_size_bytes", result.DatabaseSizeBytes,
)
if len(result.RestartRequired) > 0 {
slog.WarnContext(ctx, "postgres restart required to finish applying auto-tuned settings", "component", "app",
"settings", strings.Join(result.RestartRequired, ","),
)
}
if len(result.Reset) > 0 {
slog.InfoContext(ctx, "postgres auto-tuning reset stale settings", "component", "app",
"settings", strings.Join(result.Reset, ","),
)
}
}
// runCredentialBackfills sweeps any plaintext server-owned credential to
// ciphertext on the primary (migration-running) node. All passes are
// best-effort: a failed row leaves the prior plaintext (no new exposure) and
// still reads via the read-path pass-through, so a backfill error must never
// block boot. The sensitive-settings pass runs first so the arr
// resolve-then-encrypt pass sees consistent referenced settings.
// librarySettingsCleaner wires the per-user canonical settings cleanup the
// library delete job runs, or nil when the user store is unavailable — the
// executor treats a nil cleaner as "skip".
func librarySettingsCleaner(pool *pgxpool.Pool, stores userstore.UserStoreProvider) adminjob.LibrarySettingsCleaner {
if pool == nil || stores == nil {
return nil
}
return userstore.NewSettingValuesCleaner(auth.NewUserRepository(pool), stores)
}
func runCredentialBackfills(ctx context.Context, pool *pgxpool.Pool, cipher *secret.Cipher, settings *catalog.EncryptedSettingsRepo) {
settingsN, err := settings.BackfillSensitiveSettings(ctx)
if err != nil {
slog.ErrorContext(ctx, "secret backfill: sensitive settings", "component", "app", "error", err)
}
columnsN, err := secret.BackfillColumns(ctx, pool, cipher, secret.ColumnBackfillTargets())
if err != nil {
slog.ErrorContext(ctx, "secret backfill: credential columns", "component", "app", "error", err)
}
historyServersN, err := historyimport.NewRepository(pool, cipher).BackfillSessionServerSecrets(ctx)
if err != nil {
slog.ErrorContext(ctx, "secret backfill: history import session server credentials", "component", "app", "error", err)
}
// The arr resolver is the encrypting settings decorator: it decrypts a
// sensitive target (e.g. requests.radarr.api_key) or passes through a
// plaintext custom key, exactly replicating the deleted resolveAPIKey.
arrN, err := secret.BackfillReferencedColumns(ctx, pool, cipher, settings.Get, secret.ArrKeyBackfillTargets())
if err != nil {
slog.ErrorContext(ctx, "secret backfill: arr api keys", "component", "app", "error", err)
}
pluginConfigsN, err := plugins.NewRuntimeConfigStore(pool, cipher).BackfillEncryptedConfigs(ctx)
if err != nil {
slog.ErrorContext(ctx, "secret backfill: plugin runtime configs", "component", "app", "error", err)
}
if total := settingsN + columnsN + historyServersN + arrN + pluginConfigsN; total > 0 {
slog.InfoContext(ctx, "secret backfill: encrypted plaintext credentials at rest", "component", "app",
"settings", settingsN, "columns", columnsN, "history_session_servers", historyServersN,
"arr_keys", arrN, "plugin_configs", pluginConfigsN, "total", total)
}
}
func runCompatWebCommand(ctx context.Context, args []string) error {
if len(args) == 0 {
return fmt.Errorf("usage: silo compat-web {status|install|update|remove}")
}
command := args[0]
flags := flag.NewFlagSet("compat-web "+command, flag.ContinueOnError)
flags.SetOutput(io.Discard)
root := flags.String("dir", config.DefaultJellyfinWebInstallDir, "Jellyfin Web component install root")
version := flags.String("version", config.DefaultJellyfinWebVersion, "Jellyfin Web version without leading v")
source := flags.String("source", jellycompat.DefaultWebSourceURL, "upstream jellyfin-web git repository")
if err := flags.Parse(args[1:]); err != nil {
return err
}
switch command {
case "status":
status := jellycompat.WebComponentStatusForConfig(&config.Config{
JellyfinCompat: config.JellyfinCompatConfig{
Enabled: false,
WebVersion: *version,
WebInstallDir: *root,
WebDir: filepath.Join(*root, "current"),
},
}, map[string]string{
"jellyfin_compat.web_source_url": *source,
})
return json.NewEncoder(os.Stdout).Encode(status)
case "install", "update":
status, err := jellycompat.InstallWebComponent(ctx, jellycompat.WebComponentInstallOptions{
InstallRoot: *root,
SourceURL: *source,
Version: *version,
})
_ = json.NewEncoder(os.Stdout).Encode(status)
return err
case "remove":
return jellycompat.RemoveWebComponent(*root)
default:
return fmt.Errorf("unknown compat-web command %q", command)
}
}
func main() {
if len(os.Args) > 1 && os.Args[1] == "compat-web" {
if err := runCompatWebCommand(context.Background(), os.Args[2:]); err != nil {
log.Fatalf("compat-web: %v", err)
}
return
}
envFile := flag.String("env", ".env", "path to .env bootstrap file")
migrateOnly := flag.Bool("migrate-only", false, "apply database migrations and exit")
migrateStatus := flag.Bool("migrate-status", false, "show database migration status and exit")
migrateDownTo := flag.Int64("migrate-down-to", -1,
"roll back every migration newer than this version and exit (the version to KEEP)")
flag.Parse()
ctx := context.Background()
// Step 0: Validate the embedded settings contract before anything can
// depend on it. A malformed or self-inconsistent manifest is a build defect,
// not a runtime condition, so failing here — loudly, before the first
// request — is the whole point: the alternative is shipping an image whose
// contract disagrees with the clients that vendored it.
contract, err := settingscontract.Load()
if err != nil {
log.Fatalf("settings contract: %v", err)
}
contractETag, err := settingscontract.ETag()
if err != nil {
log.Fatalf("settings contract: %v", err)
}
slog.Info("settings contract loaded",
"revision", contract.Revision,
"definitions", len(contract.Definitions),
"etag", contractETag)
// Step 1: Bootstrap from .env
bc, err := config.LoadBootstrap(*envFile)
if err != nil {
log.Fatalf("bootstrap: %v", err)
}
// Construct the at-rest credential cipher from SECRET_KEY immediately after
// bootstrap, before any settings repo is built. It is threaded explicitly as
// a dependency into every repo that stores a server-owned secret — never a
// package-level global.
dataCipher, err := secret.New(bc.SecretKey)
if err != nil {
log.Fatalf("secret cipher: %v", err)
}
// Step 2: Connect to PostgreSQL (bootstrap pool with default max connections)
bootstrapDBCfg := config.DatabaseConfig{URL: bc.DatabaseURL, MaxConnections: 20}
pool, err := database.NewPool(ctx, bootstrapDBCfg)
if err != nil {
log.Fatalf("database pool: %v", err)
}
defer pool.Close()
slog.Info("connected to PostgreSQL")
if *migrateStatus {
migCtx, migCancel := database.MigrationContext(ctx)
statuses, statusErr := database.MigrationStatuses(migCtx, pool, migrations.FS, "sql")
migCancel()
if statusErr != nil {
log.Fatalf("failed to read migration status: %v", statusErr)
}
fmt.Printf("%-8s %8s %-25s %s\n", "STATE", "VERSION", "APPLIED_AT", "MIGRATION")
for _, status := range statuses {
appliedAt := "-"
if !status.AppliedAt.IsZero() {
appliedAt = status.AppliedAt.UTC().Format(time.RFC3339)
}
source := status.Source
if source != "" {
source = filepath.Base(source)
} else {
source = "-"
}
fmt.Printf("%-8s %8d %-25s %s\n", status.State, status.Version, appliedAt, source)
}
return
}
if *migrateDownTo >= 0 {
// Deliberately its own flag rather than a mode of --migrate-only: this
// discards data, and several of the migrations it reverses are Go ones
// the goose CLI cannot reach, so it is the only way to undo them
// short of restoring a backup.
migCtx, migCancel := database.MigrationContext(ctx)
migErr := database.MigrateDownTo(migCtx, pool, migrations.FS, "sql", *migrateDownTo)
migCancel()
if migErr != nil {
log.Fatalf("failed to roll back migrations: %v", migErr)
}
slog.Info("database migrations rolled back", "kept_through_version", *migrateDownTo)
return
}
if *migrateOnly {
migCtx, migCancel := database.MigrationContext(ctx)
migErr := database.RunMigrations(migCtx, pool, migrations.FS, "sql")
migCancel()
if migErr != nil {
log.Fatalf("failed to run migrations: %v", migErr)
}
slog.Info("database migrations applied")
return
}
// Run migrations only for integrated/api modes. Proxy and transcode nodes
// should never alter the schema — they may scale independently and would
// race or apply migrations before the primary node is deliberately upgraded.
// The same gate decides whether this node runs the credential-encryption
// backfills: only the primary (migration-running) node sweeps plaintext to
// ciphertext; secondary nodes read whatever the primary encrypted.
isPrimaryNode := bc.Mode == "integrated" || bc.Mode == "api" || bc.Mode == ""
if isPrimaryNode {
migCtx, migCancel := database.MigrationContext(ctx)
if migErr := database.RunMigrations(migCtx, pool, migrations.FS, "sql"); migErr != nil {
migCancel()
log.Fatalf("failed to run migrations: %v", migErr)
}
migCancel()
slog.Info("database migrations applied")
}
// Step 3: Load settings from DB. settingsRepo is the encrypting decorator so
// every consumer (config.LoadFromDB, admin, ABS, watchers) transparently sees
// plaintext while sensitive keys rest as ciphertext. The settings backfill
// (run after migrations, before this GetAll) is wired further below.
settingsRepo := catalog.NewEncryptedSettingsRepo(catalog.NewServerSettingsRepo(pool), dataCipher)
if isPrimaryNode {
runCredentialBackfills(ctx, pool, dataCipher, settingsRepo)
}
settings, err := settingsRepo.GetAll(ctx)
if err != nil {
log.Fatalf("loading settings: %v", err)
}
// Step 4: YAML import (one-time)
yamlPath := "silo.yaml"
if _, yamlErr := os.Stat(yamlPath); yamlErr == nil {
if settings["_yaml_imported"] == "" {
yamlSettings, importErr := config.YAMLToSettingsMap(yamlPath)
if importErr != nil {
log.Printf("WARN: could not import YAML config: %v", importErr)
} else {
for k, v := range yamlSettings {
if err := settingsRepo.Set(ctx, k, v); err != nil {
log.Printf("WARN: failed to import setting %s: %v", k, err)
}
}
if err := settingsRepo.Set(ctx, "_yaml_imported", "true"); err != nil {
slog.Warn("failed to set yaml import flag", "error", err)
}
log.Println("Imported config from silo.yaml — this file is no longer used")
settings, _ = settingsRepo.GetAll(ctx)
}
}
}
// Step 5: Auto-generate secrets
if settings["auth.jwt_secret"] == "" {
secret := make([]byte, 32)
if _, err := rand.Read(secret); err != nil {
log.Fatalf("generating jwt secret: %v", err)
}
encoded := base64.StdEncoding.EncodeToString(secret)
if err := settingsRepo.Set(ctx, "auth.jwt_secret", encoded); err != nil {
slog.Warn("failed to persist generated JWT secret", "error", err)
}
settings["auth.jwt_secret"] = encoded
}
if settings["jellyfin_compat.server_id"] == "" {
serverID := uuid.NewSHA1(uuid.NameSpaceURL, []byte("https://silo.local/jellycompat")).String()
if err := settingsRepo.Set(ctx, "jellyfin_compat.server_id", serverID); err != nil {
slog.Warn("failed to persist generated server ID", "error", err)
}
settings["jellyfin_compat.server_id"] = serverID
}
// Step 6: Build config from DB
cfg, err := config.LoadFromDB(settings)
if err != nil {
log.Fatalf("building config: %v", err)
}
// Step 7: Apply bootstrap overrides
cfg.Server.Listen = bc.Listen
cfg.Server.Mode = bc.Mode
cfg.Database.URL = bc.DatabaseURL
cfg.JellyfinCompat.Listen = bc.JFListen
if bc.RedisURL != "" {
cfg.Redis.URL = bc.RedisURL
}
// Step 8: Recreate pool if max_connections differs from bootstrap default
if cfg.Database.MaxConnections != bootstrapDBCfg.MaxConnections {
pool.Close()
pool, err = database.NewPool(ctx, cfg.Database)
if err != nil {
log.Fatalf("recreating pool with configured max_connections: %v", err)
}
}
// Re-wrap with the encrypting decorator so the recreated pool's settings repo
// still encrypts/decrypts — no raw settings repo may escape into later wiring.
settingsRepo = catalog.NewEncryptedSettingsRepo(catalog.NewServerSettingsRepo(pool), dataCipher)
nodeID := resolveNodeIdentity()
catalogSearchStartupSettings, err := catalog.CatalogSearchSettingsFromMap(settings)
if err != nil {
slog.Warn("catalog search: failed to load settings for startup wiring; using postgres", "err", err)
catalogSearchStartupSettings = catalog.DefaultCatalogSearchSettings()
}
activeCatalogSearchProvider := catalog.ActiveCatalogSearchProvider(catalogSearchStartupSettings)
// Step 9: Validate
if err := cfg.Validate(); err != nil {
log.Fatalf("config validation: %v", err)
}
// Step 10: Configure log level. The level var and quiet filter are
// shared with the operational-logging handler chain and hot-reloaded by
// the config watcher in integrated mode.
logLevelVar := new(slog.LevelVar)
logLevelVar.Set(parseLogLevel(cfg.Server.LogLevel))
// Bootstrap OpenTelemetry (logs + traces) before installing the log handler
// chain. Setup depends only on OTEL_* / SILO_OTEL_ENABLED env (not the DB),
// so it is safe to call here. When disabled, this is fully dormant: no
// providers are installed and telemetryShutdown is a no-op.
telemetryCfg := telemetry.LoadConfig(nodeID)
telemetryProviders, telemetryShutdown, err := telemetry.Setup(ctx, telemetryCfg)
if err != nil {
// Telemetry is best-effort: a malformed OTEL_* environment must not
// crash-loop the server. Setup installed no globals and returned no-op
// providers, so continue with telemetry disabled.
slog.ErrorContext(ctx, "telemetry setup failed; continuing with telemetry disabled", "component", "app", "error", err)
telemetryCfg.Enabled = false
}
defer func() {
shutdownCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if err := telemetryShutdown(shutdownCtx); err != nil {
slog.WarnContext(shutdownCtx, "telemetry shutdown error", "component", "app", "error", err)
}
}()
var otelLogHandler slog.Handler
if telemetryCfg.Enabled {
otelLogHandler = telemetry.NewOTelHandler(telemetryProviders.LoggerProvider)
}
baseHandler := buildBaseHandler(cfg.Server.LogFormat, logLevelVar, otelLogHandler)
quietFilter := logfilter.New(baseHandler, cfg.Server.LogQuiet)
slog.SetDefault(slog.New(quietFilter))
mode := cfg.Server.Mode
maybeApplyPostgresTuning(ctx, pool, cfg.Database.MaxConnections, mode)
slog.Info("silo starting", "mode", mode, "listen", cfg.Server.Listen, "log_level", cfg.Server.LogLevel, "node_id", nodeID)
appCtx, appCancel := context.WithCancel(ctx)
defer appCancel()
restartReqCh := make(chan struct{}, 1)
var restartRequested atomic.Bool
eventBus := cache.NewEventBus(cfg.Redis.URL)
logStreamHub := logstream.NewHub(nodeID, eventBus)
if err := logStreamHub.Start(appCtx); err != nil {
log.Fatalf("log stream hub start: %v", err)
}
realtimeHub := notifications.NewHub(nodeID, eventBus)
if err := realtimeHub.Start(appCtx); err != nil {
log.Fatalf("realtime hub start: %v", err)
}
eventsHub := realtimeHub.EventsHub()
scanRegistry := evt.NewScanRegistry()
operationalWriter, opsRepo, opsPM := configureOperationalLogging(appCtx, pool, settingsRepo, cfg.Redis, logStreamHub, quietFilter, nodeID)
defer func() {
if err := eventBus.Close(); err != nil {
slog.Warn("event bus close error", "error", err)
}
}()
// Proxy and transcode modes run with DB + Redis for hot-reload.
if mode == "proxy" || mode == "transcode" {
redisClient, err := cache.NewRedisClient(cfg.Redis)
if err != nil || redisClient == nil {
slog.Error("redis is required for this mode", "mode", mode, "error", err)
os.Exit(1)
}
bootstrap := nodeconfig.BootstrapOverrides{
Listen: cfg.Server.Listen,
Mode: cfg.Server.Mode,
DatabaseURL: cfg.Database.URL,
JFListen: cfg.JellyfinCompat.Listen,
RedisURL: bc.RedisURL,
}
watcher := nodeconfig.NewWatcher(pool, dataCipher, eventBus, bootstrap)
if err := watcher.Start(appCtx); err != nil {
slog.Error("config watcher start failed", "error", err)
os.Exit(1)
}
nodeURL := os.Getenv("NODE_URL")
nodeName := os.Getenv("NODE_NAME")
if nodeURL == "" {
nodeURL = "http://localhost" + cfg.Server.Listen
slog.Warn("NODE_URL not set, using listen address — session keys may collide across nodes")
}
if nodeName == "" {
nodeName = mode
}
tracker := nodesessions.NewTracker(redisClient, nodeURL, nodeName, mode)
tracker.StartRefresh(appCtx)
defer func() {
cleanupCtx, cleanupCancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cleanupCancel()
tracker.Cleanup(cleanupCtx)
}()
var handler http.Handler
if mode == "proxy" {
srv := proxy.NewServer(watcher, tracker)
handler = srv.Handler()
} else {
srv := transcodenode.NewServer(watcher, tracker)
srv.SetFFmpegLogSink(playback.NewSlogFFmpegLogSink(slog.Default(), nodeID))
// Read jellycompat reconstruction recipes central wrote at transcode
// start, so this node can rebuild a Jellyfin transcode after its own
// restart (the node hop token is recipe-less). Shares the offload Redis.
srv.SetRecipeStore(noderecipe.NewStore(redisClient, 0))
// Reclaim orphaned transcode dirs at boot and hourly thereafter, bound
// to appCtx so it stops on shutdown.
srv.StartOrphanSweeper(appCtx)
handler = srv.Handler()
}
_ = operationalWriter
_ = opsRepo
startStandaloneServer(cfg.Server.Listen, handler)
return
}
// Hot-reload config watcher for integrated/api mode. Reloads on
// EventSettingsChanged (Redis) with a 60s poll fallback, so settings
// changes apply without restart even on Redis-less deployments. The
// watcher's config supersedes the startup snapshot from here on.
configWatcher := nodeconfig.NewWatcher(pool, dataCipher, eventBus, nodeconfig.BootstrapOverrides{
Listen: bc.Listen,
Mode: bc.Mode,
DatabaseURL: bc.DatabaseURL,
JFListen: bc.JFListen,
RedisURL: bc.RedisURL,
})
if err := configWatcher.Start(appCtx); err != nil {
log.Fatalf("config watcher start: %v", err)
}
cfg = configWatcher.Config()
// Apply server.log_level / server.log_quiet changes live. Both feed the
// shared level var and quiet filter inside the default logger chain.
configWatcher.OnChange(func(_, updated *config.Config) {
logLevelVar.Set(parseLogLevel(updated.Server.LogLevel))
quietFilter.SetQuiet(updated.Server.LogQuiet)
})
// Determine which components to initialize based on mode.
needsS3 := mode == "integrated" || mode == "api"
needsScanner := mode == "integrated" || mode == "api"
needsUserDB := mode == "integrated" || mode == "api"
needsWorkers := mode == "integrated" || mode == "api"
bootstrapSensitiveConfigured := map[string]bool{}
bootstrapSensitiveValues := map[string]string{}
if bc.RedisURL != "" {
bootstrapSensitiveConfigured["redis.url"] = true
bootstrapSensitiveValues["redis.url"] = bc.RedisURL
}
if rawTrustedProxies := strings.TrimSpace(os.Getenv(clientip.EnvTrustedProxies)); rawTrustedProxies != "" {
normalizedTrustedProxies, normalizeErr := clientip.NormalizeCIDRList(rawTrustedProxies)
if normalizeErr != nil {
log.Fatalf("invalid %s: %v", clientip.EnvTrustedProxies, normalizeErr)
}
bootstrapSensitiveConfigured[clientip.SettingTrustedProxies] = true
bootstrapSensitiveValues[clientip.SettingTrustedProxies] = normalizedTrustedProxies
}
// Shared Redis client for components needing raw Redis beyond the event
// bus (websocket handshake tickets, session listing). Nil on Redis-less
// deployments; consumers fall back to in-process implementations.
apiRedisClient, apiRedisErr := cache.NewRedisClient(cfg.Redis)
if apiRedisErr != nil {
slog.Warn("redis client init failed; multi-node websocket tickets disabled", "error", apiRedisErr)
} else if apiRedisClient != nil {
defer func() { _ = apiRedisClient.Close() }()
}
// Assigned below once the trusted-proxy config is seeded; captured by the
// OnServerSettingUpdated closure, which only runs on admin requests after
// startup completes.
var ipResolver *clientip.Resolver
normalizedBootstrapRedisURL, bootstrapRedisURLErr := config.NormalizeRedisURL(bc.RedisURL)
redisBootstrapAvailable := (normalizedBootstrapRedisURL != "" && bootstrapRedisURLErr == nil) ||
(strings.TrimSpace(cfg.Redis.SentinelMaster) != "" && len(cfg.Redis.SentinelAddresses) > 0)
deps := api.Dependencies{
Config: cfg,
LiveConfig: configWatcher.Config,
OnConfigChange: configWatcher.OnChange,
BootstrapSensitiveConfigured: bootstrapSensitiveConfigured,
BootstrapSensitiveValues: bootstrapSensitiveValues,
RedisBootstrapAvailable: redisBootstrapAvailable,
AppContext: appCtx,
DB: pool,
SecretCipher: dataCipher,
EventBus: eventBus,
RedisClient: apiRedisClient,
LogStreamHub: logStreamHub,
RealtimeHub: realtimeHub,
EventsHub: eventsHub,
ScanRegistry: scanRegistry,
OpsLogRepo: opsRepo,
FFmpegLogSink: playback.NewSlogFFmpegLogSink(slog.Default(), nodeID),
PublicURL: os.Getenv("SILO_PUBLIC_URL"),
RequestServerRestart: func(context.Context) error {
if !restartRequested.CompareAndSwap(false, true) {
return handlers.ErrServerRestartAlreadyRequested
}
restartReqCh <- struct{}{}
return nil
},
OnServerSettingUpdated: func(_ context.Context, key, _ string) {
// Key-scoped reload for the client-IP trust boundary: unlike the
// whole-config watcher reload below, this cannot be blocked by an
// unrelated malformed setting failing config.LoadFromDB. Uses a
// fresh context — the setting is already persisted, so the reload
// must not be skipped because the admin request was canceled.
if key == clientip.SettingTrustedProxies && ipResolver != nil {
if cidrs, loadErr := clientip.LoadTrustedCIDRs(context.Background(), settingsRepo); loadErr != nil {
slog.WarnContext(context.Background(), "clientip config reload failed", "component", "app", "error", loadErr)
} else {
ipResolver.UpdateTrustedCIDRs(cidrs)
}
}
// Nudge the hot-reload watcher so same-process settings changes
// apply immediately even without Redis (the event bus is a no-op
// then, leaving only the 60s poll).
configWatcher.RequestReload()
},
}
accessGroupStore := access.NewGroupStore(pool)
audiobooksService := audiobooks.New(&audiobooksSettingsAdapter{repo: settingsRepo})
absCompatEnabled, err := audiobooksService.ABSCompatEnabled(appCtx)
if err != nil {
slog.Warn("Audiobookshelf compatibility disabled; failed to read setting", "err", err)
absCompatEnabled = false
}
adminJobCancelRegistry := adminjob.NewCancelRegistry()
deps.AdminJobCancelRegistry = adminJobCancelRegistry
if needsWorkers && deps.DB != nil {
deps.IntroRepository = intromarkers.NewRepository(deps.DB)
deps.IntroAnalyzer = intromarkers.NewAnalyzer(
deps.IntroRepository,
intromarkers.DefaultConfig(cfg.Playback.FFmpegPath),
slog.Default(),
)
}
if deps.DB != nil {
markerRegistry := markers.NewRegistry(slog.Default())
markerProviderConfig := markers.NewProviderConfigStore(deps.DB)
if err := markerProviderConfig.Reload(appCtx); err != nil {
slog.Warn("load marker provider config failed; falling back to registration-order fetch",
"error", err)
} else {
markerRegistry.UseConfigStore(markerProviderConfig)
if deps.EventBus != nil {
if err := deps.EventBus.Subscribe(appCtx, cache.ChannelAdmin, func(event cache.Event) {
if event.Type != cache.EventMarkerProviderConfigChanged {
return
}
if err := markerProviderConfig.Reload(appCtx); err != nil {
slog.Warn("reload marker provider config failed", "provider", event.Payload, "error", err)
}
}); err != nil {
slog.Warn("subscribe marker provider config reload failed", "error", err)
}
}
}
deps.MarkerProviderConfig = markerProviderConfig
deps.MarkerRegistry = markerRegistry
markerResolver := markers.NewDBExternalIDResolver(deps.DB)
deps.MarkerResolver = markerResolver
markerContributionStore := markers.NewContributionStore(deps.DB)
deps.MarkerContributionStore = markerContributionStore
deps.MarkerContributionService = markers.NewContributionService(
markerRegistry, markerResolver, markerProviderConfig, markerContributionStore, slog.Default(),
)
}
var watchProviderService *watchsync.Service
if deps.DB != nil {
watchProviderRegistry := watchsync.NewRegistry()
if err := watchProviderRegistry.Register(trakt.NewProvider(nil, "")); err != nil {
log.Fatalf("register watch provider: %v", err)
}
if err := watchProviderRegistry.Register(simkl.NewProvider(nil, "")); err != nil {
log.Fatalf("register watch provider: %v", err)
}
if err := watchProviderRegistry.Register(watchmdblist.NewProvider(nil, "")); err != nil {
log.Fatalf("register watch provider: %v", err)
}
watchProviderService = watchsync.NewService(
watchsync.NewPostgresRepository(deps.DB, deps.SecretCipher),
watchProviderRegistry,
)
deps.WatchProviderService = watchProviderService
}
// Initialize node pools for integrated/api modes.
if mode == "integrated" || mode == "api" {
nodeRepo := nodepool.NewRepository(pool)
deps.NodeRepo = nodeRepo
proxyPool := nodepool.NewProxyPool()
transcodePool := nodepool.NewTranscodePool()
proxyNodes, _ := nodeRepo.ListEnabled(context.Background(), nodepool.NodeTypeProxy)
transcodeNodes, _ := nodeRepo.ListEnabled(context.Background(), nodepool.NodeTypeTranscode)
proxyPool.SetNodes(proxyNodes)
transcodePool.SetNodes(transcodeNodes)
deps.ProxyPool = proxyPool
deps.TranscodePool = transcodePool
deps.NodePlanner = nodepool.NewPlanner(proxyPool, transcodePool)
healthChecker := nodepool.NewHealthChecker(proxyPool, transcodePool, nodeRepo)
healthChecker.Start(appCtx)
slog.Info("node pools initialized", "proxy_nodes", len(proxyNodes), "transcode_nodes", len(transcodeNodes))
// Subscribe to node pool change events for multi-instance reload.
_ = eventBus.Subscribe(appCtx, cache.ChannelAdmin, func(event cache.Event) {
if event.Type == cache.EventNodePoolChanged {
pNodes, pErr := nodeRepo.ListEnabled(context.Background(), nodepool.NodeTypeProxy)
tNodes, tErr := nodeRepo.ListEnabled(context.Background(), nodepool.NodeTypeTranscode)
if pErr != nil || tErr != nil {
slog.Warn("node pool reload from event failed, keeping current pools",
"proxy_err", pErr, "transcode_err", tErr)
return
}
proxyPool.SetNodes(pNodes)
transcodePool.SetNodes(tNodes)
slog.Info("node pools reloaded from event", "proxy", len(pNodes), "transcode", len(tNodes))
}
})
}
// Step 3: Create S3 clients (if needed).
if needsS3 {
configureS3Clients(cfg, &deps)
}
var literaryWorkService *literaryworks.Service
if deps.DB != nil {
literaryWorkService = literaryworks.NewService(literaryworks.NewRepository(deps.DB))
}
// Step 4: Create scanner (if needed).
if needsScanner && deps.DB != nil {
folderRepo := catalog.NewFolderRepository(deps.DB)
fileRepo := scanner.NewFileRepository(deps.DB)
deps.FolderRepo = folderRepo
deps.FileRepo = fileRepo
ffprobePath := scanner.FFprobePathFromFFmpeg(cfg.Playback.FFmpegPath)
s := scanner.NewScanner(fileRepo, ffprobePath, deps.S3Public, cfg.Scanner.Workers, cfg.Scanner.EmptyTrashAfterScan, cfg.Scanner.FileRemovalGrace)
s.SetSearchIndexProvider(activeCatalogSearchProvider)
configWatcher.OnChange(func(_, updated *config.Config) {
s.SetWorkers(updated.Scanner.Workers)
})
s.SetLiteraryWorkLinker(literaryWorkService)
s.SetEbookEnrichmentQueue(ebooks.NewEnrichmentQueue(deps.DB))
deps.Scanner = s
deps.ProbeEnsurer = scanner.NewPlaybackProbeEnsurer(fileRepo, ffprobePath, cfg.Playback.FFmpegPath, 10*time.Second)
slog.Info("scanner initialized")
}
var chapterThumbService *chapterthumbs.Service
if deps.FileRepo != nil && deps.FolderRepo != nil && deps.S3Public != nil {
chapterThumbService = chapterthumbs.NewService(
deps.FileRepo,
deps.FolderRepo,
deps.ProbeEnsurer,
settingsRepo,
deps.S3Public,
nil,
deps.TranscodePool,
cfg.Playback.FFmpegPath,
cfg.Playback.HWAccel,
cfg.Playback.HWDevice,
cfg.Playback.ChapterThumbnailWorkers,
)
if chapterThumbService != nil {
chapterThumbService.Start(appCtx)
deps.ChapterThumbnailQueuer = chapterThumbService
}
}
var pluginHost *pluginhost.Host
var pluginService *plugins.Service
var pluginInstallationStore *plugins.InstallationStore
var pluginRuntimeConfigStore *plugins.RuntimeConfigStore
var pluginHTTPProxy *plugins.HTTPProxy
pluginAutoUpdateDone := make(chan struct{})
var pluginAutoUpdater *plugins.AutoUpdateService
if deps.DB != nil {
pluginCacheDir := resolvePluginCacheDir()
repositoryStore := plugins.NewRepositoryStore(deps.DB)
installationStore := plugins.NewInstallationStore(deps.DB)
runtimeConfigStore := plugins.NewRuntimeConfigStore(deps.DB, deps.SecretCipher)
catalogService := plugins.NewCatalogService(repositoryStore, plugins.CatalogServiceOptions{
SiloAPIVersion: plugins.DefaultSiloAPIVersion,
})
installer := plugins.NewInstaller(installationStore, plugins.InstallerOptions{
BaseDir: pluginCacheDir,
})
libDataSource := pluginhost.LibraryDataSourceFunc(
func(ctx context.Context, _ string) ([]pluginhost.LibraryRecord, error) {
// TODO: scope by userID when the requests plugin needs it (Plan B).
// For now, all callers see admin-scope.
if deps.FolderRepo == nil {
return nil, nil
}
folders, err := deps.FolderRepo.List(ctx)
if err != nil {
return nil, err
}
out := make([]pluginhost.LibraryRecord, 0, len(folders))
for _, f := range folders {
out = append(out, pluginhost.LibraryRecord{
ID: strconv.Itoa(f.ID),
Name: f.Name,
MediaType: mapFolderTypeToMediaType(f.Type),
})
}
return out, nil
},
)
presenceItemRepo := catalog.NewItemRepository(deps.DB)
catalogPresence := pluginhost.NewCatalogPresence(
func(ctx context.Context, mediaType string, tmdbIDs []string) ([]pluginhost.LibraryPresenceRecord, error) {
rows, err := presenceItemRepo.LookupTMDBIDs(ctx, mediaType, tmdbIDs)
if err != nil {
return nil, err
}
out := make([]pluginhost.LibraryPresenceRecord, 0, len(rows))
for _, r := range rows {
out = append(out, pluginhost.LibraryPresenceRecord{
ExternalID: r.TMDBID,
MediaID: r.MediaID,
LibraryID: r.LibraryID,
Title: r.Title,
})
}
return out, nil
},
)
pluginHost = pluginhost.NewHost(pluginhost.Config{
EventPublisher: eventsHub,
LibraryLister: pluginhost.NewLibraryLister(libDataSource),
CatalogPresence: catalogPresence,
InstalledPlugins: pluginhost.InstalledPluginListerFunc(
func(ctx context.Context) ([]pluginhost.InstalledPluginRecord, error) {
installations, err := installationStore.List(ctx)
if err != nil {
return nil, err
}
out := make([]pluginhost.InstalledPluginRecord, 0, len(installations))
for _, installation := range installations {
// The reserved builtin row is not a plugin; keep it out
// of the host's installed-plugin listing.
if installation.IsBuiltin() {
continue
}
capabilities, err := installationStore.ListCapabilities(ctx, installation.ID)
if err != nil {
return nil, err
}
descriptors := make([]*pluginv1.CapabilityDescriptor, 0, len(capabilities))
for _, capability := range capabilities {
descriptor, err := plugins.DecodeCapability(capability)
if err != nil {
return nil, err
}
descriptors = append(descriptors, descriptor)
}
out = append(out, pluginhost.InstalledPluginRecord{
InstallationID: installation.ID,
PluginID: installation.PluginID,
Version: installation.Version,
Enabled: installation.Enabled,
Capabilities: descriptors,
})
}
return out, nil
},
),
GlobalConfigSetter: pluginhost.GlobalConfigSetterFunc(
func(ctx context.Context, installationID int, key string, value map[string]any) error {
return runtimeConfigStore.PutGlobalConfig(ctx, installationID, key, value)
},
),
Logger: hclog.New(&hclog.LoggerOptions{
Name: "plugin-host",
Level: hclog.Info,
Output: os.Stderr,
}),
})
pluginService = plugins.NewService(
repositoryStore,
installationStore,
runtimeConfigStore,
catalogService,
installer,
plugins.NewHostAdapter(pluginHost),
)
if deps.MarkerRegistry != nil && deps.MarkerProviderConfig != nil {
markerPluginResolver := markers.NewPluginResolverAdapter(pluginService)
pluginService.AddLifecycleHook(func(ctx context.Context) {
if err := reloadMarkerPluginProviders(
ctx,
deps.MarkerRegistry,
deps.MarkerProviderConfig,
installationStore,
runtimeConfigStore,
settingsRepo,
markerPluginResolver,
); err != nil {
slog.WarnContext(ctx, "reload marker plugin providers failed", "component", "app", "error", err)
}
})
}
if err := pluginService.PreloadEnabled(appCtx); err != nil {
log.Fatalf("preload enabled plugins: %v", err)
}
slog.Info("plugin cache initialized", "base_dir", pluginCacheDir)
pluginAutoUpdater = plugins.NewAutoUpdateService(
repositoryStore,
installationStore,
catalogService,
installer,
pluginHost,
slog.Default(),
// Auto-updates rewrite installation rows (new version-specific
// InstallPath/Version) and delete the old install dir without going
// through pluginService. Wire OnLifecycleChange so the service's
// installation cache is invalidated and later plugin RPCs re-read
// the fresh row instead of a stale one.
pluginService.OnLifecycleChange,
)
go func() {
defer close(pluginAutoUpdateDone)
if err := pluginAutoUpdater.Run(appCtx); err != nil {
slog.Error("plugin auto-update failed", "error", err)
}
}()
pluginInstallationStore = installationStore
pluginRuntimeConfigStore = runtimeConfigStore
pluginHTTPProxy = plugins.NewHTTPProxyWithTypedResolver(pluginService, pluginInstallationStore)
if deps.DB != nil {
pluginHTTPProxy = pluginHTTPProxy.WithUserThemeLookup(plugins.NewPgUserThemeLookup(deps.DB))
pluginHTTPProxy = pluginHTTPProxy.WithUserIdentityLookup(plugins.NewPgUserIdentityLookup(deps.DB))
}
deps.PluginService = pluginService
deps.PluginHTTPProxy = pluginHTTPProxy
defer func() {
if pluginHost == nil {
return
}
shutdownCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
if err := pluginHost.Shutdown(shutdownCtx); err != nil {
slog.Warn("failed to shut down plugin host", "error", err)
}
}()
} else {
close(pluginAutoUpdateDone)
}
if pluginService != nil && pluginInstallationStore != nil {
dispatcher := plugins.NewEventDispatcherWithTypedResolver(deps.EventBus, deps.EventsHub, pluginInstallationStore, pluginService, 4)
pluginService.SetEventDispatcher(dispatcher)
if err := dispatcher.Start(appCtx); err != nil {
log.Fatalf("plugin event dispatcher: %v", err)
}
defer dispatcher.Stop()
// Backfill the capability-subscriber index from the already-preloaded
// installations. PreloadEnabled ran earlier (before the dispatcher
// existed), so its rebuildDispatcherIndex was a no-op. Without this
// call, capability-scoped subscriptions never fire until the next
// lifecycle mutation.
pluginService.OnLifecycleChange(appCtx)
}
// backgroundInit collects non-critical startup work (catalog-size-dependent
// seeding, network-bound reconciliation) that must not block the HTTP
// listener. The steps run sequentially in a background goroutine once the
// server is ready to serve. Failures are logged, never fatal.
var backgroundInit []func(context.Context)
// Step 4b: Create metadata service and match worker (if needed).
var metadataService *metadata.MetadataService
var metadataImageCacheProcessor *metadata.ImageCacheProcessor
var personRefreshService *metadata.PersonRefreshService
var matchWorker *metadata.MatchWorker
var libraryIngestExecutor *libraryingest.Executor
var libraryScanQueue *scanqueue.Service
var itemRefreshExecutor *adminjob.ItemRefreshExecutor
var libraryRefreshExecutor *adminjob.LibraryRefreshExecutor
var itemRepo *catalog.ItemRepository
var skippedRootRepo *metadata.SkippedRootRepository
var movieQueueRepo *metadata.MovieMatchQueueRepository
var seriesQueueRepo *metadata.SeriesRootMatchQueueRepository
var matchQueueCoordinator *metadata.MatchQueueCoordinator
var rootClaimRepo *catalog.RootClaimRepository
var groupClaimRepo *catalog.GroupClaimRepository
var seasonRepo *catalog.SeasonRepository
var episodeRepo *catalog.EpisodeRepository
var audiobookEnricher *audiobooks.Enricher
var ebookEnricher *ebooks.Enricher
var mangaEnricher *manga.Enricher
if needsWorkers && deps.DB != nil && deps.FileRepo != nil {
chainRepo := metadata.NewChainRepository(deps.DB)
// Make every existing library chain aware of the built-in providers
// before serving: materialize legacy content_level='' chains per level,
// then append registered builtins disabled (idempotent; also the repair
// path after a stale chain-editor save drops a builtin row). Runs before
// the metadata service exists, so no chain cache to invalidate here.
syncCtx, syncCancel := context.WithTimeout(appCtx, 30*time.Second)
syncErr := metadata.SyncBuiltinProviderChains(syncCtx, chainRepo)
syncCancel()
if syncErr != nil {
log.Fatalf("sync builtin provider chains: %v", syncErr)
}
skippedRootRepo = metadata.NewSkippedRootRepository(deps.DB)
itemRepo = catalog.NewItemRepository(deps.DB).WithActiveSearchProvider(activeCatalogSearchProvider)
episodeRepo = catalog.NewEpisodeRepository(deps.DB)
seasonRepo = catalog.NewSeasonRepository(deps.DB)
personRepo := catalog.NewPersonRepository(deps.DB)
libraryRepo := catalog.NewLibraryItemRepository(deps.DB)
// Wait for plugin auto-update to finish before registering image resolvers.
<-pluginAutoUpdateDone
imageResolver := metadata.NewPluginImageResolver()
if pluginService != nil && pluginInstallationStore != nil {
reloadImageResolvers := func(ctx context.Context) {
if err := reloadPluginImageResolvers(ctx, pluginInstallationStore, imageResolver, pluginService); err != nil {
slog.WarnContext(ctx, "failed to reload plugin image resolvers", "component", "app", "error", err)
}
}
pluginService.AddLifecycleHook(reloadImageResolvers)
reloadImageResolvers(appCtx)
}
if deps.S3Public != nil {
presignTTL := cfg.S3.MetadataPresignExpiry
if presignTTL <= 0 {
presignTTL = 4 * time.Hour
}
imageResolver.SetS3Presigner(deps.S3Public, deps.S3Public.EffectivePresignTTL(presignTTL))
}
deps.ImageResolver = imageResolver
deps.PluginImageResolver = imageResolver
staleIDRepo := metadata.NewStaleMediaIDRepository(deps.DB)
providerIDRepo := catalog.NewProviderIDRepository(deps.DB)
movieQueueRepo = metadata.NewMovieMatchQueueRepository(deps.DB, deps.FileRepo)
seriesQueueRepo = metadata.NewSeriesRootMatchQueueRepository(deps.DB)
deps.MovieMatchQueueRepo = movieQueueRepo
deps.SeriesRootMatchQueueRepo = seriesQueueRepo
matchQueueCoordinator = metadata.NewMatchQueueCoordinator(movieQueueRepo, seriesQueueRepo)
backgroundInit = append(backgroundInit, func(ctx context.Context) {
if err := matchQueueCoordinator.WakeForChangedInputs(ctx); err != nil {
slog.WarnContext(ctx, "refresh metadata match queue inputs at startup failed", "component", "app", "error", err)
}
})
if pluginService != nil {
matchInputChanged := make(chan struct{}, 1)
go func() {
for {
select {
case <-appCtx.Done():
return
case <-matchInputChanged:
if err := matchQueueCoordinator.WakeForChangedInputs(appCtx); err != nil {
slog.WarnContext(appCtx, "wake metadata matches after plugin lifecycle change failed", "component", "app", "error", err)
}
}
}
}()
pluginService.AddLifecycleHook(func(context.Context) {
// Queue fingerprint reconciliation may touch thousands of parked
// rows. Coalesce lifecycle bursts and keep plugin admin requests
// independent of that background database work.
select {
case matchInputChanged <- struct{}{}:
default:
}
})
}
rootClaimRepo = catalog.NewRootClaimRepository(deps.DB)
groupClaimRepo = catalog.NewGroupClaimRepository(deps.DB)
pluginResolver := metadata.NewPluginResolverAdapter(pluginService)
// Serve the metadata chain's plugin-installation enabled-check from the
// plugins service's in-memory installation cache. Declared as the
// interface type and only assigned when pluginService is non-nil so a
// nil *plugins.Service is passed as a genuine nil interface (not a
// typed-nil), letting buildProviders fall back to the pool query.
var installationEnabledChecker metadata.InstallationEnabledChecker
if pluginService != nil {
installationEnabledChecker = pluginService
}
metadataService = metadata.NewMetadataService(
chainRepo, pluginResolver, installationEnabledChecker,
itemRepo, providerIDRepo, episodeRepo, seasonRepo, libraryRepo, deps.FolderRepo,
personRepo,
deps.FileRepo, skippedRootRepo, staleIDRepo, rootClaimRepo,
)
// Drop the resolved-chain cache whenever a plugin is installed, enabled,
// disabled, updated, or uninstalled. The installation-enabled check is
// served from the plugins service's in-memory cache (invalidated on the
// same events), but resolveChainCached would otherwise keep serving a
// stale provider chain for up to chainCacheTTL after a provider's
// availability changes.
if pluginService != nil {
pluginService.AddLifecycleHook(func(context.Context) {
metadataService.InvalidateChainCache()
})
}
personRefreshService = metadata.NewPersonRefreshService(deps.DB, pluginResolver, personRepo)
personRefreshService.SetImageResolver(imageResolver)
// Wire the audiobook enricher. It uses the same plugin resolver and chain
// repo as the movie/TV pipeline, but resolves providers at
// content_level='audiobook' and sweeps items directly rather than via a queue.
audiobookEnricher = audiobooks.NewEnricher(
deps.DB,
chainRepo,
pluginResolver,
itemRepo,
personRepo,
providerIDRepo,
)
ebookEnricher = ebooks.NewEnricher(
deps.DB,
chainRepo,
pluginResolver,
itemRepo,
personRepo,
providerIDRepo,
)
audiobookEnricher.SetLiteraryWorkLinker(literaryWorkService)
ebookEnricher.SetLiteraryWorkLinker(literaryWorkService)
mangaEnricher = manga.NewEnricher(
deps.DB,
chainRepo,
pluginResolver,
itemRepo,
personRepo,
providerIDRepo,
)
// Always wire the image resolver so plugin-prefixed URLs (e.g.
// metadb://) can be resolved to presigned HTTP URLs in API responses.
metadataService.SetImageResolver(imageResolver)
// Wire the image cacher whenever object storage is available so explicit
// admin image applies can succeed even if automatic metadata caching is off.
if deps.S3Public != nil {
imageCacher := imagecache.New(deps.S3Public)
imageCacher.SetArtworkRevisionTracker(catalog.NewArtworkRevisionTracker(deps.DB))
metadataService.SetImageCacher(imageCacher)
imageCacheJobs := metadata.NewImageCacheJobRepository(deps.DB)
metadataService.SetImageCacheJobEnqueuer(imageCacheJobs)
metadataImageCacheProcessor = metadata.NewImageCacheProcessorWithTargets(
imageCacheJobs,
imageCacher,
imageResolver,
metadata.ImageCacheProcessorTargets{
Items: itemRepo,
Seasons: seasonRepo,
Episodes: episodeRepo,
ItemLocalizations: catalog.NewMediaItemLocalizationRepository(deps.DB),
SeasonLocalizations: catalog.NewSeasonLocalizationRepository(deps.DB),
People: personRepo,
},
)
// Local file:// artwork (NFO sidecars): confine reads to the owning
// library's roots and sweep stale hashed local/ prefixes on re-cache.
// The processor host must mount the libraries, like the metadata worker.
metadataImageCacheProcessor.SetLibraryRootResolver(deps.FolderRepo)
metadataImageCacheProcessor.SetImagePrefixDeleter(deps.S3Public)
metadataService.SetAutoCacheImages(cfg.Metadata.CacheImages)
metadataImageCacheProcessor.SetEnabled(cfg.Metadata.CacheImages)
configWatcher.OnChange(func(_, updated *config.Config) {
metadataService.SetAutoCacheImages(updated.Metadata.CacheImages)
metadataImageCacheProcessor.SetEnabled(updated.Metadata.CacheImages)
})
if deps.Scanner != nil {
deps.Scanner.SetImageCacher(imageCacher)
}
if cfg.Metadata.CacheImages {
personRefreshService.SetImageCacher(imageCacher)
personRefreshService.SetImageCacheJobEnqueuer(imageCacheJobs)
slog.Info("metadata image caching enabled")
}
if audiobookEnricher != nil {
audiobookEnricher.SetImageCacher(imageCacher)
audiobookEnricher.SetImageCacheJobEnqueuer(imageCacheJobs)
audiobookEnricher.SetFFmpegPath(scanner.FFmpegPathFromFFprobe(scanner.FFprobePathFromFFmpeg(cfg.Playback.FFmpegPath)))
}
if ebookEnricher != nil {
ebookEnricher.SetImageCacher(imageCacher)
ebookEnricher.SetImageCacheJobEnqueuer(imageCacheJobs)
}
if mangaEnricher != nil {
mangaEnricher.SetImageCacher(imageCacher)
mangaEnricher.SetImageCacheJobEnqueuer(imageCacheJobs)
}
}
matchWorker = metadata.NewMatchWorker(metadataService, deps.FileRepo, cfg.Matcher.Workers, cfg.Matcher.BatchSize, 30*time.Second)
mwForReload := matchWorker
configWatcher.OnChange(func(_, updated *config.Config) {
mwForReload.SetConcurrency(updated.Matcher.Workers, updated.Matcher.BatchSize)
})
matchWorker.SetRealtimeHub(deps.RealtimeHub)
if movieQueueRepo != nil {
matchWorker.SetMovieFileClaimer(movieQueueRepo)
}
if seriesQueueRepo != nil {
matchWorker.SetSeriesRootClaimer(seriesQueueRepo, cfg.Matcher.TVSeriesRootQueueEnabled())
backgroundInit = append(backgroundInit, func(ctx context.Context) {
if cleaned, err := seriesQueueRepo.CleanupLegacySeriesGroupQueue(ctx); err != nil {
slog.WarnContext(ctx, "failed to clean legacy series group queue rows", "component", "app", "error", err)
} else if cleaned > 0 {
slog.InfoContext(ctx, "cleaned legacy series group queue rows", "component", "app", "count", cleaned)
}
})
}
if deps.FolderRepo != nil {
backgroundInit = append(backgroundInit, func(ctx context.Context) {
start := time.Now()
enabledFolders, err := deps.FolderRepo.GetEnabled(ctx)
if err != nil {
slog.WarnContext(ctx, "failed to seed metadata queues", "component", "app", "error", err)
return
}
seedMovieQueue := func(folderID int) {
if movieQueueRepo == nil {
return
}
if err := movieQueueRepo.SyncForFolder(ctx, folderID); err != nil {
slog.WarnContext(ctx, "failed to seed movie match queue", "component", "app", "folder_id", folderID, "error", err)
}
}
seedSeriesQueue := func(folderID int) {
if seriesQueueRepo == nil {
return
}
if err := seriesQueueRepo.SyncForFolder(ctx, folderID); err != nil {
slog.WarnContext(ctx, "failed to seed series root queue", "component", "app", "folder_id", folderID, "error", err)
}
}
for _, folder := range enabledFolders {
if folder == nil {
continue
}
switch strings.ToLower(strings.TrimSpace(folder.Type)) {
case "movie", "movies":
seedMovieQueue(folder.ID)
case "series", "tv", "show", "tvshows":
seedSeriesQueue(folder.ID)
case "mixed":
seedSeriesQueue(folder.ID)
seedMovieQueue(folder.ID)
}
}
slog.InfoContext(ctx, "deferred init: metadata match queues seeded", "component", "app", "folders", len(enabledFolders), "duration", time.Since(start))
})
}
deps.SkippedRootRepo = skippedRootRepo
deps.StaleIDRepo = staleIDRepo
deps.PersonRepo = personRepo
deps.PersonRefreshQueue = worker.NewPersonRefreshWorker(
personRefreshService,
worker.DefaultPersonRefreshWorkerConfig(),
)
deps.PersonRefresher = personRefreshService
deps.Refresher = metadataService
deps.MetadataService = metadataService
slog.Info("metadata service initialized and running")
}
if deps.Scanner != nil {
if matchQueueCoordinator != nil {
deps.Scanner.SetMetadataQueueProducer(matchQueueCoordinator)
}
if movieQueueRepo != nil {
deps.Scanner.SetMovieQueueSyncer(movieQueueRepo)
}
if seriesQueueRepo != nil {
deps.Scanner.SetSeriesQueueSyncer(seriesQueueRepo)
}
}
if deps.Scanner != nil && matchWorker != nil && deps.FolderRepo != nil && skippedRootRepo != nil {
libraryIngestExecutor = libraryingest.NewExecutor(
deps.Scanner,
matchWorker,
deps.FolderRepo,
skippedRootRepo,
deps.EventBus,
deps.RealtimeHub,
)
deps.LibraryIngester = libraryIngestExecutor
if deps.DB != nil {
libraryScanQueue = scanqueue.NewService(
scanqueue.NewRepository(deps.DB),
deps.FolderRepo,
libraryIngestExecutor,
deps.EventsHub,
appCtx,
cfg.Scanner.MaxConcurrentLibraries,
cfg.Scanner.MaxConcurrentScoped,
)
// Started below, after the notification system has attached its
// availability detector to the executor: a scan resumed by the
// workers before that wiring would complete without recording
// episode availability, silently losing release notifications.
deps.LibraryScanQueue = libraryScanQueue
}
if deps.DB != nil && deps.FileRepo != nil && metadataService != nil {
itemRefreshResolver := adminjob.NewItemRefreshResolver(
itemRepo,
seasonRepo,
episodeRepo,
deps.FolderRepo,
deps.FileRepo,
)
libraryRefreshExecutor = adminjob.NewLibraryRefreshExecutor(
adminjob.NewPGLibraryRefreshItemLister(deps.DB),
deps.FolderRepo,
itemRefreshResolver,
libraryIngestExecutor,
metadataService,
deps.EventBus,
deps.RealtimeHub,
)
}
if metadataService != nil && deps.FileRepo != nil {
itemRefreshExecutor = adminjob.NewItemRefreshExecutor(
deps.FolderRepo,
deps.FileRepo,
rootClaimRepo,
groupClaimRepo,
skippedRootRepo,
seasonRepo,
episodeRepo,
libraryIngestExecutor,
metadataService,
deps.EventBus,
deps.RealtimeHub,
)
}
}
// Ensure PersonRepo is available for the router's DetailService.
if deps.DB != nil && deps.PersonRepo == nil {
deps.PersonRepo = catalog.NewPersonRepository(deps.DB)
}
// Step 5: Create user store provider (if needed).
var userStoreProvider userstore.UserStoreProvider
if needsUserDB {
switch cfg.UserDB.Backend {
case "sqlite":
poolConfig := userdb.PoolConfig{
MaxOpen: cfg.UserDB.PoolMaxOpen,
IdleTimeout: cfg.UserDB.IdleTimeout,
DataDir: "/var/lib/silo/userdb",
}
pool := userdb.NewUserDBPool(poolConfig)
userStoreProvider = userdb.NewSQLiteProvider(pool)
slog.Info("user store initialized", "backend", "sqlite", "max_open", poolConfig.MaxOpen)
default: // "postgres"
userStoreProvider = pgstore.NewPostgresProvider(deps.DB)
slog.Info("user store initialized", "backend", "postgres")
}
defer userStoreProvider.Close()
}
var policySystem *policy.System
if mode == "integrated" || mode == "api" {
policyDecisionLogger := policy.NewDecisionLogger(
deps.DB,
nodeID,
policy.WithDecisionLogLogger(slog.Default()),
)
policyDecisionLogger.SetVerbosity(cfg.Policy.DecisionLogVerbosity)
policyDecisionLogger.SetScopeSampleRate(cfg.Policy.DecisionLogScopeSampleRate)
policySystem = policy.NewSystem(
policy.NewPolicyStore(deps.DB),
deps.EventBus,
slog.Default(),
policy.WithSystemEvalTimeout(time.Duration(cfg.Policy.EvalTimeoutMS)*time.Millisecond),
policy.WithSystemDecisionLogger(policyDecisionLogger),
)
if err := policySystem.Start(appCtx); err != nil {
log.Fatalf("policy system start: %v", err)
}
deps.PolicySystem = policySystem
configWatcher.OnChange(func(_, updated *config.Config) {
policySystem.SetEvalTimeout(time.Duration(updated.Policy.EvalTimeoutMS) * time.Millisecond)
if logger := policySystem.DecisionLogger(); logger != nil {
logger.SetVerbosity(updated.Policy.DecisionLogVerbosity)
logger.SetScopeSampleRate(updated.Policy.DecisionLogScopeSampleRate)
}
})
defer policySystem.Stop()
}
// User-facing release notifications. The system reads user state through
// the raw store provider; the provider handed to everything downstream is
// wrapped so every favorites/watchlist/progress mutation (REST handlers,
// jellycompat, imports, playback) feeds the interest index.
var notificationSystem *notifications.System
if deps.DB != nil && userStoreProvider != nil {
userRepo := auth.NewUserRepository(deps.DB)
profileTokens := access.NewProfileTokenService(cfg.Auth.JWTSecret, 0)
var notificationScopes notifications.ScopeResolver
if policySystem != nil {
notificationScopes = policy.NewViewerResolver(userRepo, userStoreProvider, profileTokens, policySystem.PDP(), accessGroupStore)
} else {
// Legacy resolver: proxy/test wiring without a policy system. Production integrated/api modes always take the policy path. Removed with the legacy cleanup phase.
notificationScopes = access.NewResolver(userRepo, userStoreProvider, profileTokens, accessGroupStore)
}
notificationSystem = notifications.NewSystem(
deps.DB,
settingsRepo,
userStoreProvider,
notificationScopes,
userRepo,
deps.EventsHub,
deps.RedisClient,
deps.SecretCipher,
mail.NewSMTPSender(settingsRepo),
)
userStoreProvider = notifications.WrapUserStoreProvider(userStoreProvider, notificationSystem)
deps.Notifications = notificationSystem
if libraryIngestExecutor != nil {
libraryIngestExecutor.SetAvailabilityDetector(notificationSystem.Detector)
}
if needsWorkers {
notificationSystem.Start(appCtx)
defer notificationSystem.Wait()
}
}
// Start the scan queue only now that the availability detector (when
// notifications are enabled) is attached to the ingest executor, so scans
// resumed at startup cannot complete before the detector exists.
if libraryScanQueue != nil {
libraryScanQueue.Start()
defer libraryScanQueue.Stop()
}
if userStoreProvider != nil && pluginService != nil {
deps.PluginUserConfig = plugins.NewUserConfigStore(userStoreProvider, pluginService)
}
// Step 6: Create playback session manager and wire into dependencies.
sessionMgr := playback.NewSessionManager(6, 2) // defaults from plan: max_streams=6, max_transcodes=2
var compatTerminalRecoveryReady <-chan struct{}
if userStoreProvider != nil {
deps.UserStoreProvider = userStoreProvider
}
if watchProviderService != nil {
historyRepo := historyimport.NewRepository(deps.DB, deps.SecretCipher)
historyIdentity := watchstate.NewStableIdentityResolver(itemRepo, episodeRepo, catalog.NewProviderIDRepository(deps.DB))
watchProviderService.
WithMatcher(historyimport.NewMatcher(historyRepo)).
WithWatchState(watchstate.NewService(userStoreProvider).WithStableIdentityResolver(historyIdentity)).
WithUserStoreProvider(userStoreProvider)
backgroundInit = append(backgroundInit, func(ctx context.Context) {
if compatTerminalRecoveryReady != nil {
select {
case <-compatTerminalRecoveryReady:
case <-ctx.Done():
return
}
}
if err := watchProviderService.SweepOpenScrobbles(ctx); err != nil {
slog.WarnContext(ctx, "failed to sweep open watch provider scrobbles", "component", "app", "error", err)
}
})
}
// Auto-remove fully-watched movies from the watchlist (standalone behavior,
// default-on per profile), propagating removals to connected providers.
// Series are never removed; watchlist read paths hide fully-watched ones
// (catalog.WatchlistVisibility) so newly added episodes bring them back.
if itemRepo != nil && userStoreProvider != nil {
maintainer := watchlist.NewMaintainer(userStoreProvider, itemRepo)
if watchProviderService != nil {
maintainer.WithListEventDispatcher(watchProviderService)
}
deps.WatchCompletionObserver = maintainer
}
deps.SessionMgr = sessionMgr
deps.PlaybackRealtimeHub = playback.NewRealtimeHub()
if chapterThumbService != nil && deps.S3Public != nil {
chapterThumbService.SetNotifier(
playback.NewChapterThumbnailNotifier(sessionMgr, deps.PlaybackRealtimeHub, deps.S3Public, 0),
)
}
// Build the reconciler early enough that playback handlers can trigger
// immediate session syncs after start/stop events.
nodeIdentity := resolveNodeIdentity()
var reconciler *worker.Reconciler
var heartbeatWriter *worker.HeartbeatWriter
if needsWorkers && deps.DB != nil {
sessionProvider := func() []worker.SessionSync {
sessions := sessionMgr.AllSessions()
syncs := make([]worker.SessionSync, len(sessions))
for i, s := range sessions {
syncs[i] = buildLiveSessionSync(s, nodeIdentity)
}
return syncs
}
reconciler = worker.NewReconciler(deps.DB, nodeIdentity, sessionProvider)
reconciler.EventBus = deps.EventBus
reconciler.EventsHub = deps.EventsHub
reconciler.PreSync = func() {
// Retire sessions that have not shown real playback activity
// recently enough to count as live. This keeps the in-memory
// limiter, transcode teardown, and synced admin view aligned.
if expired := sessionMgr.CleanStale(); len(expired) > 0 {
slog.Info("expired idle sessions", "count", len(expired))
}
}
deps.SessionSyncer = reconciler
nodeURL := fmt.Sprintf("http://%s%s", nodeIdentity, cfg.Server.Listen)
heartbeatWriter = worker.NewHeartbeatWriter(deps.DB, nodeIdentity, mode, nodeURL)
}
if deps.DB != nil {
adminStatsProvider, statsErr := handlers.NewAdminStatsProvider(appCtx, deps.DB, deps.EventBus)
if statsErr != nil {
log.Fatalf("failed to create admin stats provider: %v", statsErr)
}
defer adminStatsProvider.Close()
deps.AdminStatsProvider = adminStatsProvider
}
// Wire recommendations engine, worker, and ratings repo if enabled.
var recEngine *recommendations.Engine
var recWorker *recommendations.Worker
if cfg.Recommendations.Enabled && deps.DB != nil {
deps.RatingsRepo = catalog.NewRatingsRepo(deps.DB)
recEngine = recommendations.NewEngine(
deps.DB,
deps.RatingsRepo,
catalog.NewItemRepository(deps.DB),
catalog.NewPersonRepository(deps.DB),
userStoreProvider,
cfg.Recommendations,
)
deps.Recommender = recEngine
deps.CatalogSearchVectorizer = recEngine
var err error
recWorker, err = recommendations.NewWorker(
recEngine,
cfg.Recommendations.EmbeddingsCron,
cfg.Recommendations.TasteProfilesCron,
cfg.Recommendations.CowatchCron,
cfg.Recommendations.RecommendationsCron,
cfg.Recommendations.EmbeddingsJobTimeout,
)
if err != nil {
slog.Error("failed to create recommendation worker", "error", err)
} else {
deps.RecWorker = recWorker
}
}
// Client IP resolver with trusted proxy config.
if err := clientip.SeedDefaults(ctx, settingsRepo); err != nil {
log.Fatalf("seed clientip defaults: %v", err)
}
trustedCIDRs, err := clientip.LoadTrustedCIDRs(ctx, settingsRepo)
if err != nil {
log.Fatalf("load trusted CIDRs: %v", err)
}
ipResolver = clientip.NewResolver(trustedCIDRs)
deps.ClientIPResolver = ipResolver
// Hot-reload trusted proxies on settings changes via two complementary
// paths. The direct event-bus subscription re-reads only the clientip key,
// so a malformed unrelated setting (which fails the whole-config reload)
// cannot leave stale trust CIDRs on Redis-backed multi-instance deploys.
_ = eventBus.Subscribe(appCtx, cache.ChannelAdmin, func(event cache.Event) {
if event.Type != cache.EventSettingsChanged {
return
}
cidrs, loadErr := clientip.LoadTrustedCIDRs(context.Background(), settingsRepo)
if loadErr != nil {
slog.WarnContext(context.Background(), "clientip config reload failed", "component", "app", "error", loadErr)
return
}
ipResolver.UpdateTrustedCIDRs(cidrs)
})
// The config watcher covers the Redis-less poll/RequestReload path, so
// admin UI edits apply without a restart on single-node deployments too.
configWatcher.OnChange(func(old, updated *config.Config) {
if old != nil && old.ClientIP.TrustedProxies == updated.ClientIP.TrustedProxies {
return
}
raw := updated.ClientIP.TrustedProxies
if raw == "" {
raw = clientip.DefaultTrustedProxies
}
cidrs, parseErr := clientip.ParseCIDRs(raw)
if parseErr != nil {
slog.WarnContext(context.Background(), "clientip config reload failed", "component", "app", "error", parseErr)
return
}
ipResolver.UpdateTrustedCIDRs(cidrs)
})
// Step 6b: Create rate limiter.
if cfg.RateLimit.Enabled && deps.DB != nil {
var perKeyLimiter, globalLimiter ratelimit.RateLimiter
isMemory := true
if cfg.RateLimit.Backend == "redis" {
redisClient, redisErr := cache.NewRedisClient(cfg.Redis)
if redisErr != nil {
log.Fatalf("failed to create Redis client for rate limiting: %v", redisErr)
}
if redisClient != nil {
perKeyLimiter = ratelimit.NewRedisLimiter(redisClient)
globalLimiter = ratelimit.NewRedisLimiter(redisClient)
isMemory = false
defer redisClient.Close()
}
}
if isMemory {
perKeyLimiter = ratelimit.NewMemoryLimiter()
globalLimiter = ratelimit.NewMemoryLimiter()
}
defer perKeyLimiter.Close()
defer globalLimiter.Close()
rateLimitMW := ratelimit.NewMiddleware(perKeyLimiter, globalLimiter, settingsRepo, isMemory)
if err := rateLimitMW.Init(context.Background()); err != nil {
log.Fatalf("failed to init rate limiter: %v", err)
}
// Subscribe for multi-instance reload (only fires if EventBus is Redis-backed)
_ = eventBus.Subscribe(appCtx, cache.ChannelAdmin, func(event cache.Event) {
if event.Type == cache.EventSettingsChanged {
if reloadErr := rateLimitMW.Reload(context.Background()); reloadErr != nil {
slog.Warn("rate limit config reload from event failed", "error", reloadErr)
}
}
})
deps.RateLimitMW = rateLimitMW
}
// Activity log writer + consumer.
if err := activitylog.SeedDefaults(ctx, settingsRepo); err != nil {
log.Fatalf("seed activitylog defaults: %v", err)
}
// Seed default page sections for home and existing libraries.
sectionRepo := sections.NewRepository(pool)
var folders []*models.MediaFolder
if deps.FolderRepo != nil {
var listErr error
folders, listErr = deps.FolderRepo.List(ctx)
if listErr != nil {
log.Fatalf("list libraries for section defaults: %v", listErr)
}
}
if err := sectionRepo.SeedDefaults(ctx, "home", nil, sections.DefaultHomeSections(folders)); err != nil {
log.Fatalf("seed home section defaults: %v", err)
}
if deps.FolderRepo != nil {
for _, f := range folders {
id := f.ID
if seedErr := sectionRepo.SeedDefaults(ctx, "library", &id, sections.DefaultLibrarySectionsForType(&id, f.Type)); seedErr != nil {
slog.Warn("seed library section defaults", "library_id", id, "error", seedErr)
}
}
}
activityPM := partman.NewManager(pool, "activity_log", partman.Weekly, 2)
if err := activityPM.EnsureFuturePartitions(appCtx); err != nil {
// Non-fatal: see the operational_logs partition incident. Writes fall
// back to the default partition and periodic cleanup retries.
slog.Warn("ensure activity log partitions; continuing in degraded mode", "error", err)
}
policyPM := partman.NewManager(pool, "policy_decisions", partman.Daily, 3)
if err := policyPM.EnsureFuturePartitions(appCtx); err != nil {
// Non-fatal: decision logs fall back to the default partition and
// periodic cleanup retries partition creation.
slog.Warn("ensure policy decision log partitions; continuing in degraded mode", "error", err)
}
var activityWriter activitylog.Writer
activityConsumer := activitylog.NewConsumer(pool, nil, logStreamHub)
if cfg.Redis.URL != "" {
actRedisClient, actRedisErr := cache.NewRedisClient(cfg.Redis)
if actRedisErr == nil && actRedisClient != nil {
activityWriter = activitylog.NewRedisWriter(actRedisClient)
activityConsumer = activitylog.NewConsumer(pool, actRedisClient, logStreamHub)
go activityConsumer.RunRedis(appCtx)
defer actRedisClient.Close()
}
}
if activityWriter == nil {
memWriter := activitylog.NewMemoryWriter(10000)
activityWriter = memWriter
go activityConsumer.RunMemory(appCtx, memWriter.Chan())
}
deps.ActivityLogWriter = activityWriter
deps.ActivityLogRepo = activitylog.NewRepo(pool)
deps.NodeID = nodeID
// Create refresh worker early so the task manager can use it for FindCandidates.
var refreshWorker *worker.RefreshWorker
var personRefreshWorker *worker.PersonRefreshWorker
if needsWorkers && deps.DB != nil {
refreshWorker = worker.NewRefreshWorker(deps.DB)
if deps.PersonRefreshQueue != nil {
personRefreshWorker, _ = deps.PersonRefreshQueue.(*worker.PersonRefreshWorker)
}
}
// Construct collection service for both the router and the collection sync scheduler.
var collectionSyncScheduler *catalog.CollectionSyncScheduler
var userCollectionScheduler *usercollections.Scheduler
var trendingRefresher *sections.TrendingRefresher
if needsWorkers && deps.DB != nil {
collectionRepo := catalog.NewLibraryCollectionRepository(deps.DB)
collItemRepo := catalog.NewItemRepository(deps.DB)
libraryItemRepo := catalog.NewLibraryItemRepository(deps.DB)
collectionService := catalog.NewLibraryCollectionService(collectionRepo, collItemRepo, libraryItemRepo, nil)
collectionService.TMDBCollections = api.NewTMDBCollectionFetcher(cfg.TMDBAPIKey)
deps.CollectionService = collectionService
collectionSyncScheduler = catalog.NewCollectionSyncScheduler(collectionRepo, collectionService, slog.Default())
// The trending refresher reuses the section repo (to find used source/
// window combos), a snapshot repo, an item repo (external-ID matching),
// and the TMDB fetcher. The Trakt fetcher needs settingsRepo and is
// propagated onto deps.TrendingRefresher later in router.go.
trendingRefresher = sections.NewTrendingRefresher(
sectionRepo,
sections.NewTrendingSnapshotRepository(pool),
catalog.NewItemRepository(deps.DB),
collectionService.TMDBCollections,
collectionService.TraktCollections,
)
deps.TrendingRefresher = trendingRefresher
if deps.UserStoreProvider != nil {
userSync := usercollections.NewService(deps.UserStoreProvider, collItemRepo, libraryItemRepo, nil, slog.Default())
userSync.TMDBCollections = collectionService.TMDBCollections
// Trakt fetchers are wired in router.go (they need settingsRepo);
// router.go propagates them onto userSync once configured.
userCollectionScheduler = usercollections.NewScheduler(deps.DB, userSync, slog.Default())
deps.UserCollectionSync = userSync
deps.UserCollectionScheduler = userCollectionScheduler
deps.MDBListClient = mdblist.NewClient(cfg.MDBListAPIKey, nil)
mdblistForReload := deps.MDBListClient
configWatcher.OnChange(func(_, updated *config.Config) {
mdblistForReload.SetAPIKey(updated.MDBListAPIKey)
})
}
}
// White-label branding: one service shared by the API (public read + admin
// upload), the frontend handler (index.html title, favicon, manifest), and
// the artwork reconcile task. S3 is optional — pass a nil AssetStore (not
// the typed-nil *s3client.Client) when it isn't configured so text branding
// still works without it.
var brandingStore branding.AssetStore
if deps.S3Public != nil {
brandingStore = deps.S3Public
}
brandingSvc := branding.NewService(settingsRepo, brandingStore)
// Wire up task manager for admin task API.
if needsWorkers && deps.DB != nil {
triggerRepo := taskrepository.NewPgTriggerRepository(deps.DB)
historyRepo := taskrepository.NewPgExecutionRepository(deps.DB)
taskMgr := taskmanager.New(triggerRepo, historyRepo, triggers.New, slog.Default())
if deps.EventsHub != nil {
taskMgr.AddObserver(evt.NewTaskObserver(deps.EventsHub))
}
if deps.FolderRepo != nil && deps.LibraryScanQueue != nil {
taskMgr.Register(tasks.NewScanLibrariesTask(deps.FolderRepo, deps.LibraryScanQueue, deps.EventBus))
}
taskMgr.Register(tasks.NewCleanupOrphanedMediaItemsTask(catalog.NewOrphanedProvisionalCleaner(deps.DB)))
taskMgr.Register(tasks.NewBackfillMediaItemAliasesTask(catalog.NewItemAliasRepository(deps.DB)))
if deps.S3Public != nil {
taskMgr.Register(tasks.NewCleanupArtworkRevisionsTask(
metadata.NewArtworkRevisionGarbageCollector(deps.DB, deps.S3Public),
))
}
catalogSearchIndexer := catalog.NewCatalogSearchIndexer(deps.DB, settingsRepo)
taskMgr.Register(tasks.NewSyncCatalogSearchIndexTask(catalogSearchIndexer))
taskMgr.Register(tasks.NewRebuildCatalogSearchIndexTask(catalogSearchIndexer))
if deps.IntroAnalyzer != nil {
taskMgr.Register(tasks.NewDetectIntroMarkersTask(deps.IntroAnalyzer, settingsRepo))
}
if deps.MarkerContributionService != nil && deps.MarkerProviderConfig != nil && deps.MarkerContributionStore != nil && deps.FileRepo != nil {
taskMgr.Register(tasks.NewContributeMarkersTask(
deps.MarkerContributionService, deps.MarkerProviderConfig, deps.MarkerContributionStore, deps.FileRepo,
))
}
if chapterBackfiller, ok := deps.ChapterThumbnailQueuer.(*chapterthumbs.Service); ok {
taskMgr.Register(tasks.NewChapterThumbnailBackfillTask(chapterBackfiller, 25))
}
taskMgr.Register(tasks.NewActivityLogCleanupTask(deps.DB, settingsRepo, activityPM))
taskMgr.Register(tasks.NewOperationalLogCleanupTask(deps.DB, settingsRepo, opsPM))
var diagnosticsStore diagnostics.ObjectStore
if deps.S3Private != nil {
diagnosticsStore = diagnostics.NewS3ObjectStore(deps.S3Private)
}
taskMgr.Register(tasks.NewClientDiagnosticsCleanupTask(
diagnostics.NewPostgresRepository(deps.DB),
settingsRepo,
diagnosticsStore,
))
taskMgr.Register(tasks.NewPolicyDecisionLogCleanupTask(deps.DB, settingsRepo, policyPM))
if deps.FileRepo != nil {
// Download prepare-to-file pipeline (Phase 3): a durable, leased encode
// queue hosted on the task manager. Built here (before Start) and shared
// with the API via deps so the download service can enqueue jobs.
artifactMgr := downloads.NewArtifactManager(
downloads.NewArtifactRepository(deps.DB),
downloads.NewRepository(deps.DB),
deps.FileRepo,
downloads.NewPlaybackPreparer(),
deps.NodeID,
func() *config.Config {
if deps.LiveConfig != nil {
if c := deps.LiveConfig(); c != nil {
return c
}
}
return deps.Config
},
func(ctx context.Context, d *downloads.Download) {
if deps.EventsHub == nil {
return
}
_ = deps.EventsHub.PublishJSON(ctx, evt.ChannelUserState, "download", map[string]any{
"download_id": d.ID,
"status": d.Status,
"media_item_id": d.ContentID,
"format": d.Format,
}, evt.PublishOptions{UserID: d.UserID, ProfileID: d.ProfileID})
},
)
encodeTask := tasks.NewEncodeDownloadArtifactsTask(artifactMgr)
artifactMgr.SetKick(func() { _ = taskMgr.RunTask(appCtx, encodeTask.Key()) })
taskMgr.Register(encodeTask)
deps.ArtifactManager = artifactMgr
}
if notificationSystem != nil {
taskMgr.Register(tasks.NewSeedContentAvailabilityTask(notificationSystem))
taskMgr.Register(tasks.NewRebuildReleaseInterestTask(notificationSystem))
taskMgr.Register(tasks.NewNotificationsRetentionTask(notificationSystem))
}
if userStoreProvider != nil {
taskMgr.Register(tasks.NewSettingMutationsRetentionTask(userstore.NewSettingMutationSweeper(
auth.NewUserRepository(deps.DB), userStoreProvider,
)))
}
if matchWorker != nil {
taskMgr.Register(tasks.NewMatchMediaTask(matchWorker))
}
if refreshWorker != nil && metadataService != nil {
taskMgr.Register(tasks.NewRefreshMetadataTask(refreshWorker, metadataService))
}
if metadataImageCacheProcessor != nil {
taskMgr.Register(tasks.NewCacheMetadataImagesTask(metadataImageCacheProcessor))
}
if deps.S3Public != nil {
identity := tasks.ArtworkStorageIdentity(cfg.S3.Public.Endpoint, cfg.S3.Public.Bucket, cfg.S3.Public.KeyPrefix)
// Seed the fingerprint on first boot so an unchanged storage
// identity never triggers a sweep. On the boot after a provider
// change the stored (old) identity survives this call and the
// startup trigger runs the reconcile.
if _, err := settingsRepo.SetIfAbsent(appCtx, tasks.ArtworkStorageIdentityKey, identity); err != nil {
slog.Warn("artwork reconcile: seeding storage identity failed", "error", err)
}
var brandingReconciler tasks.BrandingAssetReconciler
if brandingSvc != nil && brandingSvc.HasStorage() {
brandingReconciler = brandingSvc
}
taskMgr.Register(tasks.NewReconcileArtworkCacheTask(
metadata.NewArtworkCacheReconciler(deps.DB, deps.S3Public),
settingsRepo,
brandingReconciler,
identity,
))
}
if pluginAutoUpdater != nil {
taskMgr.Register(tasks.NewCheckPluginUpdatesTask(pluginAutoUpdater))
}
if collectionSyncScheduler != nil {
taskMgr.Register(tasks.NewSyncCollectionsTask(collectionSyncScheduler))
}
if trendingRefresher != nil {
taskMgr.Register(tasks.NewRefreshTrendingDiscoverTask(trendingRefresher))
}
if userCollectionScheduler != nil {
taskMgr.Register(tasks.NewSyncUserCollectionsTask(userCollectionScheduler))
}
if watchProviderService != nil {
taskMgr.Register(tasks.NewSyncWatchProvidersTask(watchProviderService))
}
requestReconcileSvc := mediarequests.NewService(
mediarequests.NewRepository(deps.DB, deps.SecretCipher),
nil,
mediarequests.NewCatalogPresence(
catalog.NewItemRepository(deps.DB),
catalog.NewProviderIDRepository(deps.DB),
),
)
requestReconcileSvc.SetRequesterIdentityResolver(plugins.RequesterIdentityFromLookup(plugins.NewPgUserIdentityLookup(deps.DB)))
api.AttachRequestRouter(requestReconcileSvc, pluginService)
requestReconcileSvc.SetGroupPolicyProvider(accessGroupStore)
if userStoreProvider != nil {
userRepo := auth.NewUserRepository(deps.DB)
profileTokens := access.NewProfileTokenService(cfg.Auth.JWTSecret, 0)
var reconcileResolver scopeResolver
if policySystem != nil {
reconcileResolver = policy.NewViewerResolver(userRepo, userStoreProvider, profileTokens, policySystem.PDP(), accessGroupStore)
} else {
// Legacy resolver: proxy/test wiring without a policy system. Production integrated/api modes always take the policy path. Removed with the legacy cleanup phase.
reconcileResolver = access.NewResolver(userRepo, userStoreProvider, profileTokens, accessGroupStore)
}
requestReconcileSvc.SetEntitlementResolver(scopeEntitlementResolver{resolver: reconcileResolver})
}
if notificationSystem != nil {
requestReconcileSvc.SetFulfillmentNotifier(notifications.NewRequestFulfillmentNotifier(notificationSystem))
}
taskMgr.Register(tasks.NewReconcileRequestsTask(requestReconcileSvc, 100))
if deps.FolderRepo != nil && deps.LibraryScanQueue != nil && pluginService != nil && pluginInstallationStore != nil {
autoscanRepo := autoscan.NewRepository(deps.DB, deps.SecretCipher)
if err := autoscanRepo.MarkInterruptedEvents(appCtx); err != nil {
slog.Warn("autoscan: failed to mark interrupted polls", "err", err)
}
autoscanSvc := api.BuildAutoscanService(
autoscanRepo,
pluginService,
pluginInstallationStore,
mediarequests.NewRepository(deps.DB, deps.SecretCipher),
deps.FolderRepo,
deps.LibraryScanQueue,
deps.RedisClient,
)
// The poll task's default interval seeds the schedule from the stored
// settings (DefaultPollIntervalSeconds); per-cycle gating still runs
// off the live settings inside PollOnce. Seed in MILLISECONDS as
// seconds*1000 — the SAME computation HandleUpdateSettings uses to
// reschedule — so startup and reschedule agree for sub-minute and
// non-60-multiple intervals (the old seconds/60 minutes path diverged).
var intervalMs int64 = 10 * 60 * 1000
if settings, serr := autoscanRepo.GetSettings(appCtx); serr == nil && settings.DefaultPollIntervalSeconds > 0 {
intervalMs = int64(settings.DefaultPollIntervalSeconds) * 1000
}
taskMgr.Register(tasks.NewAutoscanPollTask(autoscanSvc, intervalMs))
taskMgr.Register(tasks.NewAutoscanWebhookRetryTask(autoscanSvc))
}
reconcileProviderIDRepo := catalog.NewProviderIDRepository(deps.DB)
reconcileEpisodeRepo := catalog.NewEpisodeRepository(deps.DB)
historyResolver := watchstate.NewStableIdentityResolver(nil, reconcileEpisodeRepo, reconcileProviderIDRepo)
historyReconciler := watchstate.NewHistoryReconciler(deps.DB, historyResolver)
taskMgr.Register(tasks.NewRepairProviderIDIntegrityTask(metadata.NewProviderIDIntegrityRepairer(deps.DB), historyReconciler))
taskMgr.Register(tasks.NewReconcileWatchHistoryTask(historyReconciler))
taskMgr.Register(tasks.NewSyncPodcastFeedsTask(podcastfeed.New(), podcastfeed.NewDBStore(deps.DB)))
if audiobookEnricher != nil {
taskMgr.Register(tasks.NewSyncAudiobookMetadataTask(audiobookEnricher))
}
if ebookEnricher != nil {
taskMgr.Register(tasks.NewSyncEbookMetadataTask(ebookEnricher))
taskMgr.Register(tasks.NewBackfillEbookMetadataTask(ebookEnricher))
}
if mangaEnricher != nil {
taskMgr.Register(tasks.NewSyncMangaMetadataTask(mangaEnricher))
}
if pluginInstallationStore != nil && pluginRuntimeConfigStore != nil && pluginService != nil {
pluginTasks, err := plugins.NewTaskRegistryWithTypedResolver(pluginInstallationStore, pluginRuntimeConfigStore, pluginService).Tasks(appCtx)
if err != nil {
log.Fatalf("plugin task registry: %v", err)
}
for _, pluginTask := range pluginTasks {
taskMgr.Register(pluginTask)
}
}
taskMgr.Start(appCtx)
defer taskMgr.Stop()
deps.TaskManager = taskMgr
slog.Info("task manager started")
}
// Build the ABS-compatible REST + Socket.io handler when a DB pool is
// available. Routes are mounted at the root level by NewRouter (not under
// /api/v1/) so ABS clients resolve /login, /api/*, /abs/api/*, and
// /abs/socket.io/* without path prefix hacks.
if absCompatEnabled && deps.DB != nil {
absUserRepo := auth.NewUserRepository(deps.DB)
absSessionRepo := auth.NewSessionRepository(deps.DB)
absJWTService := auth.NewJWTService(
cfg.Auth.JWTSecret,
cfg.Auth.AccessTokenExpiry,
cfg.Auth.RefreshTokenExpiry,
)
configWatcher.OnChange(func(_, updated *config.Config) {
absJWTService.SetExpiries(updated.Auth.AccessTokenExpiry, updated.Auth.RefreshTokenExpiry)
})
absAuthSvc := auth.NewService(
auth.NewLocalProvider(absUserRepo, absSessionRepo),
absJWTService,
absSessionRepo,
absUserRepo,
nil, // invite codes: not needed for ABS compat
nil, // settings: not needed here
nil, // user store: not needed here
)
absItemRepo := catalog.NewItemRepository(deps.DB)
absEpisodeRepo := catalog.NewEpisodeRepository(deps.DB)
absSeasonRepo := catalog.NewSeasonRepository(deps.DB)
absPersonRepo := catalog.NewPersonRepository(deps.DB)
var absFileFetcher catalog.FileVersionFetcher
if deps.FileRepo != nil {
absFileFetcher = deps.FileRepo
}
absDetailSvc := catalog.NewDetailService(absItemRepo, absEpisodeRepo, absSeasonRepo, absPersonRepo, absFileFetcher)
if deps.ImageResolver != nil {
absDetailSvc.SetImageResolver(deps.ImageResolver)
}
var absScopeResolver scopeResolver
if policySystem != nil {
absScopeResolver = policy.NewViewerResolver(absUserRepo, userStoreProvider, nil, policySystem.PDP(), accessGroupStore)
} else {
absScopeResolver = access.NewResolver(absUserRepo, userStoreProvider, nil, accessGroupStore)
}
absHDeps := audiobooks.ABSHandlerDeps{
Pool: deps.DB,
Items: absItemRepo,
Files: deps.FileRepo,
Settings: settingsRepo,
Auth: &audiobooks.SiloCredValidator{
Auth: absAuthSvc,
Pool: deps.DB,
},
AccessResolver: audiobooks.NewABSAccessResolver(absUserRepo, userStoreProvider, absScopeResolver, accessGroupStore),
Recs: recommendations.NewRepo(deps.DB),
Detail: absDetailSvc,
SessionMgr: sessionMgr,
SessionSyncer: deps.SessionSyncer,
}
absH := audiobooksService.BuildABSHandler(absHDeps)
deps.ABSHandler = absH
}
_ = audiobooksService
if deps.DB != nil && pluginInstallationStore != nil && pluginRuntimeConfigStore != nil && deps.PluginService != nil {
userRepo := auth.NewUserRepository(deps.DB)
sessionRepo := auth.NewSessionRepository(deps.DB)
authBindings, err := pluginRuntimeConfigStore.ListAuthBindings(appCtx)
if err != nil {
log.Fatalf("list plugin auth bindings: %v", err)
}
for _, binding := range authBindings {
if binding == nil || !binding.Enabled {
continue
}
installation, err := pluginInstallationStore.GetByID(appCtx, binding.InstallationID)
if err != nil {
log.Fatalf("load plugin auth installation %d: %v", binding.InstallationID, err)
}
if !installation.Enabled {
continue
}
displayName := binding.CapabilityID
mode := "credentials"
iconURL := ""
capabilities, err := pluginInstallationStore.ListCapabilities(appCtx, binding.InstallationID)
if err == nil {
for _, capability := range capabilities {
if capability != nil && capability.Type == "auth_provider.v1" && capability.ID == binding.CapabilityID {
if name, ok := capability.Metadata["display_name"].(string); ok && strings.TrimSpace(name) != "" {
displayName = name
}
// auth_modes ["oauth2"] flips the login button into
// an OAuth-style "Sign in with X" path. Mode is "oauth"
// when oauth2 is the only declared mode; "credentials"
// when password is supported alongside or alone.
if rawModes, ok := capability.Metadata["auth_modes"].([]any); ok {
hasPassword := false
hasOAuth := false
for _, m := range rawModes {
switch m {
case "password":
hasPassword = true
case "oauth2":
hasOAuth = true
}
}
if hasOAuth && !hasPassword {
mode = "oauth"
}
}
if url, ok := capability.Metadata["icon_url"].(string); ok {
iconURL = url
}
break
}
}
}
// Generic OIDC and similar multi-instance plugins ship one binary
// but install once per IdP. Their admin SPA writes display_name
// + icon_url_path to runtime config so each install renders its
// own brand on the login page. Manifest values are the fallback.
if runtimeConfigs, err := pluginRuntimeConfigStore.ListGlobalConfigs(appCtx, binding.InstallationID); err == nil {
for _, rc := range runtimeConfigs {
switch rc.Key {
case "display_name":
if v, ok := rc.Value["value"].(string); ok && strings.TrimSpace(v) != "" {
displayName = v
}
case "icon_url_path":
if v, ok := rc.Value["value"].(string); ok && strings.TrimSpace(v) != "" {
iconURL = fmt.Sprintf("/api/v1/plugins/%d/assets/%s", binding.InstallationID, strings.TrimLeft(v, "/"))
}
}
}
}
deps.AuthProviders = append(deps.AuthProviders, auth.RegisteredProvider{
Info: auth.LoginProviderInfo{
ID: fmt.Sprintf("plugin:%d:%s", binding.InstallationID, binding.CapabilityID),
DisplayName: displayName,
Mode: mode,
Default: binding.DefaultLogin,
IconURL: iconURL,
InstallationID: binding.InstallationID,
},
Provider: auth.NewPluginProvider(
auth.PluginProviderConfig{
InstallationID: binding.InstallationID,
CapabilityID: binding.CapabilityID,
DisplayName: displayName,
AutoProvision: binding.AutoProvision,
},
sessionRepo,
userRepo,
deps.DB,
deps.PluginService,
),
})
}
}
// Step 7: Build HTTP router with all dependencies.
// compatServer is populated after the compat server is constructed below;
// the closure captures the pointer so revocation calls reach the live instance.
var compatServer *jellycompat.Server
deps.OnUserSessionsRevoked = func(ctx context.Context, userID int) {
if compatServer != nil {
compatServer.SessionStore().DeleteByUserID(userID)
}
}
distFS, fsErr := fs.Sub(siloweb.DistFS, "dist")
if fsErr != nil {
log.Fatalf("failed to create frontend FS: %v", fsErr)
}
deps.FrontendFS = distFS
server.WebDistFS = distFS
// Expose the branding service (constructed before the task manager) to the
// API and the frontend handler.
deps.BrandingService = brandingSvc
server.Branding = brandingSvc
router := api.NewRouter(deps)
// Step 8: Expose Prometheus metrics endpoint (not behind auth).
metricsMux := http.NewServeMux()
metricsMux.Handle("/metrics", promhttp.Handler())
metricsMux.Handle("/api/", router)
// ABS-compat is NOT mounted on the main listener — see the "ABS compat
// listener" block below. It binds its own port so the discovery probes
// (/ping, /healthcheck, /status, /init, /login, /socket.io) own the URL
// space without collision with silo's SPA fallback. Mirrors how the
// Jellyfin compat server is set up at :8096.
metricsMux.Handle("/", server.FrontendHandler())
// Step 9: Start background workers (if needed).
var sessionCleaner *worker.SessionCleaner
var adminJobRunner *adminjob.Runner
if needsWorkers && deps.DB != nil {
if reconciler == nil {
log.Fatal("reconciler must be initialized before starting workers")
}
reconciler.Start()
defer reconciler.Stop()
if heartbeatWriter != nil {
heartbeatWriter.Start()
defer heartbeatWriter.Stop()
}
// RefreshWorker is kept as a RefreshCandidateFinder for the task manager's
// RefreshMetadataTask but no longer runs its own background loop.
// Scanning is handled exclusively by the task manager's ScanLibrariesTask.
if personRefreshWorker != nil {
personRefreshWorker.Start()
defer personRefreshWorker.Stop()
}
sessionCleaner = worker.NewSessionCleaner(deps.DB, cfg.UserDB.StaleGraceSeconds)
sessionCleaner.EventBus = deps.EventBus
sessionCleaner.EventsHub = deps.EventsHub
sessionCleaner.Start()
defer sessionCleaner.Stop()
var templateBundleApplyExecutor interface {
ExecuteTemplateBundleApply(context.Context, adminjob.TemplateBundleApplyRequest, func(int, int, string)) (any, error)
}
if deps.CollectionService != nil {
collectionRepo := catalog.NewLibraryCollectionRepository(deps.DB)
itemRepo := catalog.NewItemRepository(deps.DB)
collectionHandler := handlers.NewLibraryCollectionHandler(
collectionRepo,
deps.CollectionService,
itemRepo,
4*time.Hour,
nil,
deps.S3Public,
)
collectionHandler.FrontendFS = deps.FrontendFS
collectionHandler.SectionRepo = sectionRepo
collectionHandler.FolderRepo = deps.FolderRepo
if collectionHandler.FolderRepo == nil {
collectionHandler.FolderRepo = catalog.NewFolderRepository(deps.DB)
}
templateBundleApplyExecutor = collectionHandler
}
adminJobRunner = adminjob.NewRunner(
adminjob.NewRepository(deps.DB),
catalogseed.NewService(deps.DB, catalog.NewPersonRepository(deps.DB), recommendations.NewRepo(deps.DB)),
deps.S3Private,
itemRefreshExecutor,
libraryRefreshExecutor,
adminjob.NewLibraryDeleteExecutor(deps.FolderRepo, sectionRepo,
librarySettingsCleaner(deps.DB, userStoreProvider)),
adminjob.NewImageCacheCleanupExecutor(deps.S3Public),
templateBundleApplyExecutor,
deps.RealtimeHub,
)
adminJobRunner.SetCancelRegistry(adminJobCancelRegistry)
adminJobRunner.Start()
defer adminJobRunner.Stop()
// Start recommendation worker if enabled (reuse worker created above).
if recWorker != nil {
recWorker.Start()
defer recWorker.Stop()
// Check if this is first run (no embeddings yet).
embCount, _ := recommendations.NewRepo(deps.DB).EmbeddingCount(appCtx)
if embCount == 0 {
slog.Info("first run detected, triggering initial embedding")
recWorker.RunEmbeddingsNow()
}
}
slog.Info("background workers started")
}
// Step 10: Create and start the HTTP server.
srv := &http.Server{
Addr: cfg.Server.Listen,
Handler: metricsMux,
ReadTimeout: 30 * time.Second,
WriteTimeout: 120 * time.Second,
IdleTimeout: 120 * time.Second,
}
var compatSrv *http.Server
if (mode == "integrated" || mode == "api") && cfg.JellyfinCompat.Enabled && cfg.JellyfinCompat.Listen != "" {
compatDeps := jellycompat.Dependencies{
Config: cfg,
AppContext: appCtx,
LiveConfig: configWatcher.Config,
DB: deps.DB,
SecretCipher: dataCipher,
ClientIPResolver: ipResolver,
NodePlanner: deps.NodePlanner,
JWTSecret: cfg.Auth.JWTSecret,
RecWorker: recWorker,
FrontendFS: deps.FrontendFS,
// Hand remote-transcode recipes to the shared recipe store so a dedicated
// transcode node that restarts can rebuild a jellycompat session.
RecipeNodeStore: noderecipe.NewStore(apiRedisClient, 0),
SessionSyncer: deps.SessionSyncer,
}
// Wire direct dependencies when DB is available.
if deps.DB != nil {
browseRepo := catalog.NewBrowseRepository(deps.DB)
itemRepo := catalog.NewItemRepository(deps.DB)
seasonRepo := catalog.NewSeasonRepository(deps.DB)
episodeRepo := catalog.NewEpisodeRepository(deps.DB)
providerIDRepo := catalog.NewProviderIDRepository(deps.DB)
personRepo := catalog.NewPersonRepository(deps.DB)
folderRepo := deps.FolderRepo
var fileFetcher catalog.FileVersionFetcher
if deps.FileRepo != nil {
fileFetcher = deps.FileRepo
}
detailSvc := catalog.NewDetailService(itemRepo, episodeRepo, seasonRepo, personRepo, fileFetcher)
detailSvc.SetFolderRepository(folderRepo)
detailSvc.SetGroupClaimRepository(catalog.NewGroupClaimRepository(deps.DB))
detailSvc.SetProbeEnsurer(deps.ProbeEnsurer)
detailSvc.SetChapterThumbnailQueuer(deps.ChapterThumbnailQueuer)
if deps.ImageResolver != nil {
detailSvc.SetImageResolver(deps.ImageResolver)
}
compatDeps.BrowseRepo = browseRepo
compatDeps.ItemRepo = itemRepo
compatDeps.SeasonRepo = seasonRepo
compatDeps.EpisodeRepo = episodeRepo
compatDeps.ProviderIDRepo = providerIDRepo
compatDeps.StableIdentityResolver = watchstate.NewStableIdentityResolver(itemRepo, episodeRepo, providerIDRepo)
compatDeps.DetailSvc = detailSvc
compatDeps.FolderRepo = folderRepo
compatDeps.SessionMgr = sessionMgr
compatDeps.UserStoreProvider = userStoreProvider
compatDeps.WatchCompletionObserver = deps.WatchCompletionObserver
compatDeps.SettingsRepo = settingsRepo
compatDeps.PersonRepo = personRepo
if watchProviderService != nil {
compatDeps.WatchScrobbler = watchProviderService
}
compatSearchService := catalog.NewCatalogSearchService(
appCtx,
settingsRepo,
itemRepo,
catalog.NewSearchIndexEventRepository(deps.DB),
deps.CatalogSearchVectorizer,
)
if compatSearchService != nil {
compatSearchService.StartCoverageRefresh(appCtx)
compatDeps.CatalogSearchProvider = compatSearchService.Provider()
// Latch the resolved provider for the package-level enqueue
// helpers (idempotent with the API router's latch; this also
// covers modes that wire jellycompat without the router).
activeSearchProvider := catalog.SearchProviderPostgres
if _, ok := compatSearchService.Provider().(*catalog.MeilisearchSearchProvider); ok {
activeSearchProvider = catalog.SearchProviderMeilisearch
}
catalog.SetActiveSearchIndexProvider(activeSearchProvider)
}
if deps.S3Public != nil {
compatDeps.PosterPresigner = deps.S3Public
compatDeps.S3Client = deps.S3Public
compatDeps.S3Bucket = deps.S3Public.Bucket()
}
if deps.FileRepo != nil {
compatDeps.FileResolver = deps.FileRepo
}
compatDeps.SubtitleRepo = subtitles.NewPgRepository(deps.DB, deps.SecretCipher)
// Construct auth service for jellycompat login.
userRepo := auth.NewUserRepository(deps.DB)
compatDeps.APIKeyValidator = auth.NewAPIKeyRepository(deps.DB)
compatDeps.APIKeyUserLoader = userRepo
compatDeps.ScanQueue = deps.LibraryScanQueue
sessionRepo := auth.NewSessionRepository(deps.DB)
jwtService := auth.NewJWTService(
cfg.Auth.JWTSecret,
cfg.Auth.AccessTokenExpiry,
cfg.Auth.RefreshTokenExpiry,
)
configWatcher.OnChange(func(_, updated *config.Config) {
jwtService.SetExpiries(updated.Auth.AccessTokenExpiry, updated.Auth.RefreshTokenExpiry)
})
provider := auth.NewLocalProvider(userRepo, sessionRepo)
compatDeps.AuthService = auth.NewService(provider, jwtService, sessionRepo, userRepo, nil, nil, nil)
// Access filter resolver for viewer-scoped library access.
// Backed by the shared access.Resolver so account-level library
// restrictions (users.library_ids), profile restrictions,
// user-disabled libraries, and rating/quality ceilings apply to
// the compat API exactly as they do to the native API.
if userStoreProvider != nil {
var compatScopeResolver jellycompat.ScopeResolver
if policySystem != nil {
compatScopeResolver = policy.NewViewerResolver(
userRepo,
userStoreProvider,
nil, // profile tokens unused: compat login already verifies PINs
policySystem.PDP(),
accessGroupStore,
)
} else {
// Legacy resolver: proxy/test wiring without a policy system. Production integrated/api modes always take the policy path. Removed with the legacy cleanup phase.
compatScopeResolver = access.NewResolver(
userRepo,
userStoreProvider,
nil, // profile tokens unused: compat login already verifies PINs
accessGroupStore,
)
}
compatDeps.AccessFilterFn = jellycompat.NewScopeAccessFilter(compatScopeResolver)
}
}
compat := jellycompat.NewServerWithDependencies(compatDeps)
compatServer = compat
compatTerminalRecoveryReady = compat.StartBackgroundTasks(context.Background())
compatSrv = compat.HTTPServer()
compatSrv.ReadTimeout = 30 * time.Second
compatSrv.WriteTimeout = 0
compatSrv.IdleTimeout = 120 * time.Second
}
// ABS-compat listener — dedicated http.Server bound to its own port
// (default :13378) that hosts the Audiobookshelf-compatible API.
// Mirrors the Jellyfin compat layout above. The ABS handler mounts
// onto a fresh chi router here so /ping, /healthcheck, /status, /login,
// /socket.io, etc. own the URL space at the root — no SPA fallback,
// no collision with silo's /api/v1.
var absSrv *http.Server
if (mode == "integrated" || mode == "api") && deps.ABSHandler != nil && cfg.AudiobookshelfCompat.Listen != "" {
absRouter := chi.NewRouter()
absRouter.Use(chimiddleware.Recoverer)
absRouter.Use(chimiddleware.Compress(5))
deps.ABSHandler.Mount(absRouter)
absSrv = &http.Server{
Addr: cfg.AudiobookshelfCompat.Listen,
Handler: absRouter,
ReadHeaderTimeout: 10 * time.Second,
ReadTimeout: 60 * time.Second,
WriteTimeout: 0,
IdleTimeout: 120 * time.Second,
}
}
// Run non-critical startup work in the background so it doesn't delay the
// HTTP listener from accepting connections. Steps run sequentially and stop
// early if the app context is cancelled (shutdown).
if len(backgroundInit) > 0 {
go func() {
start := time.Now()
for _, step := range backgroundInit {
if appCtx.Err() != nil {
return
}
func() {
defer func() {
if p := recover(); p != nil {
slog.Error("deferred startup init step panicked; continuing",
"panic", p, "stack", string(debug.Stack()))
}
}()
step(appCtx)
}()
}
slog.Info("deferred startup init completed", "steps", len(backgroundInit), "duration", time.Since(start))
}()
}
errCh := make(chan error, 3)
go func() {
slog.Info("HTTP server listening", "addr", cfg.Server.Listen)
if listenErr := srv.ListenAndServe(); listenErr != nil && listenErr != http.ErrServerClosed {
errCh <- fmt.Errorf("HTTP server error: %w", listenErr)
}
}()
if compatSrv != nil {
go func() {
slog.Info("Jellyfin compat server listening", "addr", compatSrv.Addr)
if listenErr := compatSrv.ListenAndServe(); listenErr != nil && listenErr != http.ErrServerClosed {
errCh <- fmt.Errorf("jellyfin compat server error: %w", listenErr)
}
}()
}
if absSrv != nil {
go func() {
slog.Info("ABS compat server listening", "addr", absSrv.Addr)
if listenErr := absSrv.ListenAndServe(); listenErr != nil && listenErr != http.ErrServerClosed {
errCh <- fmt.Errorf("abs compat server error: %w", listenErr)
}
}()
}
// Step 11: Wait for termination signal.
sigCh := make(chan os.Signal, 1)
signal.Notify(sigCh, syscall.SIGTERM, syscall.SIGINT)
defer signal.Stop(sigCh)
select {
case sig := <-sigCh:
appCancel()
slog.Info("received signal, shutting down", "signal", sig)
case <-restartReqCh:
appCancel()
slog.Info("server restart requested, shutting down")
case serverErr := <-errCh:
appCancel()
slog.Error("server error, shutting down", "error", serverErr)
}
// Step 12: Graceful shutdown sequence.
slog.Info("beginning graceful shutdown")
shutdownCtx, shutdownCancel := context.WithTimeout(context.Background(), 30*time.Second)
defer shutdownCancel()
// 1. Stop accepting new requests.
if shutdownErr := srv.Shutdown(shutdownCtx); shutdownErr != nil {
slog.Error("HTTP shutdown error", "error", shutdownErr)
}
if compatSrv != nil {
if shutdownErr := compatSrv.Shutdown(shutdownCtx); shutdownErr != nil {
slog.Error("jellyfin compat shutdown error", "error", shutdownErr)
}
}
if absSrv != nil {
if shutdownErr := absSrv.Shutdown(shutdownCtx); shutdownErr != nil {
slog.Error("abs compat shutdown error", "error", shutdownErr)
}
}
// 2. Clean up stale sessions.
if sessionCleaner != nil {
cleaned, cleanErr := sessionCleaner.CleanStale(shutdownCtx)
if cleanErr != nil {
slog.Error("stale session cleanup error", "error", cleanErr)
} else if cleaned > 0 {
slog.Info("cleaned stale sessions", "count", cleaned)
}
}
// 2b. Remove this node's heartbeat and sessions from shared state.
if heartbeatWriter != nil {
if err := heartbeatWriter.CleanupSelf(shutdownCtx); err != nil {
slog.Error("heartbeat cleanup error", "error", err)
}
}
// 3. Close user store provider.
if userStoreProvider != nil {
if closeErr := userStoreProvider.Close(); closeErr != nil {
slog.Error("user store provider close error", "error", closeErr)
}
}
// 4. (match worker is now managed by the task manager — no separate cancel needed)
// Suppress unused variable warnings for workers used only in deferred calls.
_ = reconciler
_ = heartbeatWriter
_ = refreshWorker
_ = adminJobRunner
slog.Info("server stopped")
}
// startStandaloneServer runs a standalone HTTP server for proxy/transcode modes.
// It listens on the given address, handles graceful shutdown on SIGTERM/SIGINT.
func startStandaloneServer(addr string, handler http.Handler) {
srv := &http.Server{
Addr: addr,
Handler: handler,
ReadTimeout: 30 * time.Second,
WriteTimeout: 0, // no timeout for long streams
IdleTimeout: 120 * time.Second,
}
errCh := make(chan error, 1)
go func() {
slog.Info("HTTP server listening", "addr", addr)
if err := srv.ListenAndServe(); err != nil && err != http.ErrServerClosed {
errCh <- fmt.Errorf("HTTP server error: %w", err)
}
}()
sigCh := make(chan os.Signal, 1)
signal.Notify(sigCh, syscall.SIGTERM, syscall.SIGINT)
select {
case sig := <-sigCh:
slog.Info("received signal, shutting down", "signal", sig)
case serverErr := <-errCh:
slog.Error("server error, shutting down", "error", serverErr)
}
shutdownCtx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := srv.Shutdown(shutdownCtx); err != nil {
slog.Error("HTTP shutdown error", "error", err)
}
slog.Info("server stopped")
}
// newS3ClientIfConfigured creates an S3 client only if the bucket name is
// configured. Returns nil if the bucket is empty (not configured).
func newS3ClientIfConfigured(cfg s3client.BucketConfig) *s3client.Client {
if cfg.Bucket == "" {
return nil
}
return s3client.NewClient(cfg)
}
func configureS3Clients(cfg *config.Config, deps *api.Dependencies) {
if s3Public := newS3ClientIfConfigured(s3client.BucketConfig{
Endpoint: cfg.S3.Public.Endpoint,
PublicEndpoint: cfg.S3.Public.ReadEndpoint,
Region: cfg.S3.Public.Region,
Bucket: cfg.S3.Public.Bucket,
KeyPrefix: cfg.S3.Public.KeyPrefix,
AccessKey: cfg.S3.Public.AccessKey,
SecretKey: cfg.S3.Public.SecretKey,
PathStyle: cfg.S3.Public.PathStyle,
URLAuth: cfg.S3.Public.URLAuth,
TokenSecret: cfg.S3.Public.TokenSecret,
TokenParam: cfg.S3.Public.TokenParam,
TokenTTL: cfg.S3.Public.TokenTTL,
}); s3Public != nil {
deps.S3Public = s3Public
slog.Info("S3 public assets client configured", "bucket", s3Public.Bucket())
// Allow browsers to fetch presigned client-facing assets directly from S3.
// Skip for public/token auth (e.g. Cloudflare R2) where CORS is managed externally.
if !s3Public.UsesExternalAuth() {
corsCtx, corsCancel := context.WithTimeout(context.Background(), 10*time.Second)
if corsErr := s3Public.SetBucketCORS(corsCtx, s3Public.Bucket(), []string{"*"}); corsErr != nil {
slog.Warn("failed to set CORS on public assets bucket", "error", corsErr)
}
corsCancel()
}
}
if s3Private := newS3ClientIfConfigured(s3client.BucketConfig{
Endpoint: cfg.S3.Private.Endpoint,
Region: cfg.S3.Private.Region,
Bucket: cfg.S3.Private.Bucket,
KeyPrefix: cfg.S3.Private.KeyPrefix,
AccessKey: cfg.S3.Private.AccessKey,
SecretKey: cfg.S3.Private.SecretKey,
PathStyle: cfg.S3.Private.PathStyle,
}); s3Private != nil {
deps.S3Private = s3Private
slog.Info("S3 private internal client configured", "bucket", s3Private.Bucket())
if !s3Private.UsesExternalAuth() {
corsCtx, corsCancel := context.WithTimeout(context.Background(), 10*time.Second)
if corsErr := s3Private.SetBucketCORS(corsCtx, s3Private.Bucket(), []string{"*"}); corsErr != nil {
slog.Warn("failed to set CORS on private assets bucket", "error", corsErr)
}
corsCancel()
}
}
if s3UserDB := newS3ClientIfConfigured(s3client.BucketConfig{
Endpoint: cfg.S3.UserDB.Endpoint,
Region: cfg.S3.UserDB.Region,
Bucket: cfg.S3.UserDB.Bucket,
KeyPrefix: cfg.S3.UserDB.KeyPrefix,
AccessKey: cfg.S3.UserDB.AccessKey,
SecretKey: cfg.S3.UserDB.SecretKey,
PathStyle: cfg.S3.UserDB.PathStyle,
}); s3UserDB != nil {
deps.S3UserDB = s3UserDB
slog.Info("S3 user-db client configured", "bucket", s3UserDB.Bucket())
}
}
type pluginImageResolverCapabilityStore interface {
ListEnabled(ctx context.Context) ([]*plugins.Installation, error)
ListCapabilities(ctx context.Context, installationID int) ([]*plugins.Capability, error)
}
func reloadPluginImageResolvers(
ctx context.Context,
store pluginImageResolverCapabilityStore,
resolver *metadata.PluginImageResolver,
service *plugins.Service,
) error {
if resolver == nil {
return nil
}
if store == nil || service == nil {
resolver.ReplaceSources(nil)
return nil
}
installations, err := store.ListEnabled(ctx)
if err != nil {
return fmt.Errorf("list enabled plugin installations: %w", err)
}
sort.Slice(installations, func(i, j int) bool {
if installations[i] == nil {
return false
}
if installations[j] == nil {
return true
}
return installations[i].ID < installations[j].ID
})
var registrations []metadata.PluginImageResolverSourceRegistration
for _, installation := range installations {
if installation == nil {
continue
}
// Builtin installations resolve in-process metadata providers only;
// registering them here would claim their capability id as a gRPC
// image-resolver scheme with no binary behind it.
if installation.IsBuiltin() {
continue
}
capabilities, err := store.ListCapabilities(ctx, installation.ID)
if err != nil {
return fmt.Errorf("list image resolver capabilities for installation %d: %w", installation.ID, err)
}
sort.Slice(capabilities, func(i, j int) bool {
if capabilities[i] == nil {
return false
}
if capabilities[j] == nil {
return true
}
if capabilities[i].Type != capabilities[j].Type {
return capabilities[i].Type < capabilities[j].Type
}
return capabilities[i].ID < capabilities[j].ID
})
for _, capability := range capabilities {
if capability == nil {
continue
}
switch capability.Type {
case sdkcapability.ImageResolver:
schemes, priority := imageResolverCapabilityConfig(capability)
if len(schemes) == 0 {
slog.WarnContext(ctx, "plugin image resolver capability has no valid schemes", "component", "app",
"installation_id", installation.ID,
"capability_id", capability.ID)
continue
}
for _, scheme := range schemes {
source := metadata.NewPluginClientSource(installation.ID, capability.ID, func(
ctx context.Context, installationID int, capabilityID string,
) (metadata.PluginMetadataClient, error) {
return service.ImageResolverClient(ctx, installationID, capabilityID)
})
registrations = append(registrations, metadata.PluginImageResolverSourceRegistration{
Scheme: scheme,
Source: source,
Kind: metadata.PluginImageResolverSourceExplicit,
Priority: priority,
InstallationID: installation.ID,
CapabilityID: capability.ID,
})
}
case sdkcapability.MetadataProvider:
scheme := strings.TrimSpace(capability.ID)
if !metadata.ValidImageResolverScheme(scheme) {
slog.WarnContext(ctx, "skipping legacy metadata image resolver with invalid scheme", "component", "app",
"installation_id", installation.ID,
"capability_id", capability.ID)
continue
}
source := metadata.NewPluginClientSource(installation.ID, capability.ID, func(
ctx context.Context, installationID int, capabilityID string,
) (metadata.PluginMetadataClient, error) {
return service.MetadataProviderClient(ctx, installationID, capabilityID)
})
registrations = append(registrations, metadata.PluginImageResolverSourceRegistration{
Scheme: scheme,
Source: source,
Kind: metadata.PluginImageResolverSourceLegacy,
InstallationID: installation.ID,
CapabilityID: capability.ID,
})
}
}
}
resolver.ReplaceSources(registrations)
slog.InfoContext(ctx, "reloaded plugin image resolvers", "component", "app", "sources", len(registrations))
return nil
}
func imageResolverCapabilityConfig(capability *plugins.Capability) ([]string, int) {
if capability == nil {
return nil, 0
}
meta := capabilityMetadataFields(capability.Metadata)
return metadataStringList(meta["schemes"]), metadataInt(meta["priority"])
}
func capabilityMetadataFields(raw map[string]any) map[string]any {
if raw == nil {
return nil
}
if nested, ok := raw["metadata"]; ok {
switch typed := nested.(type) {
case map[string]any:
return typed
}
}
return raw
}
func metadataStringList(value any) []string {
var out []string
switch typed := value.(type) {
case []string:
for _, item := range typed {
if scheme := strings.TrimSpace(item); metadata.ValidImageResolverScheme(scheme) {
out = append(out, scheme)
}
}
case []any:
for _, item := range typed {
text, ok := item.(string)
if !ok {
continue
}
if scheme := strings.TrimSpace(text); metadata.ValidImageResolverScheme(scheme) {
out = append(out, scheme)
}
}
}
return out
}
func metadataInt(value any) int {
switch typed := value.(type) {
case int:
return typed
case int32:
return int(typed)
case int64:
return int(typed)
case float64:
return int(typed)
case json.Number:
n, _ := typed.Int64()
return int(n)
default:
return 0
}
}
type markerPluginCapabilityStore interface {
ListEnabled(ctx context.Context) ([]*plugins.Installation, error)
ListCapabilities(ctx context.Context, installationID int) ([]*plugins.Capability, error)
}
type markerPluginRuntimeConfigStore interface {
ListGlobalConfigs(ctx context.Context, installationID int) ([]*plugins.RuntimeConfig, error)
PutGlobalConfig(ctx context.Context, installationID int, key string, value map[string]any) error
}
type markerLegacySettingsStore interface {
Get(ctx context.Context, key string) (string, error)
}
func reloadMarkerPluginProviders(
ctx context.Context,
registry *markers.Registry,
configStore *markers.ProviderConfigStore,
store markerPluginCapabilityStore,
runtimeConfigs markerPluginRuntimeConfigStore,
legacySettings markerLegacySettingsStore,
resolver *markers.PluginResolverAdapter,
) error {
if registry == nil {
return nil
}
var providers []markers.Provider
if store == nil || resolver == nil {
return registry.SetProviders(providers)
}
installations, err := store.ListEnabled(ctx)
if err != nil {
return fmt.Errorf("list enabled plugin installations: %w", err)
}
sort.Slice(installations, func(i, j int) bool {
if installations[i] == nil {
return false
}
if installations[j] == nil {
return true
}
return installations[i].ID < installations[j].ID
})
nextPriority := 1000
for _, installation := range installations {
if installation == nil {
continue
}
// Builtin installations expose no marker providers; defense in depth
// alongside the capability-type filter below.
if installation.IsBuiltin() {
continue
}
capabilities, err := store.ListCapabilities(ctx, installation.ID)
if err != nil {
return fmt.Errorf("list marker provider capabilities for installation %d: %w", installation.ID, err)
}
sort.Slice(capabilities, func(i, j int) bool {
if capabilities[i] == nil {
return false
}
if capabilities[j] == nil {
return true
}
return capabilities[i].ID < capabilities[j].ID
})
for _, capability := range capabilities {
if capability == nil || capability.Type != sdkcapability.MarkerProvider {
continue
}
descriptor, err := plugins.DecodeCapability(capability)
if err != nil {
return fmt.Errorf("decode marker provider capability %d/%s: %w", installation.ID, capability.ID, err)
}
metadataMap := markerCapabilityMetadata(descriptor)
provider, err := markers.NewPluginProvider(markers.PluginProviderOptions{
InstallationID: installation.ID,
CapabilityID: capability.ID,
DisplayName: firstNonEmptyMarkerText(descriptor.GetDisplayName(), capability.ID),
PluginID: installation.PluginID,
RequiredExternalIDs: markers.PluginRequiredExternalIDsFromMetadata(metadataMap),
}, resolver)
if err != nil {
return err
}
providers = append(providers, provider)
priority := nextPriority
nextPriority++
if configuredPriority, ok := markers.PluginDefaultFetchPriorityFromMetadata(metadataMap); ok {
priority = configuredPriority
}
if configStore != nil {
defaultConfig := markers.ProviderConfig{
Provider: provider.ID(),
FetchEnabled: true,
FetchPriority: priority,
ContributeEnabled: false,
ContributeAutoLocal: false,
ContributeMinConfidence: 0.95,
}
if legacy, ok := legacyIntroDBProviderConfig(configStore, installation, capability, provider.ID()); ok {
defaultConfig = legacy
}
if err := configStore.Ensure(ctx, defaultConfig); err != nil {
return err
}
}
if err := copyLegacyIntroDBPluginConfig(ctx, runtimeConfigs, legacySettings, installation, capability); err != nil {
return err
}
}
}
return registry.SetProviders(providers)
}
func legacyIntroDBProviderConfig(
configStore *markers.ProviderConfigStore,
installation *plugins.Installation,
capability *plugins.Capability,
providerID string,
) (markers.ProviderConfig, bool) {
if configStore == nil ||
installation == nil ||
capability == nil ||
installation.PluginID != "silo.theintrodb" ||
capability.ID != "introdb" {
return markers.ProviderConfig{}, false
}
if _, exists := configStore.Get(providerID); exists {
return markers.ProviderConfig{}, false
}
legacy, ok := configStore.Get("introdb")
if !ok {
return markers.ProviderConfig{}, false
}
legacy.Provider = providerID
return legacy, true
}
func copyLegacyIntroDBPluginConfig(
ctx context.Context,
runtimeConfigs markerPluginRuntimeConfigStore,
legacySettings markerLegacySettingsStore,
installation *plugins.Installation,
capability *plugins.Capability,
) error {
if runtimeConfigs == nil ||
legacySettings == nil ||
installation == nil ||
capability == nil ||
installation.PluginID != "silo.theintrodb" ||
capability.ID != "introdb" {
return nil
}
configs, err := runtimeConfigs.ListGlobalConfigs(ctx, installation.ID)
if err != nil {
return fmt.Errorf("list TheIntroDB plugin config: %w", err)
}
for _, config := range configs {
if config != nil && config.Key == "account" {
return nil
}
}
apiKey, err := legacySettings.Get(ctx, "introdb.api_key")
if err != nil {
return fmt.Errorf("load legacy introdb.api_key: %w", err)
}
if strings.TrimSpace(apiKey) == "" {
return nil
}
if err := runtimeConfigs.PutGlobalConfig(ctx, installation.ID, "account", map[string]any{
"api_key": strings.TrimSpace(apiKey),
}); err != nil {
return fmt.Errorf("copy legacy introdb.api_key to plugin config: %w", err)
}
return nil
}
func markerCapabilityMetadata(descriptor *pluginv1.CapabilityDescriptor) map[string]any {
if descriptor == nil || descriptor.GetMetadata() == nil {
return nil
}
return descriptor.GetMetadata().AsMap()
}
func firstNonEmptyMarkerText(values ...string) string {
for _, value := range values {
if strings.TrimSpace(value) != "" {
return strings.TrimSpace(value)
}
}
return ""
}
// mapFolderTypeToMediaType maps silo's MediaFolder.Type values
// ("movies", "series", "mixed") to the SDK's MediaType values
// ("movie", "tv", "mixed"). Unknown values map to "mixed".
func mapFolderTypeToMediaType(t string) string {
switch t {
case "movies":
return "movie"
case "series":
return "tv"
default:
return "mixed"
}
}
type scopeResolver interface {
Resolve(ctx context.Context, input access.ResolveInput) (access.Scope, error)
}
type scopeEntitlementResolver struct {
resolver scopeResolver
}
func (r scopeEntitlementResolver) MaxPlaybackQuality(ctx context.Context, userID int, profileID string) (string, error) {
scope, err := r.resolveScope(ctx, userID, profileID)
if err != nil {
return "", err
}
return scope.MaxPlaybackQuality, nil
}
// MaxContentRating implements mediarequests.ContentRatingResolver so request
// discovery honors the profile's parental rating ceiling.
func (r scopeEntitlementResolver) MaxContentRating(ctx context.Context, userID int, profileID string) (string, error) {
scope, err := r.resolveScope(ctx, userID, profileID)
if err != nil {
return "", err
}
return scope.MaxContentRating, nil
}
func (r scopeEntitlementResolver) resolveScope(ctx context.Context, userID int, profileID string) (access.Scope, error) {
return r.resolver.Resolve(ctx, access.ResolveInput{
UserID: userID,
ProfileID: profileID,
SkipPINVerification: true,
})
}
// audiobooksSettingsAdapter bridges catalog.ServerSettingsRepo (which
// exposes Get) to the audiobooks.SettingsReader interface (which
// requires GetString). The two signatures are identical modulo name.
type audiobooksSettingsAdapter struct {
repo catalog.SettingsStore
}
func (a *audiobooksSettingsAdapter) GetString(ctx context.Context, key string) (string, error) {
return a.repo.Get(ctx, key)
}