Commit Graph
85 Commits
Author SHA1 Message Date
881c96864b feat(playback): finalize platform-neutral protocol v3 (#567)
* docs(playback): add v3 neutral-contract finalization plan

Supersedes the wire-contract sections of the 2026-07-12 v3 plan: server-owned
attempt keys, delivery-keyed negotiation without Media3 engine names, tiered
capability evidence, neutral device/output context, track/quality replan
operations, audio-only planning, and coordinated no-back-compat rollout
across server, Android, Apple, and web.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(playback): make v3 attempt keys server-owned and replace engines with deliveries

Contract core of the platform-neutral v3 finalization (plan sections 3.1
and 3.2), breaking on purpose — v3 is dark and all clients move together:

- Every PlanV3 now carries plan_attempt_key, an opaque server-computed
  token clients store and echo in attempted_plan_keys; ReplanRequestV3
  gains bounded local_mutations that the replan handler folds into the
  failed plan's key. Clients never hash anything.
- KotlinName() is deleted from DeliveryV3, StreamProtocolV3 and
  SubtitleModeV3; the attempt-key canonical string now uses lowercase
  wire tokens, and PlanRecipeVersionV3 bumps to v3.3 so no key or plan
  ID computed under the old canonicalization can collide.
- EngineV3 leaves the wire: ClientPlaybackContextV3.Engines (media3_*)
  becomes Deliveries keyed original_http|progressive|hls, with
  EngineCapabilityV3 renamed DeliveryCapabilityV3. PlanV3.Engine is
  removed; the planner, subtitle policy and quirk registry re-key on
  delivery class, and the media3_only feature token is deleted.
- Validated-claim strings drop the prefix: media3_h264_decode ->
  h264_decode, media3_audio_decode -> audio_decode.
- Golden fixtures in testdata/protocol_v3 are regenerated by Go and are
  now the cross-repo source of truth.

Part of the playback protocol v3 neutral-contract train (steps 2-3 of
docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(playback): add v3 evidence tiers and neutral device/output context

Implement plan sections 3.3 and 3.4 of the v3 neutral-contract pass:

- ClientCodecCapabilitiesV3 gains required video_evidence and
  audio_evidence closed enums (exact | platform_attested | declared).
  Planner strictness follows the tier: exact keeps the strict decode-entry
  validation, platform_attested validates codec/resolution/bit-depth/
  frame-rate but skips profile/level matching, declared grants copy routes
  from the flat codec lists. Only exact audio evidence earns passthrough
  claims. The detailed_decode_capabilities feature token is deleted
  (subsumed by video_evidence=exact), and evidence-blocked direct routes
  carry the new evidence_insufficient_for_direct reason/warning.

- DeviceContextV3 is now platform/os_version/manufacturer/model plus a
  bounded platform_details map (<=16 entries, <=128 chars); the Android
  Build dump fields are gone. Fire TV quirks keep matching on
  manufacturer/model (brand fallback removed with the field).

- output_route_generation (int64, dual-location) becomes an optional
  opaque output_context_id string on the output context; the dual-location
  consistency validation is deleted. Attempt keys, plan invalidation,
  route events, and the planstore column follow (new Goose migration).

- Feature advertisement collapses to the top-level client_features list
  only; ClientPlaybackContextV3.Features is deleted and ReplanRequestV3
  gains an optional client_features refresh.

- PlanRecipeVersionV3 bumped v3.3 -> v3.4; fixtures re-keyed.

Part of the playback protocol v3 neutral-contract finalization plan
(docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(playback): add v3 intent replans, quality menu, and audio-only routes

Protocol v3 could only replan after a failure, so changing the audio track
or the quality still required the legacy audio PATCH and the client-recipe
transcode start — the two endpoints v3 is meant to replace. Clients also had
to own a resolution ladder to render a quality menu, and a source with no
video track was terminaled by the video/HDR gates, keeping audiobooks on the
legacy path.

Add track_change and quality_change replan operations. They carry no failure
classification and route through the existing replan transaction, so they
inherit its idempotency, capacity reservation, and staged-successor commit
for free. Because nothing failed, the previous route stays eligible: neither
the attempted-key history nor the failed-plan exclusion applies to them.

Publish the server ladder on the plan as available_qualities so the quality
menu is server-owned; the rungs come from the same resolutionLabelV3 and
ladderBitrateKbpsV3 helpers the planner itself uses, not a parallel table.

Plan audio-only sources through their own reduced route family: original_http
when the client decodes the codec, otherwise a progressive AAC conversion.
The plan advertises audio/mp4 for that remux and the transport now serves the
same value, because a declared-tier client probes the advertised MIME with
isTypeSupported before attaching a source buffer, and "video/mp4" on a stream
with no video track is exactly the mismatch that makes the probe lie.

Name the protocol's string vocabulary (dynamic ranges, transformations,
executors, validated claims, terminal reasons) as constants while touching
these lines, so the wire values have one definition.

Part of #135

* docs(playback): publish the v3 protocol contract and fix subtitle ordinals

Protocol v3 exists only as Go code today, so the Android and Apple ports have
no authority to implement against other than reading this repository. Publish
the contract as a normative document, machine-checkable schemas, and generated
golden fixtures, and fix the one place where the server's own wire output
disagreed with the ordinal space it publishes.

- docs/architecture/playback-protocol-v3.md is self-contained enough for a
  third-party client: endpoints and status codes, evidence tiers and their
  bound-matching rules, delivery classes, the timeline model, replan
  semantics, registries, track identity, plan identity, quality, and
  transformations.
- docs/design/schemas/playback-v3/ carries JSON Schemas for the five wire
  shapes plus valid and invalid fixtures, following the client-diagnostics
  layout. internal/playback/contract validates every fixture against its
  schema, so a schema that drifts from the Go types fails the Go suite.
- cmd/playbackfixtures generates internal/playback/testdata/protocol_v3 from
  the production planner. `make playback-fixtures` writes them and
  `make verify-playback-fixtures` (wired into CI) fails when they are stale.
  These files are what the client ports consume, so drift would otherwise
  surface as a playback bug on three platforms at once.

The subtitle fix: combined ordinals are one dense space over externals, then
embedded tracks, then downloaded ones, but the legacy URL builder skipped
burn-in-only tracks while assigning indices, so every track after a DVD/DVB
track was numbered one too low and resolved to its neighbour. Ordinal
assignment now lives in playback.BuildSubtitleInventoryV3 and both the plan
inventory and the legacy `subtitle_urls` shape project from it; the legacy
shape still filters burn-in-only entries but keeps each track's real index.

Part of #135

* feat(web): migrate the players to the neutral playback v3 contract

The web player was the last client still speaking the legacy start
protocol: it picked its own file version from a codec probe, posted an
ffmpeg recipe to start a transcode, PATCHed an endpoint to change audio
tracks, and derived its own quality ladder. None of that survives a
server-owned plan, and none of it produced telemetry the apps could be
compared against.

Video player: starts with a v3 request that advertises `declared`
evidence from `isTypeSupported` probes and the three delivery classes,
then consumes the returned plan for its URL, timeline, tracks and
warnings. Quality and track changes become replans (`quality_change`,
`track_change`), the quality menu renders `available_qualities` instead
of computing rungs, and playback failures emit `route-events` so web
failures land in the same diagnostics as Android and Apple. The
duration comes from `source.duration_seconds` rather than the playback
engine, and the "how was this delivered" overlay reads the plan's
delivery and server transformations instead of comparing codec strings.

Audiobook player: starts against the audio-only planner path with a
single `original` rung, and takes its seek anchor from
`timeline.player_start_seconds` so the progressive-remux route (which
anchors the stream and restarts the player clock at zero) does not seek
twice.

Server side, `disable_progress_persistence` left the wire, so the rule
it encoded is now derived. Resume state is keyed on the item, but every
part of a multipart presentation shares that key while carrying its own
file-local clock — persisting part 4's position would store "12 minutes
in" as the book's resume point. `PresentationPartTotal > 1` expresses
that directly and generalizes to multipart movies and split episodes,
and a client can no longer forget to ask or lie about it.

`useTranscodeQuality` and the legacy response types are deleted, and
`WEBTEST_KNOWN_FAILURES` loses the audiobook entry along with its fix.

Part of #135

* feat(playback)!: make v3 the only playback protocol

Protocol v3 shipped behind a flag, alongside the legacy start path it was
designed to replace. Running both meant every planner change had to be made
twice, in two shapes that disagree about who decides the route: the legacy
body carried a decision the client had already made, while v3 asks the server
to make it. This deletes the legacy half.

Removed:

- `handleStartPlaybackLegacy` and its request/response bodies. The
  `POST /playback/start` route stays, but the protocol-version dispatch
  envelope is now a strict v3 decode — a body that does not declare
  `protocol_version: 3` gets `426 client_upgrade_required` so an outdated app
  can render a clear "update required" state instead of misreading a plan.
  Deliberately not a `400`: the request may be well-formed for the protocol it
  was written against.
- `POST /playback/transcode/start`, superseded by the `quality_change` replan
  operation, and `PATCH /playback/{session_id}/audio`, superseded by
  `track_change`. Both mutated a session without re-planning.
- The shadow planner and both rollout settings rows. With v3 the only
  protocol, `playback.protocol_v3_enabled` would mean "no playback at all";
  `playback.protocol_v3_shadow_enabled` gated a comparison against a path that
  no longer exists. `409 protocol_disabled` on route-events goes with them, and
  capability `enabled` is now constant `true` (the field stays — clients
  feature-detect against it).
- Version-selection helpers in `internal/playback/resolver.go` that only legacy
  start reached. `Resolve`/`ClientCapabilities`/`PlayDecision` stay: downloads
  consumes them. `internal/jellycompat` has its own resolution surface and is
  untouched.

Behaviour the legacy handlers owned and v3 now owns explicitly: series version
and audio-track preferences are persisted on start and on a `track_change`
replan (not on failure recovery, whose forced route is not a user choice); an
omitted `start_position` resolves to the profile's saved resume point; and an
omitted audio track resolves through the series preference, the profile audio
language, then the library override. Both are settled before planning, because
the plan's timeline is cut at the start position. Spec §2.2 documents this as
"omission is a request, not a default".

The encode-target clamp that lived in the deleted transcode handler is already
enforced in the planner, twice — `availableQualitiesV3` omits rungs at or above
the source height, and the encode path clamps `targetHeight` to it.

Unchanged: progress, stop, HLS manifest and segment delivery, the realtime
control socket, stream tokens and restart reconstruction, watch together,
downloads, jellycompat.

Every removal is recorded in the pre-lock removals table in
docs/architecture/v1-scope.md.

Part of #135

* fix(scanner): stop recording embedded cover art as a video track

ffprobe reports embedded cover art as a video stream carrying
disposition.attached_pic. convertProbeData appended every "video" stream
to VideoTracks without consulting isMainVideoStream, the predicate that
already existed for duration decisions, so the picture was persisted as a
playable track. That misreports the file twice:

  - An audio file with a cover picks up a video track, so it no longer
    satisfies MediaFile.IsAudioOnly and the v3 planner routes an
    audiobook through the video path instead of planAudioOnlyV3.
  - When the picture is ordered ahead of the real stream, the flat
    codec_video/resolution/hdr columns describe the poster: a 954x720
    h264 episode was stored as mjpeg 480x480.

Filter attached_pic streams out of the track loop. The guard is the
disposition flag, not the codec name, so a genuine MJPEG video is still
probed as video — the library has one.

Already-probed rows self-heal on the next playback: NeedsCriticalProbeRepair
already reprobes tracks missing color_range, which covers 21 of the 23
affected rows, and applyProbeData overwrites VideoTracks wholesale. The
remaining two need a rescan; nothing persisted records attached_pic, and
keying repair off still-image codec names would reprobe the genuine MJPEG
file on every playback forever.

Part of the playback v3 neutral-contract work: it is what lets Android
drop AUDIOBOOK_COVER_ART_CODECS, which fabricated decode support the
client cannot honestly claim under video_evidence: "exact".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(playback): publish subtitle URLs even when playback starts with subtitles off

The v3 plan's subtitle inventory is the authoritative track list a client builds
its subtitle menu from, but the handler only rewrote it with session-scoped URLs
when a track was actually selected. A start or replan that resolved to
`subtitle.mode: "off"` therefore returned the planner's URL-less inventory, so a
client whose picker reads the inventory had a menu it could not fetch anything
from. The Cast path hits this every time: it starts with subtitles off and needs
the receiver's text tracks up front.

attachSubtitleArtifactV3 now scopes and publishes the inventory unconditionally
and gates only the artifact stamping on the selection. Spec §8 records that the
`url` on a sidecar entry does not depend on the current selection.

Part of the v3 neutral-contract finalization.

* chore(playback): reconcile neutral v3 with main

* fix(playback): preserve subtitle intent across replans

* fix(playback): retain subtitle inventory on adapted routes

* fix(playback): software-decode High10 AVC for QSV

* fix(playback): scale High10 frames before QSV upload

* fix(playback): preserve empty subtitle inventories

* fix(playback): freeze terminal attempt contract

* chore(playback): name fixture contract tokens

* fix(playback): close v3 conformance review gaps

* chore(playback): name conformance category

* fix(playback): complete v3 conformance contract

* fix(playback): keep schema fixtures generated

* fix(playback): emit schema-valid conformance arrays

* fix(playback): omit empty replan failures

* fix(web): omit empty replan failures

* fix(playback): close neutral v3 contract gaps

* fix(playback): harden v3 replan, transcode, and quality-ladder edge cases

Review remediation for the neutral v3 cutover, server side:

- A failed replan no longer overwrites the durable StartResponse with a
  terminal or advances the replan request ID; an idempotent start replay
  of a still-healthy session returns the original plan.
- SoftwareVideoDecode is now derived inside the transcode layer from
  source facts (codec/profile/bit depth) carried on TranscodeOpts, so
  jellycompat, downloads, recipe-card reconstruction, and transcode
  nodes get the High10 software-decode fix, not just the v3 handler.
  video_to_h264 recipe version bumps to 2 so mixed-version node pools
  that would silently drop the flag fail validation instead.
- Local transport startup shares the 30s ManifestStartupTimeout; a
  timeout with the process still running stays retryable and is no
  longer persisted as a durable terminal against the attempt.
- Sparse replan bodies (failure_recovery et al) no longer reset a
  user-selected quality preference to auto; the empty-value guard now
  covers every operation.
- availableQualitiesV3 publishes no fixed rungs when the source height
  is unknown, keeping the no-upscaling ladder contract.
- The proxy remux path serves audio-only fMP4 as audio/mp4 via a new
  additive AudioOnly token claim, matching the integrated path.
- Plain text subtitle sidecars accept any requested extension again
  (served as VTT), restoring the permissive v1 behavior; ASS and bitmap
  handling is unchanged.
- The 4K-disallowed terminal message discloses when a lower-resolution
  alternate exists but was pinned away by quality "original".

Part of #135.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): keep playback alive through failed replans and honest audio claims

Review remediation for the neutral v3 cutover, web player:

- A failed or refused replan no longer unmounts the player: the fatal
  error screen is reserved for loads with no adopted plan, and replan
  failures surface through the existing non-fatal replanError path.
- changeQuality rolls its optimistic preference back when the replan is
  refused or errors, so a failed switch is not silently applied by the
  next unrelated replan and the menu shows the real active rung.
- The capability probe now tests mp3/vorbis codecs and mp3/flac/ogg
  containers (MediaSource with a canPlayType fallback), restoring
  direct play for mp3 audiobooks instead of per-part AAC re-encodes.
- Reanchor seeks issued while a replan is in flight coalesce and run
  when it settles instead of being silently dropped with the scrubber
  pinned to a phantom position.
- Subtitle refresh/translation replans use the resume anchor while the
  media element has no metadata, so a subtitle_ready broadcast during
  startup no longer restarts a resumed stream at 0:00.
- An exhausted failure-recovery chain sets a visible error instead of
  returning silently.

Part of #135.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): accept video-only and VP9 probe metadata

Treat audio and video probe completeness independently so legitimate video-only assets converge without repeated ffprobe repair. Allow unknown codec profile/level metadata to fall through to server adaptation while preserving exact direct-decode constraints.

Fixes #574

* fix(playback): address protocol v3 review findings

* fix(playback): harden lease and probe repair decisions

* fix(playback): close remaining v3 review gaps

* fix(playback): recover failed transcode starts

* fix(playback): address remaining review-bot findings on v3 replan and audio planning

Server:
- The deferred replan lease release is bounded by a 3s timeout so a
  saturated pool or DB outage cannot wedge a handler goroutine that
  holds the per-session store lock on an uncancellable context.
- planAudioOnlyV3 honors the request bandwidth cap: an over-cap source
  skips the original_http direct route and converts to AAC with the
  same bandwidth_cap_applied warning and decision reason the video
  ladder uses. Unknown source bitrate never triggers the cap.
- A copy-audio progressive plan rejected only by a per-delivery
  audio_decode_codecs subset retries as an AAC conversion instead of
  returning adaptation_unavailable, and the AAC recipe respects the
  delivery's max_channels.

Web:
- failure_recovery replans issued while another replan is in flight
  queue (superseding a pending seek reanchor) instead of being
  silently dropped with the fatal overlay already suppressed.
- A terminal response to a fresh non-preserving start clears the
  previous plan and stops its session, so episode navigation cannot
  keep rendering the prior item under the new title.
- A refused recovery replan for a transport-dead plan surfaces the
  error and re-arms the plan failure key, so transient recovery
  failures no longer strand an endless spinner; the audiobook player
  gets the same guard reset.
- A track-less subtitle_translation_completed hands off to the
  refreshed persisted track once the inventory settles, clearing the
  live overlay, instead of pinning the synthetic live track forever.

Part of #135.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): reuse HLS transport for sidecar replans

* fix(playback): stabilize copy HLS remount timeline

* fix(playback): address v3 review findings

* fix(playback): satisfy player contract types

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 18:14:49 -04:00
40a9de7f26 feat(watchsync): add plugin-backed providers (#475)
* feat(watchsync): add plugin-backed providers

* fix(watchsync): address plugin review findings

* fix(watchsync): harden plugin provider failures

* feat(watchsync): complete plugin provider contract

* fix(watchsync): address provider review feedback

* fix(watchsync): keep device state host-private

* fix(watchsync): build reconciliation index concurrently

* fix(watchsync): preserve empty device state updates

* chore(deps): use released watch-sync SDK

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-08-06 10:30:49 -04:00
QuickandGitHub 3bdfc58512 feat(settings): sync navigation and card customization by client family (#538)
* test(web): use safe auth placeholders

* feat(settings): sync navigation and card customization

* fix(settings): address customization review feedback

* fix(settings): address customization review feedback

* fix(settings): harden customization capability handling
2026-08-04 08:20:41 -04:00
QuickandClaude Fable 5 a28960819d fix(settings): make device setting controls work and fit the device
The device settings screen shipped with three defects:

- Sliders (audio sync, playback speed, subtitle sync, next-up prompt,
  sleep timer) were rendered controlled with only onValueCommit, so the
  thumb never moved and no gesture could change the value. A shared
  SettingSlider now keeps a draft during the drag and commits once on
  release; the admin RegistrySettingControl had the same bug and uses it
  too.
- "Change how they look" called an onOpenPanel the page never passed.
  The screen now opens the shared SubtitleAppearancePanelView scoped to
  the selected device and profile, with a reset that clears the row.
- Every setting was shown for every device, so a browser offered
  "Screen orientation" and native-only toggles. settingsgen now emits
  the manifest's advisory platforms field into the TypeScript contract,
  and the screen hides settings whose platforms exclude the target
  device — unless the device stores a value, which must stay clearable.
  audio_sync_ms, dolby_vision_enabled and dv_profile7_hdr10_fallback
  gain native-only tags (no web consumer exists); platforms is advisory
  UI metadata, so no revision bump.

Kotlin/Swift generators deliberately unchanged: the native clients
hardcode their own applicability today, and their vendored bindings are
regenerated from their own repos.

Part of #215

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 16:15:51 +00:00
QuickandGitHub d70b291bb8 feat(settings): add shared language option catalogs (#521) 2026-07-30 18:17:49 -04:00
dc4b9a0909 feat(settings): add the cross-platform settings contract and its manifest (#479)
* docs(settings): define the cross-platform settings contract

Turns the audit in #376 into a decision-complete design for how user settings
work across the server, bundled web client, Apple clients, and Android clients.

Today there are three partial contracts - the server registry, the web client's
own manifest, and independently owned key constants in each native client - and
they have measurably drifted. The root enabler is that keyUsesUserScope returns
true for any unregistered key, so a client can invent a production setting
unilaterally and the server stores it as an unvalidated string.

The design decides:

Ownership. Every production user-facing setting needs a server-owned manifest
entry, even when the value is stored only on one client. The single exception is
private local.<client>.* diagnostics, bounded by five conditions.

Types and scopes. Native JSON values instead of strings. Five remote scopes plus
client_local, and each definition declares its own resolution order rather than
inheriting a global precedence.

Preferences versus restrictions. internal/policy already resolves
max_playback_quality and metadata-language limits over the same controls this
contract resolves preferences for. Definitions declare constrained_by, the
effective response reports the permitted value alongside the user's stored one,
and a mutation exceeding a restriction is stored rather than rejected - a capped
4K preference should take effect the day the cap lifts, not be destroyed by it.

Compatibility. Widening a scope, adding an enum member, or widening a range is
additive and revision-tagged; narrowing anything needs a new key. introduced_in
is a manifest revision attached to individual enum members and scopes, not just
whole definitions, so a newer client never offers a choice an older server will
reject.

Rollout. One coordinated breaking release, with no compatibility shim,
projection, or client fallback. After the cutover no future setting requires
coordination. No settings version check goes in the authenticated middleware and
nothing returns 426: deleting the old routes already produces the break, and a
gate would be more code in four repos for the same outcome while permanently
coupling every endpoint to one subsystem's versioning.

Scope placement. Appearance and date/time move from account to profile scope.
Account scope was an artifact of pre-profile storage; leaving it there means a
household shares one theme and text size, and any non-child profile can restyle
everyone else.

Read path. Batched context resolution, index requirements, a session-snapshot
rule, and a no-regression benchmark gating storage consolidation - profile_series
resolution is per-item, so a season view would otherwise issue one request per
episode.

Verified against the current server, Apple, and Android implementations. Two
findings shape it: the unknown-key extension bag is real, and v1 scope reads NOT
LOCKED, so removing the legacy surface needs no amendment if it lands before
lock.

Related to #376.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the canonical settings contract manifest

First implementation step for the cross-platform settings contract (#376).
Adds the artifact everything else depends on: the manifest, its JSON Schema,
the object value schemas, and a Go loader that validates the whole thing at
load time. No routes, no storage, no behavior change — nothing reads this yet.

contracts/settings/v1/ holds the artifact at a stable path because clients
vendor it and generate bindings from it. The embed directive has to sit beside
it (go:embed cannot reach outside its own directory), so that directory is a
tiny Go package containing nothing else; loading and validation live in
internal/settingscontract.

38 definitions: 35 remote, 3 contract-known client_local. That covers every key
the legacy registry accepts, every unregistered key the extension bag was
silently accepting from the web client, every unregistered device key Android
writes, and the profile preference columns that become settings.

Registering the previously-unregistered keys is where the drift shows up, and
the manifest records each case in a notes field:

- ui_theme, ui_text_scale, ui_text_weight, ui_high_contrast,
  ui_custom_theme_vars, and ui_custom_css reached the server only because
  keyUsesUserScope returns true for any unregistered key. They are now typed,
  renamed to the dotted convention every other key uses, and moved to profile
  scope per the design.
- player.match_frame_rate and player.sleep_timer_default_minutes are written by
  Android against a server that does not register them, so every write and reset
  is currently rejected. Registered.
- player.next_up_prompt_seconds is Android's alias for
  playback.next_up_prompt_seconds and does not become a definition; the test
  matrix pins it as a migration alias.
- player.playback_speed is capped at 3.0, matching the server rather than
  Android's 4.0.
- subtitle_appearance becomes playback.subtitle_appearance. Every other
  canonical key carries a domain prefix, and preserving accidental key names is
  an explicit non-goal of the design.

Validation is deliberately stricter than the schema can express. Beyond shape,
it enforces that a resolution order ends in "default", that it only resolves
scopes the definition allows, and — the one most likely to bite — that every
writable scope is actually read, so a setting cannot accept writes at a scope it
will never honor. Defaults are validated against their own value schema, so a
default that violates its own range or enum fails at load. Revision tags are
checked to never run ahead of the manifest revision, which is what makes
revision-aware client filtering trustworthy. Ceiling and floor policy
constraints are rejected on unordered types, where capping would silently do
nothing; playback.preferred_quality's enum is therefore ordered ascending.

ValidateValue is the single validation path, so the mutation endpoint, the
migration, and the manifest's own default checks cannot diverge later. Numbers
decode through json.Number so an integer setting rejects 30.5 rather than
truncating, and object values validate against their referenced JSON Schema
instead of accepting arbitrary JSON the way validateJSONSetting does today.

Canonicalization implements RFC 8785 over the value domain the contract uses:
sorted keys, no insignificant whitespace, ECMAScript number formatting. The
digest is the ETag, and PublicBytes strips maintainer notes so the served
manifest never carries internal commentary.

Promotes santhosh-tekuri/jsonschema/v6 from indirect to direct.

Verification: 124 tests pass across 16 cases; golangci-lint clean;
make verify-local-paths passes. Two failures in internal/api/handlers
(TestRemoveJellyfinCompatWebDisablesWebSetting, the playback v3 seek recovery
test) reproduce unchanged on main and are unrelated.

Part of #376.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): give ui.theme a device override

Theme joins text scale, text weight, and high contrast as a profile default
with an optional per-device override, resolving profile_device -> profile ->
default. The right theme is partly a function of the screen and the room — a
light theme on a phone in daylight, a dark one on a TV at night — which is the
same reasoning the other three appearance keys already used.

All four appearance settings now cascade consistently, which also means one
rule to explain in the UI rather than "these three follow the device, that one
does not".

ui.custom_theme_vars and ui.custom_css stay profile-wide. They are authored
styling rather than a contextual preference, so a profile's custom tokens still
apply on top of whichever theme a device resolves to. Recorded in the
definition notes because it is a visible consequence: vars tuned against a dark
theme will sit on top of a light one if a device overrides the theme. Widening
those to profile_device later is an additive revision bump if it turns out to
matter.

Part of #376.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web): tag local appearance caches with their owning account

The theme, text scale, text weight, high contrast, custom theme variable
and custom CSS caches in localStorage were untagged, so on a shared
browser a second account inherited the first account's appearance: with
no server value of its own, every fallback resolved to whatever the
previous account had stored, and the leftover `silo-theme` key also
suppressed the admin-configured default theme for the new account.

DateTimeFormatProvider already solved this by stamping its cache with the
authenticated user id and refusing another account's values. Extract that
mechanism into `createOwnedCache` in utils/storage.ts (where key
namespacing lives) and put all three groups behind it, so appearance and
custom theme get the same protection instead of a third copy of the rule.

- Each group carries its own owner stamp. A shared stamp would be unsafe:
  the groups are written by hooks nested inside each other, and effects
  run inner-first, so whichever hook stamped first would vouch for the
  other's still-stale values.
- A null owner (auth bootstrapping, or signed out) still trusts the
  cache, which keeps the warm start and the login screen's last look.
- An unstamped cache is not trusted once an account is known, so existing
  users take a one-time appearance reset on first load rather than a
  chance of seeing someone else's settings.
- When a foreign cache is detected the values are dropped and the empty
  cache is handed to the new account, so a later single save cannot
  re-trust the rest of the previous account's state.

Owner is the user id because /settings is user-scoped server side; it
lives in one helper (`appearanceCacheOwner`) so it can be widened if
appearance moves to profile scope. `shouldLoadApiTheme` is gone: it had
become a synonym for `appearanceCacheOwner(...) !== null` with no callers
left.

Part of #376

AI-use disclosure: implemented with Claude Code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): make the settings contract enforceable and fix the appearance cache

The contract manifest landed as a document nothing checked. This makes it a
mechanism, and fixes the one defect in the change set that hurt users on merge
rather than at cutover.

Web appearance cache. useTheme cleared the cache for any account whose stamp
did not match and never repopulated it — the only writers were the four
user-action setters — so every upgrading user lost their warm start on every
load, not once, and x-large-text and high-contrast users lost theirs too. The
owner-stamp protocol is replaced with per-account key namespacing
(`silo-theme:7`): a foreign value is absent rather than present-and-distrusted,
so nothing has to be deleted, the first account keeps its warm start, and there
is no shared stamp for a second tab, a stale debounce timer, or an out-of-order
effect to race on. Widening ownership to profile scope, which this manifest
requires, is now a change to appearanceCacheOwner alone. Adds the API-to-cache
mirror useTheme was missing, cancels pending debounced writes across an account
change, and re-seeds provider state during render so no frame paints the
previous account's look.

Canonicalization. writeCanonical used json.Marshal, which HTML-escapes < > and
&, and canonicalNumber used Go's 'g' format — both diverge from RFC 8785, so
the first label containing an ampersand or bound below 1e-4 would have forked
the server's ETag from every conforming client. Output is now byte-identical to
ECMAScript String() across the edge cases, verified against node. The ETag also
covers the value schemas, which decide what the server accepts and previously
could change while the tag stood still. All four derived representations are
memoized; a conditional GET no longer costs a full parse and re-serialize.

Validation. strictUnmarshal's decoder.More() answered false for a stray ] or },
so `true]` validated as a boolean. Enum matching compared fmt.Sprintf tokens, so
the string "3" satisfied an integer member. Declared steps were never enforced.
The language pattern rejected tags both mobile platforms emit unprompted
(en_US, ca-ES-valencia, ar-EG-u-nu-latn) and never normalized case, so en-US and
en-us were two rows for one preference; NormalizeValue now canonicalizes on the
shared path.

Manifest. show_forced_subtitles defaulted false where the server column is NOT
NULL DEFAULT true, which would have turned forced subtitles off for every
profile that never touched it. preferred_quality declared 13 members where the
planner speaks 6 and collapses the rest to auto. metadata_language's allowlist
was bound to the very column it migrates from. subtitle-appearance pinned
fontFamily to three families while Apple stores any installed system font.
Registers five user-facing settings the clients already ship, and corrects three
notes that described Android behaviour that was not true.

Enforcement. The package had no non-test callers, so MustLoad never ran; it now
loads and logs at startup. The inventory test compared the manifest against a
hand-copied map and could not see the drift it named; it now iterates
settingsRegistry and checks defaults too — both verified to fail on injected
drift. Adds .github/workflows/ci.yml, the repo's first CI that runs go test,
go vet, gofmt, and the frontend suite. Known pre-existing failures are named
individually in the Makefile so everything else stays gated and the list can
only shrink.

Part of #135

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): align the sleep timer default and range with the shipped client

Android is the only client that implements this setting. It clamps to 0..240
and defaults to 30. The manifest said 0..480 with a default of 0, so a
manifest-driven UI would have offered durations no client can store, and every
user who never opened the picker would have had the preset silently turned off
at cutover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci: give the new workflow the deps it actually needs

The first run exposed two gaps in the workflow itself. go build ./... fails
without libvips headers, because h2non/bimg binds libvips through cgo and
pkg-config; the Dockerfile installs the same package. And pnpm/action-setup
resolves its version from package.json, but there is no package.json at the
repo root — the packageManager field lives in web/package.json, and a job's
defaults.run.working-directory does not apply to an action's inputs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(web): stop the diagnostics download test depending on the Node version

new Response(blob) reads the body through blob.stream(), which jsdom's Blob
does not implement on Node 22 — the version the Dockerfile builds with. The
test passed locally on Node 24 and threw "object.stream is not a function" in
CI. Nothing in it asserts on the body, only that the object URL and filename
reach the anchor, so a string body is equivalent and works on both.

Surfaced by the CI workflow added in this branch, which is the first thing in
this repo to run the frontend suite anywhere but a developer's machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(build): copy the settings contract into the container build context

Both Dockerfiles copy cmd/, internal/, migrations/ and web/embed.go, but the
manifest lives in contracts/settings/v1 — an embedded Go package that sits
outside internal/ because clients vendor those files. The image build therefore
fails with "no required module provides package .../contracts/settings/v1".

Caught deploying to the dev box. Nothing had built an image since the manifest
landed: the Docker workflow only runs on pushes to main and workflow_dispatch,
and CI's go build runs against a full checkout, so neither gate covers the
container context. This would have broken the published image on merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): enforce the language-tag and step constraints the manifest declares

A sweep of all 43 manifest definitions against the running server (160 checks:
declared default, both boundaries, and deliberate violations for each remote
key) found two places where the live registry accepts what the contract
forbids. Both are fixed by calling the contract's own validators rather than
adding a second implementation.

playback.audio_language was checked as "32 characters or fewer", so the server
stored "!!!" for a field the manifest declares as language_tag — a value track
matching would then silently never match. It now requires a well-formed tag via
settingscontract.NormalizeLanguageTag. The empty string is still accepted: the
string-only endpoint has no way to send null, and both Android and web send ""
to clear the choice, so rejecting it would break clearing the preference.

player.playback_speed declared step 0.05 and nothing enforced it, so 0.26 was
stored — a value no client's stepper can represent and that every client would
silently snap on the next write. settingscontract.StepAligned is now exported
and used by both the contract validator and the registry, so there is one
definition of "on step" rather than two that can drift.

This gives the contract its first production consumer beyond the startup load,
which is the direction Phase 2 continues in.

Also fixes a genuinely flaky test that the new CI gate would have hit
intermittently: TestRemoveJellyfinCompatWebDisablesWebSetting used t.TempDir as
the install root, but the endpoint returns 202 and its goroutine keeps writing
there after the test body returns, so cleanup tripped "directory not empty"
roughly one run in four. Confirmed pre-existing and unrelated to settings; the
suite now passes six consecutive full-package runs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): keep widened numeric bounds resolvable at older revisions

A bound was one scalar plus the revision that introduced it, which discards
the value it replaced. Widening a maximum from 240 to 480 at revision 3 left
a revision-3 client with no correct answer against a revision-1 server:
honoring 480 offers values that server rejects, and filtering the tagged
bound out leaves the setting unbounded. Since clients are specified to filter
their pinned contract against the server's advertised revision, the bound has
to carry what it used to be.

Bounds now hold their full history, oldest first, and AtRevision hands back
the limit a given peer actually enforces. A bound nobody has widened still
serializes as a bare number, so the manifest reads the same and untouched
entries do not churn the ETag.

Validation gains the rules the representation makes checkable: a maximum may
only grow and a minimum may only shrink, history is strictly ordered, later
entries must say when they arrived, and the first entry cannot predate the
definition. That last rule is the lower bound allowed_scopes already
enforced; the same gap is closed for enum members, which could previously
claim to predate the definition containing them.

Reported by Codex review on #479.

* fix(settings): accept the partial subtitle appearance objects already stored

The schema required all nine properties, but the current API accepts and
round-trips sparse objects — settings_device_test.go stores
{"fontSize":"xxlarge"} and reads it back — and the web client has always
merged whatever it gets over DEFAULT_SUBTITLE_APPEARANCE. Requiring the full
object would have made the cutover migration quarantine preferences users
really set, or block on them.

Every property is now optional and a stored value is documented as a sparse
override merged over the definition's complete default. An empty object is
still rejected: an override that overrides nothing is the same state as no
override, which the contract represents as unset.

Cross-scope resolution is deliberately unchanged. A device override still
replaces the profile's object rather than merging into it, because a device
override means "draw subtitles this way on this screen", not "amend the
profile" — and that is what the server does today.

Reported by Codex review on #479.

* fix(jellycompat): scan the parent directory when a sidecar changes

Autoscan matched scantrigger rejections by comparing RequestError.Message
against literal strings. One of those messages became "Unsupported media file
extension for library type" and the copy in handlers_autoscan.go did not, so
the comparison silently stopped matching.

The effect is user-visible: a Jellyfin client posting a change for Movie.nfo
or poster.jpg gets a 400 and the batch is abandoned, when the sidecar should
have resolved to a scan of the directory containing it. Three tests covered
exactly this and had been excluded rather than read.

RequestError now carries a Reason the caller can switch on. Message stays
prose for the client reading the response — it is meant to be reworded, and
nothing should break when it is.

Also makes two tests honest about asynchronous work. The Jellyfin Web
teardown deleted its install root while the operation goroutine was still
writing to it, where a late write recreates a path RemoveAll already walked
past; it now waits for the operation's terminal state, which required
exporting CurrentWebOperation. And the direct-play If-Range test pinned size
and mtime so ctime was the only remaining validator, then read it back inside
a single coarse-clock tick — it failed about 85% of the time on main for a
reason unrelated to what it tests, and now rewrites until the stamp moves.

With those fixed, GOTEST_KNOWN_FAILURES is empty and gone: make test-go runs
the whole Go suite. The one test that cannot pass yet —
TestHandleReplanPlaybackV3SeekFailureRecoveryNeverChangesMediaVersion, which
has failed since the commit that introduced it and describes unimplemented v3
planner behavior — carries a t.Skip explaining that where the test is, rather
than a regex in the Makefile.

Reported by CodeRabbit review on #479.

* fix(settings): reject JSON the decoder would otherwise rewrite

Two cases where encoding/json accepts input by quietly changing it, which is
the one thing a contract promising byte-identical agreement between peers
cannot tolerate.

Duplicate object properties. jsonschema.UnmarshalJSON keeps the last
occurrence, so {"fontSize":"small","fontSize":"large"} validated and stored
"large". Which one wins is a property of the parser, not of the contract: a
client generated against a different JSON library can disagree about what it
just sent, and the canonical form cannot represent the duplicate at all.

Lone surrogates. An unpaired \ud800 became U+FFFD and canonicalization
reported success, so the server would issue canonical bytes and an ETag for
an artifact a conforming implementation must refuse — RFC 8785 requires
terminating here. Substitution also means the value read back is not the
value written.

Both checks run before the decode that would hide them, on the shared
decodeJSON path that the manifest, its public projection and every value
schema go through, and again on the object branch of ValidateValue, which
uses a different decoder.

Reported by Codex review on #479.

* ci: gate Go lint on the lines a branch changes

AGENTS.md told contributors CI ran the same checks as `make lint`, and the Go
job ran only gofmt and vet. A change failing the documented Go lint gate
passed all three jobs.

Running the linter as-is is not an option: the tree has ~296 findings today,
which is why this half of `make lint` was never enforced. Blocking every PR
on a cleanup nobody has scheduled gets the gate deleted again, so CI runs
with --new-from-merge-base and only the lines a branch touches have to be
clean. The count can then only fall.

golangci-lint is built from source at a pinned version rather than
downloaded. A released binary refuses to run against a Go newer than the one
it was built with, and go.mod here tracks Go closely enough that the current
release already fails that way on 1.26.4.

.golangci.yml declared version 2 while still using v1's issues.exclude-rules
key. Current golangci-lint ignores it, so the "allow repeated strings and
unchecked cleanup errors in tests" exclusions silently did not apply — 16
findings in test files that the config says to skip. Moved to
linters.exclusions, which `golangci-lint config verify` accepts.

The four lines this surfaced in scantrigger are fixed rather than excluded:
its repeated status codes and messages are now named constants, so one
condition cannot end up worded two ways.

Also drops the workflow token to contents:read and stops persisting
credentials in the three checkouts, neither of which any job needs.

Reported by CodeRabbit and Codex review on #479.

* docs(v1): record the settings removal as a pre-lock exception

The design removes the legacy /api/v1/settings routes and the profile DTO
preference fields, while AGENTS.md states /api/v1 is additive-only and
removals go through Deprecation/Sunset. Read together those contradict.

They do not actually conflict: v1-scope.md scopes the additive-only rule to
"when the scope locks", and the scope is still open, so a removal taken now
is in scope and there is no amendment process to invoke yet. But that
reasoning lived only in the settings design, where nobody checking the API
policy would find it.

v1-scope.md now carries a pre-lock removals table naming what goes and why
waiting is worse, and states the deadline the argument depends on: a removal
listed there must ship before lock or fall back to Deprecation/Sunset.
AGENTS.md points at the table and says to treat an unlisted removal as a
mistake.

Reported by CodeRabbit review on #479.

* fix(settings): clear the remaining review findings

Small, unrelated except that each was raised on #479.

compileObjectSchemas parsed every non-directory file under schemas/ as a JSON
Schema, so a stray editor backup or .DS_Store would panic the server at
startup through MustLoad. schema_ref can only name a .json file; anything
else is skipped.

cmd/silo used MustLoad while the ETag check beside it and every other startup
failure use log.Fatalf. It now fails the same way, so a bad contract prints
an error instead of a stack trace.

TestRegistryDefaultsMatchTheContract called scalarDefault before handling
null, and scalarDefault rejects null as non-scalar — so the subtest skipped
and the comparison after it was unreachable. A nullable contract default
could disagree with a non-empty registry default and nothing failed.
Confirmed by injecting that drift, which now reports it.

The three appearance providers each adapted the auth context to
AppearanceAuth with identical code, putting the shape of auth back in three
places that widening cache ownership would have to find. useAppearanceCacheOwner
now does it once.

useTheme.test.ts cleared storage.KEYS between cases, but appearanceCache
writes namespaced keys and an owner pointer that are not in that list, so
both survived and the suite was order-dependent. It clears the store, as
storage.test.ts already did.

The abs_smart_collection_store comment is reworded rather than given back its
SQL quotes: gofmt folds a pair of apostrophes in a doc comment into a
typographic quote, which is how it became one in the first place.

Reported by CodeRabbit review on #479.

* feat(settings): add canonical typed storage for the settings contract

The cross-platform settings contract needs one typed store behind it before a
resolver, routes or a migration can exist. This adds that storage to both
user-store backends and holds them to identical behavior.

PostgreSQL gets user_setting_values with the scope CHECK constraints, the five
partial unique indexes that enforce one explicit value per identity, and the
covering indexes the one-query read path needs, plus user_setting_mutations for
mutation_id idempotency and the inert user_setting_migration_rejects audit
table. The per-user SQLite store gets the same shape minus user_id, since that
database is already user-scoped.

The UserStore interface grows the typed operations: read one explicit value at
one scope, collect every candidate row for a resolution request in a single
query, upsert with a revision increment, unset, and the idempotency receipt
operations. The resolution read deliberately returns unranked candidates so the
resolver can rank in Go — one query per request, never one per scope, which the
pgx query-count test pins.

Delete behavior is application-enforced. Neither backend can inherit it from
constraints: the SQLite store declares no foreign keys, and library, series and
device columns are not FK targets in Postgres either. Profile deletion cascades
to profile-anchored values while account scope survives, forgetting a device
clears its profile_device values alongside the legacy overrides, and the
library/series purges remove only what is scoped to that entity.

The shared conformance suite covers all of it, including the set-versus-unset
distinction for false, 0, "" and null, so a divergence between the two backends
fails a test rather than reaching a client.

Part of #376

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(settings): pin the settings-value schema constraints in both backends

Completes the storage track. The conformance suite exercises the store API,
which validates identities in Go before any SQL runs — so nothing noticed
whether the CHECK constraints and partial unique indexes actually existed.
The one-time migration writes these rows in bulk without going through the
per-request path, so the schema is the only thing guarding it.

Adds constraint tests to both backends covering every scope's column
requirements, rejection of an unknown scope, a profile that does not exist,
non-JSON values, and each of the five partial unique indexes.

Also clears the lint the storage commit did not get to: sql.ErrNoRows and
pgx.ErrNoRows compared with == rather than errors.Is (which fails on a
wrapped error), an unchecked rows.Close, and repeated fixture literals in the
shared suite now named so a backend that confuses two scope columns fails on
the assertion rather than on a typo.

* fix(settings): close the review findings in the validator and the theme cache

Four defects the existing tests did not reach.

The web theme resolver compared the server's value against the appearance
cache and fell back when they agreed, but the mirroring effect writes the
server's value into that same cache — so the comparison held on the first
render and stopped holding on the second, reverting an explicitly chosen
theme to the default. The server's value is this account's own stored
choice, so it now simply wins. The regression test re-renders rather than
asserting on the first paint, which is why the original one passed.

golangci-lint's exclusions.paths is a path regex, not a directory list, so
a bare `web` also excluded internal/jellycompat/web_component.go,
internal/webhooksync/, internal/notifications/webhook*.go and eleven other
non-test files that were being linted before. Anchored.

json.Number is a string kind, so `"1.5"` unmarshalled into it happily and
Float64 parsed the quoted digits: a numeric setting validated as a JSON
string and NormalizeValue stored the quoted form into jsonb. Rejected.

The lone-surrogate check ran only on the object branch, so a lone surrogate
in ui.custom_css decoded to U+FFFD on SQLite and was refused outright by
Postgres jsonb — the two backends disagreeing about whether the same value
could be stored. Hoisted to cover every type.

The strict language-tag validation this branch added is correct, but it
rejects what the shipped Android client sends; the companion fix is
silo-android 4aeb78b4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(auth): stop TestJWT_TamperedToken passing a valid signature

The test overwrote the last character of the signature with "X". An
HMAC-SHA256 signature is 32 bytes, so its base64url encoding is 43
characters and the final one carries only four significant bits — U, V, W
and X all decode to the same trailing byte. Roughly one token in sixteen
was therefore left byte-identical and validly signed, and the test failed
because ValidateToken correctly accepted it.

Measured at 3098/50000 (6.2%) over distinct signatures; it just failed the
Go job on this branch for reasons unrelated to the branch. Flipping a
character in the middle of the signature is 0/50000.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): reject raw invalid UTF-8, not just escaped surrogates

The previous commit hoisted the lone-surrogate check to cover every value
type, but that only closes the escaped path. A raw 0xff byte inside a
quoted string — what an HTTP body carries when a client encodes text in the
wrong charset — is not an escape, so the surrogate scan never sees it, while
encoding/json still substitutes U+FFFD and reports success. NormalizeValue
then stores the original bytes, which SQLite's json_valid accepts and
Postgres jsonb refuses: the same backend divergence, reached the other way.

Found by the Codex review bot on the previous commit's own diff.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(settings): size the library page state bound to what the web client writes

ui.library_page_state's `search` was bounded at 256 characters. The web
client serializes an advanced library view as URLSearchParams, encoding each
filter rule as three groups[i][rules][j][field|op|value] keys — measured at
216 characters for one rule, 518 for three, 820 for five.

The current endpoint validates this key by checking only that it parses, so
those oversized values are already stored in production. Typing them at the
declared bound would have failed the migration for anyone who had saved a
view with more than one filter rule, and rejected the equivalent write
afterwards.

Raised to 4096, which clears ten rules with room to spare while staying a
real bound. The test pins it against the key shapes
libraryPageSearchParams.ts actually emits rather than a round number.

Reported by the Codex review bot; the lengths above were measured by calling
serializeLibraryPageSearchParams, not estimated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): split quality into two axes and register the orphan keys

Two manifest changes the cutover needs.

**Quality becomes resolution + bitrate.** The legacy ladder values
(1080p-high, 720p-medium, 1080p-8, 420p, 328p) were never a third dimension
— they are a bitrate spelled into the resolution string. The web player
already decomposes them: useTranscodeQuality.ts defines 1080p-high as
{resolution: 1080p, bitrate: 10000} and sends the two separately, so the
compound form never reached the wire. Downloads went further and kept only
a bitrate ladder.

So playback.preferred_quality keeps the six clean resolutions and
playback.max_bitrate_kbps becomes the second axis, nullable because
"uncapped" is a real answer and a numeric sentinel would need widening
every time hardware improves. Clients compose their own presets from the
pair, which means retuning what "High" means is a client release rather
than a contract break. Migration decomposes each legacy value losslessly,
so none of them lands in the rejects table.

**The five extension-bag keys are now definitions.** card_overlays,
next_up_mode, sidebar_pins, disabled_library_ids and library_order reached
the server only through the unknown-key path, stored as unvalidated
strings. Two of them the server reads back — next_up_mode decides home
section assembly and card_overlays falls back to an admin default — so
they cannot be demoted to client-local. Registering them is what lets the
extension bag close.

Adds three schemas for their shapes and a test that exercises every
schema_ref against a real value: each of these is nullable with a null
default, so the existing default-validation test returns at the null branch
without ever compiling the reference.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the canonical resolution engine

One answer to "what is this setting, for this profile, on this device, for
this content". Before this, each caller carried its own ladder:
catalog/detail.go resolved subtitles across four levels by hand and audio
across three, handlers/settings.go had a two-level device/user resolution
with a lazy write-back inside a GET, and jellycompat read profile columns
directly. Those disagreed about precedence, which is the drift the contract
exists to remove.

Resolution is one batched read regardless of how many keys, libraries, or
series are in play — ranking happens in Go against each definition's
declared resolution_order. Five sequential index lookups per key per item
is the implementation the design rejects, and a season view is exactly
where it would have shown up.

An absent identity drops its scope rather than erroring, so one code path
serves an identified client, an anonymous jellycompat seed, and a batch
spanning many series. Rows for a foreign profile, device, library or series
are ignored even though the batched read returns them.

Constraints narrow without destroying: a capped 4K preference resolves to
the cap, reports itself constrained, and keeps the authored value so it
takes effect the day the cap lifts. Two cases needed care — null on a
nullable numeric means unbounded, so a ceiling must cap it rather than rank
it equal and let the value that most needs capping slip past; and an
allowlist falls back to a permitted member rather than the definition's
default, which may itself be outside the list.

Adds ValueSchema.CompareValues to the contract package, since ordering
values is what makes a ceiling or floor mean anything and value semantics
belong with the schema that declares them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the one-time migration planner

The conversion rules from legacy settings storage to canonical values, as
ordinary Go rather than twice in two SQL dialects. Both backends read their
own rows, hand them to Plan, and write what comes back — so the decisions
are testable without a database and SQLite and Postgres cannot drift apart
in what they decide.

The rules that needed care, each pinned by a test:

Column defaults are not choices. quality_preference is NOT NULL DEFAULT
'1080p' while the contract defaults to auto, so migrating the column
unconditionally would pin every profile in the install to 1080p having
never chosen it — and that stored value would then outrank the contract
default forever. Same for language 'en', subtitle_mode 'auto', and
show_forced_subtitles true.

The empty string is unset, not a value. The legacy string API had no way to
send null, so both Android and web spell "clear my choice" as "". Storing
that would make a cleared setting outrank the default.

Legacy quality decomposes rather than rejects. Every compound value maps to
a resolution and a bitrate from the ladder in useTranscodeQuality.ts, so
nothing lands in the rejects table.

Account rows fan out to every profile, which is the account-to-profile move
the contract makes for appearance and search scope: a household that shared
one theme each end up owning theirs.

Legacy strings become typed JSON — "true" to true, "30" to 30 — or every
generated binding would fail to decode what the migration wrote.

Nullability differs per backend, so profile columns arrive as pointers and
the caller resolves "chose the default" versus "never written" when it
reads. jellycompat's DisplayPreferences blobs ride the same table under
synthetic keys and are left alone; they are that subsystem's storage.

Everything that cannot convert is recorded with a reason rather than
dropped, and a final test asserts every planned row would be accepted by
the mutation endpoint's own validation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): run the one-time migration on the SQLite backend

Wires the planner to real storage as userdb migration V15. V14 created the
tables; this fills them.

It runs inside runMigrations' existing transaction, so a database either
comes out fully migrated or untouched — a partial migration is the one
state neither the operator's backup nor a rollback covers. Pinned by a test
that rolls back and asserts nothing was left behind.

Two things the wiring had to get right that the planner could not see:

Reject identities are JSON. Postgres declares that column jsonb NOT NULL
and SQLite guards it with a json_valid CHECK, so the free-form
"profile=p1 device=d1" the planner emitted would have failed to insert — on
exactly the rows the table exists to record. They are structured documents
now, which is also queryable.

Subtitle and audio preferences are two tables keyed the same way, so they
merge into one per-series record before planning. Converting them
independently would have produced two rows racing for the same identity.

Every legacy read tolerates a missing table, since this runs against
databases created at any schema version, and preferred_metadata_language is
deliberately absent: that column exists only in the Postgres schema.

Tested end to end against a real database rather than only through the
planner — the rows land, satisfy the scope CHECK and the partial unique
indexes, and hold valid JSON. Also covers the empty-install case and
asserts a second run fails rather than silently doubling every value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): run the one-time migration on the Postgres backend

The mirror of userdb V15, registered with goose as a Go migration rather
than SQL: the conversion validates every value against its own definition
and re-encodes it as typed JSON, and one legacy quality string becomes two
rows — neither is expressible in SQL without duplicating the manifest. The
rules stay in internal/settingsmigrate, so the two backends cannot disagree.

RunTx, so the whole backfill lands in goose's transaction. The down
migration empties the canonical tables; the legacy ones are never touched
by the up, which is what keeps the cutover reversible until the follow-up
migration drops the superseded columns.

preferred_metadata_language is read here and only here — the column exists
in this schema and not in SQLite's, so this is the sole source for
catalog.metadata_language.

Verified against a real Postgres: the full goose chain runs, 1080p-high
decomposes to ("1080p", 10000), values land as typed jsonb rather than
strings (jsonb_typeof reports number), rejects carry a queryable jsonb
identity, and the composite profile foreign key refuses a row naming a
profile that does not exist.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): add the canonical settings API

The routes that make the typed storage reachable. Until now the manifest,
the resolver and the migration all existed with nothing able to call them.

GET /settings/contract serves the public manifest behind an ETag — clients
vendor a pinned copy and generate bindings from it, so the common request
asks "still the same contract?" rather than transferring it. Its
capabilities sibling reports revision and supported scopes for feature
detection instead of version sniffing.

/settings/values/{key} reads, writes and clears an explicit value at one
named scope, which is what a reset affordance needs: "did I set this here"
is a different question from "what applies", and the old endpoint could
only answer a blurred version of both. Scope comes from the query while
profile and device come from session headers, so one profile cannot address
another's settings by naming it.

/settings/values/effective resolves any number of keys in one request, with
the resolution ladder and the source of each answer reported so a client can
offer "reset this device's override" against the exact row holding it.
Asking for no keys returns every remote setting, which is what a settings
screen wants.

Writes are idempotent when a client sends X-Silo-Mutation-Id: a retry after
a dropped response replays the receipt, and reusing an id with different
content is a conflict rather than a silent overwrite of the wrong thing.

Three things the string-only endpoint could not do, each pinned by a test:
an unknown key is refused rather than stored in the extension bag, values
are checked against their declared type and range, and a write to a scope
the definition does not allow is rejected.

Registered before the catch-all /{key} routes, which would otherwise
swallow "contract" and "values" as setting names. The legacy endpoints stay
live for now; deleting them is the next commit, once their consumers move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): generate typed bindings for all four languages

One generator rather than one per repo. The point of the contract is that
four codebases agree on keys, types, scopes and defaults, and four
independently written generators would be four chances to disagree.

Go and TypeScript land in this repo; Kotlin and Swift are written into the
sibling client checkouts, skipped with a note when they are not present so
a server-only developer can still run it. Output is sorted by key so an
unrelated manifest edit does not produce spurious diffs.

The Kotlin output is the interesting one: it generates the DeviceSettings
allowlist Android maintained by hand, plus the BOOLEAN_KEYS/INT_KEYS/
DOUBLE_KEYS classification it kept as a *second* hand-maintained table that
had to agree with the first. Both are manifest questions now, so the whole
class of "wrote a local key to the server" and "flushed a value the store
could not parse" bugs stops being possible by construction.

The TypeScript output carries the full definition table — labels, controls,
enum members, bounds — so web/src/lib/settingsManifest.ts can be deleted
rather than kept in sync: it declared 17 definitions against the contract's
49, with its own two-scope model that does not match the contract's five.

make verify-settings-bindings fails when the committed output disagrees
with the manifest, wired into CI, so a manifest change cannot merge leaving
every client reading stale keys.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(web): add the two-axis quality picker and typed settings hooks

Quality becomes one picker over two stored values.

The server holds a resolution cap and a bandwidth cap independently, which
is what the player has always sent on the wire — useTranscodeQuality.ts has
decomposed 1080p-high into {resolution, bitrate} for as long as it has
existed. Presets live in the client rather than the contract so retuning
what "High" means is a one-line edit here instead of a contract change four
codebases have to agree on, and an older server keeps working because it
only ever sees the two axes it already understands.

A combination no preset covers still gets a truthful label rather than a
picker showing the wrong entry: reachable by setting the axes separately
through the API, or from a legacy value whose bitrate is off this ladder.
Choosing an uncapped preset clears the bitrate rather than storing a
sentinel, so "no cap" stays the absence of a value at every layer.

Adds hooks over the canonical API alongside the legacy ones rather than
replacing them wholesale — a key that is not in the manifest cannot be
expressed, because SettingKey is generated from it, and the default for an
unset value comes from the generated table rather than a literal at the
call site. That last part is what stops the flip-off bug the Apple client
carries a hand-written guard for.

A test asserts every preset composes values the contract actually accepts,
so a preset naming a resolution outside the enum or a bitrate outside the
declared bounds fails here rather than 400ing when a user picks it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): resolve catalog playback preferences through the contract

catalog/detail.go held the two hardest ladders in the codebase: subtitles
resolved across four levels by hand, audio across three, each partially
overriding the last through Has* flags. Both now call the canonical
resolver, so the precedence lives in the manifest and this file cannot
disagree with the contract about which override wins. Adding a scope is a
manifest change rather than another branch here.

The subtitle track signature stays on its specialized table — it identifies
a concrete track rather than expressing a preference, so it is not a
setting.

Resolution keeps the memoization the old lookups had: the audio resolver
still reads once per profile and once per library rather than once per
file, which is what kept a many-track audiobook detail page fast. The test
that guards it now counts resolver reads instead of GetProfile calls, since
the guarantee is about scaling with file count rather than about which
method does the reading.

Four tests seeded the profile column directly. That column is a migration
source now, not a read path, so they seed the canonical value instead —
they were passing against storage nothing reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): close the unknown-key extension bag

keyUsesUserScope returned true for any key the registry did not know, so a
client could invent a production setting unilaterally and the server stored
it as an unvalidated string. That is how six ui.* settings and five orphan
keys reached production untyped, and it is the root enabler the design
names.

An unknown key is no longer a user setting, so the legacy write path
rejects it and the canonical API — which validates every value against its
own definition — is the only way to store something new.

jellycompat's DisplayPreferences blobs ride the same table under synthetic
keys and keep working: they are that subsystem's storage rather than user
settings, and they move to dedicated storage in the follow-up rather than
being dropped here.

Also repoints the DisplayPreferences seed at the canonical resolver.
Resolved at profile scope with no device on purpose — Jellyfin clients do
not carry Silo's device identity, so a device override leaking into the
seed would hand one device's settings to every Jellyfin client on the
account.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(settings): enforce viewer quality caps through resolver constraints

constraintsFor was the unwired half of the preferences-versus-restrictions
seam: it returned nil, so a profile capped at 1080p by policy still resolved
its stored 2160p preference at face value through the effective endpoint.

The settings routes are mounted inside RequireViewerAccess, so the resolved
access scope is already on the request context. Scope.MaxPlaybackQuality
holds a literal member of the contract's quality enum ("1080p"/"2160p"),
which is exactly what the manifest binds playback.preferred_quality's
ceiling to under policy_input "max_playback_quality" — so the wiring is a
direct map with no translation table. An empty value means the policy sets
no cap, expressed by returning nil so the resolver leaves the preference
alone.

catalog.metadata_language deliberately stays unconstrained: the manifest
notes record that the allowlist draft was circular (the policy input it
would bind to is populated from the very preference it would narrow).

The handler test covers both halves of the seam: a 2160p preference under a
1080p cap resolves to the cap with constrained:true/ceiling and the authored
value reported in stored_value, the stored row itself is not rewritten, and
an uncapped viewer gets the preference unchanged with no constraint noise.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): publish user_settings change events

Add a user_settings realtime channel so clients learn when a setting
changed on another device without polling. The channel is modeled on
user_state: non-admin subscribable, per-user addressed envelopes, null
snapshot.

SettingValuesHandler gains an EventsHub and publishes
user_settings.changed after every successful PUT and DELETE on
/settings/values/{key}. The payload carries only key, scope and
profile_id — never the value. Admins receive every user's user-scoped
events, so a value in the payload would leak private settings to
admins; interested clients re-fetch over the scoped REST API instead.
The payload is always non-empty because an empty Data falls back to a
null snapshot in the hub. A nil hub (tests) skips publishing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): sweep expired mutation receipts daily

Setting-mutation idempotency receipts were written with an expires_at that
nothing enforced, so the table grew forever. Add a hidden daily system task
(05:00) that walks every login account, opens its user store, and calls
DeleteExpiredSettingMutations. A user whose store fails to open or sweep is
logged and skipped so one broken store cannot stall retention for everyone
else; the delete is idempotent, so the next run repairs anything missed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(settings): resolve metadata language canonically in access and policy

Repoint the last legacy column readers onto canonical contract resolution
(settings cutover task A4a):

- access.Resolver and policy.ViewerResolver now resolve
  catalog.metadata_language through settingsresolve (profile scope ->
  contract default) via a shared access.PreferredMetadataLanguage helper,
  instead of reading user_profiles.preferred_metadata_language. Resolution
  is deliberately unconstrained: the policy input this preference feeds is
  the one a constraint would have to reference, which is circular — see the
  key's manifest notes.
- playback start now resolves playback.audio_language canonically for the
  profile default instead of reading user_profiles.language, matching the
  catalog detail path. Series and library override handling is unchanged.
- items.go needed no change: it already consumes the resolver-produced
  scope.PreferredMetadataLanguage.

The legacy columns keep their values but are no longer read on these
paths; a profile with only a column value now resolves to the contract
default, and a stored canonical value wins. Tests pin both directions in
access, policy (including scope parity, where the column is now a decoy),
and the playback handler. Read cost is one batched store read per
resolution, same as the profile-row read it replaces.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(jellycompat): give DisplayPreferences its own table

The Jellyfin DisplayPreferences blobs rode the legacy user_settings
key/value table under synthetic jellycompat:* keys, which forced the
legacy settings API to carry a prefix carve-out in its otherwise-closed
unknown-key gate. They are the compat subsystem's storage, not user
settings: the contract neither validates nor resolves them.

Move them to a dedicated jellycompat_displayprefs table in both
backends, keyed by (prefs id, client) per user, with the blob stored as
opaque text served back byte-for-byte (deliberately not jsonb, which
would re-serialize it). The data-copy migrations — per-user SQLite V16
and a paired SQL + Go goose migration for Postgres — are transactional
and harmless to re-run, and both drive their key parsing and row
classification from the new internal/jellycompat/displayprefs package
so the backends cannot diverge, following the internal/settingsmigrate
precedent. A jellycompat:* row that does not parse as a DisplayPrefs
key (only ever writable through the removed carve-out) is recorded in
user_setting_migration_rejects rather than silently deleted.

With the last non-settings tenant gone, the jellycompatSettingPrefix
carve-out is deleted: the legacy settings endpoints now refuse
jellycompat:* keys like any other unknown key and never surface them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): serve admin user-settings through the canonical API

Replace the ten string-registry /admin/users/{id}/settings* and
device-settings* routes with the canonical contract surface: one list of
every explicit value the target user has stored across all scopes, and
set/delete at an explicit scope named in the query string.

The admin handlers live on SettingValuesHandler and share the session
routes' implementation rather than duplicating it — the same key/scope
parsing, identity validation, contract scope allowance, value
normalization and mutation-receipt idempotency, factored into
keyedScopeFromRequest/completeIdentity and setValueAt/deleteValueAt.
The only admin-specific parts are the target user coming from the path,
profile and device ids coming from the query (an admin holds no session
claim to the user being inspected, so its named profile is checked to
exist), and change events attributed to the target user so their
clients refresh.

The list is a new UserStore read, ListAllSettingValues, implemented in
both backends and pinned by the shared storetest conformance suite:
the admin surface wants the stored truth (which overrides exist, for a
per-row reset affordance), which no resolution-shaped read answers.

The ten removed routes are recorded in the pre-lock removals table in
docs/architecture/v1-scope.md per the v1 API rules; the web admin
device-overrides page moves onto the new surface in the Phase B
rewrite inside this same unmerged PR.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(settings): add the cross-platform conformance fixture and its Go and web runners

contracts/settings/v1/conformance.json is the spec's named drift gate: 21
hand-authored cases of {keys, stored rows, context, constraints, expected
effective value + source}, every one executable against the shipped manifest.
They pin the semantics most likely to drift across four resolver
implementations: the full resolution ladder (series > library > device >
profile > default), an absent identity dropping its scopes, foreign-identity
rows never resolving, ceiling caps that report the authored value with
constrained:true, the ordered-enum sentinels (auto below every cap, original
above), null-on-a-nullable-numeric meaning unbounded and being brought down by
a ceiling but ignored by a floor, allowlist falling back to the first allowed
member rather than the (possibly forbidden) default, and
playback.subtitle_appearance resolving device > profile only with the sparse
device object replacing, not merging. Cases may inject a constraint binding
onto a copy of a real definition so constraint kinds no shipped definition
carries stay testable.

The Go runner (internal/settingsresolve/conformance_test.go) resolves each
case through the real resolver against the embedded manifest. The web runner
(web/src/lib/settingsConformance.test.ts) runs the same cases through a new
client-side resolver, web/src/lib/settingsResolve.ts, which mirrors the
server's semantics; the TypeScript bindings now carry each definition's
ordered flag and constrained_by binding so that resolver derives constraint
behavior from the contract instead of hardcoding it. Both runners reject
unknown fixture fields — schema drift in the fixture itself is drift — and
both refuse a fixture authored against a different manifest revision.

The fixture travels with the bindings: make settings-bindings vendors the copy
the web runner reads, and make verify-settings-bindings fails CI when that
copy goes stale. The Kotlin and Swift copies land together with their runners
in the client repos, which will pick their own test-resource paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): review pass over the phase A stack

Fixes the eight adversarially-confirmed defects the review of the
unpushed phase A stack (40e0f77a..1f2c7fe4) found, each with a test
that fails without its fix.

Writers left behind by the language cutover (high). 22e9d7f1 made
access, policy and playback start resolve catalog.metadata_language and
playback.audio_language exclusively from user_setting_values, but
POST/PUT /profiles — the write path the shipped web UI uses — still
wrote only the legacy columns, so a language change after the one-time
backfill never took effect (a stale backfilled row, or the contract
default, won forever). Profile mutations now mirror their preference
fields into the canonical profile-scope rows through the same contract
validation /settings/values applies (audio, subtitle and metadata
language, subtitle mode, forced subtitles; the empty string clears the
row, matching the migration's unset spelling), publish
user_settings.changed for each row moved, and 400 on a value the
canonical endpoint would refuse. quality_preference is deliberately not
mirrored: the server never resolves the legacy column and the two-axis
picker already writes canonically.

Web admin settings 404s (high + medium). facad78d removed the ten
/admin/users/{id}/settings* and device-settings* routes but shipped no
web changes, so the user-detail settings and device-overrides tabs and
the devices-page override editor were dead. The seven admin hooks now
speak the canonical values API: one list across all scopes feeds both
tabs, mutations address an explicit scope identity, values re-type
through the generated contract (display stringifies for the
registry-era controls), device rows are enriched with device and
profile names client-side, and the removed bulk device reset becomes
per-key deletes that treat 404 as already-reset.

Silent metadata-language degrade (medium). PreferredMetadataLanguage
now logs a warning with the profile and error when contract load or
store resolution fails, so pool exhaustion is distinguishable from "no
preference"; the healthy paths stay quiet.

Displayprefs move data loss (medium). Under READ COMMITTED the blanket
pattern DELETEs in moveDisplayPrefs/unmoveDisplayPrefs could destroy a
row an old-binary instance committed between the SELECT and the DELETE
during a rolling deploy — reproduced against real Postgres. Both
directions now delete only the exact rows they read (rejects restore by
primary key), leaving a late row stranded for a re-run to pick up.

Coverage the review proved missing (medium x3): admin mutations are now
tested to attribute change events to the target user, not the acting
admin (the exact regression passed the whole suite before); the
user_settings websocket channel is subscribed through the real events
websocket, failing if the channel is dropped from either
allowedChannelsForRole or AllChannels; and the conformance fixture
gains three locked-constraint cases (replace, equal-value pass-through,
locked default) so the Go and TypeScript locked branches — previously
executable by no test on either platform — are pinned by the shared
drift gate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): read and write appearance and format preferences through the settings contract

Move the four identity-sensitive preference hooks — useTheme,
useCustomTheme, useDateTimeFormat, useSearchMediaScope — off the legacy
string-only /settings endpoints and onto the canonical settings API.
Each surface now reads through one batched useEffectiveSettings call and
writes via useSetSettingValue at scope "profile", matching what the
generated manifest declares: ui.theme / ui.text_scale / ui.text_weight /
ui.high_contrast are profile-scoped with a profile_device override the
effective read already resolves (no device-override UI exists, so writes
stay profile-wide), and ui.custom_theme_vars / ui.custom_css /
ui.date_format / ui.time_format / search.media_scope are profile-wide.
Keys come from the generated SETTING_KEYS table, so a typo'd or
unmanifested key can no longer be expressed.

Because the canonical effective endpoint always answers — resolving
unset keys to the contract default with source "default" — the hooks now
use the source to distinguish "the profile chose this" from "nobody
stored anything". That preserves the admin-default theme layering and
keeps resolved-but-unchosen values out of the warm-start mirror.

ui.theme moving account→profile scope means the appearance warm-start
cache must not be shared by sibling profiles on one account, so
appearanceCacheOwner widens its token from the user id to user id plus
active profile id. Every cache read/write already resolves through that
one function, so no call site could be left behind; the API→cache
mirror, the render-time re-seed on identity change, and the debounced
write cancellation all follow automatically. The ownership tests now
cover profile switches within one account: no theme/text-scale/CSS leaks
between profiles, each profile's warm start survives the switch, and a
debounce armed by one profile never persists under its sibling.

Part of the Phase B settings-contract cutover; the legacy hooks in
queries/settings.ts keep their remaining callers until B4 deletes them.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): store library, sidebar, and overlay preferences through the settings contract

Phase B2 of the settings-contract cutover: the query-layer preference
stores — sidebar pins, library page state, disabled libraries, library
order, and card overlay prefs — move off the legacy string-valued
/settings endpoints onto the canonical values API, using generated
SETTING_KEYS and each definition's declared scope (profile for pins,
visibility, order, and overlays; profile_device for page state and the
remember toggle).

Values are now written as typed JSON matching the contract schemas
(sidebar-pins.json, library-page-state.json, library-id-list.json,
card-overlays.json) instead of JSON-encoded strings, so the encoding the
migration produced keeps validating. Every parser accepts both the
canonical object value and the legacy string encoding, so nothing breaks
while caches or older rows still hold strings.

Semantics preserved deliberately:
- Sidebar pin toggles keep their optimistic update with the
  revision-guarded rollback, now layered on the effective-settings cache
  entry (effectiveSettingsQueryKey is exported for exactly this).
- The remember-library-pages toggle clears the device override to
  inherit again rather than storing the default, via
  useClearSettingValue; the canonical DELETE's 404 for "nothing stored"
  is treated as already-done, matching the legacy delete's idempotency.
- Overlay prefs keep the admin default / kill-switch layering: the
  contract default null means "no preference expressed", which is what
  lets /settings/overlay-config defaults apply, and only a stored value
  overrides them.
- Library visibility/order keep their optimistic local state with
  rollback on error; ids are normalized client-side with the same rules
  library-id-list.json enforces.

parseDisabledLibraryIDs/parseLibraryOrder collapse into one
parseLibraryIDList (they were byte-identical), and the serialize helpers
disappear with the string encoding. Legacy hooks in queries/settings.ts
stay for the remaining consumers until B4.

Part of the settings-contract cutover (see
docs/superpowers/specs/2026-07-10-cross-platform-user-settings-contract-design.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): write playback, subtitle, and library preferences at canonical scopes

The settings screens and the player panels were the last web surfaces still
speaking the legacy string API, and each carried its own idea of where a
preference lives. Playback and subtitle behavior wrote profile columns through
PUT /profiles; auto-play and next-up wrote untyped strings; subtitle appearance
went through three bespoke routes that existed only because the string API had
no way to express an object-valued setting per device. All of them now read one
batched effective resolution and write typed JSON at an explicit scope.

Where each preference lands follows the manifest rather than the endpoint that
happened to hold it:

  - Playback and subtitle defaults, and next-up mode, write at profile.
  - Subtitle appearance writes playback.subtitle_appearance at profile_device,
    replacing /settings/subtitle_appearance/effective and the PUT/DELETE pair on
    /settings/device/subtitle_appearance. One hook now owns that value for the
    settings screen, the in-player panel, and the cue renderer, which before
    each parsed the effective response separately.
  - Per-library edits write at profile_library with the library identity, one
    key at a time. The legacy endpoint replaced a composite row, so clearing one
    field meant re-sending the other three and losing any concurrent change to
    them; independent per-key writes have no such coupling, and "inherit" is a
    delete rather than a sentinel.
  - The in-player series choice splits along the line the contract draws:
    language and mode are preferences and move to profile_series, while the
    track index and signature stay on /subtitle-prefs because they identify a
    concrete track rather than expressing a preference.

Controls render from the generated SETTING_DEFINITIONS. The hand-written
registry beside it had drifted — it declared several profile-only keys as device
overrides, and disagreed with the manifest about the bounds of two sliders — so
the display helpers now derive control shape, options, bounds, and the
device-overridable key list from the contract. (Deleting settingsManifest.ts
itself is B4; nothing outside its own test imports it any more.)

Two follow-on fixes fell out of reading the contract rather than the registry.
playback.auto_skip_recap and playback.auto_play_next_preview are declared at
profile_device but only the intro override was ever consulted, so a device
override on either silently did nothing; the player resolves all three now.
And per-library "Original Language" is gone: the contract types these as BCP 47
tags, and the phase-A migration already rejects "original" at profile_library,
so offering it would have written a value the server refuses.

Risk worth naming: LibrarySettings decides "overrides" from the resolved source
rather than by comparing values, which is what keeps three distinct cases apart
— a library row holding the same value as the profile is still an override, and
a library row holding null is an explicit "no subtitles" rather than an absent
choice. A screen that compared values would collapse the first into "inherits"
and the second into "unset".

Part of #135

AI-assisted: authored with Claude Code; reviewed and verified by the committer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): refresh settings from the user_settings channel

A canonical settings write reached only the tab that made it. The server
already publishes user_settings.changed on every write and delete, but no
web client subscribed, so a preference changed on a phone or by an admin
sat stale here until a manual reload or the 5-minute staleTime expired.

Subscribe the channel and treat the frame purely as an invalidation
signal. The payload carries the key, the scope and the profile — never a
value, because admins receive other accounts' user-scoped events and a
value there would leak private settings. Marking the value queries stale
lets react-query refetch only what a mounted screen is reading, and a
burst of writes coalesces into one fetch per key rather than one per
event.

A profile-addressed change to a profile other than the signed-in one is
dropped: it cannot alter what this tab resolves. Account-scoped changes
carry no profile and always invalidate.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): render settings from the generated contract

web/src/lib/settingsManifest.ts was a hand-written table of labels,
controls, defaults and bounds sitting beside the generated contract, and
it had already drifted: it declared profile-scoped keys as device
overrides, disagreed with the server on the type and range of several
keys, and enumerated a language subset narrower than the one the player
speaks. lib/settingsDisplay.ts has derived all of that from
SETTING_DEFINITIONS since the contract landed, and nothing but the
manifest's own test still imported it.

Delete the manifest and its test. The one piece it owned that the
contract cannot express is the language list — language settings are
typed as BCP 47 rather than as an enum, so there is no member list to
render — which moves to lib/languageOptions.ts and is now derived from
the shared player language list. Two shapes ship: NAMED_LANGUAGE_OPTIONS
for a control that spells its own unset entry, and LANGUAGE_OPTIONS with
the leading "no preference" row for a nullable setting.

The per-library editor's LANGUAGE_OPTIONS re-export goes with it, so
every language dropdown in settings now iterates one list in one shape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): show canonical device overrides in admin devices

The device detail panel read its override rows from
GET /admin/devices/{user}/{device}, whose `settings` array still comes
out of the legacy user_device_settings table. The settings-contract
migration folded that table into user_setting_values and nothing writes
to it any more, so an override created since the cutover — including one
the admin had just saved through this very panel — was invisible here,
while the migrated rows stayed visible. The panel's own writes go to the
canonical route, which made the list look like it silently dropped
edits.

Read the overrides from the canonical values API instead, filtered to
device scope and to this device. Both storage generations show, because
the migration moved the legacy rows into the same table. The detail
endpoint is still the source for registration metadata — device name,
owner, which profiles have used it — which is not a setting and has no
canonical equivalent.

The override count and last-updated readouts move to the canonical rows
for the same reason: override_count is computed over the legacy table and
would disagree with the rows rendered underneath it. "Reset all for
device" has no bulk canonical route, so it keeps issuing one delete per
key, now over the keys that actually exist. The reset button also takes
the profile id from the tab rather than from its first row, which a
profile registered on the device with no override yet does not have.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): delete the legacy settings hooks

hooks/queries/settings.ts spoke the string-only registry API: every
value a string, scope implied by which function you called, and an
unknown key silently accepted. Phase B moved every consumer onto the
canonical value hooks, and the last importer left was the file's own
test — so both go together, along with the client functions they were
the only callers of.

hooks/queries/libraryPlaybackPreferences.ts goes with them. It wrapped
GET/PUT/DELETE /library-playback-prefs, which LibrarySettings replaced
with profile_library-scoped canonical writes; nothing in web has called
it since. The server route stays for now — the Android and Apple clients
may still use it — but the web type and query keys have no reason to
linger.

settingsKeys keeps only `all` (the prefix the canonical invalidation
targets) and the plugin entries, which are a different system. The
list/detail/deviceDetail/effective builders described the registry's
cache layout and had no remaining callers; effectiveSettingsQueryKey in
settingValues.ts owns the canonical shape.

hooks/useSettingsForm.ts is deliberately untouched: it edits admin
server_settings through /admin/settings, which is a separate surface
from the per-user contract, and has more than twenty live consumers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): review pass over the phase B adoption

Phase B moved the web client onto the canonical settings surface. Three
scope mistakes slipped in, all of the same shape: a value written at a
scope no UI can reach, shadowing the one the user can edit.

Auto-play next. The post-roll toggle wrote profile_device while Settings →
Playback wrote profile, and the contract resolves the device row above the
profile row. Turning auto-play off in the player therefore made the
settings switch permanently inert — it saved a profile value the device row
kept shadowing and snapped straight back, with no web affordance able to
clear the device row. Both surfaces now share useAutoPlayNextSetting, which
writes the profile and clears any device row (also the only way a migrated
per-device override becomes reachable). Before Phase B both writers used
useSetDeviceSetting, so they could not disagree; this restores that
invariant at the scope the rest of the Playback screen edits.

In-player subtitle picks. handleSubtitleChanged wrote three canonical keys
at profile_series, the top of the resolution ladder, while "Auto" on the
item page still deleted only the legacy /subtitle-prefs row — so the reset
silently stopped working and the abandoned language kept resolving for
every episode of the series, forever. One of the three,
show_forced_subtitles, was worse: the player has no forced-subtitle
control, so the value it wrote back was the *resolved* one, which for a
viewer who never expressed a preference is the contract default. That
pinned the default above the profile-scope toggle on the Subtitles screen.
The written set now comes from SERIES_SUBTITLE_SETTING_KEYS — language and
mode only, both derived from the user's actual choice — and
useDeleteSubtitlePreference clears exactly that list, so the writer and the
reset cannot drift. show_forced_subtitles still rides the legacy composite
row, which is keyed to a concrete track selection and is not part of the
canonical ladder.

Admin user settings. The tab now lists every non-device canonical row,
which includes the object-valued profile settings (sidebar pins, card
overlays, disabled libraries, library order, custom theme vars). It gated
only on `definition`, and controlKindFor has no `object` branch, so those
fell through to RegistrySettingControl's select — rendering a user's pins
as a one-entry "Unset" dropdown whose only option nulls them. It now uses
the same isStructuredSetting guard the device tab got, routing them to a
raw JSON editor.

Tests: each fix has a test that fails without it, verified by reverting the
fix in place. The auto-play and subtitle tests resolve through
lib/settingsResolve rather than a canned answer, so the scope-precedence
assertions exercise the real ladder.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): serve profile preference fields from canonical resolution

PUT /settings/values?scope=profile writes only user_setting_values, but
GET /profiles still served the legacy user_profiles columns. A preference
saved through the canonical API was therefore invisible in every profile
DTO reader on every platform — Apple's shipped build reads exactly those
fields — while profiles_settings_sync.go mirrored one way only, legacy
column write to canonical row.

Serve those five fields (language, preferred_metadata_language,
subtitle_language, subtitle_mode, show_forced_subtitles) by resolving
their canonical keys through the settingsresolve seam at profile scope,
falling back to the contract default rather than to the stale column.
This matches the cutover direction taken everywhere else: the legacy
columns stay written but stop being read, so "clear this preference"
cannot resurface a pre-cutover value the one-time backfill already
converted. The write paths that accept these fields and mirror them are
unchanged; this is read-side only, and the DTO's field names and types
are untouched.

Resolution is batched. A profile list serves the whole household, so
SettingResolutionQuery.ProfileID becomes ProfileIDs and the new
Resolver.ResolveProfiles ranks every profile against one candidate set —
one store read per list request instead of one per profile. Both backends
carry the widened predicate and the shared storetest conformance suite
gains a household case, so they cannot drift on it.

quality_preference stays column-backed: the legacy column is one compound
value while the contract splits it across playback.preferred_quality and
playback.max_bitrate_kbps, so there is no lossless read. The auto_skip_*
and auto_play_next_preview fields stay column-backed too — the sync path
never mirrored them, so their canonical rows can lag the columns.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): repair CI findings after the main merge

CI runs checks the local loop does not: golangci-lint (not installed
here) flagged two unchecked Close errors in the new websocket test, and
tsc -b (the tests were only vitest-run locally) rejected strict
indexed-access in four test files touched by the review passes. The
merge also brought main's onboarding tour, whose SettingControl wrote
through the legacy useSetSetting hook this branch deletes — it now
writes the canonical scoped mutation, re-typing the tour's string values
through the generated contract like the admin surface does.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(settings): satisfy the incremental lint pass

golangci-lint reports findings incrementally, so these three surfaced
only after the previous fix: errors.Is for the pgx.ErrNoRows compare
(wrapped errors), and named constants for the repeated "values"
response key and the "usersettings" prefs id goconst flagged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(database): pin the read value when deleting moved displayprefs rows

Under READ COMMITTED the move's DELETE takes its own snapshot, so during
a rolling deploy an old-binary instance could update a jellycompat row
between the migration's SELECT and its delete — and the (user_id, key)
predicate would destroy the newer value after copying only the older
one. Naming the value the transaction actually read makes such a row
survive as a stranded legacy row instead, the same disposition a
late-inserted row already had.

Extends the concurrent-write migration test to commit an update to an
already-read row during the stall and assert the newer value survives.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(api): make canonical mutation writes honest under failure

Three review findings on the canonical settings endpoints:

- Idempotency receipts were recorded via defer, so a failed upsert
  still left a receipt and the client's retry replayed a success for a
  write that never happened. The receipt is now written only after the
  upsert lands, and it stores the actual response — revision and
  updated_at included — so a replay is byte-identical instead of a
  reconstruction of the input with revision 0.

- The mutation envelope accepted trailing JSON after the first
  document, leaving the interpreted mutation parser-dependent. The
  decoder now requires EOF after the envelope.

- Resolving a device-aware key without X-Silo-Device-Id silently
  skipped every stored device override and passed the profile fallback
  off as the effective value. The effective endpoint now fails closed
  with 400, matching the write path's existing requirement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): stop the contract rejecting values shipped clients store

Four bounds in the contract were narrower than what a shipped client
already produces, so real stored preferences would fail validation or
be quarantined at migration:

- The BCP 47 grammar rejected extlang tags (zh-cmn) and private-use-only
  tags (x-private) the legacy length-only validator accepted, turning an
  existing 204 into a 400. The pattern now covers both, and
  NormalizeLanguageTag cases a script correctly after an extlang and
  leaves private-use content lowercase.

- subtitle_appearance.fontFamily allowlisted ASCII, contradicting its
  own description: Apple clients store CTFontManager family names
  verbatim and those are routinely CJK. The pattern now excludes unsafe
  characters instead of allowlisting ASCII.

- theme-var-overrides capped CSS values at 128 characters, which real
  multi-stop gradients exceed; the web importer stores them unchecked.
  Raised to 1024.

Plus one tightening the review asked for: card-overlays.order now
declares uniqueItems, matching library-id-list, so an overlay cannot be
rendered twice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(userstore): reject non-canonical identities and bound resolution batches

Three review findings on the canonical settings storage layer:

- SettingIdentity.Validate trimmed ids only to check emptiness, so a
  padded id like " p1 " validated, persisted verbatim, and was then
  invisible to resolution queries, which bind trimmed forms — a
  silently orphaned row. Validation now rejects any id that is not in
  canonical trimmed form, pinned in the shared conformance suite so
  both backends hold the line.

- The effective-values endpoint accepted unbounded library_ids and
  series_ids lists; the SQLite backend expands each id into a bound
  parameter, so a crafted batch could exhaust the host-parameter budget
  and fail the whole resolution. The request boundary now caps the
  combined content ids at 200.

- pickForScope's doc comment promised ties broken "by the most
  specific id in the request order" while the implementation sorts by
  ascending library then series id; the comment now describes the
  actual (deliberately deterministic-only) behavior.

Plus: the pgstore conformance cleanups now assert the ON DELETE CASCADE
they rely on instead of discarding the delete error, so a dropped FK
can no longer leak seeded rows into the shared test database silently.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): close the discovery gaps around the canonical API

Three review findings:

- Canonical profile_device writes never touched the device registry, so
  a device that only ever wrote through /settings/values was invisible
  to ListDevices and the admin device surfaces — undiscoverable and
  unforgettable. Device-scope writes now refresh the registry from the
  request's device headers, throttled the same way the legacy route is.

- The contract spec tells clients to probe GET /settings/manifest (and
  /settings/capability), and to read a 404 as "pre-contract server";
  the router only exposed /settings/contract*. The documented paths now
  alias the same handlers.

- The plugin proxy's X-Silo-Theme header came from the legacy
  account-level user_settings.ui_theme row, so a profile's theme change
  through the canonical API never reached plugins and profiles sharing
  an account were indistinguishable. The lookup now resolves the
  canonical profile-scoped ui.theme row (falling back to the legacy row
  for stores the backfill has not covered) using the request's active
  profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(settings): emit revision metadata in generated bindings and verify the TS one

Two review findings on the generator surface:

- The bindings dropped every introduced_in tag, so a client generated
  from revision N could not filter its pinned contract down to an older
  server's advertised revision — the promised negotiation had no data.
  The TypeScript definitions now carry introducedIn per definition,
  per scope, per enum member, and the full history of any widened
  numeric bound. (Go/Kotlin/Swift emit keys, not definition tables, so
  they only need the Revision constant they already have.)

- make verify-settings-bindings compared only the generated Go file and
  the conformance fixture, so a manifest change could merge with a
  stale web/src/lib/settingsContract.ts. The target now regenerates and
  diffs the TypeScript binding too, through the same prettier config
  the bindings target applies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: drop the accidentally committed settingsgen binary

24ee9952 checked in a 5.5 MB compiled settingsgen alongside its source.
The binary is a local build artifact — cmd/settingsgen is the source of
truth and make settings-bindings runs it with go run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): close the migration planner's data-loss and crash findings

Four review findings on the one-time legacy-to-canonical migration:

- A device holding both player.next_up_prompt_seconds and its
  playback.* rename canonicalized to one identity, and both backends
  insert bare — a unique violation that failed NewUserDB (SQLite) or
  aborted the goose migration (Postgres). Plan now ends with a
  deterministic dedup keyed on the canonical identity; a canonically
  keyed row beats a renamed alias, since the runtime writes the
  canonical spelling first and only best-effort-deletes the alias.

- The four auto-skip profile columns (auto_skip_intro/credits/recap,
  auto_play_next_preview) were never read, so an explicit true silently
  became the contract default false. They now migrate — explicit true
  only, so an untouched false column does not become a choice.

- Profiles with language 'en' emitted no playback.audio_language row
  because the column default was suppressed, but that default WAS the
  effective behavior: the old playback path preferred English, while
  the canonical null default skips language matching entirely. English
  now migrates as an explicit row. The other suppressed defaults stay
  suppressed — their empty-string defaults already meant unset.

- Stored v1 card_overlays documents were quarantined because the
  planner validated them against the v2-only schema; the web parser has
  upgraded v1 at read time all along. The planner now applies the same
  v1-to-v2 upgrade before validation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): delete canonical library-scoped values with the library

The canonical settings schema deliberately has no FK on library_id or
series_id, and the migration comment promised the owning delete paths
would clean these rows up — but nothing called
DeleteSettingValuesForLibrary/-Series outside stores and tests, so a
deleted library left orphaned profile_library preferences in every
user's store forever.

Adds userstore.SettingValuesCleaner, a per-user best-effort sweep in the
mutation-sweeper's mold, and wires it into the library delete job. The
series-side cleanup is exposed on the same cleaner for the scanner's
orphan pruning to adopt; series have no single delete executor today.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): keep off-step playback speeds working until cutover

The legacy device endpoint gained step enforcement mid-branch, turning
an existing in-range PUT of 0.26 from 204 into 400 — a behavior change
on a live /api/v1 endpoint before the coordinated break, which the v1
rules forbid. The legacy validator is back to range-only; the typed
mutation endpoint keeps enforcing the manifest's step.

The migration planner now snaps stored off-step numbers onto their
definition's step grid instead of quarantining them: a stored 0.26 is a
real preference, and every client's stepper was going to snap it on the
next write anyway.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): serve canonical values to the readers the cutover stranded

Four P1 review findings where the web writes canonical rows the server
never reads — and the legacy keys those readers use are now unwritable,
so the values are frozen and user edits silently do nothing:

- access.DisabledLibraryIDs and the policy viewer resolver read the
  legacy account key while the library screen writes profile-scoped
  ui.disabled_library_ids. Both now resolve the canonical profile row,
  falling back to the legacy key only when no canonical row exists.

- The sections fetcher and handler read the legacy next_up_mode account
  key while the playback screen writes ui.next_up_mode. Same ladder,
  behind one shared sections.NextUpMode helper.

- Profile creation committed the profile and then synced settings
  non-atomically, so a mid-sync failure left a profile the retry could
  not recreate (name conflict) with preferences that read as contract
  defaults forever. The create path now compensates by deleting the
  profile it created.

- The mounted legacy PUT /subtitle-prefs/{series_id} wrote only
  user_subtitle_preferences, but item detail resolves those three keys
  canonically, so a post-upgrade client's "subtitles off" returned 204
  and was ignored. The legacy handler now dual-writes the canonical
  profile_series rows, and its delete clears them.

Plus the migration's disposition for stranded Apple device-scope audio
language rows: nothing read them before the contract, so promoting them
to real overrides would change track selection at upgrade. They are
recorded in the rejects table instead of copied.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): make canonical settings writes take effect

Three P1 review findings on the web half of the cutover:

- Every profile-default editor reads the resolved value but writes the
  profile row, so a device override — left by the migration converting
  legacy user_device_settings, or written by another client — kept
  shadowing the save and snapped the control back with no affordance to
  remove it. useAutoPlayNextSetting already solved this for one key;
  that logic is now a shared useProfileDefaultWriter used by the
  playback screen, the quality picker, subtitle behavior, and the four
  appearance setters. It only clears when the key is device-scopable
  and the resolved value actually came from a device row.

- The appearance cache only ever grew: when the effective response
  resolved a key to "default" — because another client deleted it —
  the namespaced entry and local state survived and kept winning the
  fallback, so a removal never reached this browser. The mirror now
  runs both ways, clearing only on an explicit default answer (silence
  is not a deletion) and only within the current identity's namespace.
  Custom theme vars and CSS do the same, except while a local draft is
  unsaved.

- The quality picker wrote the canonical two-axis keys while playback
  still derived its cap from currentProfile.quality_preference, a
  legacy compound column the canonical write deliberately does not
  mirror — so choosing a quality changed nothing about what played. The
  watch route and both item-detail pages now read
  playback.preferred_quality, falling back to the profile column until
  the settings read resolves so playback never blocks on it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ci): repair the type error and the bindings gate's job placement

Two breaks from the previous commits:

- useTheme referenced storage.StorageKey, but storage is a value, not a
  namespace — the Web job's tsc caught what the local incremental
  typecheck had already cached past. Imported the type properly.

- verify-settings-bindings gained a prettier step, and the Go job that
  runs it has no pnpm, so the check failed on its own tooling rather
  than on a stale binding. Split the web half into
  verify-settings-bindings-web and moved it to the Web job, which has
  pnpm; verify-settings-bindings-all runs both locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): close the second-round review findings

Three from the review of the pushed work:

- The live profile sync omitted auto_skip_intro/credits/recap and
  auto_play_next_preview, which my own change made load-bearing: the
  player now resolves those keys canonically, so a legacy PUT /profiles
  moved the columns, returned 200, and changed nothing about playback.
  All four now mirror on write. The DTO read block keeps its shape —
  clients pin it — and its columns are what the sync keeps current.

- The effective endpoint dropped unknown keys silently, letting a
  client fill the gap with its own vendored default and present a value
  this server would refuse to store. Unknown keys now 404 by name.

- Two sidebar-pin toggles in flight at once could commit in either
  order, and the server upsert is last-write-wins, so the first request
  landing second restored the pre-toggle document. The writes are now
  chained, and each link reads the document when it runs, so a queued
  toggle sends the newest state rather than the one it was queued with.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(database): make the settings-contract deploy reversible

Rolling back this release meant restoring a backup, for a reason that
was not obvious: the DisplayPreferences move deletes the jellycompat
rows from user_settings once it has copied them, and the previous
binary reads exactly those rows. An older server therefore starts
cleanly and silently serves defaults, so every Jellyfin client's saved
view preferences look reset.

The down functions were already written and correct — nothing could
invoke them. The backfill and the DisplayPreferences move are Go
migrations registered in-process, so the standalone goose CLI in the
Makefile cannot see them, and the server exposed only --migrate-only
and --migrate-status.

Adds MigrateDownTo, the --migrate-down-to flag, and a make target, plus
a rehearsal test that seeds a legacy row the way the old binary wrote
it, applies the move, rolls back, and asserts the row returns
byte-for-byte.

Documents the ordering in the spec's cutover section, including the two
caveats an operator needs beforehand: take a backup, and the per-user
SQLite backend cannot be rolled back at all — its migrations have no
down path and an older binary refuses to open a newer database, so
those installs restore from backup rather than degrade.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): skip legacy rows whose profile was deleted

The dev-server migration aborted on real data:

  writing playback.subtitle_appearance at profile_device for user 1:
  violates foreign key constraint user_setting_values_profile_fkey

user_device_settings carries an ON DELETE CASCADE on (user_id,
profile_id) today, but rows written before that constraint outlived the
profiles they belonged to — that install had 46 such rows across 14
deleted profiles. The planner copied them faithfully and the canonical
table, which declares the same foreign key, refused them; because the
backfill runs in one transaction, the whole migration failed and the
server could not start.

An override belonging to a profile nobody can select is not a preference
anyone can be shown or reset, so Plan now drops those rows rather than
repairing them, recording each in user_setting_migration_rejects so an
operator can see what was left behind. Account-scope rows carry no
profile and pass through untouched.

Verified by replaying that install's 514 device rows through the
planner: 9 rows would have hit the constraint before, 0 after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(settings): address canonical cutover review findings

* fix(settings): address latest review findings

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:52:41 -04:00
CoffeeKnyteandQuick 9281ae61d9 fix(playback): transcode H.264 sources with conflicting in-band PPS
H.264 streams that redefine the same pic_parameter_set_id in-band with
different content cannot be safely stream-copied into an avc1/fMP4 HLS
segment: the avcC advertises a single parameter set, so VideoToolbox
(Safari/Chrome on macOS) decodes with the wrong PPS and desyncs mid-GOP,
surfacing as PIPELINE_ERROR_DECODE / kVTVideoDecoderBadDataErr (-12909).

At playback start, a bitstream scan (DetectMultiplePPSH264) runs for H.264
files on the probe-ensure path, grouping in-band PPS by id and flagging any
id carrying more than one distinct definition. The result is a runtime-only
flag (VideoTrack.MultiplePPS, json:"-") memoized per process — never written
to the database, no schema change, recomputed on the first play after a
restart.

The v3 planner and legacy resolver disqualify a copy-unsafe source from the
video stream-copy / remux ladder, routing it to a real transcode. Direct
play of the original container is left intact: decoders that reparse in-band
parameter sets (ExoPlayer, VLC, native) handle the source fine.

Part of #135.
2026-07-29 09:49:34 -04:00
c52ca7dd7a feat(admin): identify compat sessions and Android devices in the live session view (#495)
* feat(admin): identify Android devices by model in live session view

Android clients that send a bare default User-Agent (e.g.
"Dalvik/2.1.0 (Linux; U; Android 11; AFTKRT Build/RS8180.3729N)")
showed up as "Dalvik" in the admin live-session view, which tells an
operator nothing about the device.

Parse the model code out of the UA (the token between the last ';' and
"Build/") and map the Amazon Fire TV family and NVIDIA Shield to product
names. Unknown but parseable models fall back to "Android · <MODEL>"
instead of "Dalvik"; multi-word models like "Pixel 7" are preserved
whole. This is display-only: the session still stores the raw model code
in its user agent, and no response field or contract changes.

* feat(admin): mark Jellyfin-compat sessions with the JF pill by origin

The admin "JF" pill was derived at read time by substring-matching a
token list against the client name / user agent. A real Jellyfin
client that authenticates through the compat surface but sends a bare
User-Agent and no MediaBrowser client name (e.g. a Fire TV app) got no
pill, even though it plainly came through the Jellyfin API.

Stamp compat origin as immutable identity at session creation and carry
it through to the admin view:

- ClientInfo.IsCompat is set true in the jellycompat auth path; newSession
  copies it onto Session.IsJellyfinCompat.
- The flag rides the durable RecipeCard (next to the client metadata that
  already exists so the pill survives reconstruction) and is restored in
  ReconstructSession, so a server restart keeps the pill.
- buildLiveSessionSync -> worker.SessionSync -> a new compat_origin column
  on playback_sessions_sync (added migration); the reconciler upserts,
  reloads, and compares it so origin changes still publish and unchanged
  rows do not churn.
- The handler ORs the stored origin with the existing name/UA heuristic,
  which stays as a fallback for rows written before this column existed.

is_jellyfin_client keeps the same name and type on the wire; it is only
sourced more accurately.

* fix(admin): correct Android device labels

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-28 21:27:58 -04:00
54e184df85 feat(requests): enforce per-profile rating limits in discovery (#505)
* feat(requests): enforce per-profile rating limits in discovery

- Resolve each profile's max content rating and filter discovery, detail, and browse results against it, failing closed on missing ratings
- Reject request submissions for titles above the viewer's ceiling
- Add TMDB GetCertification backed by release_dates/content_ratings with a long-lived cache and singleflight
- Push certification.lte to TMDB for studio/network/genre browse as a cost pre-filter
- Backfill restricted section pages from a fixed window of TMDB pages to keep carousels populated and pagination stable

* fix(requests): address discovery rating review findings

- Preserve backfill overflow: sections use plain TMDB cursor semantics
  plus an additive next_page field instead of fixed windows, so an early
  stop never drops allowed titles from unconsumed pages (bit hardest at
  permissive R/TV-MA ceilings).
- Bound cold-path cost: DiscoverAll backfills at most 2 TMDB pages per
  section (vs 5 for a direct section request), capping worst-case cold
  certification hydration at 240 lookups instead of 600.
- Keep the TMDB prefilter a superset: rank-3 ceilings now push down
  certification.lte=NC-17/TV-MA rather than R, so titles the local
  ladder allows can't vanish upstream unrecoverably.
- Fail closed on foreign certifications: enforcement-path lookups use
  new US-only pickers (a Canadian PG no longer reads as US PG), while
  the display path keeps its any-country fallback. US multi-entry
  disagreements prefer the theatrical/real rating over festival NR.
- Detach shared certification fetches from the first caller's context
  (WithoutCancel + 30s bound) so one disconnecting client can't fail
  the singleflight result for concurrent waiters.
- Advertise enforcement via rating_restrictions_enforced on
  /requests/status so clients can feature-detect instead of
  version-sniffing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(requests): harden rating enforcement per second review pass

- GetDetail gates on the US-only enforcement certification (cached
  GetCertification) instead of the display rating, whose any-country
  fallback let a foreign "PG" pass the US ladder.
- pickUSMovieCertification takes the strictest recognized US rating when
  multiple release entries disagree ([PG, R] -> R); entry order is not
  meaningful and enforcement must not admit a title on its most lenient
  certificate.
- Certification singleflight uses DoChan so a canceled caller returns
  ctx.Err() immediately instead of blocking up to 30s on the detached
  shared fetch (which still completes for surviving waiters).
- Viewer rating ceiling resolves once per request and threads through
  discover/browse/detail enrichment (enrichPageWithCeiling); DiscoverAll
  drops from 12 scope resolutions per load to 1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 22:34:30 -04:00
22fec4ed2d feat(metadata): add resilient match queue diagnostics (#463)
* feat(metadata): add resilient match queue diagnostics

* fix(metadata): harden match queue lifecycle

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-24 12:18:49 -04:00
383973ec22 feat(metadata): improve match accuracy and localized titles (#461)
* feat(metadata): improve match accuracy and localized titles

* fix(metadata): address matching review findings

* test(catalog): align empty alias snapshot scope

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-24 12:02:52 -04:00
Quick104 3c56ed606c fix(admin): enforce settings contracts end to end 2026-07-23 11:24:55 -04:00
QuickandGitHub c0b2e18173 Merge pull request #433 from RXWatcher/fix/ebook-enrichment-architecture
feat(ebooks): decouple metadata enrichment from scans via a durable queue
2026-07-23 11:18:42 -04:00
Quick104 f6615ef8d7 fix(jellycompat): confirm authoritative watch stops 2026-07-23 09:15:19 -04:00
Quick104 166c5ef32f Add reliable Jellycompat watch scrobbling
- Forward start, pause, resume, and stop events with stable media identities
- Persist and retry terminal scrobbles across teardown and restart paths
- Reject ambiguous playback-report route matches
2026-07-22 21:41:05 -04:00
rxwatcher 8d40138bdb feat(ebooks): drain enrichment backlog with progress 2026-07-22 13:43:57 +02:00
rxwatcher 1c2d422548 fix(ebooks): bound enrichment lease work 2026-07-22 13:43:57 +02:00
Quick104andClaude Fable 5 4fa84a661a feat(diagnostics): client diagnostics server foundation
Implements slice 1 of docs/design/2026-07-19-client-diagnostics.md: the
versioned contract (schemas, fixtures, Go validator), storage-validated
diagnostics.uploads_enabled gate, account-scoped status endpoint, hardened
streaming multipart ingest with quota reservation and a receiving/ready/
failed report state machine, S3 streaming puts, acting-admin report API
(list/detail/download/delete with audit events), and the retention +
orphan-reconciliation cleanup task.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XppCCycoaskCsW7ja1fZct
2026-07-20 11:13:52 -04:00
91e1164090 feat(metadata): local NFO metadata and sidecar artwork (builtin chain provider) (#390)
* feat(metadata): register builtin NFO provider and broaden parsing

Phases A and B of the #216 local-NFO work, implemented test-first.

Registration & hint-first identity (Phase A):
- Migration seeds a reserved kind='builtin' silo.builtin installation
  and an 'nfo' metadata capability (default_enabled=false, priority 1
  for movie/series) with a partial unique index and documented Down.
- In-process builtin provider registry (internal/metadata/builtin.go);
  buildProviders returns the registered provider for builtin rows.
- Guard rails keep the reserved row out of every plugin surface (user
  plugin-settings, installations list, image resolvers, preload,
  auto-update, store Delete, mutation handlers -> 409); silo.builtin is
  a reserved manifest id.
- Startup sync materializes legacy content_level='' chains per level,
  then appends builtin capabilities disabled via
  AppendProviderToAllChains (idempotent); resolveEnabledProvidersBy
  priority now respects default_enabled=false.
- NFO uniqueids seed the trusted-hint machinery via IdentityHintProvider
  with per-mode conflict policy (stored IDs win on scheduled refresh,
  NFO wins on manual refresh, Identify skips NFO); ID-less candidates
  are excluded from provider-priority tie-breaks and nfo never counts
  as corroboration.
- Web chain-editor empty-state gate is now server-derived so builtin
  providers are reachable on plugin-less servers.

Parser breadth & sidecar hardening (Phase B):
- Parser covers the practical Kodi/Jellyfin field set for <movie> and
  <tvshow>: original title, tagline, runtime, dates, content rating,
  genres/studios/countries/tags, multi-source ratings with scale
  normalization, cast with roles/order, director/credits. Empty
  collections stay nil so merge early-returns apply.
- findNFO parses candidates and falls through on read/parse failure or
  root-type mismatch, so a stray movie.nfo cannot shadow tvshow.nfo;
  GetMetadata gains the same ContentType guard Search has.
- New FieldReleaseDates lock gates Year/ReleaseDate/First+LastAirDate
  in merge (Go) and the edit-metadata dialog (web), closing the gap
  where a manual refresh re-applied NFO dates over admin corrections.
- Merge-contract tests pin NFO fill semantics, genres whole-list
  first-provider-wins, and NFO edits propagating on manual refresh only.
- Docs: new admin wiki page (supported fields, merge semantics,
  naming-supplies-structure contract), index bullet, sidecar wording
  revision, v1-scope feature-detection note.

Zero behavior change while the provider is disabled (default); pinned
by CI-mode and DB-gated test suites.

Part of #216

AI-use disclosure: implemented with Claude Code (Fable 5) via
spec-driven TDD and agent-assisted implementation.

* feat(metadata): ingest local sidecar artwork and read series-depth NFO

Phases C and D of the #216 local-NFO work, implemented test-first, plus
the mixed-library use-case pins. Together these deliver the headline
case: a series absent from every remote database (e.g. a fitness
library) scans into a fully presented show -> named seasons -> titled
episodes tree from NFO files and sidecar art alone.

Local sidecar artwork through the S3 image cache (Phase C):
- The NFO provider implements ImageProvider: poster/backdrop/logo
  sidecar discovery with a fixed precedence map, symlink/non-regular
  rejection, an 8 MiB cap, and file:// source URLs at rating 0. Generic
  filenames apply only via the sidecar search paths, so a shared
  folder.jpg in a flat multi-movie directory applies to none.
- file:// becomes a live local source scheme: routed into *_source_path
  (never *_path), accepted by every image enqueue gate, attributed as
  provider "local", excluded from cached-path detection.
- The image-cache processor caches local files with lexical-on-logical
  confinement to the library roots, open-handle reads with re-checks,
  the same variant widths as remote art, and stable (7-day) failure
  classification. Keys land under
  local/{contentType}/{contentID}/{hash8}/{imageType}; superseded
  prefixes are cleaned on re-cache and item deletion.
- applyIfBetter gains a local exemption so rating-0 local art can fill
  matched items without being stickily displaced; ImageRequest carries
  additive sidecar path context.

Series depth (Phase D):
- SeasonsRequest/EpisodesRequest carry additive local path context
  (series roots, per-season directories, per-episode file paths),
  derived from naming at match time and reconstructed on refresh.
- season.nfo supplies season name/plot; NFO season numbers are advisory
  (directory-derived number wins with a Warn - naming owns structure).
  <episodedetails> gains aired/runtime/ratings; <basename>.nfo titles
  episodes and <basename>-thumb.ext supplies thumbs; filename SxxEyy
  wins over NFO numbers.
- Episode NFOs work without a season.nfo (provider seasons unioned with
  on-disk seasons); SynthesizeFallbackEpisodes always runs after persist
  so NFO-less episodes keep synthesized rows. Season/episode file:// art
  rides the Phase C pipeline unchanged.
- Migration adds season:1/episode:1 to the builtin NFO capability's
  default_priority (still default_enabled=false).

Mixed sports-library use case (tests only, no product change):
- Pins the classification contract for one library holding movie-shaped
  and show-shaped content (WWE PPV events as movies next to a "WWE
  SmackDown" show, NASCAR/F1/FIFA with partial TVDB/TMDB data): naming
  decides movie-vs-series per file before any provider runs; the NFO
  supplies metadata/identity but never flips type (ContentType guard);
  the per-root Type override is the correction path.
- NFO-driven type classification at scan time is recorded as an explicit
  deferred open question.

Part of #216

AI-use disclosure: implemented with Claude Code (Fable 5) via
spec-driven TDD and agent-assisted implementation.

* docs(metadata): document local NFO metadata architecture

Add a single as-built architecture page
(docs/architecture/local-nfo-metadata.md) for the #216 local-NFO
feature: the builtin registration model, hint-first identity semantics,
the file:// -> S3 artwork pipeline and its deployment constraint, series
depth, the mixed-library classification contract, and known limitations.

This replaces the working implementation plan, the per-phase specs, and
the narrow sidecar-artwork note, which were planning drafts and are left
untracked; admin-facing behavior remains in the wiki.

Part of #216

AI-use disclosure: planned, drafted, and consolidated with Claude Code
(Fable 5) using multi-agent exploration and adversarial review.

* fix(metadata): address PR review findings on NFO builtin provider

Fold in the valid, low-risk fixes surfaced by automated review on #390:

- imagecache: extract validateCacheRequest so CacheBytes (the local
  sidecar season/episode path) enforces the same episode-requires-season
  guard as Cache, preventing distinct episodes' art from colliding under
  one S3 key.
- image_cache_processor: close the sidecar symlink-swap window by
  rejecting the opened handle unless os.SameFile matches the Lstat'd
  file, so a leaf swapped to a symlink can't pull an out-of-root target
  into the public cache.
- plugins: guard the reserved builtin installation row in the store's
  Update, matching Delete, so its version/enabled/capabilities can never
  be rewritten even if a mutation slips past the HTTP layer.
- cmd/silo: bound SyncBuiltinProviderChains with a 30s timeout so a stuck
  DB round-trip fails fast at startup instead of hanging.
- metadata: panic instead of silently no-op'ing on an invalid
  RegisterBuiltinProvider call (init-time programmer error).
- docs: correct the media-folder-and-naming NFO paragraph to state
  season/episode NFOs and sidecar artwork are actively read.

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-16 17:55:36 -04:00
1664c60425 fix(metadata): publish artwork revisions atomically (#399)
* fix(metadata): publish artwork revisions atomically

* fix(metadata): harden artwork revision cleanup

* fix(metadata): address artwork revision review findings

- restore image applies for all media_items types and reject unsupported
  target/image combinations with 400 before uploading; episodes coerce to
  stills and the web dialog no longer offers image tabs episodes can't use
- add WHEN clauses to displacement triggers and hoist to_jsonb so bulk
  catalog upserts that assign unchanged artwork columns skip the trigger
- make artworkkey the single variant-ladder owner: imagecache derives its
  widths from it and triggers store image_type instead of hardcoded
  variant arrays, expanded by the collector at deletion time
- sweep dormant registry rows periodically so references lost through
  untriggered surfaces degrade to slow cleanup instead of leaking
- park just-published revisions dormant, keep dormant rows dormant on
  re-cache, and batch the GC reference pre-check per run
- heal rows re-referencing a just-deleted revision via reconciler-style
  resets after the deletion commits
- share a per-URL image-loaded hook across DetailHero, ItemCard,
  SectionItemCard, GlobalSearch, and CollectionPosterCard
- deduplicate Cache/CacheBytes finalization and drop unused VariantPaths
  plumbing

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(catalog): cast reused timestamp parameter in revision upsert

Postgres cannot deduce one type for $3 used both as a plain value and
inside a CASE arm; the dev deploy surfaced it as SQLSTATE 42P08 on every
publication. Cast both uses and cover the arm/park/track upserts with
database-backed tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(metadata): address artwork revision review comments

- keep a durable heal path: deletion marks deleted_at instead of removing
  the registry row, so a failed post-delete heal retries with backoff and
  broken references never park; trackers clear the marker on re-upload
- never treat bare existence as an immutable-content match; backends
  without content verification rewrite the object
- exercise revisioned cover keys in scanner/enrichment fakes, compare the
  tracked manifest exactly, and honor cancellation in the blocking test
  deleter

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 17:33:21 -04:00
CoffeeKnyteandGitHub b3722dac58 fix(playback): open the listener before sweeping stale transcode dirs (#413)
* fix(playback): run orphaned-transcode cleanup in the background at startup

The native and Jellyfin-compat routers swept stale per-session transcode
dirs synchronously during NewRouter, before the listener bound. On a slow
network filesystem this blocked startup for 80+s (64 leftover dirs on the
last deploy), so restart-reconnect clients were turned away and the health
check reported the server unhealthy the whole time.

Move both sweeps into a background goroutine (StartBackgroundOrphanCleanup)
so the listener comes up immediately and the cleanup runs concurrently. The
delete logic is unchanged: same active-session snapshot and MaxTokenTTL
age-sparing, only later. A package-level mutex serializes concurrent sweeps
of the shared transcode root so the two background sweeps can't race on
os.RemoveAll.

Part of #412

* fix(transcode): background the node boot-time transcode-dir sweep

A dedicated transcode node swept leftover transcode dirs synchronously in
NewServer, before startStandaloneServer bound its listener. On a slow
network filesystem that delete blocked the node from coming online at boot,
the same startup-stall class as the main server.

Move the sweep into the shared StartBackgroundOrphanCleanup goroutine so the
node's listener binds immediately. Backgrounding required an age guard: the
sweep previously ran as a full wipe (minAge=0) with an empty active-set,
which was only safe because it completed before any request could arrive.
Run concurrently that would race a token-carried reconstruct writing into
TranscodeDir/<sessionID>, deleting segments a fresh ffmpeg is producing.
Passing MaxTokenTTL spares any dir younger than the max token lifetime —
exactly the ones a still-valid reconnect could reconstruct — while dirs
older than any surviving token (never reconstructable) are still reclaimed.

Part of #412

* feat(playback): reclaim orphaned transcode dirs periodically, not just at boot

The orphaned-transcode sweep only ran at startup on both the central server
and transcode nodes, so it only ever reclaimed dirs left by an ungraceful
prior shutdown. During a long uptime the in-memory session reapers delete the
dirs of sessions they still track, but a dir whose owning session was dropped
without its RemoveAll succeeding becomes an "untracked orphan" with no runtime
GC — on a box that runs for weeks these accumulate until the next restart.

Add StartPeriodicOrphanCleanup: an immediate background sweep followed by an
hourly re-run bound to a lifecycle context. Wire it on all three surfaces —
native API and Jellyfin-compat (via deps.AppContext) and the transcode node
(via a new Server.StartOrphanSweeper(appCtx), replacing its boot-only sweep).
When no context is supplied (tests) it degrades to a single boot-time sweep so
no ticker goroutine outlives the caller. The sweep stays age-guarded at
MaxTokenTTL, so nothing reconstructable is ever reaped.

Because the node sweep now runs during live traffic, it snapshots the live
job set (Server.activeSessionIDs) and spares those dirs by id rather than by
age alone — a long-lived session that only re-serves already-written segments
stops advancing its dir mtime, which age could otherwise misclassify. In
integrated mode the native and compat sweeps share one TranscodeDir but each
snapshots only its own manager's live set; the resulting cross-manager reap of
a >24h idle dir is bounded (rebuilds from token/recipe) and documented at both
call sites.

Part of #412
2026-07-16 14:48:00 -04:00
28c6ddc237 feat(playback): add per-user transcoding controls (#375)
* feat(playback): add per-user transcoding controls

* fix(playback): enforce forced video transcode permission

* chore: address transcode control review feedback

* fix(playback): recheck transcode permission on audio switch

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-10 22:30:04 -04:00
d68e70bb47 feat(autoscan): Sonarr/Radarr webhook intake without arr API keys (#353)
* docs(autoscan): add arr webhook intake spec and implementation plan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoscan): add webhook intake schema migration

Adds delivery_mode to autoscan_sources, the autoscan_webhook_endpoints
table, and delivery_mode/provider_event_type on autoscan_events.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoscan): add built-in arr-webhook source identity

Host-discovered scan-source entry so webhook-mode sources need no
plugin installation; composite lister appends it to plugin discovery.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoscan): persist delivery mode, webhook endpoints, event metadata

Sources carry delivery_mode; autoscan_webhook_endpoints CRUD with
SHA-256 token lookup and AAD-bound encrypted redisplay; events record
delivery_mode/provider_event_type; CreateEvent gains SkipRunningCheck
so webhook deliveries are never dropped by the poll exclusion.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoscan): share the consume path and add webhook IngestChanges

Extracts consumeSourceChanges from PollOnce (marker semantics
preserved, existing poll tests unchanged); PollOnce skips webhook
sources; IngestChanges feeds deliveries through the shared pipeline
without markers and without the running-event exclusion.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoscan): add Sonarr/Radarr webhook payload parser

Host-side arrwebhook package: provider inference, import/rename/delete
path extraction with vanished-path-friendly previous paths, subtree
fallback, exact-path dedupe, and no-op unknown events. Fixture-backed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(autoscan): add public webhook delivery route and admin endpoint management

Public POST /api/v1/autoscan/webhooks/{token} with per-IP rate
limiting, 256KiB body cap, 202-for-noop semantics, and token/body kept
out of logs; admin create/rotate/delete endpoint routes; source
responses carry delivery mode + webhook status/URL; create/update
validate delivery mode against source identity.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add webhook delivery mode to Autoscan admin UI

Webhook sources get a generate/copy/rotate webhook URL section,
provider selector, delivery status, and a connection-free Add-source
flow; activity rows badge webhook deliveries with the arr event type.
Path rewrites stay editable in both modes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(api): redact secret path params from request and activity logs

The request logger and activity-log middleware recorded raw URLs, so
bearer credentials in secret path segments (autoscan webhook {token},
webhook-sync {secret}) were persisted to app logs and activity_log.
Redact the secret segment via the chi route params in both sinks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(autoscan): make webhook delivery reliable

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 14:13:31 -04:00
04c4344f52 feat(metadata): reconcile artwork cache after public S3 provider changes (#349)
* feat(metadata): reconcile artwork cache after public S3 provider changes

Changing the public S3 provider previously broke every cached image
permanently: the DB keeps bucket-relative keys, the image cache pipeline
treats a cached path as its durable dedup marker and never re-enqueues,
and clients eat the 404s straight from S3 so the server never notices.

Add a storage identity fingerprint (s3.public_storage_identity, seeded
via SetIfAbsent at boot) and a reconcile_artwork_cache task whose
startup trigger only fires when the identity changed; manual runs
always sweep, doubling as bucket-data-loss recovery. The task probes a
random sample of cached objects, then either bulk-resets (near-total
miss) or per-row verifies. Missing provider-sourced artwork is reset to
its *_source_path so the existing enqueue loop re-caches it; surfaces
without a re-downloadable source (chapter thumbnails, collection
artwork, library posters, branding refs, embedded book covers) are
cleared so their owning pipelines refill them. Small upload-holding
tables are always per-row verified so bulk mode cannot blind-clear an
upload that survived migration, and transport errors never reset rows.

Users never see broken images during the transition: reset rows serve
the provider's original URL via the existing absolute-URL pass-through
and thumbhashes are preserved. The storage settings page now warns that
uploads cannot be re-downloaded when the identity fields are edited.

Part of #348

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(metadata): harden artwork reconcile per code review

Address the confirmed findings from the PR review:

- Fingerprint the key prefix case-sensitively and slash-trimmed exactly
  as s3client applies it (new exported NormalizeKeyPrefix): a case-only
  prefix edit is a real storage move and must reconcile; a slash-only
  edit is not and must not.
- Certify the storage fingerprint immediately after the artwork sweep
  succeeds and make the 4-object branding check non-fatal (reported in
  the task message), so a transient branding error cannot discard a
  completed catalog sweep and force it to repeat every boot.
- Fail closed on conditional-task preflight errors in the task manager
  (previously fail-open ran the task), and retry transient settings
  reads in ShouldRun since the startup trigger fires once per process.
- Track probe HEAD errors against a separate baseline so a flaky probe
  cannot consume the sweep's error budget.
- Probe before counting: bulk mode skips the per-surface count(*)
  full scans entirely, and probe sampling drops ORDER BY random()
  (plain LIMIT answers "is the cache in this bucket" just as well).
- Verify chapter thumbnails across a whole 500-file batch in one HEAD
  fan-out instead of per file, keeping the worker pool saturated.
- Replace the 10 inline non-provider-scheme ARRAY literals in the
  enqueue query with the shared nonProviderImageSchemesSQL constant.

Part of #348

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(metadata): guard bulk reset against degraded probes, certify only clean sweeps

Address bot review feedback on the reconcile hardening:

- A probe where more than half the HEAD requests error aborts the run:
  errored requests are excluded from the sample, so a partial outage
  could otherwise present a handful of surviving 404s as a ~100% miss
  rate and bulk-reset the catalog. Bulk mode additionally requires a
  minimum number of successful samples; thinned probes and tiny
  catalogs take the safe per-row verify path.
- Track sweep errors separately from probe/branding errors
  (stats.sweep_errors) and certify the storage fingerprint only when
  the sweep completed with zero of them — skipped rows were never
  verified, so the next startup retries. Applied resets stay durable.
- Give each ObjectExists attempt its own timeout so a stalled HEAD
  fails that attempt instead of pinning the retry loop to the run
  context.
- Report branding assets checked (not just cleared) in stats.Checked.
- Drop the dead settingsRepo/brandingSvc nil guards in cmd/silo and
  sync spec numbers with the implementation constants.

Part of #348

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 12:00:24 -04:00
203a18ae83 feat(observability): OpenTelemetry logs+traces with secret redaction and slog standardization (#290)
* feat(observability): OpenTelemetry logs+traces with secret redaction

Part of #265. Adds opt-in OpenTelemetry (logs + traces) alongside the existing
stderr + opslog pipeline, plus secret redaction on all sinks. Default-off: with
no OTEL_* / SILO_OTEL_ENABLED config, behavior is unchanged.

Bootstrap (internal/telemetry):
- Setup() builds one shared resource, a TracerProvider (parent-based trace-id
  ratio sampler), a LoggerProvider, and the W3C TraceContext+Baggage propagator
  from env. It installs NO MeterProvider — metrics stay on Prometheus, and the
  built-in no-op global MeterProvider keeps the trace instrumentation libs from
  double-emitting. Shutdown is deferred with a flush timeout.
- Logs are bridged via otelslog fan-out (slog.MultiHandler), level-gated by the
  shared LevelVar and best-effort so a failing collector can't break the console
  or DB branches. stderr + opslog stay untouched.

Secret redaction (internal/logredact):
- A slog.Handler masks secret-keyed attributes (password, token, api_key,
  authorization, cookie, ...) — including .With-bound attrs, nested groups,
  secret-keyed group subtrees, and values behind a LogValuer — on the console
  and OTLP sinks, with a no-op fast path when a record has no secret keys.
  opslog.shouldRedact delegates to logredact.SecretKey so all sinks share one
  marker list.

Rotation is infra-managed (no custom file sink): container runtime for stderr,
collector/backend for OTLP, opslog partition-pruning for the DB. Documented in
docs/architecture/observability.md.

Verification: go build ./..., go vet, gofmt -l — clean; go test
./internal/telemetry/ ./internal/logredact/ -race pass.

AI-use disclosure: implemented with AI assistance (Claude Code), including
adversarial reviews that hardened the bootstrap and fixed two redaction leak
paths; reviewed by the author.

* refactor(observability): slog context+component sweep, sloglint gate (phase 3)

Part of #265. Builds on the OTel bootstrap + redaction commit.

Standardizes every log call site onto the context-carrying slog variants so
records correlate with the active OpenTelemetry trace, and locks the standard
in with a machine gate so future code (human- or AI-authored) can't drift back.

- Call-site sweep: converted the remaining slog.<Level>(...) calls to the
  slog.<Level>Context(ctx, ...) form wherever a context.Context is in scope
  (background/init calls with no ctx are left as-is), across 183 files. Applied
  via a type-aware AST codemod. Log levels and message strings are preserved
  verbatim; a component attr (canonical per-package name) is added to direct
  package-level slog calls. Bound-logger calls keep their existing .With
  bindings. The main.go and telemetry package conversions rode with their file
  in the previous commit to keep each file within a single commit.
- Enforcement (.golangci.yml): enable sloglint with context=scope, static-msg,
  key-naming-case=snake, no-mixed-args. After the sweep all four report zero
  violations repo-wide (tests included), so make lint / CI now blocks any
  regression to the non-context form. The gate ships with the sweep because it
  cannot be green until the legacy sites are converted.

Metrics remain on Prometheus; no behavior change to /metrics or Grafana.

Verification: go build ./..., go vet ./..., gofmt -l — clean; sloglint (all 4
rules) 0 violations repo-wide; log levels verified unchanged.

AI-use disclosure: implemented with AI assistance (Claude Code), including the
codemod; reviewed by the author.

* fix(observability): honor per-signal OTLP protocol and secret WithGroup names

Two Codex review findings on PR #290:

- telemetry: OTEL_EXPORTER_OTLP_{TRACES,LOGS}_PROTOCOL now override the
  generic OTEL_EXPORTER_OTLP_PROTOCOL per signal, so mixed collector
  setups (e.g. HTTP logs + gRPC traces) build the right exporter.
- logredact: entering a group whose name is secret-bearing (e.g.
  WithGroup("authorization")) now masks every leaf in that subtree,
  matching how slog.Group("authorization", ...) is masked as a whole.

* fix(observability): address review feedback on telemetry bootstrap

- Telemetry setup failure no longer kills boot: Setup returns usable
  no-op providers alongside the error and main logs and continues with
  telemetry disabled, honoring the best-effort contract.
- Honor OTEL_TRACES_SAMPLER (always_on/off, traceidratio, parentbased_*
  variants); unsupported values fall back to parentbased_traceidratio.
- Attach node identity as semconv service.instance.id instead of the
  non-semconv node.name.
- Rename opslog retention-scope log attrs to target_component/target_level
  so they no longer collide with the canonical component routing key, and
  tag those lines with component=opslog.
- Fix stale levelGated comment casing; use WarnContext in the telemetry
  shutdown defer; document the LogValuer double-resolve on the redaction
  slow path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 08:53:52 -04:00
QuickandClaude Fable 5 28340ad6f5 feat(watchlist): hide fully-watched series instead of removing them
Removing a series from the watchlist on full watch stranded it once new
episodes aired: nothing ever re-added it. Split the behavior by type:

- watchlist.Maintainer now auto-removes only fully-watched movies (still
  propagating removals to connected providers).
- Series stay on the watchlist; the new catalog.WatchlistVisibility
  filter hides series whose available episodes are all completed on the
  display surfaces (sections rail, catalog watchlist source, GET
  /watchlist). A newly added episode makes the series reappear on the
  next fetch, and nothing is synced upstream since the entry never
  leaves the list.

Sync, recommendations, notifications, and the watchlist check endpoint
intentionally keep seeing the full list. The filter honors the existing
per-profile remove-watched preference and uses batch lookups only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 19:04:42 -04:00
42602b7896 feat(policy): access groups + embedded OPA policy engine with decision audit log (#282)
* docs(policy): add OPA policy engine design spec and implementation plan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* build(deps): add OPA v1.18.2 SDK for the policy engine

Pulls github.com/open-policy-agent/opa v1.18.2 (policy engine core for
the upcoming internal/policy subsystem) and the transitive upgrades go
mod tidy applied (otel 1.44, grpc 1.81.1, prometheus/common 0.67.5).
Full build verified.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add OPA engine core, vendor scope policy, and parity suite

New internal/policy package (dead code — nothing wires into request paths
yet): prepared-query Engine with 25ms eval timeout and fail-closed decode,
typed PDP.ResolveViewerScope, go:embed vendor bundle, capabilities lockdown
for future admin-authored Rego, and vendor scope.rego reproducing
access.Resolver.Resolve (library intersection, disabled-library handling,
quality/rating ceilings) with a narrowing-only silo_custom.scope.override
extension hook.

Parity proven by 1368 dual-execution subtests against the real
access.Resolver, including the nil-vs-empty AllowedLibraryIDs battery and
quality/rating variation; rank tables are test-pinned to internal/access.
Rego unit tests run via opa/v1/tester inside go test. Bench:
~106µs/op per scope decision incl. input marshaling.

Also restores the OPA requirement to go.mod (the earlier deps commit ran
go mod tidy before any import existed, so tidy dropped it).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed,
corrected (quality.allowed raw-file-rank divergence), and verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add policy document store, foundation schema, and compile-check

policy_foundation migration: policy_documents (one enabled doc per domain
via partial unique index — two enabled docs would define override twice
and conflict at eval), immutable policy_document_versions, single-row
policy_generation counter, and the partitioned policy_decisions log table
(daily range partitions, no FK, denial partial index).

PolicyStore: transactional version numbering (FOR UPDATE), activation
that verifies compiled_ok and bumps the generation in the same tx,
enable/disable with typed ErrDomainAlreadyEnabled, and a delete guard for
documents with an active version. CompileCheck sandboxes admin Rego:
locked capabilities (no http.send/net.*/opa.runtime), enforced
silo_custom.<domain> package path, vendor+stub layering, 2s budget,
structured row/col errors. Engine gains NewEngineWithCustom /
NewEngineFromStore with WARN-and-skip for invalid custom rows.

DB-backed tests verified against a migrated Postgres (concurrent version
numbering, atomic generation bumps, activation guards).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here (domain constants extracted).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add policy System lifecycle with hot reload and cross-node invalidation

policy.System owns one long-lived Engine and reloads it in place when
policy documents change: EventPolicyChanged on the existing ChannelAdmin
bus (new cache event constant) plus a 60s generation-poll fallback for
Redis-less deployments, with a generation-consistent snapshot read.
Vendor compile failure is startup-fatal; store/custom failures degrade
to vendor-only and the poll loop heals them; runtime reload failures
keep the last known-good engine. NotifyChanged gives the future admin
handlers synchronous local reload + cross-node publish.

Wiring: constructed in integrated/api modes only, PolicySystem field on
api.Dependencies (unused by routes yet), policy.eval_timeout_ms setting
(hot-reloaded via configWatcher.OnChange; default 25ms). Verified by a
full server boot smoke and DB-backed convergence tests (event + poll
paths, degraded boot, last-known-good).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add async decision logging with sampling, retention, and query repo

DecisionLogger batch-inserts each node's policy decisions straight to
the partitioned policy_decisions table via a non-blocking buffered
channel (drop-and-count on overflow — logging never adds latency to or
fails a decision). Scope decisions sample 1-in-N (default 50, setting
policy.decision_log_scope_sample_rate); denials and eval errors always
log; input/result JSON samples only at policy.decision_log_verbosity=
verbose. Cursor-paginated DecisionRepository backs the upcoming admin
log viewer. Retention via partman (daily partitions) and a
PolicyDecisionLogCleanupTask honoring policy.decision_log_retention_days
(default 14). PDP emits entries per evaluation; the System owns the
logger lifecycle and settings hot-reload.

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here (removed an unused, unsynchronized PDP setter).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(api): add admin policy management API and capability endpoint

/api/v1/policy/capability (authenticated feature detection) plus the
acting-admin /api/v1/admin/policy surface: vendor Rego viewer, document
CRUD with the one-enabled-per-domain conflict mapped to 409, immutable
version creation (compile-checked; failed versions persist as audit
history with structured row/col errors and can never activate),
activate/rollback with synchronous reload + cross-node invalidation via
System.NotifyChanged, stateless validate, throwaway-bundle simulate
(never touches the live engine, never logs decisions), and
cursor-paginated decision-log queries. Routes mount only when the
policy system is wired, keeping proxy/transcode modes untouched.

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here (seeded the FK'd test user; replaced an unchecked
fmt.Sscanf with strconv).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add /admin/policy workspace with Rego editor, simulate, and decision log

New Policy admin page (System nav group): documents list with
one-enabled-per-domain conflict handling, CodeMirror 6 Rego editor
(hand-rolled StreamLanguage mode) with server compile issues rendered as
inline lint diagnostics, explicit Save-version vs Activate flow with
confirm, read-only vendor module viewer, simulate panel with seeded
example inputs, version history with rollback, and a cursor-paginated
decision-log browser. Capability-gated via /policy/capability. Adds the
three decision-log settings to Log Retention. First code-editor
dependency in web/ (@uiw/react-codemirror + @codemirror/*), decided in
the design spec.

Implementation drafted by Codex (GPT-5.5) via codex exec; verified here
(lint, format:check, tsc --noEmit, vitest policy suites).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): make OPA authoritative for viewer scope resolution

policy.ViewerResolver implements the ViewerResolver interface backed by
PDP.ResolveViewerScope and replaces access.Resolver at all five
construction sites: router viewer middleware, notifications scopes, the
reconciler, jellycompat's scope filter, and the ABS resolver (which now
accepts a pre-built resolver, preserving its PIN-at-login semantics).
PIN/profile-token verification and disabled-library loading are
extracted into shared exported helpers used by both implementations, so
the legacy resolver stays compiled as the parity reference with
identical behavior. The adapter lives in internal/policy (which already
depends on internal/access transitively) — direct typed PDP calls, no
new import cycle. Sites without a policy system (proxy modes, bare test
routers) keep the legacy resolver until the cleanup phase.

Verified: full test suite green (jellycompat TestBeginWebOperation* and
one playback GPU test are pre-existing failures, confirmed identical on
main), 1368-case parity suite, dedicated ViewerResolver parity/PIN/
nil-vs-empty/fail-closed tests, and a full server boot smoke.

Implementation drafted by Codex (GPT-5.5) via codex exec; a first-pass
reflection-based adapter was rejected and reworked into the typed
in-policy adapter; reviewed line-by-line and verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): make OPA authoritative for acting-admin and permission gates

vendor/permission.rego reproduces the acting-admin rule (admin role +
primary-profile-or-none), HasEffectivePermission semantics for
marker_edit, and the metadata-curation rule including the subtle
admin-past-refused-bypass case that requires the explicitly ASSIGNED
permission. Policy-backed middleware in policy_gates.go keeps all Go-side
lookups (declared-profile primary check, item->library resolution, the
404-on-unknown-item path) and preserves the legacy status/body taxonomy
exactly — proven by dual-execution middleware tests that run every
scenario through both implementations and assert byte-equal responses.
Permission decisions always log (allowed flag populated); simulate and
the capability endpoint gain the permission domain automatically via the
domain registry. Router swaps behind single constructor choice points
with the legacy gates retained for policy-less wiring.

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): make OPA authoritative for download and playback admission decisions

vendor/action.rego decides download eligibility (downloads enabled +
user allowed), download-transcode eligibility (transcode enabled + user
allowed + artifacts available), and playback admission (stream/transcode
counts vs limits, zero = unlimited), with a tightening-only
silo_custom.action override that can also clamp a quality ceiling (never
widen — merged via quality.min). Go keeps everything stateful: config
loading, preset-ladder enumeration, and live session counting.

Downloads consult an optional ActionDecider (nil = legacy logic) mapped
back to the existing sentinel errors and capability response. Playback
gains a minimal AdmissionDecider hook at the exact point of the legacy
limit comparison: counts snapshot under the session mutex, PDP evaluated
OUTSIDE the lock, then revalidated under lock before insert (retry on
count drift) — no admission ever decided on stale counts and no eval
under the mutex. Deny reasons map to the legacy ErrTooManyStreams /
ErrTooManyTranscodes sentinels, pinned by tests.

Parity: combination tables driven against the real PresetsFor /
ensureTranscodeAllowed / SessionLimits math; full suite green (known
pre-existing jellycompat flakes only).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed
(locking design verified line-by-line) and verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): satisfy tsc -b strict return typing in the Rego stream tokenizer

The production build (tsc -b) rejects assigning CodeMirror's
string | void next() result to string | undefined; tsc --noEmit did not
catch it. Restructured the string-literal loop.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): clearer error when a decision is undefined for partial input

Vendor policies index required input fields directly, so a hand-written
simulate payload missing fields yields an undefined decision. Surface
that as 'decision X is undefined for this input (missing required input
fields?)' instead of 'empty result' — found while exercising the
simulate API against a live server.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(web): set changeOrigin automatically when the API proxy target is remote

Remote dev backends sit behind vhost-routing proxies that reject a
localhost Host header; local targets keep the existing pass-through
behavior. Enables pointing the Vite dev server at a hosted backend via
VITE_API_PROXY_TARGET in web/.env.local.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): redesign the policy workspace around the decision pipeline

The first-pass UI was structurally generic: a five-column document table
squeezed beside the editor, three equal-weight action buttons with
hidden preconditions, raw version IDs, and jargon copy — nothing taught
the model. The page now teaches it:

- A pipeline strip states the mental model up front: Silo decides the
  baseline -> your overrides narrow it -> every decision is logged. Tabs
  renamed to Overrides / Baseline / Decision Log (ids stay stable for
  bookmarked URLs).
- The document table becomes one card per domain (Library visibility /
  Admin & permissions / Downloads & playback) with plain-language
  descriptions, example rules, status pills (Live vN / Draft / Disabled),
  inline creation, and the enable kill-switch in place.
- Selecting an override drills into a full-width editor with a visible
  lifecycle rail (Draft -> Validated -> Saved -> Live) and one contextual
  primary action per step; the unedited live source shows no actions
  until edited. Version comments appear only at the save step.
- Simulate is reframed as 'Test before going live' with a human verdict
  chip (Allowed / Denied — reason / ceiling summary) above the raw JSON;
  internal generation counters no longer surface.
- History uses 'Make live' with plain go-live copy; authors read
  'User N'; the baseline tab explains that upgrades never touch
  overrides.

Hand-written redesign (no Codex); verified via vitest, tsc, eslint,
prettier, and a production build.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): present the policy baseline as readable rules, not raw Rego

The Baseline tab dumped five Rego modules into read-only editors. It now
leads with what the rules actually do: one card per domain with
plain-language statements of the shipped behavior and a note on what an
override may change, plus content-rating and playback-quality tier
ladders parsed live from the lib module sources (so the tiers shown are
the ones the server enforces, not a hardcoded copy). The Rego source
stays one click away behind a per-module accordion and remains the
stated source of truth; unrecognized modules fall back to source-only.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(policy): add access-groups design addendum

Groups with permission toggles become the everyday admin surface; the
Rego editor is demoted behind policy.editor_enabled (default off).
Restriction-only composition: group grants are an upper bound, per-user
settings tighten further — same rule as the existing account/profile
merge, one layer up.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): add access groups — group defaults with restriction-only composition

New access_groups table + users.access_group_id (one group per user, NULL
= today's behavior). Group grants are an upper bound composed with the
user's own settings by strictest-wins rules — library intersection,
MinQuality, AND'd booleans, strictest positive stream/transcode limits,
permission-mask intersection, and a requests toggle gating CreateRequest.
The merge happens in Go (access.ApplyGroupPolicy /
EffectivePolicyForUser) before policy inputs are built, so vendor Rego,
the parity suites, and the decision log are untouched; every enforcement
surface (viewer scope in both resolvers, permission gates, downloads,
playback admission, requests) consumes the effective policy and fails
closed on provider errors. Changing a group's quality ceiling bumps its
members' access_policy_revision, mirroring the per-user rule.

Additive admin API: /admin/access-groups CRUD with member counts;
PUT /admin/users/{id} + user DTOs gain access_group_id.

Also demotes the Rego editor: policy.editor_enabled (default off,
hot-reloaded) drives the capability endpoint's editor_available and
403-gates editor endpoints while the engine and decision logging keep
running.

Design: docs/superpowers/specs/2026-07-02-access-groups-design.md.
Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed
(composition core + fail-closed call-site audit) and verified here.
DB-backed group-store tests pending local Postgres recovery.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add Access Groups admin page and gate the policy editor

New /admin/access-groups: a card grid summarizing each group (member
count + key restrictions), drilling into an editor that reuses the same
LibraryAccessSelector and quality presets as the user editor, with
toggles for downloads/transcoded-downloads/requests, concurrent-stream
and transcode limits, and a permissions mask (all-assignable by default,
narrowable to specific permissions). Delete warns how many members fall
back to the built-in defaults. Copy states the composition rule up front:
a group grants the most a member can do; their own restrictions still
apply on top.

The user editor gains a Group picker and read-only row; the Policy nav
entry is now hidden unless the capability reports the editor enabled.
Plumbing (types, hooks, user-editor picker, nav gating) drafted by Codex
(GPT-5.5); the Groups page hand-built. Verified: 25 tests across the
touched suites, tsc, eslint, prettier, and a production build.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): seed a Default Group and auto-assign newly created users

Adds access_groups.is_default with a partial unique index (one default
at most — the profiles is_primary pattern) and seeds a permissive
'Default Group' whose ceiling is a no-op, so assignment never changes
anyone's effective access until an admin edits it. The seed is guarded
against pre-existing defaults and name collisions; the Down migration
only removes the row if it is still untouched.

Assignment happens at the single INSERT INTO users choke point
(UserRepository.Create): when no explicit group is given, access_group_id
is filled by a scalar subquery on the default flag — NULL when no default
exists. Every creation path (setup, signup, invites, OAuth, admin create)
is covered by construction. Setting a new default via the API atomically
clears the previous one in the same transaction.

Deleting or unsetting the default is legal: new users then start with no
group, which is pre-feature behavior.

Implementation drafted by Codex (GPT-5.5); migration guards and the
choke-point subquery reviewed line-by-line here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): surface the default access group

Cards show a Default badge; the group editor gains a 'Default for new
users' toggle (with copy noting existing users are never moved); the
delete dialog warns when removing the default that new accounts will
start with no group.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): ship the Default Group with house-rule ceilings

Seed values per product decision: 5 concurrent streams, 5 transcodes,
transcoded downloads off, and a permission mask of marker_edit only
(metadata curation excluded). Plain downloads and requests stay on. The
Down guard matches the new values so it still only removes an untouched
seed row. Only newly created users are affected; existing users are
never assigned.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): retire per-user defaults — the Default Group is the sole default policy

Removes both legacy 'user defaults' mechanisms now that the seeded
Default Group owns new-user policy:

- users.max_streams / max_transcodes column defaults drop from 6/2 to 0
  (= unrestricted at the user layer), so group ceilings apply to new
  signups/invites/OAuth users instead of fighting stale per-user
  numbers. Existing rows keep their stored values — nobody is silently
  uncapped on upgrade.
- The dead defaults.max_playback_quality / defaults.max_profiles
  settings validation goes away with its only writer (the User Defaults
  dialog, removed on the web side).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): replace the User Defaults dialog with group-governed creation

The Users page's 'User Defaults' dialog (defaults.* server settings)
duplicated what access groups now do properly, and its values were only
ever form prefill — no backend path applied them. The button now links
to Access Groups, and the create-user form seeds unrestricted user-layer
values (0 streams/transcodes, any quality, downloads allowed) so the
member's group governs; per-user fields remain for tightening individual
users.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): migrate existing non-admin users into the Default Group

Existing users join the seeded Default Group on upgrade so one policy
source governs the whole instance. Their per-user limits still holding
the retired 6/2 column defaults are normalized to 0 in the same
statement so the group's ceilings actually apply; deliberately
customized values are preserved. Admin accounts stay ungrouped —
scope/action decisions are role-blind, so grouping an admin would cap
the server owner on upgrade.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(access): keep admins out of the Default Group and treat group moves as policy changes

New-user creation now mirrors the migration's admin exclusion: the
default access group is only auto-assigned to non-admin roles, so a
fresh server owner no longer inherits the starter group's transcode
denial and stream caps.

Changing a user's access group now bumps access_policy_revision (the
group carries permissions, quality, and limits, exactly like the
per-user fields that already bump it) and triggers admin session
revocation when the group actually changes.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): enforce marker_edit through the PDP on marker write routes

The Rego permission policy owned marker_edit but no Go caller ever
consulted it: PUT/DELETE /markers went through a handler-local check
that short-circuited admins and read only the user's own permissions,
so group permission masks and custom policy overrides were ignored.

Marker writes are now gated by router middleware like the other
permission surfaces: a PDP-backed RequireMarkerEdit that evaluates the
group-merged effective permissions (plus the legacy variant for
proxy/test wiring without a policy system). The handler-local check and
its user loader are gone.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): assert device/quality policy facts and honor the quality ceiling

The download_transcode action check hard-coded an empty device ID and
never asserted the requested quality, and no caller consumed
ActionDecision.QualityCeiling — custom download policies keyed on those
inputs were silently ineffective.

Resolve now threads the request's device ID and requested quality into
the action input, and a returned quality ceiling downscales the
prepared transcode target (the ceiling applies to what is served,
matching the serve-time rule in serveDownloadBytes). FileQuality and
the content-rating pair stay intentionally empty for downloads —
documented on downloadActionInput: those ceilings are enforced against
the served artifact by the scope-derived access filter, and asserting
the source's quality would wrongly deny capped transcodes.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(access): align the default-group seed assertions with the migration

The DB test still asserted the earlier no-op seed (transcode allowed,
unlimited streams/transcodes, null permissions); the shipped migration
seeds transcode denied, 5/5 limits, and marker_edit-only permissions,
so the test failed on any database with the migration applied.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): lock the Rego sandbox by builtin purity and bound compile work

Exclude every nondeterministic builtin from the admin sandbox instead of
denylisting names, so OPA upgrades cannot silently expose impure builtins
while pure helpers like net.cidr_contains stay usable. Apply the same
capabilities to the runtime engine, cap concurrent compile checks, and
reject oversized sources before they reach the uncancelable compiler.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): require literal booleans in vendor override and input checks

Bare object.get truthiness treated any non-false value as satisfied, so a
malformed override 'allowed' value could fail to tighten a base grant and
hand-crafted simulate input could flip flag predicates. Compare against
literal true so anything else denies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): surface decision log cleanup failures to the task manager

CleanupDecisionLogsOnce now returns the first error alongside the deleted
count so a broken partition manager or DB outage marks the scheduled task
failed instead of reporting 100% success while policy_decisions grows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): log admission decider errors before failing closed

A policy-evaluation failure was silently mapped to the too-many-streams
denial, making an engine outage indistinguishable from a real limit hit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(access): nil-guard the downloads user and restore the ABS legacy resolver

effectiveDownloadUser dereferenced policy state before its nil-user check,
and the ABS handler lost viewer-scoped filtering entirely when the policy
system was unavailable because no legacy access.NewResolver fallback was
wired like the other resolver paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): address admin policy review feedback

- invalidate the version query by version_number, the key usePolicyVersion
  actually caches under
- keep the goPrevious cursor-stack updater pure (Strict Mode double-invoke)
- make version history rows keyboard-selectable like the document list
- clamp download_transcode_allowed when downloads are disabled so groups
  cannot save a contradictory record

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(api): cap policy endpoint request bodies at 1 MiB

The policy write endpoints (create document/version, set enabled,
validate, simulate) decoded JSON bodies without a size limit, so an
oversized payload buffered fully in memory before CompileCheck's
256 KiB source cap could reject it. Route all five through a shared
decodePolicyRequest helper that wraps the body in http.MaxBytesReader
and returns 413 with the repo's standard too_large error shape.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UtnZ2Uewzo959hpneLrtRN

* fix(access): forbid deleting or demoting the default access group

Deleting the default group (or unsetting its is_default flag) left the
server with no default: new non-admin users were then created ungrouped
with max_streams/max_transcodes of 0 — unlimited — because the legacy
per-user column defaults were retired in favor of the group's ceilings.

The store now rejects both operations with ErrDefaultGroupRequired
(mapped to 409); promoting another group remains the supported way to
move the default, and atomically clears the previous one. The admin UI
disables the delete button and the default toggle on the default group
and explains the promote-another-group flow.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UtnZ2Uewzo959hpneLrtRN

* fix(web): keep unsaved policy drafts when a newer version activates elsewhere

The editor state was keyed on the active version's id/sha, so a
background refetch after another admin (or another tab) activated a
version remounted the editor and silently discarded the dirty draft.

PolicyEditorPanel now pins the seed it is editing against and only
adopts an incoming seed when nothing can be lost: the editor is clean,
the draft already equals the incoming source (the same-admin activate
flow), or the selection moved to a different document. Otherwise the
pinned editor stays mounted and an inline notice offers an explicit
"Load live version" action.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UtnZ2Uewzo959hpneLrtRN

* fix(policy): fail reloads on invalid custom sources and surface degraded/apply state

A stored custom source that stops compiling used to be silently skipped on
reload: the bundle widened to vendor-only for that domain while the generation
reported fully applied. Reload is now strict — a bad enabled source fails the
reload and the last known-good engine keeps serving. Boot keeps its vendor
fallback for availability, but skips are recorded on the engine and exposed
(with store-outage reasons) through System.DegradedState and additive
degraded fields on GET /policy/capability. Activate/SetEnabled re-run
CompileCheck instead of trusting the stored compiled_ok flag.

Mutation endpoints also no longer conflate persistence with live apply:
activation/enable responses carry additive applied/failed_step/
loaded_generation fields and return 202 when the store change persisted but
the local reload failed.

Addresses review findings C1, C2, and the degraded-signal gap (6.1).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): type deny reasons across the contract and enforce profile_verified

Deny handling used to branch on exact free-text reason strings in three Go
consumers, and playback reported ANY unrecognized reason — including custom
override free text and engine failures — as a stream-limit error. Decisions
now carry a stable reason_code (custom overrides always get custom_denial);
downloads, the metadata-curation gate, and playback admission switch on codes,
with a new ErrPlaybackNotAllowed -> 403 playback_not_allowed mapping for
non-limit denials. Rego tests pin every vendor code.

The scope contract's tighten-only profile_verified output was also emitted but
never consumed; a policy revocation now surfaces as ErrProfileUnverified (403
profile_unverified) instead of silently proceeding.

Addresses review findings 6.2 and C4.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(catalog): close the dual-library disabled-scope bypass in direct item authorization

EnsureAccessible, EnsureAccessibleIDs, and FilterAccessibleContentIDs gated
library access with allow/deny predicates over a single joined
media_item_libraries row, so an item linked to BOTH a passing library and a
disabled one satisfied the disabled check via the passing row — a direct-ID
bypass of disabled-library scope on the detail, media-file, playback, and
download paths. All library access predicates now share one helper
(libraryAccessConditions) emitting independent EXISTS / NOT EXISTS subqueries,
the semantics GetByIDsWithAccess already used, including the orphan-item
membership guard for disabled-only scopes. SQL-shape tests pin every builder
and a DB-gated regression test covers the dual-library item end to end.

Addresses review finding C3 (plus the same shape in
buildFilterAccessibleContentIDsSQL, which the review did not flag).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): serialize quota check and row creation under a per-user advisory lock

The concurrent-download quota was check-then-insert with nothing serializing
the pair: parallel creates could all observe free quota before any row
existed, bypassing the cap and stacking artifact encode jobs. All four
check->insert spans (ephemeral original, artifact-backed, series batch,
managed batch) now run inside Repository.WithUserQuotaLock — a
pg_advisory_xact_lock keyed by user, so the serialization holds across nodes.
The artifact path keeps the limiter-before-Ensure ordering (a rejected request
must not leave an encode job behind) by holding the lock across Ensure.
Managed-entry replacement stays quota-exempt and lock-free. A DB-gated
barrier test races 8 creates against a cap of 1.

Addresses review finding C5.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): assert served quality at create time for original and remux downloads

Direct-original and remux downloads serve the source resolution unchanged, but
create-time policy checks left file_quality empty — an over-ceiling source
registered a row serveDownloadBytes could never satisfy. Resolve now runs a
final download action check with FileQuality populated on those two paths
(capped transcodes keep the ceiling-on-artifact behavior), a custom override
ceiling below the served resolution denies, and quality_ceiling_exceeded maps
to ErrQualityUnavailable. The ActionInput contract now documents exactly when
file_quality and the rating facts are supplied so custom policy authors are
not misled.

Addresses review finding C6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): guard activation against slow overrides and make eval timeouts observable

A custom scope override that exceeds the 25ms eval budget compiled fine,
activated fine, and then converted to 500s on every authenticated request —
server-wide lockout authored in the admin editor. Activation and enable now
run GuardEvalCost: the candidate source is evaluated on a throwaway engine
against a canned representative input under the live budget, and a source
that cannot complete is rejected 422 with ErrPolicySlowEval before it goes
live. Runtime timeouts keep failing closed but now carry a distinct
ErrPolicyEvalTimeout sentinel, an Error log, and a per-engine counter exposed
as eval_timeouts on GET /policy/capability so intermittent near-budget
policies are attributable.

Addresses review finding C7.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: gofmt remediation files

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 17:38:19 -04:00
5f59f8e952 feat(clientip): expose trusted proxy CIDRs in the Admin UI and via SILO_TRUSTED_PROXIES (#310)
* feat(clientip): expose trusted proxy CIDRs in the admin UI and via env var

Trusted reverse-proxy CIDRs (clientip.trusted_proxies) previously required
hand-editing server_settings via SQL and a restart. Now:

- Admin UI: a Network > Trusted Proxies field on the General settings page,
  with server-side CIDR validation and normalization on save.
- Env var: SILO_TRUSTED_PROXIES is validated at startup and persisted to
  server_settings (re-applied on every boot while set), so Docker operators
  never touch the database and the UI shows the effective value.
- Hot reload: the setting now rides the nodeconfig watcher snapshot, so
  changes apply without restart on Redis-less deployments too (previously
  reload only worked via the Redis event bus, and only when rate limiting
  was enabled).

Closes #300

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(clientip): keep key-scoped event-bus reload alongside the config watcher

A malformed unrelated setting fails the whole-config watcher reload; the
direct subscription re-reads only clientip.trusted_proxies so the trust
boundary still updates on Redis-backed multi-instance deployments.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(clientip): key-scoped same-process reload in OnServerSettingUpdated

Covers the Redis-less path: an unrelated malformed setting that fails the
whole-config watcher reload can no longer leave stale trusted-proxy CIDRs
after a successful admin save.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(clientip): reload with a fresh context in OnServerSettingUpdated

The setting is already persisted when the hook runs; a canceled admin
request must not skip the trust-boundary reload.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style(web): wrap long trusted-proxies hint to the 100-char width

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add guidance tip for trusted proxy ranges

Explains that the setting replaces the private-network defaults, the
recommended /32 pattern, CDN multi-range caveats (Cloudflare), and why
0.0.0.0/0 is unsafe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 16:44:36 -04:00
430224a1b9 perf: cut home-screen, Continue Watching, and Latest latency; cache shared home rails (#292)
* perf(jellycompat,sections): bound resume scan, batch leaf detail progress, widen section concurrency

Three low-risk fixes from the section-fetch performance investigation
(docs/superpowers/plans/2026-07-03-section-fetch-performance.md):

- jellycompat: bound loadProgressPage at resumeScanMaxRows=300 so a single
  request never pages through more than that many in-progress rows. The cap is
  unconditional: it also covers the sparse-visible-set case (a heavy watcher
  whose recent rows are mostly dismissed/superseded, or a Series/Season-only
  request that matches no leaf in-progress row), where the page never fills and
  the loop would otherwise scan the entire history — previously an O(history)
  scan reaching tens of seconds. In the common case the loop exits far earlier,
  so the cap only bounds the pathological worst case; 300 leaves ample headroom
  to fill a ~20-item Continue Watching page. Beyond the cap the reported total
  is a clamped lower bound. Covered by TestLoadProgressPage_BoundsScanForSparseVisibleSet.
- jellycompat: batch the leaf-item (movie/episode) progress lookup in
  GetItemDetailsByIDs via ListProgressWithCompletedHistory instead of a
  per-item GetProgressWithCompletedHistory (~100 sequential queries for a
  50-item detail page). Series keep the per-item episode-rollup path (they own
  no progress row). Output is unchanged; a batch-lookup failure is now logged
  rather than silently dropping played state for the whole page.
- sections: raise fetchAllMaxConcurrency 4 -> 6 to cut FetchAll wave count for
  large home layouts, staying within the default 20-conn pool.

Part of the home/continue-watching latency work.

AI-use disclosure: implemented with AI (Claude) assistance.

* perf(jellycompat): keep Latest browse on the cross-library fast path under isPlayed

/Items/Latest with isPlayed=false is the highest-frequency compat browse
(~10.8k calls/day). The played overlay can't be pushed into SQL, so browse
over-fetches and filters locally. The cross-library recently_added fast path
(BrowseRecentlyAddedAcrossLibraries: one ~1ms index walk per library) was
gated on Offset==0, so a heavy watcher who had already seen the newest items
needed a 2nd chunk and fell through to BrowsePage — a whole-catalog
MIN(first_seen_at) + GROUP BY HashAggregate over ~147k movies measured at
~755ms per call (0.8-1.6s observed end-to-end).

Fetch the entire over-fetch budget (maxScannedRows) in a single merged
fast-path walk instead of paging into BrowsePage, so the loop fills from one
call. The clamp caveat (MaxLimit=1000 leaves a fall-through only for
requestedLimit>200, off the Latest hot path) is documented inline.

Part of the home/browse latency work.

AI-use disclosure: implemented with AI (Claude) assistance.

* fix(jellycompat): scope resume scan cap to resume path and bound the fast-path loop

Addresses PR #292 review feedback:

- Codex (P2): the resumeScanMaxRows cap was applied unconditionally in the
  general loop, which also paginates the completed (watched-items) list. Gate it
  on resumeFiltered so the completed path keeps exact TotalRecordCount and deep
  StartIndex pagination. Covered by TestLoadProgressPage_CompletedScanNotCapped.
- CodeRabbit (Critical): the earlier raw-offset fast-path loop — the default
  Continue Watching shape and the sections-fallback route — had the same
  unbounded-scan bug and was not covered by the cap (the existing test forces
  EnableTotalRecordCount=true, routing around it). Bound it with the same
  resumeScanMaxRows guard. Covered by
  TestLoadProgressPage_BoundsFastPathScanForSparseVisibleSet.
- CodeRabbit (Minor): tag the doc's fenced example blocks as text to satisfy
  markdownlint MD040.

AI-use disclosure: implemented with AI (Claude) assistance.

* perf(sections): cache shared user-agnostic home rails per access scope

Home-screen rails that are identical for everyone who can see the same
libraries (recently added, recently released, genre, trending on server,
most watched, new to library, critically acclaimed, award winners, format
showcase, seasonal, mood, trending discover, admin-curated lists, and
library collections) were rebuilt from Postgres once per request, per user.
Only the overlay on top of each row (watched flags, play position, presigned
poster URLs) is actually per-user.

Insert a process-global resolved-list cache at the FetchOne choke point in
internal/sections. Each cacheable row is built once per access scope, held
with a 15m TTL, and refreshed in the background 3m before expiry;
singleflight collapses cold-miss stampedes into a single build. The per-user
overlay still runs fresh in buildSectionsResponse, so no profile state is
ever shared. Random and per-user rows (continue watching, next up,
recommendations, hidden gems, forgotten favorites, activity feed, user
collections) bypass the cache.

The access-scope key captures every access boundary the fetch path enforces
-- section identity (type + id + config hash) + item limit + accessible and
disabled libraries + max content rating + excluded media types + name prefix
+ allowed-content-id allowlist -- and nothing per-user, so entries are safely
shared. Empty membership is never cached (avoids freezing a transiently empty
rail); background refreshes are bounded by a timeout.

Scale (analytical, derived from the cache behavior -- not a measured
latency): for the user-agnostic rows, Postgres section-query volume collapses
from O(rows x concurrent requests) to O(rows x distinct access scopes) per
15m refresh window, because most users share a handful of access scopes.
Illustrative -- 40 cacheable rows on a home screen, 1000 concurrent users
falling into ~5 distinct access scopes:

  - before: ~40 x 1000 = ~40,000 section queries per wave of home loads
  - after:  ~40 x 5    = ~200 builds per 15m window (plus one background
            refresh per row per scope), i.e. a warm home load runs zero
            section queries for these rows.

That is a ~99% reduction in shared section-query load at that concurrency;
the win grows with concurrency and shrinks as access-scope diversity rises.

Design/plan doc added under docs/superpowers/plans/.

* perf(jellycompat): serve per-library Latest via the cached recently-added section

A jellyfin-compat per-library /Items/Latest rail is the same user-agnostic
list as the native "recently added" library rail -- both order by
mil.first_seen_at DESC. It was rebuilt on every request through
directContentService.BrowseItems, missing the resolved-list cache entirely.

Route per-library Latest for movies and series libraries through the native
section fetch instead, so it reuses the shared cache. HandleLatest resolves
the library's type once, and for a movies/series library builds a synthetic
SectionRecentlyAdded with the same type + config + limit + access scope the
native rail uses and calls FetchOne; the per-user overlay (favorites,
progress, episode targets, presign) is extracted into buildLatestItemDTOs and
shared by both the native and BrowseItems paths, so no overlay logic is
duplicated. Cached *models.MediaItem values are read-only -- LocalizeItemModels
deep-copies before any presign mutation.

To let the two surfaces share one entry, resolvedListCacheKey no longer
includes the arbitrary section ID: every cacheable section type derives its
membership from type + config + limit + scope, never from its own ID (audited
all 14 cacheable types plus the library-collection path; the sole s.ID read
lives in the non-cacheable user-collection branch). A native recently-added
rail and the compat Latest for the same library + scope now collapse to ONE
cache entry, built once and reused. Access-scope isolation is unchanged --
the removed ID never carried access information, and every access boundary
(libraries, rating cap, excluded types, content allow-list, name prefix)
still keys the entry.

Guardrails: the native path is restricted to movies and series libraries;
every other library type (ebook, music, manga, mixed) is ignored and keeps
its exact BrowseItems behavior -- important because an unfiltered
recently-added fetch would otherwise surface non-video items to Jellyfin
clients that only expect video. Deeper pages, played-filter and
backdrop-required requests, a client asking for a type other than the
library's own, and any FetchOne error also fall back to BrowseItems.

Chosen over an alternative that gave the synthetic section a deterministic ID
(which kept two separate cache entries): both returned identical data with
similar complexity, so the shared-entry design won.

* fix(sections,jellycompat): post-review fixes for the shared-list cache and Latest path

Consolidates fixes from the branch's adversarial review and PR #292 review
comments into one commit:

- Latest fast path: fall back to BrowseItems when a request carries a genre,
  name-prefix, or person filter (the synthetic recently-added section cannot
  express these, so serving it unfiltered would return a wrong, broader set).
  Eligibility is decided by latestFastPathEligible and covered by a test.
- Clamp the /Items/Latest page size to compatBrowseMaxLimit before building the
  section, matching the BrowseItems fallback, so a large client Limit can't drive
  an oversized recently-added fetch or explode the shared cache key with unbounded
  ItemLimit values.
- Evict expired entries from the process-global resolvedListCache: resolvedListSet
  sweeps expired keys at most once per minute, bounding the map to scopes seen
  within one TTL window. Covered by TestResolvedListCacheEvictsExpiredEntries.
- Log a short digest of the cache key (resolvedListLogKey) instead of the raw key
  in the background-refresh panic/error paths, since the key embeds
  user-controlled access-scope fields such as NamePrefix.

Skipped review comments (verified already fixed or stale against current code):
the resume fast-path scan bound and watched-items cap (04d2e795) and the docs
fence-language tags (already addressed).

Build, vet, and go test -race pass for internal/sections and internal/jellycompat.

* perf(plugins): cache plugin installations in-memory, invalidated on lifecycle change

## Problem
Every poster/image on a warm home rail re-read plugin_installations from
Postgres to answer "is this plugin enabled?" and to acquire the plugin client
(Source A: metadata chain buildProviders enabled-check; Source B: ensureClient
-> loadInstallation). Plugin-resolved image URLs are never URL-cached, so the
plugin source and the DB read behind it fired again on every identical warm
request; 100% of images in the target library are plugin-backed.

## Solution
- Guarded in-memory installation cache (map[int]*Installation + RWMutex) in
  plugins.Service. loadInstallation reads through it; the requireEnabled gate
  stays after the cache read so ErrInstallationDisabled semantics are unchanged.
  invalidateInstallationCache clears it and is self-registered as a lifecycle
  hook, so Service.OnLifecycleChange wipes it on install/enable/disable/update/
  uninstall.
- A generation counter closes an invalidate-vs-repopulate race: captured before
  installations.GetByID and re-checked under the write lock, so a row fetched
  before a lifecycle mutation is never written into a freshly cleared cache
  (would otherwise resurrect a just-disabled plugin).
- Route the metadata chain enabled-check through the same cache via a structural
  InstallationEnabledChecker interface (nil-safe: falls back to the pool query
  when no checker is injected), wired in cmd/silo/main.go.

## Post-review fix (auto-update reliability blocker)
AutoUpdateService mutated installations (new InstallPath, old dir deleted) on
the default auto update policy without firing OnLifecycleChange, leaving the
cache stale and breaking plugins with "stored plugin manifest mismatch" until
restart. It now takes an onChange callback wired to Service.OnLifecycleChange
and fires it once per Check run that mutated a row.

## Verification
go build/vet, go test ./internal/plugins/... ./internal/metadata/... (-race).
Tests: cache hit/invalidation, racing-invalidation guard, IsInstallationEnabled,
auto-update fires onChange.

## AI-use disclosure
Implemented with AI assistance (Claude).

* perf(jellycompat): batch per-item presign, and enrich series on the cached Latest path

## Problem
List rails presigned each item's poster/backdrop/logo/still image individually
(~160 singular resolver calls for a 40-item page where 4 batched calls suffice),
and ItemsHandler carried a near-verbatim duplicate of the batch presigner.

## Solution (batching)
Promote the batch presigner to a shared package-level presignCompatListItems
(presign_list.go) with a generic collectImagePaths[T]; convert the per-item
loops (cached home/Latest rail, favorites, batch loaders, userdata favorites) to
one batched PresignImageURLsWithExpiry per image type per page; batch the
season/episode collections; delete the three duplicate presign helpers. URL
output is unchanged (verified byte-for-byte).

## Post-review fix (series Latest data-parity regression)
The native cached Latest fast path built items via compatListItemsFromModels +
buildLatestItemDTOs and never ran the series watch-state rollup, so a series
library's Latest lost Played / UnplayedItemCount and page 1 disagreed with the
BrowseItems fallback. enrichSeriesUserData is promoted to the ContentService
interface and called on the native path (reused, not duplicated).

## Verification
go build/vet, go test ./internal/jellycompat/... ./internal/catalog/...
Tests: bounded presign invocation counts + per-item URL mapping; series rollup
populated on the native Latest path.

## AI-use disclosure
Implemented with AI assistance (Claude).

* perf(sections): gate personalized rails out of the shared cache; widen refresh lead

## Problem
1. The shared home-rail cache whitelisted custom_filter/genre sections by TYPE
   alone, but those route through fetchFiltered -> ParseQueryDefinition and can
   carry personalized (per-profile) rules/sorts (watched, favorited,
   in_watchlist, in_progress, last_watched; sorts progress/date_viewed/plays).
   Their membership is per-profile yet the cache key excludes userID/profileID,
   so a personalized rail built for one profile was served to others in the same
   access scope for up to 15m -- a cross-profile watchlist/watch-state leak.
2. The background-refresh lead was tuned so steady traffic is served a warm
   entry from a longer soft window.

## Solution
- Add QueryDefinition.IsPersonalized() (reusing the existing
  QueryFieldRequiresProfile/QuerySortRequiresProfile helpers).
  isCacheableSectionType now parses the section QueryDefinition and refuses to
  cache custom_filter/genre when personalized; non-personalized definitions stay
  cacheable. Seasonal/mood/trending build their definitions server-side and stay
  unconditionally cacheable.
- resolvedListRefreshLead 3m -> 10m (soft threshold builtAt+5min instead of
  builtAt+12min).

## Verification
go build/vet, go test ./internal/sections/... ./internal/catalog/... (-race).
Test: personalized custom_filter/genre not cacheable; non-personalized are.

## AI-use disclosure
Implemented with AI assistance (Claude).

* fix(sections,metadata): post-review fixes for shared cache and plugin chain staleness

Addresses three review findings on PR #292:

- sections: canonicalize section config JSON before hashing so configs
  differing only in whitespace/field order share a cache entry (native +
  jellycompat rail sharing). Added TestHashSectionConfigCanonicalizes.
- metadata: invalidate the resolved-chain cache on plugin lifecycle
  changes; the installation-enabled check already reads the invalidated
  plugin cache, but resolveChainCached could serve a stale provider chain
  for up to chainCacheTTL after a provider's availability changed.
- jellycompat: move ctx to the first parameter of presignCompatListItems
  for consistency with the other presign helpers.

Skipped the episode-image presign batching nitpick: the resolver already
dedupes+singleflights, so it is a Minor perf-only item not worth the
two-pass refactor risk in this pass.

* fix(sections,jellycompat): harden shared rail cache and Latest fast path per review

Addresses the eight findings from the deep review of this PR:

- Detach the blocking cold-miss rebuild from the singleflight leader's
  request context (context.WithoutCancel + the shared 30s build timeout)
  so one client disconnect no longer fails every collapsed waiter and
  leaves the entry uncached.
- Stop client-controlled values minting unbounded cache entries: the
  compat Latest fast path now always fetches a fixed 100-row budget and
  slices to the requested limit (one entry per scope+library instead of
  one per Limit value), and an unrecognized MaxOfficialRating string
  disqualifies the fast path instead of entering the global cache key.
- Add release_date to the sections item projection/scan so movies served
  via the Latest fast path keep PremiereDate (Jellyfin default-set field)
  in parity with the BrowseItems fallback.
- Fall back to per-item progress lookups when the batched leaf progress
  query fails, restoring one-item-at-a-time degradation instead of
  blanking played state for the whole page.
- Derive cache eligibility from a single source of truth: fetchSection
  and isCacheableSectionType now share the userAgnosticSectionFetcher
  table, whose no-userID/profileID signature makes a fetcher drop out of
  the cacheable set at compile time if it ever gains per-profile inputs.
- Decide Latest fast-path eligibility off the actual browse params the
  fallback would receive, so any filter later added to buildBrowseParams
  automatically disqualifies the cached path; share one
  compatDefaultBrowseLimit constant between both paths.
- Extract AccessFilter.WriteAccessScopeCacheKey as the shared, security-
  critical serializer for all access-scoped caches (resolved-list,
  editorial candidates, audiobook groups); the editorial key now captures
  ExcludedMediaTypes, which its loaders already applied in SQL.
- Strip leaked agent-transcript markup from the section-fetch plan doc.

go build ./..., go vet, gofmt clean; go test -race on
internal/sections, internal/catalog, internal/jellycompat passes
(TestBeginWebOperation* failures are the known pre-existing flakes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 02:01:30 -04:00
ce09d2b6fc fix(catalog): harden Meilisearch search integration (#291)
Findings from a full review of the Meilisearch implementation:

- Give the indexer its own 2m HTTP timeout instead of reusing the 800ms
  search-path timeout_ms, so large document uploads to a non-loopback
  Meilisearch stop timing out.
- Delete superseded indexes after a rebuild (previous active + leftover
  <prefix>_rebuild_* partials); every rebuild previously leaked a full
  copy of the catalog on the Meilisearch instance.
- Cache the index state row + pending count for 3s on the search hot
  path (was two Postgres round trips per search request); a failed
  search invalidates the cache immediately.
- Swap the active-index pointer before marking outbox events processed
  so a crash between the two replays events instead of losing them.
- End pagination only on a short page; estimatedTotalHits is an
  estimate and could truncate results.
- Latch the startup-resolved provider process-wide so package-level
  enqueue helpers stop querying server_settings in write transactions.
- Surface dead-lettered outbox events in the admin status + web UI.
- Remove unwired provider config knobs, dedupe the manga-chapter
  exclusion predicate, split the vector cache onto its own mutex, real
  rebuild progress percentages, and expand client/coalesce test
  coverage.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 00:18:14 -04:00
2155261706 fix(recommendations): make embedding backfill job timeout configurable (#229)
The embedding backfill (TriggerEmbeddings + the scheduled runEmbeddings)
ran under a hardcoded 30-minute context. That is fine for a fast hosted
embedding API, but local/self-hosted embedders (e.g. Ollama on CPU) are
far slower — on a large catalog they embed only a few thousand items
before the context deadline aborts the run with "context deadline
exceeded". The job is idempotent and resumable, so progress is not lost,
but it never finishes without repeatedly re-triggering it.

Make the per-run timeout configurable via a new
`recommendations.embeddings_job_timeout` setting (default 24h), threaded
through RecommendationsConfig -> NewWorker and applied to both the manual
trigger and the cron-scheduled run. A non-positive value falls back to
24h. Default behavior is unchanged for hosted users (a full backfill
comfortably fits in 24h); local LLM users can now complete a one-shot
backfill instead of stalling.

AI-use disclosure: implemented with assistance from Claude Code.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-05 00:17:22 -04:00
604bbf1a0f feat(playback): unified restart-resilient playback (native + jellycompat) (#174)
* feat(playback): unified restart-resilient playback via shared TranscodeManager

Make direct, remux, and native HLS transcode sessions survive a server
restart through one shared flow instead of per-method paths. A missing
in-memory session becomes a reconstruct trigger, not a 404: the server
rebuilds the session from a tiny durable recipe card plus the position the
client re-supplies on its next request.

- internal/playback/transcode_manager.go: shared TranscodeManager owning the
  transcodes map, recipe-card lifecycle, reconstruct single-flight +
  concurrency cap, LoadOrReconstructSession front door, ReconstructSession /
  ReconstructTranscode, and orphan cleanup. ~90% is logic moved out of the
  native handler (no behavior change), not new surface.
- internal/playback/recipecard.go + recipecard_postgres.go: RecipeCard with a
  PlayMethod discriminator (direct/remux/transcode; empty decodes as transcode
  for back-compat) behind a swappable, nil-safe RecipeStore interface backed by
  transcode_recipes.
- internal/playback/session.go: RegisterReconstructed inserts a rebuilt Session
  under its existing id (no UUID mint, no limit double-count, race-yielding).
- internal/playback/transcode.go: CloseProcess keeps the output dir so a
  reconstruct winner keeps serving; Close removes it.
- internal/api/handlers: drain the transcode lifecycle into the manager; wire
  reconstruct into the stream/segment serve paths; re-bind ownership to the live
  caller (refuse userID==0/mismatch); card-aware orphan cleanup.
- migrations: add transcode_recipes (expires_at TTL, filter-on-read, indexed).

Ownership stays two-factor: an authenticated caller AND a session.UserID that
matches; the card stores no secrets and identity is re-resolved per request.

Tests: recipe-card round-trip/legacy-decode/disabled-noop, RegisterReconstructed
insert/race/concurrency, close-vs-close-process dir semantics, the
LoadOrReconstructSession status matrix, and the reconstruct concurrency cap.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* feat(jellycompat): reconstruct transcodes across restart via shared manager

Bring Jellyfin (jellycompat) HLS playback onto the same restart-resilient flow
as the native path. Previously jellycompat owned a separate PlaybackHandler with
a private transcodes map and a duplicated transcode lifecycle that never grew the
reconstruct half, so an in-flight Jellyfin transcode died on restart and the next
segment request 404'd.

- Embed the shared playback.TranscodeManager and delete the duplicate lifecycle,
  so jellycompat gets reconstruct, the concurrency cap, the node-affinity rule,
  and the card lifecycle for free.
- internal/jellycompat/playback_sessions_postgres.go: DurableCompatPlaybackStore,
  a write-through cache over jellycompat_playback_sessions behind the new
  CompatPlaybackStore interface (nil pool degrades to cache-only). This persists
  the load-bearing PlaySessionId -> UpstreamSessionID mapping (plus media sources,
  route item id, seek) so it survives a restart instead of vanishing with the map.
- Write a recipe card on compat transcode start keyed by the upstream session id,
  using the native StreamAppUserID so the ownership re-bind matches; reconstruct
  the upstream session and the transcode seeked to the requested seg_NNNNN.
- migrations: add jellycompat_playback_sessions (expires_at TTL + compat_token
  index, full PlaybackSession in data JSONB).

Auth is mapped to the native user id before reconstruct so the same two-factor
ownership check and userID==0/mismatch refusal apply unchanged.

Tests: DB-gated (SILO_TEST_DATABASE_URL) durable-store round-trip proving a
session written by one instance reloads in a fresh one (the restart case), plus
a nil-pool cache-only path; existing handler tests updated to the manager.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* docs(playback): consolidate unified playback reconstruction design

Replace the three overlapping playback docs (the native Postgres
restart-resilience spec, the jellycompat plan, and the unification spec) with a
single self-contained design at
docs/superpowers/specs/unified-playback-reconstruct.md.

The doc leads with the unified design — the one-idea reconstruct model, a strong
visual flow of a restart mid-playback, the shared TranscodeManager + recipe card,
the two swappable durable stores, security, the concurrency cap and node-affinity
constraint, preconditions, and verification. The design history and rationale
(reconstruct-not-rehydrate, phased delivery, Redis-vs-Postgres, token-as-
descriptor, failure analysis) move to an appendix. It references no other md file.

AI-use: written with AI assistance.

* fix(playback): address review on restart-resilient playback

Four fixes from PR review of the unified reconstruction work:

- Rewrite the recipe card on audio-track change. HandleChangeAudioTrack only
  updated the in-memory session/transcode, so after a restart reconstruct
  resumed with the stale AudioTrackIndex/TranscodeAudio (and stale play method)
  from the start-time card. Re-save the card (direct/remux/transcode) with the
  switched state, mirroring the start-card pattern.
- Guard nil TranscodeManager in LoadOrReconstructSession and ReconstructSession.
  StreamHandler.TM is documented optional (tests/minimal setups); a missing
  session previously panicked in recipeEnabled instead of returning
  SessionMissing. ReconstructTranscode already guarded nil; make the two
  siblings consistent.
- Reject direct/remux cards in doReconstructTranscode before spawning ffmpeg, so
  a non-transcode card id can never enter the HLS reconstruction path.
- Log a non-success status from the remote transcode-node DELETE in
  CloseTranscodeSession; a 401/404/500 was previously silent.

AI-use: implemented with AI assistance.

* fix(playback): harden restart-resilient compat sessions

* feat(playback): token-carried reconstruction across restarts

Build on the shared TranscodeManager (introduced earlier in this branch) so a
playback session survives an API-server or transcode-node restart without the
client re-negotiating, and retire the Postgres transcode_recipes store in favor
of a recipe carried inside the signed stream token.

- RecipeCard encodes the byte-affecting encode parameters and rides inside the
  stream token; LoadOrReconstructSession rebuilds the in-memory Session (and,
  for integrated transcodes, the ffmpeg process) on a cold miss, single-flighted
  per session and paced by a spawn semaphore. Removes recipecard_postgres.go and
  the 20260617233705_add_transcode_recipes migration.
- transcodenode reconstructs a lost ffmpeg node-side from the forwarded token.
- TR-lease: proxy/streamauth enforce a revocation deny-marker on every served
  segment, with a 500ms Redis timeout, a bounded per-session "allowed" cache
  (3s TTL, expiry-first graceful eviction), and a degraded-fail-open counter.

Review hardening folded in:
- Manifest/segment handlers do the in-memory session lookup first and only
  verify the stream token on a reconstruct miss (token HMAC was per-segment).
- Copy-mode reconstruct never applies the encoded-only seg*dur seek, at spawn
  time or via the recovery path: RestartSeekTarget reports "unresolved" for a
  copy session whose manifest cannot yet map the segment, so the client retries
  instead of seeking to a fabricated source time.
- Crash teardown is a compare-and-delete (CloseTranscodeSessionIf returns
  whether it matched); the crash closure tears down the playback session only
  when it matched, so a session reconstructed under the same id is not killed.
- Reconstruct enforces the same per-user stream/transcode caps as a fresh start
  (RegisterReconstructedWithLimits), closing a token-replay slot bypass.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* feat(jellycompat): node-side transcode reconstruct via shared recipe store

Make Jellyfin-compat playback sessions survive a server or transcode-node
restart by reusing the shared TranscodeManager reconstruct path and a durable
recipe store, on top of the durable compat session store added earlier in this
branch.

- Node-side transcode reconstruct goes through the shared recipe store; the
  recipe is persisted to the control-plane store (Redis) when a dedicated
  transcode node is used so the node can rebuild ffmpeg after its own restart.
- Adopt the shared manager's API (3-arg OnFFmpegCrash carrying the dead session,
  guarded CloseTranscodeSessionIf, RegisterReconstructedWithLimits).

Review hardening folded in:
- Recipe lifecycle: noderecipe.Store gains Delete, called on deliberate
  teardown (stop, method-switch discard, node stop/force-reload) so a stopped
  session cannot be resurrected by a buffered request after a node restart;
  crash paths intentionally keep the recipe so a resume can reconstruct.
- Crash closure tears down the upstream session only when the guarded transcode
  close matched, so a reconstructed successor is never left orphaned.
- Copy-mode segment recovery surfaces a retryable not-found instead of a
  wrong-position restart, matching the native and node paths.
- Durable Update is now a SELECT ... FOR UPDATE transaction, removing the
  lost-update clobber that could silently drop a transcode recipe.
- Empty-token route resolution no longer falls back to an unbounded full-table
  scan; DB expiry filters bind the injected clock; the redundant re-Get is gone.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* docs(playback): consolidate restart-resilient playback design

Replace the superpowers spec with a single architecture record describing the
token-carried recipe card, the shared TranscodeManager reconstruct path for
direct/remux/transcode, the jellycompat durable session + node recipe store, and
the revocation-lease model with its fail-open tradeoff.

AI-use disclosure: written with AI assistance (Claude Code).

* docs(playback): correct jellycompat node-recipe rationale in comments

The noderecipe / transcode-node / jellycompat comments justified the Redis
recipe store with "a Jellyfin client cannot round-trip a token". The real
reason: the node-hop token is server-minted and could carry the recipe, but the
recipe is mutated in place under a stable session id (a /Sessions/Playing/Progress
audio switch restarts ffmpeg without re-minting the client's token) and a
third-party Jellyfin client cannot be driven to refresh a stale token, so the
node must reconstruct from a server-authoritative, node-reachable store.

Aligns the comments with docs/architecture/restart-resilient-playback.md §10.
Comment-only; no behavior change.

* refactor(playback): remove deny-lease revocation, defer to future PR

The deny-lease stream-revocation mechanism (the internal/streamauth
package, its silo:streamauth:<sid> Redis markers, the proxy Allowed()
enforcement, and the admin Stop/Terminate deny write) only ever enforced
on the offload-proxy topology and was a silent no-op on the integrated
single box and the dedicated transcode node. Rather than ship a partial
revocation feature that looks complete but isn't, remove it wholesale and
defer a uniform cross-topology revocation design to a dedicated follow-up.

Removed: internal/streamauth (package + tests); the LeaseDenier field,
StreamLeaseDenier interface, and denyStreamLease helper in playback.go;
the admin deny write; the router/main wiring; and the proxy verifyToken
Allowed() gate. The unified-reconstruct core (recipe-token,
LoadOrReconstructSession) is orthogonal and untouched.

Known limitation (now on every topology): admin Terminate and user Stop
tear down the live in-memory session and ffmpeg producer, but a still-valid
stream token can reconstruct the session until its 24h TTL expires. No
node-side byte-withholding ships in this PR.

docs/architecture/restart-resilient-playback.md is updated to mark the
revocation/deny-lease sections as deferred and to drop the overstated
"instant revocation on admin kill" claim.

* fix(playback): allow zero-caller bearer on transcode reconstruct

The authless HLS transcode delivery routes (master.m3u8 / segment) treat
the session UUID as the bearer credential, so a real request carries
requestUserID == 0. The live serve path already allows this, but
ReconstructSession hard-rejected a zero caller, so a request that worked
before a restart became SessionMissing -> 404 after the in-memory session
was gone, breaking the restart resilience these routes advertise.

Match the live-path contract in LoadOrReconstructSession: allow a zero
caller (UUID-as-bearer) and refuse only a non-zero caller that mismatches
the card owner. The reconstructed session is bound to card.UserID either
way. Adds TestReconstructSession_Ownership covering both cases.

* fix(jellycompat): re-persist recipe on local audio switch

A Jellyfin client switching audio on an integrated/local compat transcode
restarted live ffmpeg with the new track but did not re-persist
PlaybackSession.Recipe. The remote branch already re-persists via
startRemoteTranscode -> persistTranscodeRecipe. After a central restart,
reconstruct rebuilt ffmpeg from the stale Recipe.AudioTrackIndex, so the
integrated session resumed on the original audio track.

Persist the updated recipe (best-effort) after a successful Restart in the
local branch, mirroring the remote branch, so the durable
Recipe.AudioTrackIndex tracks live ffmpeg. Adds a regression test.

* fix(playback): strip stream token from proxied transcode-node URL

proxyToTranscodeNode appended the client's raw query string to the internal
transcode-node URL and logged that URL on transport failure. When a remote
transcode runs without a separate proxy node, that query carries
?st=<signed JWT> — a 24h bearer reconstruction descriptor exposing the
media path and recipe claims — placing the token into internal requests and
error logs.

Strip the "st" param before building targetURL, preserving any other query
params. The token is neither forwarded to the node nor present in the
logged URL. Header-forwarding of the token (so the node can reconstruct) is
a separate follow-up (#6).

* fix(playback): fail open on transient limit-provider error in reconstruct

During the reconstruct wave right after a restart (Postgres under peak
load), a transient limit-provider DB error was collapsed into a hard 404,
permanently stopping playback for a user within their limits. limitsForUser
wrapped any provider error, RegisterReconstructedWithLimits propagated it,
and ReconstructSession mapped every error to SessionMissing -> 404 -
indistinguishable from a genuine over-cap rejection.

Distinguish the two: tag provider errors with a new ErrLimitProviderUnavailable
sentinel and, during reconstruct, fail OPEN on a provider error (admit via
RegisterReconstructed + log a degraded warning) rather than refuse - mirroring
the reliability-first fail-open-on-dependency-error philosophy. A genuine
ErrTooManyStreams / ErrTooManyTranscodes over-cap still refuses. Adds tests
for both the fail-open and still-refused paths.

* fix(playback): forward stream token to transcode node as header

The dedicated transcode node's reconstruct path reads the stream token only
from the X-Silo-Stream-Token header, but proxyToTranscodeNode forwarded only
the node-API bearer token (and #5 now strips st from the URL). So when the
central API proxied to the node and the node self-restarted, it could not
reconstruct from the recipe-complete native token -> 404.

Capture st before stripping it from the URL, verify it at the API boundary
(streamtoken.Verify + SessionID match, mirroring the node's own check), and
forward it as X-Silo-Stream-Token. Best-effort: a missing/invalid token never
blocks the live proxy, and the token is still kept out of the forwarded URL
and logs.

* fix(playback): restart node ffmpeg on native remote audio switch

A native audio-track switch on an offloaded/remote transcode was a no-op at
the node yet returned 200 with a fresh URL: HandleChangeAudioTrack restarted
ffmpeg only when the API owned a LOCAL TranscodeSession, so for an offloaded
transcode the node kept serving the OLD audio (the node consults the token
only on a session miss). The replacement URL was also minted from identity-
only claims, so a later node restart 404'd.

For the offloaded transcode case (detected via session.TranscodeNodeURL),
POST a fresh /transcode/start to the node with the new AudioTrackIndex
(handleStart tears down and restarts ffmpeg) and mint the replacement proxy
URL from a full RecipeCard so reconstruct survives a node restart. The encode
recipe is derived from the durable session target fields plus the file,
mirroring HandleStartTranscode. A concrete SegmentDuration
(playback.DefaultSegmentDuration) is embedded rather than 0: the node's token
completeness gate treats SegmentDuration<=0 as incomplete and falls back to a
recipe store the native path never populates, which would 404 on a node
restart - the exact resilience this path provides. A failed node POST now
surfaces 502 rather than a false 200. Remux and non-offloaded (local)
transcode paths keep their prior identity-claim URLs unchanged.

Known limitation: Session does not persist the original SegmentDuration or
SubtitleTrackIndex/SubtitleBurnIn, so a remote audio switch resets subtitle
selection to none and assumes the default segment length; a client that
started with a non-default segment length will resegment on switch. Making
that state durable on the session is a follow-up.

* docs(playback): scrub stale deny-lease/revalidator comments

The deny-lease revocation mechanism and its "central revalidator" were removed
earlier in this branch, but four comments still described them as live
(transcode_manager.go, noderecipe/store.go, streamtoken/token.go,
proxy/server.go). Reword them to match the shipped behavior: ownership claims
are re-resolved at reconstruct, the noderecipe store shares Redis only with the
node-session tracker, and a sub-TTL hard cut depends on a node-side revocation
mechanism that is deferred to a future PR.

* fix(jellycompat): surface durable playback-session write failures

DurableCompatPlaybackStore.Update applied the in-memory mutation and then
swallowed every Postgres commit-failure path, returning nil. Callers that
promise restart resilience (persistTranscodeRecipe's recipe write, the
upstream-session binds in streams.go) were told the session was durably
persisted when only the cache held it, so a transient DB hiccup could leave
the next restart reloading a stale row (wrong audio track) or 404ing.

updateDB now returns the genuine DB round-trip error (begin/query/unmarshal/
marshal/exec/commit); Update propagates it while still applying the in-memory
mutation so live state stays correct. A nil pool and a genuinely absent/expired
row remain best-effort (return nil) — only real infrastructure failures
propagate, so existing rollback paths fire exactly when durability is lost.

Part of #174

* fix(playback): re-inject stream token into proxied transcode manifests

API-proxied remote transcode manifests dropped the reconstruct token from
their segment URLs, so playback died after a node or API restart. When a
remote transcode has no separate proxy node, the client loads its manifest via
the API-local path; proxyToTranscodeNode strips the signed token ("st") from
the forwarded URL (keeping it off node URLs and logs, forwarded only as the
X-Silo-Stream-Token header), and the node builds relative segment URIs from
that token-less query. The segment URLs the client received carried no token,
and the proxy only re-attached the header when an incoming segment request
already had "st" — which it never did — so a restart made those segments
non-reconstructable and they 404'd.

proxyToTranscodeNode now rewrites the manifest body at the boundary: every
segment and #EXT-X-MAP init URI gets the client-facing, API-verified token
re-appended (new playback.AppendManifestQueryParam helper), so the client's
later segment fetches carry "st" again and reconstruct after a restart. The
token still never reaches the node URL or its logs. Only 200 .m3u8 responses
are rewritten (Content-Length corrected); segments stream through untouched.

Part of #174

* fix(playback): preserve subtitle/cadence recipe across offloaded audio switch

Switching audio on a remote (offloaded) transcode with burned-in subtitles
silently dropped them, and reset a non-default segment cadence. The offloaded
audio-switch restart rebuilt the node start request from Session state, but
Session/SessionStreamState retained no subtitle or segment-duration state
(only the live local ts.Opts() and the RecipeCard did), so the branch
hard-coded SubtitleTrackIndex:-1, SubtitleBurnIn:false and
SegmentDuration:Default — signing that altered recipe into the replacement
stream token. An audio switch then changed bytes beyond audio selection, and
any later reconstruct kept the wrong no-subtitle/wrong-cadence recipe.

Persist the byte-affecting recipe on the session: SubtitleTrackIndex,
SubtitleBurnIn and SegmentDuration are added to Session/SessionStreamState,
populated at start (finalizeTranscodeStart) and on post-restart reconstruct
(ReconstructSession from the card), carried forward on every audio-switch
state update, and read back when rebuilding the offloaded node request and its
recipe card. The restart now reproduces the exact live stream. Also resolves
the M-4b non-default segment_duration reset.

Part of #174

* fix(playback): serialize transcode spawn paths with a per-session lock

Reconstruct was single-flighted only against other reconstructs, so a
restart-driven segment reconstruct racing a quality/seek/audio fresh start
could spawn two ffmpeg processes writing the same output directory at once —
segment corruption, partial-write closes, orphaned processes, and skewed
active-job accounting. The atomic register-after-spawn (GetOrRegister / the
reconstruct compare-on-register) prevented a map leak but not the concurrent
disk writers, because the losing path had already spawned. The dedicated
transcode node had the same split between handleStart and spawnReconstruct.

Add a refcounted per-session lifecycle lock to both TranscodeManager and the
node Server, held across "check existing -> spawn -> register":
- reconstruct (doReconstructTranscode / spawnReconstruct) re-checks under the
  lock and yields to any live session instead of spawning a duplicate;
- the native and jellycompat fresh-start paths take the lock around their
  spawn+register (the native path also closes any session a reconstruct rebuilt
  in the meantime so its fresh ffmpeg is the sole writer);
- the node handleStart holds it across teardown+spawn+register.
The refcount drops the map entry once no path holds/waits, keeping it bounded.
GetOrRegisterTranscodeSession is removed — the lock supersedes it and keeping a
register-after-spawn primitive would invite reintroducing the race.

Part of #174

* fix(playback): serialize restart re-spawn under the session lifecycle lock

TranscodeSession.Restart() releases s.mu across cancel -> wait-for-done ->
re-exec and spawns ffmpeg into opts.OutputDir without holding the per-session
lifecycle lock. LockSessionLifecycle's contract (fresh start, restart,
reconstruct) requires restart to hold it too, but all five callers invoked
Restart unlocked: native audio-switch and segment-recovery, compat
audio-switch and segment-recovery, and the transcode-node segment-recovery.

A restart racing another restart (audio-switch vs segment-recovery) or a
fresh-start/reconstruct could land two ffmpeg processes writing the same
segment directory -- mixed timelines, init.mp4/segment mismatch, and an
orphaned-but-still-writing ffmpeg -- the exact concurrent-writer corruption
the lifecycle lock exists to prevent.

Add RestartSessionLocked (TranscodeManager) and restartSessionLocked (node
Server) that hold LockSessionLifecycle only across the cancel->respawn
transition, re-check that the handle is still the live mapped session under
the lock, and return ErrSessionSuperseded rather than re-spawning a stale
handle. Route all five call sites through them. The lock is released before
callers wait on segments so recovery latency is unchanged.

Tests: gating (restart blocks until the lifecycle lock frees, then spawns),
concurrent-restart serialization, and superseded re-check on both the manager
(covers native + compat) and node lock owners.

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-02 14:23:14 -04:00
d1fd7c236f fix(jellycompat): sync activity dashboard immediately on compat playback start/stop (#261)
* fix(jellycompat): sync activity dashboard immediately on compat playback start/stop

Compat (Infuse/Jellyfin) sessions were only reconciled into
playback_sessions_sync on the periodic 15s reconciler tick, so stopped
streams lingered in the admin activity dashboard and overlapped with
newly started ones as ghost sessions. Native playback handlers already
trigger an immediate SessionSyncer.SyncNow on start/stop; wire the same
syncer into the jellycompat playback handler and flush after
teardownPlaySession (Stopped report + ActiveEncodings teardown) and
after a new upstream session starts.

The stop-path sync detaches from request cancellation
(context.WithoutCancel) because clients often drop the connection right
after reporting a stop.

Also adds the ListProgressSince method to the jellycompat test fake so
the package's tests compile again after the downloads-v2 interface
change.

Part of #205

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(worker): serialize SyncNow and bound compat request-path session sync

Adversarial review follow-ups: SyncNow snapshots could commit out of
order when request-path syncs raced the periodic tick (an older
snapshot committing last would resurrect stopped sessions or drop
fresh ones), and the detached stop-path sync had no deadline, letting
a stalled DB pin request goroutines. Serialize snapshot capture +
reconcile under one lock and cap request-path syncs at 5s so failures
degrade to the periodic tick.

Part of #205

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(worker): coalesce immediate session syncs instead of queueing on a lock

Review iteration 2: a stop-triggered sync queued on a bare mutex could
expire its 5s deadline waiting behind a slow tick and fail before
removing the stopped row. Coalesce instead: one owner reconciles at a
time, callers that arrive mid-flight return immediately after flagging
a follow-up pass, and the owner re-captures a fresh (post-change)
snapshot afterwards. Request goroutines never block behind another
sync, snapshots still commit in capture order, and an expired-context
owner leaves the queued pass for the next tick instead of burning it.

Part of #205

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 10:07:09 -04:00
9cae868a27 feat(downloads): offline sync for mobile — downloads v2 (#258)
* feat(downloads): offline sync for mobile (downloads v2)

Replace internal/download with a unified internal/downloads package and add
fully-offline download + watch-sync support for mobile clients, across five
independently-shippable phases:

- Phase 0: reshape the downloads table and the /downloads contract to be
  device- and format-aware; add GET /downloads/capability; extend
  DownloadConfig (default-off keys); update the web download hooks/components in
  lockstep. This is the one approved pre-lock exception to the additive-only
  /api/v1 rule (the web app is the only consumer and is updated together).
- Phase 1: managed device-library entries (create/list/PATCH/delete/serve),
  keyed on the X-Silo-Device-Id header.
- Phase 2: offline playback manifest plus artwork/subtitle proxy endpoints that
  strip every presigned URL (inline thumbhashes + authenticated proxies).
- Phase 3: prepare-to-file (remux + transcode-to-single-file) as a durable,
  leased artifact queue with startup recovery, hosted on the task manager;
  playback.PrepareFile emits one +faststart MP4. Adds the admin transcode
  toggle and per-artifact LRU cleanup.
- Phase 4: offline progress reconciliation -- a clamped event_at LWW key plus a
  server-assigned synced_seq cursor on watch_progress; an optional clamped
  updated_at on POST /sync/progress and an opaque ?since= cursor on
  GET /progress (additive; existing callers unaffected).

Security & reliability invariants, each with an acceptance test:
1. Server-owned sync ordering: ?since= delta delivery is driven only by the
   server-assigned synced_seq; the client clock is bounded (event_at, clamped
   to now+skew) and used only for last-write-wins on the caller's own profile.
2. Full profile+device authorization on every managed endpoint, with a
   per-profile content/library access re-check before serving any bytes/assets.
3. Durable artifact recovery: a transactionally-claimed (FOR UPDATE SKIP
   LOCKED), lease-heartbeat, attempt-counted queue with a startup sweep, so no
   crash strands a download in preparing and concurrent workers never
   double-encode.

Migrations are timestamped Goose files: reshape downloads (device/format);
download_artifacts (durable queue); watch_progress event_at/synced_seq.

DB-backed acceptance tests skip without SILO_TEST_DATABASE_URL and run in CI;
the invariant-1 progress test also runs against the real SQLite backend locally.

Client repos (silo-android, silo-apple) consume the reshaped /downloads/*
contract and the updated_at/?since= progress fields and require coordinated
follow-up.

Implements the maintainer-approved v1 capability proposal for offline sync
(downloads v2).

AI-use disclosure: implemented by Claude (Claude Code) from the approved design
doc under docs/superpowers/specs, with human review.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(downloads): series & season downloads + client-pull monitoring

Build season downloads and a "monitor a series" capability on top of the
downloads v2 (offline sync for mobile) work.

Season downloads:
- POST /downloads accepts season_number (with series:true) to download one
  season. CreateSeries/CreateSeason share one body via a listEpisodes closure
  and register managed entries under a shared batch_id (original-only). Episode
  files are resolved in a single batched query.

Series monitoring (auto-download), client-driven:
- New device-scoped download_subscriptions table with a Sonarr-style mode
  (all | future | latest_season | specific_seasons), a client-enforced
  delete_watched flag, and a max_storage_bytes cap. The server never deletes
  on-device files; retention and the hard cap are the client's, the server
  only soft-gates registration.
- The client calls POST /downloads/subscriptions/sync on open / background
  refresh; the server registers the in-scope, not-yet-downloaded episodes
  (idempotent via the managed-entry unique index) and the device pulls them on
  its own schedule. No background worker and no dependency on the notifications
  subsystem. latest_season follows new seasons (>= subscribe-time season);
  future excludes the back catalog via air date.
- Subscription CRUD + sync are profile+device authorized (device id from the
  X-Silo-Device-Id header only) with a per-request content-access re-check. The
  capability endpoint advertises season_download / series_monitoring /
  monitoring_modes.

Also lands the downloads-v2 work already present in the tree: durable artifact
(remux/transcode) preparation and offline watch-progress reconciliation, plus
the design-spec updates.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* WIP: epitaxy pre-switch from feat/downloads-v2-offline-sync

* test(downloads): fix deterministic ID collision in reconcile test

Artifact IDs are time-sortable, so two artifacts created in the same
moment share their first 8 chars; combined with a captured timestamp the
two preparing-download IDs collided on downloads_pkey. Use the full
artifact ID, which is unique per row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): support sqlite userdb backend for managed downloads

With the sqlite userdb backend, profiles live only in per-user SQLite
stores and public.user_profiles stays empty, so user_devices'
profile FK made every managed create/subscription/offline-sync request
fail with an FK violation. Drop the FK (shared Postgres tables must not
FK profile tables — same rule as notifications) and replace the lost
cascade with an app-level purge on profile deletion, wired through
ProfileHandler for both backends. DB-backed regression tests cover the
no-Postgres-profile-row path and the purge cascade.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): dispatch encode kick asynchronously

triggerDrain invoked the kick inline, and the kick (taskmanager RunTask)
executes the encode task on the caller's goroutine — so a POST
/api/v1/downloads with a bitrate quality blocked the HTTP request on the
entire queue drain, ffmpeg encodes included, delaying the 202 by minutes
on an idle queue. Dispatch the kick on a goroutine; the task manager
already serializes concurrent runs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): enforce per-user quota on the encode pipeline

Two gaps let a user bypass MaxConcurrentPerUser entirely for prepared
downloads: artifact-backed rows are created in 'preparing' (never
'queued'/'downloading'), which CountActiveByUser didn't count, and
createArtifactDownload enqueued the encode job before limiter.Check, so
even a 429-rejected request left a job the worker would transcode.
Count 'preparing' as active and check the limiter before Ensure; managed
replacements stay quota-exempt since they don't add a row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): protect ephemeral artifact links from LRU eviction

HasActiveLink only counted managed (device_id IS NOT NULL) rows, so
under a byte budget Cleanup could delete an artifact still referenced by
a ready-but-unfetched ephemeral web download — permanently 404ing a row
the API kept listing as ready (the artifact row is gone, so recovery
can't re-queue it). Any non-terminal link now protects the artifact;
only artifacts whose links are all cancelled/failed/revoked are
evictable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): batch manifests skip bad entries instead of failing whole batch

One deleted or access-filtered episode made GET
/downloads/batches/{id}/manifests 404 for the entire season, so a
client could no longer fetch manifests for the still-valid entries.
Report unbuildable entries in a skipped[] array (revoked | not_found |
error) alongside the delivered manifests, mirroring the create path's
skip idiom. Also cut the batch cost: the shared series detail is
resolved once per batch instead of once per episode, and buildSubtitles
reuses the already-loaded media file instead of re-querying it per
manifest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migrations): wrap DO block in StatementBegin/End markers

Under NO TRANSACTION goose splits statements on semicolons, so the
dollar-quoted DO block failed every fresh install with 'unterminated
dollar-quoted string' (SQLSTATE 42601). Already-applied databases are
unaffected. Same fix is being applied to main; identical content merges
cleanly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(api): allow season 0 (Specials) in season downloads

season_number was a plain int dispatched with '> 0', so requesting the
Specials season was indistinguishable from omitting the field and
silently broadened to a full-series download. Dispatch on pointer
presence, treat 0 as the Specials season, and reject negatives with 400.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): capability quality_presets is never JSON null

PresetsFor returned a nil slice when downloads are disabled or the user
lacks the permission, and Capability's []string{} initialization was
immediately overwritten by it — so GET /downloads/capability serialized
"quality_presets": null where the contract documents an array.
Normalize at the source so every caller inherits the guarantee.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): subscription sync correctness + batched registration

Three subscription fixes:

- A paused subscription no longer syncs: PATCHing scope (or pausing and
  changing scope in one request) registered episodes for a monitor the
  user had just stopped, inconsistently with SyncSubscriptions' guard.
- SubModeFuture compares calendar days (UTC): air_date is date-only, so
  the strict instant comparison permanently excluded episodes airing the
  same day the user subscribed; episodes with no air date now fall back
  to their ingest time instead of never registering.
- Registration is one batched fetch (GetManagedEntriesByKeys) plus one
  batched INSERT ... ON CONFLICT DO NOTHING RETURNING
  (CreateManagedEntriesBatch) instead of a SELECT+INSERT per episode —
  a 300-episode series cost ~600 sequential round trips per request and
  every no-op sync re-walked the full set. RETURNING yields exactly the
  new rows, so the sync response's 'registered' count now honestly
  reports 0 in the steady state instead of the full in-scope count on
  every app open. The now-unused InsertManagedEntryIfAbsent is removed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(userstore): stamp triggers own the event_at LWW key

MarkProgressBatch (jellycompat series mark-played) advanced updated_at
but never event_at, and both stamp triggers only defaulted event_at when
NULL — so a queued offline event with a client time between the row's
old event_at and the mark could win SetProgressIfNewer and resurrect a
stale resume position that then re-synced to every device.

Make the triggers authoritative instead of adding a tenth hand-written
SET clause: whenever an UPDATE changes updated_at without explicitly
changing event_at, the trigger advances the LWW key; writes that do set
event_at (offline sync's clamped client event time) keep their value.
Postgres gets a CREATE OR REPLACE migration; SQLite gets a v12 userdb
migration that drops and reinstalls the trigger bodies (CREATE TRIGGER
IF NOT EXISTS never replaces). Conformance tests cover both batch paths,
the preserved-client-time invariant, and the v11→v12 upgrade.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): lifecycle hygiene — squash migrations, dead status, stale-row sweeps

Migrations: fold the 20260621 corrective migration back into the base
Downloads V2 migrations (its columns/constraints already exist there)
and fix the reshape Down, which re-added the narrow status CHECK without
collapsing managed-lifecycle rows first — rollback aborted on any DB
with preparing/ready/revoked rows; validated against a live row. Branch
databases that applied the corrective migration need its version row
removed: DELETE FROM goose_db_version WHERE version_id = 20260621020459.

Code: drop the dead 'registered' status (nothing ever wrote it; the
lifecycle is preparing -> ready; 'revoked' stays reserved for the
planned admin revoke flow) along with unused KindDirect and
ErrInvalidFormat.

Sweeps: Cleanup now runs an age-based hygiene pass independent of the
byte budget — cold terminally-failed artifacts (with .part leftovers),
orphaned ready artifacts no download row references, and ephemeral web
rows older than their convenience-record lifetime (also unpinning their
artifacts and bounding GET /downloads growth). The byte budget remains
the disk quota per the limits & restrictions design.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(downloads): sync API doc with v2 fixes; HEAD on file route; Android handoff

Document the contract changes from the review fixes: batch-manifest
skipped[] shape, honest subscription 'registered' semantics, season 0 =
Specials, always-array quality_presets, bytes_sent actual behavior,
ephemeral 7-day retention, header-pairing requirement, progress-delta
deletion caveat, and the ready/failed push event schema (new §9.4).
Add an Android client handoff section (§11) mirroring the Apple one,
register HEAD on /downloads/{id}/file for download stacks that probe
before ranged GETs, and add season_number to the web create-request
type. Flag the /direct-download session-token-in-URL tradeoff; a
short-lived download-scoped URL is a follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: consolidate download/progress helpers, prune dead code, gate sweeps

Behavior-preserving consolidation from the Downloads V2 review:

- appendVideoFilterArgs: one home for the burn-in/hwaccel -vf selection,
  shared by the HLS builder and the single-file prepare builder (the
  drift pattern that already bit tone-mapping once).
- userstore.ResolveProgressState: one home for the min-resume/watched
  threshold rule, replacing five identical copies across both store
  backends and the offline-sync ingest.
- Download file selection ranks resolutions via access.CompareQuality
  (adds 4320p, agrees with playback) instead of a private switch.
- writeSubtitle uses the shared subtitles.SubtitleContentType mapping.
- config.DefaultTranscodeDir replaces three '/tmp/silo-transcode'
  literals.
- Read-side quality/revision defaulting helpers removed: insertArgs plus
  the NOT NULL/CHECK schema already guarantee the invariant.
- Dead code removed: Repository.ListByUser, SubscriptionRepository.
  ListActiveBySeries, and the stale auto-register-worker comments (the
  design is client-pull; no worker exists).
- Redundant left-prefix indexes dropped from the base migrations (their
  unique indexes serve the same prefixes).
- recover()'s disk-presence sweep and the stale-row hygiene sweep run on
  startup then hourly instead of every 30s tick (both are O(cache
  size)).
- gofmt/prettier fixes for pre-existing drift in handlers/playback.go
  and pages/Profiles.tsx.

Deferred (noted for follow-ups): quality-ladder preset table collides
with the drafted download limits & restrictions design, which specifies
its own ladder helper; Download-literal construction consolidation and
the managed-identity value object remain open.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(downloads): draft download limits & restrictions design

Design input for the follow-up v1 capability proposal (quality ceiling,
batch size cap, per-user quantity/bandwidth overrides). Committed with
downloads v2 because the remediation work explicitly defers the quality
ladder refactor and revocation wiring to this spec.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(progress): reject malformed updated_at; clamp negative progress inputs

Review findings on #258:

- A malformed (non-RFC3339) updated_at in POST /sync/progress previously
  parsed to the zero time, which clampEventAt treated as "now" — letting a
  stale offline event win LWW as a fresh server-time write. The item is now
  rejected with a per-item error instead.
- ResolveProgressState now clamps negative position/duration before
  classification so no backend can persist negative progress through
  UpdateProgress/SetProgress.
- The online-write event_at invariant test is table-driven over both
  SetProgress and UpdateProgress, which share the same contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): close review gaps — permission gates, file-access recheck, artifact-true manifests

Review findings on #258:

- UpdateSubscription now applies the same feature/DownloadAllowed gate as
  CreateSubscription and SyncSubscriptions; a PATCH could previously
  re-activate or widen a monitor and register managed rows after an admin
  disabled downloads or revoked the user.
- Serving download bytes (managed and ephemeral) and /direct-download now
  mirror playback's per-file authorization via catalog.FileAllowedByAccess:
  library scope and the profile's max playback quality are re-checked at
  serve time, with artifact-backed rows checked against the artifact's
  resolution (a 720p transcode of a 4K source stays servable under a 1080p
  ceiling).
- Offline manifests for remux/transcode entries now describe the prepared
  artifact (container, codecs, resolution, single selected audio track)
  instead of the catalog source file the client never receives.
- ArtifactRepository.Requeue reports ErrNotFound when the row was
  concurrently swept; ArtifactManager.Ensure recreates the job in that case
  instead of linking downloads to a dead artifact id.
- "No downloadable episodes" is a sentinel (mapped to 404
  no_downloadable_episodes) rather than a bare error that surfaced as 500.
- Subscription season_numbers are bounds-checked (0–9999) before the int32
  narrowing in the repo could silently wrap them.
- HandlePatchDownload reuses requireManaged instead of hand-rolling the
  same managed-identity checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 22:05:36 -04:00
QuickandGitHub a0fbe2cdda [codex] feat(jellycompat): use catalog search provider (#236)
* feat(jellycompat): use catalog search provider

* fix(jellycompat): route video search buckets to provider
2026-06-27 20:10:48 -04:00
02e62767a1 feat(watchsync): sync watchlists with Trakt/Simkl/MDBList (#227)
* feat(watchsync): sync watchlists with Trakt/Simkl/MDBList

Extend the watch-providers feature to sync a user's watchlist, generalizing
the existing favorites pipeline rather than duplicating it.

What changed
- Generalize the favorites sync into one ListKind-parameterized pipeline
  (internal/watchsync/lists.go) driving both favorites and watchlist; the
  per-favorites service methods are replaced by kind-generic ones. The shadow
  table watch_provider_favorite_items becomes watch_provider_list_items with a
  list_kind discriminator.
- Providers: Trakt gains watchlist sync (/sync/watchlist, distinct from
  favorites); Simkl gains plan-to-watch sync; MDBList is re-mapped from
  favorites to watchlist (its only list is a watchlist) — its capabilities now
  report import_favorites=false / import_watchlist=true, and the migration
  re-binds existing MDBList connections.
- Auto-remove watched items from the watchlist: a standalone, default-on
  profile preference (user_profiles.remove_watched_from_watchlist) removes a
  movie when watched and a series once every episode is watched. Implemented as
  watchstate.CompletionObserver (internal/watchlist.Maintainer), wired into the
  manual mark-watched, playback-stop, and jellycompat mark-played paths.
- Optional MDBList sort-order mirroring: an opt-in, capability-gated toggle
  mirrors MDBList's watchlist order into Silo via user_watchlist.sort_index;
  ListWatchlist orders by sort_index then added_at, so both /api/v1/watchlist
  and the catalog watchlist view inherit it.
- Real-time + scheduled: local add/remove pushes to connected providers
  immediately (removals gated by the opt-in removals toggle); the hourly job is
  the inbound/import + retry/reconcile path.
- Web: watch-provider settings gain watchlist import/export/removals and
  "mirror watchlist order" toggles plus watchlist sync stats.

Why
- The favorites and watchlist pipelines are ~90% identical; generalizing keeps
  one code path (per CLAUDE.md's anti-duplication guidance) instead of cloning.

API/compat
- All new fields on ConnectionStatus/Capabilities/ConnectionUpdate/SyncRun and
  the web types are additive (Silo v1 additive-only rule). No existing field is
  renamed, removed, or retyped.

Risks / follow-up
- MDBList capability flip is intentional and client-visible: silo-android /
  silo-apple may need to surface MDBList under the watchlist (not favorites) UI.
- MDBList existing users: their MDBList list previously mirrored Silo favorites
  and now mirrors Silo watchlist; the first post-migration sync is a union
  (removals default off), so nothing is destructively purged.
- Order mirroring reflects the order MDBList returns from /watchlist/items
  (couldn't confirm against their docs — Cloudflare-blocked); if it ever
  diverges from the UI sort, a sort param is the small follow-up.

Tests: new maintainer (auto-remove) and watchlist-order unit tests; provider +
service tests updated. go build, go test (affected pkgs), migrate-validate,
verify-local-paths, web prettier/eslint/tsc all pass.

AI-use disclosure: implemented with Claude Code (Claude Opus 4.8).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(watchsync): update list shadow table references

* fix(watchsync): address review — retry/progress + error propagation

Addresses CodeRabbit review on #227:
- maintainer: propagate transient catalog lookup errors instead of silently
  treating every items.GetByID failure as "maybe an episode".
- exportList: mark every queued item not confirmed sent (not_found, failed, or
  omitted) so the pending loop always advances; the next run's upsert clears the
  error and re-attempts, so transient failures still retry.
- removePendingListItems + realtime removal: treat Sent and NotFound as
  reconciled; leave true failures pending (no last_error, which would strand
  them from the removal query) so the scheduled run retries, using in-memory
  dedupe to terminate the loop.
- exportLocalListItems: send the normalized items (with computed
  ProviderItemKey), not the original event slice.
- UpdateConnection: clear mirrored watchlist order before persisting the disable
  and propagate failures, so a failed clear can't report "disabled" while
  sort_index ordering is still active.
- web: include favorite + watchlist removal counts in the exported "sent" total.
- test: align serviceFakeRepo list-state with Postgres (clear last_error on
  successful transitions); add maintainer error-propagation test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 16:05:53 -04:00
QuickandGitHub d83f7fadef fix(audiobooks): mirror ABS playback into live sessions (#203) 2026-06-26 12:27:12 -04:00
QuickandGitHub 4473f8c60b Merge pull request #220 from Silo-Server/codex/search-provider-interface
feat(search): add provider interface with initial Meilisearch support
2026-06-26 09:32:53 -04:00
QuickandClaude Opus 4.8 ed6c084c68 feat(search): gate index events by active provider and harden rebuild reconcile
Completes the search-provider-interface wiring that the catalog hardening
commits already call into:

- Skip the transactional search-index-event write path when Meilisearch is
  not the active provider (ItemRepository.WithActiveSearchProvider /
  SearchIndexEventRepository.disabledByActiveProvider).
- Dead-letter catalog_search_index_events after 10 attempts instead of
  retrying forever.
- Track the rebuild high-water mark (MaxEventID / MarkProcessedThrough) and
  persist last_processed_event_id in UpdateStateAfterRebuild so a rebuild
  reconciles events enqueued during the rebuild.
- Validate (read-only) the embedding lock when embedding a search query
  instead of establishing/mutating it.
- Surface total_exact on the legacy /items browse response.
- Wire the active catalog search provider into the scanner and item repo at
  startup.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 08:34:04 -04:00
Quick 40329f616d perf(search): speed up catalog query results 2026-06-25 20:46:08 -04:00
Quick 6ca427096b Add catalog search provider support 2026-06-25 16:20:14 -04:00
Quick af52ea7bcc fix(plugins): use scoped installation store for image resolver 2026-06-25 15:08:55 -04:00
Quick 2045b7a0b2 feat(plugins): add image resolver registry 2026-06-25 14:47:48 -04:00
Quick 227986094b Propagate playback client metadata through session sync 2026-06-23 14:26:14 -04:00
CoffeeKnyteandGitHub 6e1f79e8a9 feat(jellycompat): expose library collections as an auto-shown Collections view (#175)
* feat(jellycompat): expose library collections as a Collections library view

Surface server library collections as a top-level Jellyfin "Collections"
library (CollectionType "boxsets") so compat clients see them as the first
library in /UserViews and can browse them by ParentId. The BoxSet machinery
(list/detail/children) already existed; this adds the library wrapper.

- catalog: add LibraryCollectionRepository.AnyVisibleInLibraries, an
  index-only EXISTS probe (no item join/aggregation) used to gate the view so
  an empty Collections tab never shows. Mirrors collectionVisible semantics
  (multi-library scope rows or the legacy single library_id fallback).
- jellycompat: add the synthetic Collections CollectionFolder (fixed Jellyfin
  sentinel ID, stable across servers), prepend it to the user's views when a
  visible collection exists, and route ParentId/Items-by-ID for that sentinel
  to the existing BoxSet listing. ChildCount is left omitted (no per-call
  count, no unwatched badge).

AI-use disclosure: implemented with assistance from Claude.

* fix(jellycompat): display collection posters and a generated Collections tile

Library collections surfaced as Jellyfin BoxSets showed blank cards: their
poster_url is frequently a bundled frontend template path
(/images/collection-templates/x.jpg), which the compat image route passed
through unchanged and then rejected in parseRemoteImageURL (no scheme/host),
returning BadGateway. Clients probed Images/Primary and got nothing.

- images: serve app-relative artwork (bundled template posters) straight from
  the embedded frontend FS (new ImagesHandler.frontendFS), with content-type
  and cache headers. Wired through jellycompat.Dependencies.FrontendFS.
- poster_gen: on-the-fly gradient poster generator (per-title hue, centered
  white caption with black outline, gobold/opentype), memoized in a bounded
  cache. Used for the synthetic Collections library tile and as a fallback for
  collections without usable artwork, so cards are never blank.
- consolidate collection/view image routing in HandleItemImage, authorized by
  the signed tag or an authenticated, visibility-checked session.

AI-use disclosure: implemented with assistance from Claude.

* fix(jellycompat): declare 2:3 PrimaryImageAspectRatio on BoxSets and Collections tile

Clients defaulted collection cards to a square and crop the 2:3 poster to fit.
Set PrimaryImageAspectRatio (portrait 2/3) on the BoxSet DTO and the synthetic
Collections library tile so the full poster is shown, matching Jellyfin.

AI-use disclosure: implemented with assistance from Claude.
2026-06-18 10:20:48 -04:00
14ffc91dfb [codex] Expand provider image cache queue (#176)
* feat(metadata): expand provider image cache queue

* fix(metadata): harden provider image cache queue

Addresses bug-review feedback from Codex/CodeRabbit on the metadata image
cache pipeline. All findings validated against the code before fixing;
false positives (rows/connection deadlock, PhotoSourcePath merge coupling)
were confirmed non-issues and left unchanged.

- Honor metadata.cache_images for the background processor. The
  cache_metadata_images task was registered whenever S3 was configured,
  so merely enabling object storage downloaded the entire provider-artwork
  catalog even with caching disabled. Add ImageCacheProcessor.SetEnabled,
  gate RunOnce/RunUntilIdle on it, and wire it (with hot reload) from
  cfg.Metadata.CacheImages in main.go.
- Guard terminal job updates with lease ownership. EnqueueBatch can
  repurpose a running row with a new source; MarkSucceeded/MarkFailed
  keyed on id alone let a stale worker finalize the replacement job and
  drop the new artwork. Thread locked_by through and add
  status='running' AND locked_by=$n guards.
- Avoid uploading stale jobs onto the live artwork key. Verify the
  target still references the job's source (CurrentTargetSourcePath)
  before CacheImage, so a job whose source an admin/refresh already
  replaced cannot overwrite the deterministic storage object.
- COALESCE nullable external IDs in EnqueueExistingProviderArtwork. A
  NULL tmdb_id/tvdb_id/imdb_id on any candidate failed the scan and
  aborted the whole cache run; matches the existing item_repo pattern.
- Stop re-downloading the catalog every 30 days. Discovery now skips
  targets whose *_path is already a cached relative path, making the
  cached row the durable dedup marker instead of the prunable job row.
- Decouple catalog sweeps from queue draining. RunOnce no longer runs
  discovery per batch; RunUntilIdle sweeps only when the queue drains and
  throttles full sweeps to every 15m, so idle installs stop full-scanning
  every entity table each minute.
- Requeue claimed-but-unstarted jobs on cancellation. Acquire the
  semaphore before spawning workers and RequeueClaimed any jobs not yet
  started, instead of leaving them locked until the 15m lease expires.
- Skip the backoff sleep after the final upload attempt in
  putObjectWithRetry (saves ~1.5s on permanent failures).
- Add the s3/file/local/upload/generated exclusion to the seasons and
  episodes backfill in migration 20260617184537 for consistency with the
  later migration (the bad backfill was inert downstream, but the
  asymmetry is removed).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 10:07:58 -04:00
c4cbcddeae feat(manga): manga library type — series grouping, reading loop, AniList/MangaDex metadata + status badge (#138)
* docs: design spec for manga library type (host sub-project)

Forks the ebooks library type into a 'manga' type: series detected from the
folder tree as a first-class type='manga' item, .cbz/.cbr chapters stay
readable ebook items linked via a new manga_chapters table, browse shows series
cards, enrichment targets the series item at content level 'manga'. Hands off to
a follow-on plugin spec for the manga metadata source.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: implementation plan for manga library type (host)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(scanner): manga filename index/volume parser

* feat(scanner): manga series-name-from-folder detection

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(plan): align manga DB/scanner tasks to scanner pure-planner pattern (no test-DB)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(scanner): manga parser corpus regression

Add TestParseMangaIndexCorpus — 36 real-world scanlation filenames
covering bare chapter, decimal chapter, v/vol-prefix volume, and
c/ch-prefix chapter patterns; asserts <5% miss rate.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(db): manga_chapters link table

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(scanner): manga_chapters repository + pure chapter-write mapping

Adds mangaChapterWrite (pure, unit-tested), upsertMangaChapter, and
listMangaChapters following the ebook/audiobook thin-SQL pattern.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(scanner): recognize manga library type

Add isMangaLibraryType helper (unexported, matching the style of
isEbookLibraryType / isAudiobookLibraryType) with a corresponding
TestIsMangaLibraryType unit test.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(api): manga library content level

Map library type "manga" to content level ["manga"] in
metadataContentLevelsForLibraryType so that seedDefaultChain seeds a
manga-level metadata provider chain when a manga library is created.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(scanner): route manga libraries to a manga scan path

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(scanner): group manga chapters under a manga series item

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(scanner): give manga series item a library membership so it browses

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(catalog): browse manga libraries as series

Accept "manga" as a valid media_scope so a manga library browses only its
type='manga' series items; the per-chapter type='ebook' items are naturally
excluded because MediaScopeItemTypes("manga") expands to {"manga"}. Add the
manga default library sections (scoped to media_scope='manga') so the library
feed shows series cards. Refresh the two media_scope validation error messages.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(catalog): manga series detail lists chapters

For a type='manga' item, attach its chapters to the detail response via a new
MangaDetailExtension. fetchMangaChapters joins manga_chapters to media_items on
the chapter content ID, scopes to the series, and orders by chapter_index
(NULLS LAST) then sort_title — matching the scanner's chapter ordering.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(web): manga detail types + library browse scoping

Add MangaChapter/MangaDetailExtension TS types mirroring the host
catalog structs, wire manga? onto ItemDetail, and admit "manga" as a
QueryDefinition.media_scope. Scope manga libraries to media_scope=manga
in browse (host expands it to type=manga series items) while reusing the
ebook sort universe via getLibrarySortRelevanceScope. Add isMangaLibraryType.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(web): manga series detail with volume-grouped chapter list

Add MangaContent detail view: a DetailHero series header plus a chapter
list grouped by volume. groupMangaChapters (pure, unit-tested) buckets
chapters by their volume token, orders chapters within a group by
chapter_index (nulls last) and orders groups by their minimum index;
loose (volume-less) chapters collapse into a trailing "Chapters" group.
Each chapter links to the existing ebook reader by content_id alone
(file_id is optional — the reader resolves the file server-side), reusing
buildMediaPlayHref. Admit "manga" into ItemDetail.type and wire the
detail switch. Continue-reading is deferred (needs per-chapter progress
fan-out / a last-read timestamp not in the current payload).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(web): handle manga in playable-type + collection filter-scope unions

Adding "manga" to the shared ItemDetail["type"] and
QueryDefinition["media_scope"] unions leaked into consumers with narrower
local types, breaking the production tsc build. Fixes:

- mediaNavigation: admit "manga" into PlayableMediaType. Manga series are
  not directly playable (you open the detail page and read a chapter,
  itself an ebook item), so buildMediaPlayHref falls through to the item
  href for them, like series/season.
- FilterRuleEditor: add "manga" to FilterRuleMediaScope and relabel
  "watched" -> "Read" for manga as well as ebook (manga is read).
- CollectionGuidedRulesEditor: add "manga" to GuidedFormState.mediaScope,
  a "Manga" media-type option, ebook-like "Read Status" labels, and map
  manga -> ebook sort-relevance scope (manga has no dedicated sort scope).
- CatalogFilterBar (cascading leak surfaced after the above): add a
  "Manga" scope option and map manga -> ebook sort-relevance scope in both
  scope handlers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(web): offer manga as a library type in the create dialog

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(scanner): strip scene-release junk from manga series names

Add cleanMangaSeriesName which repeatedly strips trailing parenthetical
groups (year, year-range, Digital, release-group tags) then trims any
dangling dash, so folder names like "404 Demons (Digital) (Oak)" resolve
to "404 Demons". Wire it into mangaSeriesFromPath so both the series
title and the mangaSeriesGroupKey identity key use the cleaned value.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(web): flat volume/chapter manga list; nest only multi-chapter volumes

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(scanner): parse manga index after stripping series-name prefix

Numbers inside a series title (e.g. "404 Demons", "365 Days to the
Wedding") were wrongly grabbed as the chapter number because
parseMangaIndex matched the first bare number in the full filename.
mangaIndexForFile now strips the series-name prefix before delegating
to parseMangaIndex, so only the number that follows the title is used.
reconcileMangaFile in manga_scan.go is updated to call mangaIndexForFile
instead of parseMangaIndex directly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(scanner): stop missing-file reconcile from deleting manga series items

Manga series items are file-less virtual parents; the shared
ReconcileFolderMembership swept them every scan because they have no
media_file. Exclude type='manga' from file-presence membership reconciliation,
and add a manga-scan step that deletes only series with zero remaining chapters.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ebooks): exclude manga chapters from individual ebook enrichment

Manga chapters are type='ebook' parts of a series; the ebook enrichment sweep
was searching each one against book sources (Gutenberg/Anna's/etc.) and failing
in a pointless storm. Exclude items with a manga_chapters link; series-level
enrichment is handled separately.

* docs: design spec for manga metadata plugin + series enrichment (sub-project 2)

New silo-plugin-manga-metadata (AniList, high-confidence matching) + a host
MangaEnricher for type='manga' series; default-enabled metadata source for manga
libraries.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: implementation plan for manga metadata plugin + series enrichment

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(db): manga_enrichment_state table

Mirrors ebook_enrichment_state: dedicated failure counter for the manga
enrichment sweep so it does not contend with media_items.refresh_failures.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(manga): series enricher (claims type='manga', resolves manga chain)

* feat(manga): sync_manga_metadata task + enricher wiring

* feat(catalog): expose manga chapter/volume counts in browse

Add manga_chapter_count and manga_volume_count to browse cards so the
frontend can render a Vols N / Ch N chip on manga series. The counts come
from two index-backed correlated subqueries over manga_chapters in the
browse SELECT (mangaCountColumns), scanned positionally before added_at and
nilled out for non-manga rows. Threaded through models.MediaItem and exposed
on the itemListResponse JSON card.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sections): scope manga home recent sections to type=manga series

A manga library mixes type='manga' series with type='ebook' chapters, so
the auto-generated home 'Recently Added/Released in <Library>' rows surfaced
the junk chapter filenames. Add GeneratedHomeLibraryRecentConfigScoped which
emits the modern QueryDefinition shape (library_ids + media_scope) so a manga
library's generated home rows filter to type='manga' only. Other library
types are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(catalog): exclude manga chapters from browse/section/search surfaces

Manga CHAPTER items (type='ebook' rows linked into a type='manga' series
via manga_chapters) were leaking into catalog browse, section resolution,
and search as standalone items showing junk filenames. They are internal
sub-units of the series and only the series should appear.

There is no single shared item-listing chokepoint: browse, the query/preview
executor, and search each build their own WHERE. Add a shared, index-backed
anti-join predicate (manga_chapters.chapter_content_id is the PK) via
mangaChapterExclusionWhere and wire it into all three builders. By-id fetch
paths that legitimately resolve chapters (ebook reader, continue-reading,
series detail chapter list) use separate queries and are unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(scanner): use #NN as the manga volume for Vol.YYYY #NN releases

mangaVolYearIssue early-return was returning the year token (e.g. "Vol.2003")
as the volume label, which the frontend couldn't prettify to "Volume N".
Now returns "v<issue>" (e.g. "v04") so the existing frontend regex ^v?(\d+)$
renders it as "Volume 4" correctly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(web): manga count chip on posters

Add an optional manga_chapter_count / manga_volume_count to the browse
item type and render a top-right "Vols N" / "Ch N" chip on ItemCard,
strictly gated on type==='manga'. The label prefers "Vols" when the
volume count dominates, "Ch" otherwise; the chip is hidden when the
chapter count is missing or non-positive. No other card type renders it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(web): manga reader back returns to series (no loop)

The ebook reader's back action defaulted to the chapter's own item
detail (/item/<chapter>), whose back returned to the reader — an
infinite loop for manga chapters. The reader now honors an explicit
backTo search param when present, navigating there instead. Absent for
normal ebooks, so their back behavior is unchanged. Only manga chapter
rows pass backTo, keeping the fix manga-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(web): manga chapter row actions (read/mark-read/download)

Each manga chapter/volume row now offers Read (the existing reader link,
now carrying a backTo to the series), Mark-read (the shared watched-state
mutation per chapter content_id), and Download (lazily fetches the
chapter's file versions on demand and opens the shared
DownloadVersionPicker, gated on user.download_allowed). The
volume-unit / loose-chapter / section structure from buildMangaList is
unchanged. Scoped to MangaContent only; EbookContent is untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(web): validate reader backTo param is a safe in-app relative path

Prevents open-redirect / javascript:-URI XSS from a crafted ?backTo= URL.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(catalog): include per-chapter read state in manga detail

Manga chapters are ebook items, so a chapter is "read" when the viewer's
ebook_reader_progress row crosses the finished threshold. fetchMangaChapters
now LEFT JOINs that table scoped to the AccessFilter's user_id/profile_id and
exposes a per-chapter Read bool on MangaChapter, threaded through
buildMangaExtension. The detail payload previously carried no read state, so
the row toggle always started unread.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(web): manga rows reflect read state on load

MangaChapter now carries an optional read flag from the detail payload, and
MangaRow seeds its mark-read toggle from chapter.read instead of always
starting unread. The optimistic toggle + shared watched mutation are
unchanged; only the initial value is seeded.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(sections): exclude manga chapters from recently-added/released/random + other library-listing sections

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(sections): manga recently-added/released cards show the latest volume's cover

* fix(manga): keep enrichment honest about no-match vs enriched, batch 50->200

- sweep stats now separate enriched / no_match / failed: a stamped no-match
  was counted (and logged) as an enrichment, which masked a collapse of the
  real match rate during the backfill
- batch size 50 -> 200 (SILO_MANGA_ENRICH_BATCH overrides): with the plugin
  serving GetMetadata from its search cache an item costs one rate-limited
  AniList request, so a sweep still fits the 5-minute task interval

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(manga): size enrich batch to the 5-minute interval at AniList's real budget

140 items x ~2.1s/request fits the interval; an overlong sweep makes the task
manager drop the next trigger and the effective rate falls below the AniList
budget.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(catalog): manga count chip data missing from library browse

manga_chapter_count/manga_volume_count were only added to BrowseRepository,
but /library/{id}?tab=library flows through previewQuerySource ->
QueryExecutor.PreviewPage, which selects qualifiedListItemColumns and scans
with scanItems - so manga cards never carried the counts and the Vols/Ch
poster chip stayed hidden.

Append mangaCountColumns to the preview-page SELECT and scan them via a new
scanItemsWithMangaCounts (nil for non-manga rows, mirroring scanBrowseItems).
Extract listItemScanDests so the three scan variants share one destination
list instead of duplicating the 48-column scan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): manga chip reads 'X Volumes · X Chapters', menu verbs say Read

- chip: show distinct-volume and loose-chapter counts side by side instead
  of the single 'Vols N'/'Ch N' heuristic; mangaCountColumns now counts
  DISTINCT volume tokens (rows sharing a volume are one volume) and only
  un-volumed rows as chapters
- watched-state labels: type='manga' fell through to the video default, so
  the card dot menu and detail page said 'Mark Watched' - manga now uses
  the ebook reading verbs (Mark Read / Mark Unread, 'Marked as read' toast)
- format MangaContent.test.tsx (pre-existing prettier miss)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(manga): backdrop enrichment - banner hero art + backdrop-only backfill

- cache remote backdrops like posters (cacheRemoteImages generalizes the
  poster-only path; failures keep the provider URL, which still renders)
- claim arm for enriched items missing a backdrop: fetched by stored
  provider ID (search skipped - no rate spend, no re-match risk) and only
  the backdrop is written; stamping after the attempt keeps banner-less
  series from being re-claimed every sweep
- backfill = one-time SQL clearing last_refreshed for poster-set/
  backdrop-empty manga

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(manga): reading-loop UX - continue CTA, next chapter, series-aware cards, file details

Fixes the four high-priority findings from the manga UX review plus a
file-inspector request:

- H1: series hero gets a Continue / Start Reading / Read Again CTA
  targeting the first unread chapter (firstUnreadChapter over the ordered
  list), plus an overflow menu (View Details, admin Refresh Metadata)
- H2: the reader resolves its owning manga series (chapter detail now
  carries series_id/series_title) and offers next-chapter navigation: a
  header next button and an end-of-book floating CTA at >=99.5% progress;
  back defaults to the series even without a backTo param
- H3: chapter rows show a persistent read check + muted title, and the
  mark-read mutation carries series_id so the series detail cache
  invalidates (read states no longer revert on revisit)
- H4: continue-reading cards for manga chapters present the series:
  sections payload resolves chapter->series linkage, the card heading/image
  link to the series, and meta lines launch the reader
- View Details: manga series menus (card dot menu + detail overflow) open
  a file inspector showing folder paths and per-chapter file names/sizes
  via GET /catalog/items/{id}/manga-files; paths are stripped for viewers
  without file-path visibility (item-versions policy)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(manga): UX mediums - richer detail page, smarter list, manga sort scope

Second batch from the manga UX review (M1-M7):

- M1: multi-chapter volume sections are collapsible (fully read sections
  start collapsed) with sticky headers, and long series get a 'Jump to
  <next unread>' anchor above the list
- M2: the series hero shows the author line (HeroCrewLine learns Author
  credits with person links; DetailHero now renders crewLine and genre
  chips independently) and Volumes/Chapters badges
- M3: browse-card count chip abbreviates to '12 Vol - 3 Ch' so it fits
  narrow cards without occluding covers
- M4: manga gets its own sort scope: Duration/Bitrate (meaningless for
  file-less series rows) disappear, reading labels (Date Read / Reads)
  apply, Author stays
- M5: global search labels manga results 'Manga' instead of the raw type
- M6: chapters carry the viewer's reading fraction; part-read rows show an
  inline progress bar + percent
- M7: chapter rows show the extracted cover thumbnail (presigned
  poster_url on the chapters payload) instead of a generic icon

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(manga): UX lows - volume token dedupe, comic reader chrome, empty-state hint

- buildMangaList buckets volumes by canonical numeric token so mixed
  release naming (v01 + 1) yields one Volume 1 instead of duplicates
- cbz/cbr readers start with the side panel closed and hide prose-only
  chrome (reading ruler, TTS, typography/font controls, hyphenation,
  writing mode) while keeping comic-relevant settings (theme, brightness,
  margin, right-to-left, spread, flow)
- manga empty state mentions chapters appear after the library scan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(manga): publication status badge via new SDK status field

- vendor the unpublished plugin SDK (adds MetadataItem.status) under
  internal/compat/ with a relative go.mod replace, following the
  zishang520-webtransport-go convention; swap to the published module
  before the upstream PR
- map plugin status into MetadataResult.ShowStatus, persist it during
  manga enrichment, and show it as the hero status badge (show_status was
  already on the detail payload and MetadataBadges)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(manga): generalize backdrop pass to secondary fields (backdrop + status)

The backdrop-only claim arm becomes a secondary-fields pass: enriched items
missing a backdrop and/or publication status are claimed, fetched by stored
provider ID, and only the missing secondary fields are written. Lets the
new status field backfill across the already-enriched library instead of
applying only to future enrichments.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(metadata): merge ShowStatus through MergeMetadata/MergeGlobalMetadata

The new MetadataResult.ShowStatus never reached the accumulated result the
manga enricher persists from - the field-by-field merges didn't know it, so
the status backfill pass obtained nothing. Regression-tested on both paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(manga): keep scanner identity IDs out of the metadata flow

filterMangaProviderIDs passed the scanner's manga_series identity row
through, so the search-skip-when-already-matched guard saw provider IDs on
every item and never searched: unmatched items went straight to a by-ID
fetch with no usable ID and were stamped as terminal no-match without a
single provider request (and the MangaDex fallback was never consulted).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore: gitignore docker-compose.override.yml (local deployment override)

The override unpublishes the bundled redis/postgres host ports
(ports: !override []). It is a per-deployment, local-only file: ignoring it
keeps a rebase from main and git clean -fd from disturbing it, and keeps it
out of any PR. Its accidental absence once exposed Redis to the internet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(manga): code-review fixes — no-match guard, sort comparator, volume-count consistency

- enrichWithProviders: set accumulator.HasMetadata after a provider result
  merges (MergeMetadata doesn't propagate it). Without this, a confident
  match carrying only genres/authors/status/year but no cover and no overview
  failed the no-match check and was discarded + terminally stamped.
- byChapterIndex: both un-indexed chapters yield POSITIVE_INFINITY, so the
  subtraction was Infinity-Infinity=NaN (Array.sort treats NaN as 0, leaving
  order undefined). Compare explicitly for a stable order.
- MangaContent volume/chapter badges: derive counts from the rendered
  buildMangaList entries (which canonicalize v01 ≡ 1) instead of raw distinct
  volume tokens, so the badge can no longer say '2 Volumes' over one row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(manga): clarify the enrichment claim's secondary arm is admin-reset-only

The secondary arm (poster present, backdrop/status missing) requires
last_refreshed IS NULL, so it is only reachable when an operator resets
last_refreshed to backfill a newly-added field — not an automatic periodic
re-check (which would re-fetch banner-less series every sweep). Documents the
intent so it does not read as dead code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(manga): collapse continue-reading chapters per series; batch provider-id lookup

- Continue Reading now collapses multiple in-progress chapters of the same
  manga into one card (most recently read kept), mirroring the episode→series
  collapse. The reading section resolves chapter→series linkage into itemMeta
  (applyMangaChapterSeriesMeta) and runs the shared
  collapseContinueWatchingSeriesCandidates, which the reading path previously
  skipped.
- claimBatch resolves provider IDs for the whole batch in one query via the
  new ProviderIDRepository.GetByContentIDs (content_id = ANY), replacing the
  per-item GetByContentID N+1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): manga publication-status chip on browse cards + more legible chips

- Color-coded publication status pill (Ongoing/Completed/Hiatus/Cancelled/
  Upcoming) in the manga card's top-left corner, mirroring the vol/chapter
  count chip top-right. Strictly manga-gated; show_status was already on the
  browse payload.
- New .glass-chip (78% surface vs glass-subtle's 40%) for the manga count +
  status pills so the labels stay legible over busy cover art.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* build(manga): depend on published silo-plugin-sdk v0.7.0

Replace the vendored internal/compat/silo-plugin-sdk copy with a normal
dependency on the published SDK module at v0.7.0, which adds
MetadataItem.status (publication/airing status) consumed by the manga
status badge at internal/metadata/plugin_provider.go.

- go.mod: pin v0.7.0, drop the local-path replace directive
- remove the vendored internal/compat/silo-plugin-sdk tree
- Dockerfile: drop the vendored-SDK COPY
- strip the manga design docs/plans from docs/superpowers (internal)

Requires Silo-Server/silo-plugin-sdk#4 merged and tagged v0.7.0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(manga): exclude chapters from the matcher's unmatched-item lister

Manga chapters are type='ebook' items that stay status='pending' by
design - provider metadata lives on the type='manga' series item. The
scan-final RetryUnmatchedItemsByFolderAndPathPrefix listed all of them
and ran a rate-limited ebook-plugin search per chapter: 31,564 chapters
x ~1s = 8h46m appended to a 2-minute manga library scan (observed
live), every one a guaranteed no-match. Earlier runs never survived to
completion, so the library's last_scanned_at stayed NULL forever.

Add the same manga_chapters NOT EXISTS guard the ebook enricher's
claim query already uses. Verified live: the same library now scans in
27s with retried_items=0.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scanner): never probe-repair ebook/comic files (ebook+manga detail-page killer)

NeedsCriticalProbeRepair was always true for BaseType 'ebook' files (epub, pdf,
cbz, cbr — incl. manga chapters): buildEbookMediaFile leaves ProbeUpdatedAt nil
and they have no audio/video, so probeEnsurer.Ensure spawned ffprobe per file on
every detail/watch load and never converged (ffprobe errors on zip/rar, result
never persisted). Short-circuit probe-repair for ebook base type — they're read
directly and never use the transcode/playback probe pipeline.

SHARED fix: benefits both the ebooks and manga library types.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf+fix(ebooks): parallelize detail extension + preserve finished read-state

- buildEbookExtension ran its 3 related-content queries (series, also-by-author,
  similar) sequentially; run them concurrently like buildAudiobookExtension so
  ebook detail latency is the slowest query, not their sum.
- PGEbookReaderProgressStore.Upsert did an unconditional SET progress=EXCLUDED;
  a routine autosave (e.g. reopening a finished book) could drop it below the
  0.9 finished threshold and silently un-mark it read (and clear the manga
  chapter checkmark, which rides on the same row). Guard: once finished,
  progress only moves on an explicit unread (row delete); below threshold it
  tracks freely.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(manga): batch chapter presign, index volume counts, quiet scan log

- fetchMangaChapters presigned each chapter poster individually; a long-running
  series has hundreds of chapters. Batch them in one PresignImageURLs call, and
  add the missing rows.Err() check (was silently returning partial lists).
- The browse manga count chip's count(DISTINCT volume) subquery wasn't covered
  by manga_chapters_series (series_content_id, chapter_index); add
  idx_manga_chapters_series_volume (series_content_id, volume) so both count
  subqueries are index-only.
- Downgrade the per-chapter "manga scan: indexed" log from Info to Debug (one
  line per .cbz; the 500-file progress log already covers operator visibility).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(manga): address PR #138 code-review findings

Folds PR #142 into the manga branch (already done via fast-forward) and
remediates the issues surfaced in the #138 code review.

Correctness:
- Preserve the scanner's manga_series identity anchor through enrichment.
  ReplaceByContentID's DELETE was unconditional, so the first successful
  enrichment wiped the manga_series provider-id row the scanner relies on
  for idempotency, causing duplicate series + metadata loss on the next
  scan. excludedProviderIDs now also means "not deleted", and the DELETE
  preserves those rows. (internal/catalog/provider_id_repo.go)
- Fall back to the series cover when the latest chapter has no poster.
  Poster columns default to '' (not NULL), so the manga series-card poster
  override blanked cards via a plain COALESCE; wrap operands in NULLIF.
  (internal/sections/fetcher.go)
- Keep backTo a real query param on reader links when libraryId is absent.
  It was string-concatenated with '&', producing a malformed URL on
  deep-links; route it through the query helper instead.
  (web/src/lib/mediaNavigation.ts, EbookReader.tsx, MangaContent.tsx)

Quality:
- Hide manga chapters from favorites/watchlist browse, matching the
  exclusion enforced on every other listing surface.
  (internal/catalog/favorites_browse.go)
- Centralize the manga chapter exclusion predicate into a single exported
  catalog.MangaChapterExclusionWhere, removing four duplicated copies.
  (catalog, sections, ebooks)
- Skip the two manga count subqueries on browse scopes that cannot contain
  manga (non-manga type filters), substituting NULL placeholders.
  (internal/catalog/browse.go)
- Normalize provider publication status (AniList/MangaDex/SDK variants)
  into a stable label set so show_status carries one manga value-domain.
  (internal/manga/enrichment.go)

Adds unit tests for the poster NULLIF contract, browse gating, and status
normalization.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: regenerate go.sum after rebase onto main

Drops stale silo-plugin-sdk v0.6.0 and other leftover hashes from the
intermediate rebased states; go.mod is now on the published v0.7.0 tag.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(scanner): adapt manga scan to ebookFileShouldSkip 3-value signature

main changed ebookFileShouldSkip to also return the existing content ID;
the manga scan path only needs the unchanged flag, so discard the new
return. Resolves a silent semantic conflict from the rebase onto main.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Silo Server Developer <warmasterx555@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 20:13:10 -04:00
e084cdd1d6 Add unified literary works for ebooks and audiobooks (#107)
* docs: add literary works design and plan

* feat(literary): add work link schema

* feat(literary): add work domain primitives

* feat(literary): persist work links

* feat(literary): score work matches

* feat(catalog): include literary work summary on item detail

* feat(literary): expose work detail API

* feat(literary): assemble work detail

* feat(literary): add admin work linking primitives

* feat(catalog): group literary items by work

* feat(literary): auto-link works during book scans

* fix(literary): narrow work match candidates

* fix(literary): address work merge blockers

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-06-16 18:17:08 -04:00
13c5e0ba2f feat(catalog): deterministic cross-server content_id (#155)
* feat(catalog): deterministic cross-server content_id

Replace per-server Sonyflake content_id with a structured natural key
derived from provider IDs (movie:tmdb:…, series:tvdb:…, episode:…,
local:… fallback), so two servers holding the same title share one
anchor for artwork, watch history, progress, favorites and ratings.

- internal/contentid: derivation core, SeriesIDFromContentID transform,
  frozen precedence, SchemeVersion=1, embedded-series-anchor invariant.
- internal/metadata/service.go: deterministic id at every mint site.
- internal/catalog/history_source.go: resolve show via string transform
  for anchored episode ids; skip the episodes_pkey probe.
- migrations/sql/20260612130000: collision-safe value remap across the
  65-column reference graph + COLLATE "C", FK/trigger handling, audit
  map, working down.

Benchmarked against an exact-cardinality copy of cprod-postgres
(1.93M episodes, 775k history rows): 2.57x faster history page, 1.7x
throughput at 100 concurrent users, 2.7x cheaper per content_id probe.

* feat(catalog): re-ID untagged items to deterministic content_id at first match

Untagged libraries get a path-derived local: content_id at scan time and
only learn their provider IDs later, when the match worker confirms a
result. Previously that id was never folded back in, so untagged-then-matched
items kept a per-server local: placeholder forever and never converged across
servers (re-ID was deferred to a migration rerun).

mergeAndPersist now promotes a local: skeleton to its deterministic
provider-anchored id at the moment of first confirmed match, via a single
new gate (canonicalizeLocalContentID):

  - target id already taken  -> merge onto it (existing rebind machinery)
  - target id free           -> rename in place

The rename is a single SQL function (silo_rename_content_id); FK children
follow via ON UPDATE CASCADE added to the content_id family, so a fresh
skeleton moves a handful of rows rather than the full-table remap the bulk
migration does. The guard is one IsLocal prefix check, so tagged content and
all refreshes pay nothing, and the move is self-healing under retry.

Verified: gofmt/vet/build clean; migrate-validate passes; migration applies
on the real schema (up/down/up), FKs gain ON UPDATE CASCADE while keeping
ON DELETE; functional test confirms series PK move + series_id cascade +
provider-id sweep, and movie rename.

Follow-ups (noted in docs): recomposeSeriesChildIDs for a series that
accumulated episodes before matching; a lockstep test for the soft-ref list.

* fix(catalog): harden content_id parsing and merge per review

Address review feedback on the deterministic content_id work:

- history_source.go: gate the anchored-episode display-id transform on the
  full five-part episode shape (split_part parts 2-5 non-empty), not just the
  'episode:' prefix, so a malformed id can't transform to 'series:broken:' and
  vanish at the media_items join. Shared anchoredEpisodePredicate drives both
  the null-poisoned join key and the series-recovery expression.
- contentid.go: unexport the provider-precedence slices so no package can
  mutate the frozen SchemeVersion ordering at runtime.
- contentid.go: add parseAnchored to validate the exact per-kind arity and
  numeric season/episode suffixes; SeriesIDFromContentID and IsProviderAnchored
  now fail closed on truncated/malformed ids (e.g. "episode:tvdb:296762").
- canonicalize.go: distinguish catalog.ErrItemNotFound from transient lookup
  errors (a real error no longer masquerades as "target free"), and allow a
  matched local source to be consolidated onto the canonical row instead of
  orphaning a duplicate.

* refactor(contentid): URL-safe "-" separator in content_id

Use "-" instead of ":" to join content_id components
(movie-tmdb-228064, episode-tvdb-296762-1-5, local-<hex>). "-" is an RFC 3986
unreserved character, so a content_id is URL-safe verbatim: encodeURIComponent
is a no-op and the id is its own tidy path segment (/item/series-tvdb-296762)
with no %3A escaping. The stored value equals the URL value, so there is no
encode/decode boundary and an operator can grep the id straight out of a URL or
log. Every component is [a-z0-9]+ (or "tt"+digits), so "-" is unambiguous.

Pre-release format finalization: this branch is unmerged, so no deployed data
carries ":" ids — the migration mints the "-" form fresh and no re-migration is
needed. Still SchemeVersion 1.

- contentid.go: single `sep` constant drives construction and parsing so the two
  can never drift; all constructors/parsers and doc examples updated.
- history_source.go: split_part transform and the anchored-episode predicate use
  '-'; kept in lockstep with the package via a code comment.
- 20260612130000_deterministic_content_id.sql: derivation and season/episode
  composition emit '-'; LIKE filters match 'series-%'.
- docs/architecture/deterministic-content-id.md: format spec + rationale for the
  separator choice; this is the design doc the change is derived from.

Client-side: the web frontend treats content_id as an opaque string (no
splitting/regex), so no client changes are required; existing
encodeURIComponent call sites simply stop emitting %3A.

* docs(contentid): show why hash/bigint rejected in probe-cost table

Add Cross-server deterministic / Zero-join show transform / Human-readable
columns to the index-probe-cost comparison so the trade-off is legible at a
glance: the 128-bit hash and bigint surrogate are faster but each give up a
load-bearing property, and the structured key is the only all-checkmark row.

* docs(contentid): order probe-cost table to end on the structured key

* docs(contentid): label fenced blocks and drop stray EOF tags

Per CodeRabbit review: add 'text' language to three fenced code blocks
(MD040) and remove accidental </content></invoke> artifacts at EOF.

* fix(catalog): remap array-valued content_id soft references in deterministic id migration

The value-remap migration (20260612130000) enumerates the reference graph by FK
plus a scalar name+type sweep (text/varchar/bpchar). That misses
trending_discover_snapshots.content_ids: it is text[] (excluded by the type
filter), named content_ids not content_id (excluded by the name list), and
cannot carry an FK — so the bulk remap left those arrays holding stale Sonyflake
ids that resolve to nothing until the snapshot regenerates. A counterexample to
the migration's "self-protecting, cannot orphan" invariant.

Remap the array element-wise in both directions (Up old->new, Down new->old),
preserving order and leaving collision/unmatched elements untouched; a WHERE
EXISTS guard skips empty/unaffected arrays so array_agg never collapses the NOT
NULL column to NULL. Mirror the gap in silo_rename_content_id (20260614120000)
with array_replace for the single-value runtime rename so the two stay in
lockstep.

Verified on PG18: mixed/collision/empty arrays remap correctly and round-trip
clean; runtime array_replace preserves order.

Surfaced reviewing #155. The jellycompat restart-decode regression and the
atomicity-wording nit are posted as review comments, not addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(jellycompat): pack content_id into compat UUID reversibly so item ids survive restarts

Addresses the restart-decode regression raised in review of #155. With
content_id now a structured string instead of a numeric Sonyflake,
EncodeStringID sent every item/season id down the one-way SHA1 path, making
decode depend on an in-memory reverse map. That map is cold after a process
restart (the codec is a process-lifetime singleton), so a client presenting a
previously-issued item UUID — resume-from-home, deep link, detail page, image,
userdata — got "unknown compat id" until the item was re-listed.

Make the encoding reversible instead of stateful:

- internal/contentid: add Pack/Unpack, a bit-packed, fixed-budget (<=15 byte)
  binary form of a structured or local content_id. digitCount preserves
  provider-id leading zeros (e.g. imdb tt0944947); structured forms are
  self-delimiting; the local form fills the budget exactly. Provider ids that
  overflow uint64 return ok=false.
- Shrink ForLocal to a 112-bit (sha256(path)[:14]) hash so a local id packs
  losslessly into the 15-byte UUID payload. 112 bits is far beyond any single
  server's local-item count. No other code assumed the old width.
- internal/jellycompat: EncodeStringID packs item/season content_ids into the
  UUID (byte 0 = kind, bytes 1..15 = packed, non-zero tag distinguishes it from
  the numeric encoding); DecodeStringID unpacks first and re-packs to confirm,
  so an opaque id whose bytes merely parse is rejected and falls through to the
  map. Numeric ids and arbitrary names (genres, studios) are unchanged.

Net: item/season ids decode by pure computation — stable across restarts and
across instances — with no lookup table. Only the rare unpackable content_id and
non-content names still use the in-memory map.

TDD: round-trip property tests in contentid (all kinds, leading zeros, reject
cases) and a cross-instance decode test in jellycompat that fails on the old
hash+map path. Full contentid + jellycompat suites green; production code
golangci-clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(migrate): make the migration run timeout configurable (SILO_MIGRATE_TIMEOUT)

The boot-path migration runner hardcoded a 5-minute context timeout. The
deterministic-content-id value-remap (20260612130000) does a full-table COLLATE
rewrite + 65-column remap that needs ~20 min on a real dataset (615k items /
2M episodes), so it was cancelled at 5 min. Worse, Postgres keeps the orphaned
backend running (holding AccessExclusive locks) until it notices the dead client
at a statement boundary, while the goose session advisory lock releases on
disconnect — so each 5-min boot retry piled a new attempt behind the previous
one's locks. The migration never applied; the server boot-looped.

Make the timeout configurable via SILO_MIGRATE_TIMEOUT (a Go duration like
"60m"); 0 or negative disables the deadline for a one-off heavy migration. Default
stays 5m. All three entry points (migrate-status, --migrate-only, boot) honor it.

Required for the deterministic-content-id migration to apply on any real-sized
database, not just dev — the 5m cap made the PR undeployable at scale.

Follow-up (not here): on cancellation the runner should actively terminate its
backend so a future timeout cannot orphan a lock-holding statement.

TDD: MigrationTimeout parsing (default/override/zero/invalid) + MigrationContext
deadline behavior.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(contentid): require exact length for local ids in Unpack

Tighten the tagLocal branch of Unpack from `len(body) < localHashLen` to an
equality check. The local form fills the compat-UUID payload exactly (no
padding), so a body of any other length is non-canonical; matching it exactly
keeps Unpack a strict fail-closed inverse of Pack for the fixed-length branch,
which decodes client-supplied UUIDs.

Not applied to the structured branch (a review suggestion proposed the same
change there): structured ids are self-delimiting and the compat layer pads them
with trailing zeros to fill the 15-byte UUID payload, so ignoring trailing bytes
is intentional and documented. Rejecting them would make every structured id
fail to decode — the jellycompat cross-instance test guards against that.

Not a live bug today (the only caller passes u[1:] from a 16-byte UUID, so body
is always exactly localHashLen, and idcodec re-packs to verify), but it is the
correct contract and zero-risk. Adds a regression test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:34:13 -04:00
5afe56cfc0 feat(jellycompat): add runtime-managed Jellyfin Web compatibility (#77)
* feat(jellycompat): install web assets at runtime

* fix(jellycompat): recover stale web operation locks

* fix(jellycompat): harden web component management

* feat(admin): refine compat settings and restart status

* chore(dev): add hot-reload docker compose stack

* fix(dev): include npm in hot-reload backend

* feat(admin): refine Jellyfin compatibility settings

* feat(settings): improve jellyfin proxy summary

* feat(settings): improve jellyfin web controls

* fix(settings): update jellyfin web removal status

* fix(settings): enable jellyfin web after install

* feat(jellycompat): auto-select web ui version

* test(api): update rate limit handler setup

* feat(jellycompat): refine web ui install onboarding

* fix(jellycompat): address web ui install review issues

* fix(onboarding): mirror jellyfin api runtime status

* fix(admin): remove global restart banner

* fix(settings): gate restart required tracking

* fix(jellyfin): ignore live settings for restart status

* fix(jellyfin): avoid restart for live compat settings

* fix(subtitles): normalize AI language codes

* fix(catalog): support partial title search tokens

* feat(branding): add white-label customization

* Add push relay engineering plan

- Document relay API contracts, APNs/FCM behavior, auth, storage, and ops
- Capture implementation plan, provider references, decisions, and README

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-06-15 09:34:08 -04:00