Commit Graph
6 Commits
Author SHA1 Message Date
13c5e0ba2f feat(catalog): deterministic cross-server content_id (#155)
* feat(catalog): deterministic cross-server content_id

Replace per-server Sonyflake content_id with a structured natural key
derived from provider IDs (movie:tmdb:…, series:tvdb:…, episode:…,
local:… fallback), so two servers holding the same title share one
anchor for artwork, watch history, progress, favorites and ratings.

- internal/contentid: derivation core, SeriesIDFromContentID transform,
  frozen precedence, SchemeVersion=1, embedded-series-anchor invariant.
- internal/metadata/service.go: deterministic id at every mint site.
- internal/catalog/history_source.go: resolve show via string transform
  for anchored episode ids; skip the episodes_pkey probe.
- migrations/sql/20260612130000: collision-safe value remap across the
  65-column reference graph + COLLATE "C", FK/trigger handling, audit
  map, working down.

Benchmarked against an exact-cardinality copy of cprod-postgres
(1.93M episodes, 775k history rows): 2.57x faster history page, 1.7x
throughput at 100 concurrent users, 2.7x cheaper per content_id probe.

* feat(catalog): re-ID untagged items to deterministic content_id at first match

Untagged libraries get a path-derived local: content_id at scan time and
only learn their provider IDs later, when the match worker confirms a
result. Previously that id was never folded back in, so untagged-then-matched
items kept a per-server local: placeholder forever and never converged across
servers (re-ID was deferred to a migration rerun).

mergeAndPersist now promotes a local: skeleton to its deterministic
provider-anchored id at the moment of first confirmed match, via a single
new gate (canonicalizeLocalContentID):

  - target id already taken  -> merge onto it (existing rebind machinery)
  - target id free           -> rename in place

The rename is a single SQL function (silo_rename_content_id); FK children
follow via ON UPDATE CASCADE added to the content_id family, so a fresh
skeleton moves a handful of rows rather than the full-table remap the bulk
migration does. The guard is one IsLocal prefix check, so tagged content and
all refreshes pay nothing, and the move is self-healing under retry.

Verified: gofmt/vet/build clean; migrate-validate passes; migration applies
on the real schema (up/down/up), FKs gain ON UPDATE CASCADE while keeping
ON DELETE; functional test confirms series PK move + series_id cascade +
provider-id sweep, and movie rename.

Follow-ups (noted in docs): recomposeSeriesChildIDs for a series that
accumulated episodes before matching; a lockstep test for the soft-ref list.

* fix(catalog): harden content_id parsing and merge per review

Address review feedback on the deterministic content_id work:

- history_source.go: gate the anchored-episode display-id transform on the
  full five-part episode shape (split_part parts 2-5 non-empty), not just the
  'episode:' prefix, so a malformed id can't transform to 'series:broken:' and
  vanish at the media_items join. Shared anchoredEpisodePredicate drives both
  the null-poisoned join key and the series-recovery expression.
- contentid.go: unexport the provider-precedence slices so no package can
  mutate the frozen SchemeVersion ordering at runtime.
- contentid.go: add parseAnchored to validate the exact per-kind arity and
  numeric season/episode suffixes; SeriesIDFromContentID and IsProviderAnchored
  now fail closed on truncated/malformed ids (e.g. "episode:tvdb:296762").
- canonicalize.go: distinguish catalog.ErrItemNotFound from transient lookup
  errors (a real error no longer masquerades as "target free"), and allow a
  matched local source to be consolidated onto the canonical row instead of
  orphaning a duplicate.

* refactor(contentid): URL-safe "-" separator in content_id

Use "-" instead of ":" to join content_id components
(movie-tmdb-228064, episode-tvdb-296762-1-5, local-<hex>). "-" is an RFC 3986
unreserved character, so a content_id is URL-safe verbatim: encodeURIComponent
is a no-op and the id is its own tidy path segment (/item/series-tvdb-296762)
with no %3A escaping. The stored value equals the URL value, so there is no
encode/decode boundary and an operator can grep the id straight out of a URL or
log. Every component is [a-z0-9]+ (or "tt"+digits), so "-" is unambiguous.

Pre-release format finalization: this branch is unmerged, so no deployed data
carries ":" ids — the migration mints the "-" form fresh and no re-migration is
needed. Still SchemeVersion 1.

- contentid.go: single `sep` constant drives construction and parsing so the two
  can never drift; all constructors/parsers and doc examples updated.
- history_source.go: split_part transform and the anchored-episode predicate use
  '-'; kept in lockstep with the package via a code comment.
- 20260612130000_deterministic_content_id.sql: derivation and season/episode
  composition emit '-'; LIKE filters match 'series-%'.
- docs/architecture/deterministic-content-id.md: format spec + rationale for the
  separator choice; this is the design doc the change is derived from.

Client-side: the web frontend treats content_id as an opaque string (no
splitting/regex), so no client changes are required; existing
encodeURIComponent call sites simply stop emitting %3A.

* docs(contentid): show why hash/bigint rejected in probe-cost table

Add Cross-server deterministic / Zero-join show transform / Human-readable
columns to the index-probe-cost comparison so the trade-off is legible at a
glance: the 128-bit hash and bigint surrogate are faster but each give up a
load-bearing property, and the structured key is the only all-checkmark row.

* docs(contentid): order probe-cost table to end on the structured key

* docs(contentid): label fenced blocks and drop stray EOF tags

Per CodeRabbit review: add 'text' language to three fenced code blocks
(MD040) and remove accidental </content></invoke> artifacts at EOF.

* fix(catalog): remap array-valued content_id soft references in deterministic id migration

The value-remap migration (20260612130000) enumerates the reference graph by FK
plus a scalar name+type sweep (text/varchar/bpchar). That misses
trending_discover_snapshots.content_ids: it is text[] (excluded by the type
filter), named content_ids not content_id (excluded by the name list), and
cannot carry an FK — so the bulk remap left those arrays holding stale Sonyflake
ids that resolve to nothing until the snapshot regenerates. A counterexample to
the migration's "self-protecting, cannot orphan" invariant.

Remap the array element-wise in both directions (Up old->new, Down new->old),
preserving order and leaving collision/unmatched elements untouched; a WHERE
EXISTS guard skips empty/unaffected arrays so array_agg never collapses the NOT
NULL column to NULL. Mirror the gap in silo_rename_content_id (20260614120000)
with array_replace for the single-value runtime rename so the two stay in
lockstep.

Verified on PG18: mixed/collision/empty arrays remap correctly and round-trip
clean; runtime array_replace preserves order.

Surfaced reviewing #155. The jellycompat restart-decode regression and the
atomicity-wording nit are posted as review comments, not addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(jellycompat): pack content_id into compat UUID reversibly so item ids survive restarts

Addresses the restart-decode regression raised in review of #155. With
content_id now a structured string instead of a numeric Sonyflake,
EncodeStringID sent every item/season id down the one-way SHA1 path, making
decode depend on an in-memory reverse map. That map is cold after a process
restart (the codec is a process-lifetime singleton), so a client presenting a
previously-issued item UUID — resume-from-home, deep link, detail page, image,
userdata — got "unknown compat id" until the item was re-listed.

Make the encoding reversible instead of stateful:

- internal/contentid: add Pack/Unpack, a bit-packed, fixed-budget (<=15 byte)
  binary form of a structured or local content_id. digitCount preserves
  provider-id leading zeros (e.g. imdb tt0944947); structured forms are
  self-delimiting; the local form fills the budget exactly. Provider ids that
  overflow uint64 return ok=false.
- Shrink ForLocal to a 112-bit (sha256(path)[:14]) hash so a local id packs
  losslessly into the 15-byte UUID payload. 112 bits is far beyond any single
  server's local-item count. No other code assumed the old width.
- internal/jellycompat: EncodeStringID packs item/season content_ids into the
  UUID (byte 0 = kind, bytes 1..15 = packed, non-zero tag distinguishes it from
  the numeric encoding); DecodeStringID unpacks first and re-packs to confirm,
  so an opaque id whose bytes merely parse is rejected and falls through to the
  map. Numeric ids and arbitrary names (genres, studios) are unchanged.

Net: item/season ids decode by pure computation — stable across restarts and
across instances — with no lookup table. Only the rare unpackable content_id and
non-content names still use the in-memory map.

TDD: round-trip property tests in contentid (all kinds, leading zeros, reject
cases) and a cross-instance decode test in jellycompat that fails on the old
hash+map path. Full contentid + jellycompat suites green; production code
golangci-clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(migrate): make the migration run timeout configurable (SILO_MIGRATE_TIMEOUT)

The boot-path migration runner hardcoded a 5-minute context timeout. The
deterministic-content-id value-remap (20260612130000) does a full-table COLLATE
rewrite + 65-column remap that needs ~20 min on a real dataset (615k items /
2M episodes), so it was cancelled at 5 min. Worse, Postgres keeps the orphaned
backend running (holding AccessExclusive locks) until it notices the dead client
at a statement boundary, while the goose session advisory lock releases on
disconnect — so each 5-min boot retry piled a new attempt behind the previous
one's locks. The migration never applied; the server boot-looped.

Make the timeout configurable via SILO_MIGRATE_TIMEOUT (a Go duration like
"60m"); 0 or negative disables the deadline for a one-off heavy migration. Default
stays 5m. All three entry points (migrate-status, --migrate-only, boot) honor it.

Required for the deterministic-content-id migration to apply on any real-sized
database, not just dev — the 5m cap made the PR undeployable at scale.

Follow-up (not here): on cancellation the runner should actively terminate its
backend so a future timeout cannot orphan a lock-holding statement.

TDD: MigrationTimeout parsing (default/override/zero/invalid) + MigrationContext
deadline behavior.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(contentid): require exact length for local ids in Unpack

Tighten the tagLocal branch of Unpack from `len(body) < localHashLen` to an
equality check. The local form fills the compat-UUID payload exactly (no
padding), so a body of any other length is non-canonical; matching it exactly
keeps Unpack a strict fail-closed inverse of Pack for the fixed-length branch,
which decodes client-supplied UUIDs.

Not applied to the structured branch (a review suggestion proposed the same
change there): structured ids are self-delimiting and the compat layer pads them
with trailing zeros to fill the 15-byte UUID payload, so ignoring trailing bytes
is intentional and documented. Rejecting them would make every structured id
fail to decode — the jellycompat cross-instance test guards against that.

Not a live bug today (the only caller passes u[1:] from a 16-byte UUID, so body
is always exactly localHashLen, and idcodec re-packs to verify), but it is the
correct contract and zero-risk. Adds a regression test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:34:13 -04:00
6f9427d613 docs(process): v1 scope-lock process — proposal template, CODEOWNERS gate, agent instructions (#145)
Implements the process layer of the v1 feature-lock planner: capability
proposals arrive uniform via issue form; the lock artifact
(docs/architecture/v1-scope.md) is CODEOWNERS-gated; the shared
CLAUDE.md/AGENTS.md guidelines gain the scope gate, additive-only API
rules, and pre-push checklist for agent-driven contributions.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-12 18:43:23 -04:00
QuickandClaude Fable 5 a24279e3e3 docs(notifications): import notification design docs, add web push + v1.5 specs
Imports the notification system design folder (architecture overview,
release-events/inbox foundation, APNs/FCM relay specs, outbound webhooks)
and adds the Web Push spec (05, implemented in this branch), the shared
outbound-email architecture note, and the v1.5 roadmap (06) covering the
remaining work after APNs/FCM were deferred to v2.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-11 14:55:46 -04:00
0bd4f8cb3b fix(jellycompat): restore CanDownload with a real Download route for Infuse (#123)
dd81a7ef set CanDownload=false to stop Wholphin's screensaver from
404ing on the nonexistent /Items/{id}/Download route — but the flag is
load-bearing for Infuse, which refuses Direct Play (Static=true
streaming) of items it believes it cannot download. With omitempty the
field vanished from the JSON entirely and Infuse playback broke, while
PlaybackInfo-negotiating clients were unaffected.

Resolve the underlying inconsistency instead of trading one client for
the other: implement GET/HEAD /Items/{id}/Download serving the original
file (range support, Content-Disposition, optional mediaSourceId for
multi-version items) under stream-group auth, and restore
CanDownload=true now that the route exists. Fixes Infuse playback and
keeps Wholphin's download callers working.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 21:09:47 -04:00
c5f21cb10d fix(jellycompat): parse repeated Fields query params (#110)
* fix(jellycompat): parse repeated Fields query params

parseItemsQuery read the Fields parameter via q.Get("Fields"), which
returns only the first value when a client sends Fields as repeated
query params (Fields=A&Fields=B&...) instead of comma-separated in a
single param (Fields=A,B,C).

The jellyfin-sdk-kotlin (used by Wholphin) sends repeated params. When
such a request listed a detail-only field like MediaSources after other
fields — e.g. the episode-playlist request
  /Shows/{id}/Episodes?Fields=PrimaryImageAspectRatio&...&Fields=MediaSources&...
silo saw only the first value (PrimaryImageAspectRatio), so
needsDetailFields stayed false, the request took the list path, and the
response came back without MediaSources. Clients then could not start
playback of the returned episodes ("no media sources").

Join all repeated Fields values before splitting on commas so field
order and delimiter style no longer matter. Comma-separated single-param
clients (e.g. VidHub) are unaffected.

* fix(jellycompat): stop advertising CanDownload and stub ThemeSongs

Wholphin (jellyfin-sdk-kotlin) audit surfaced two reachable gaps:

- mapping.go set CanDownload=true on every playable item while no
  /Items/{id}/Download route exists, sending clients that honor the flag
  (e.g. Wholphin's screensaver/slideshow) into 404s. Advertise false until
  a download route exists.
- GET /Items/{id}/ThemeSongs 404'd, so enabling theme songs in Wholphin
  silently failed on every detail page. Stub it with an empty
  ThemeMediaResult. This cannot reuse the generic item stub: the SDK
  models OwnerId as non-nullable, so the response must include it even
  when empty.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(architecture): add Wholphin endpoint coverage audit

Cross-references every Jellyfin endpoint the Wholphin client can call
against the routes jellycompat serves, with gating evidence for each
missing-but-unreachable endpoint and prioritized recommendations.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 20:42:46 -04:00
9e29e7b330 feat(security): encrypt server-owned credentials at rest (#45) (#95)
* feat(security): encrypt server-owned credentials at rest

Introduce AES-256-GCM at-rest encryption (HKDF-derived from a required
SECRET_KEY) for server-owned credentials, with row-bound AAD, a versioned
enc:v1: envelope, and an idempotent startup backfill.

- internal/secret: cipher + RowAAD/SettingsAAD + the startup backfill engine.
- SECRET_KEY required at bootstrap; cipher threaded as an explicit dependency.
- server_settings: EncryptedSettingsRepo decorator over the audited
  SensitiveSettingKeys (also drives admin redaction); the config watcher and
  watch-sync settings reads decrypt too.
- Arr keys inline-encrypted; the ambiguous SecretResolver indirection removed
  from requests/autoscan.
- Per-table columns encrypted: subtitles, watch-sync, webhook-sync (not
  webhook_secret), history-import, and the jellycompat session's bridged Silo
  access/refresh tokens.
- Startup backfill (resolve-then-encrypt for arr refs) is best-effort and
  primary-node gated.

Equality-looked-up secrets and plugin_runtime_configs.config_value are out of
scope (need hashing / cross-repo design) — see
docs/architecture/secret-encryption.md.

Refs #45

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(compose): require SECRET_KEY in docker-compose

The server now fatals without SECRET_KEY, so the integrated service (and the
commented distributed proxy/transcode examples) pass it through with a
fail-fast guard matching the existing MEDIA_ROOT pattern. Distributed worker
nodes must use the SAME key as the primary to decrypt shared data.
Generate with: openssl rand -base64 48.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(security): encrypt history import session credentials

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 15:25:48 -04:00