Files
silo-server/internal/metadata/canonicalize.go
T
13c5e0ba2f feat(catalog): deterministic cross-server content_id (#155)
* feat(catalog): deterministic cross-server content_id

Replace per-server Sonyflake content_id with a structured natural key
derived from provider IDs (movie:tmdb:…, series:tvdb:…, episode:…,
local:… fallback), so two servers holding the same title share one
anchor for artwork, watch history, progress, favorites and ratings.

- internal/contentid: derivation core, SeriesIDFromContentID transform,
  frozen precedence, SchemeVersion=1, embedded-series-anchor invariant.
- internal/metadata/service.go: deterministic id at every mint site.
- internal/catalog/history_source.go: resolve show via string transform
  for anchored episode ids; skip the episodes_pkey probe.
- migrations/sql/20260612130000: collision-safe value remap across the
  65-column reference graph + COLLATE "C", FK/trigger handling, audit
  map, working down.

Benchmarked against an exact-cardinality copy of cprod-postgres
(1.93M episodes, 775k history rows): 2.57x faster history page, 1.7x
throughput at 100 concurrent users, 2.7x cheaper per content_id probe.

* feat(catalog): re-ID untagged items to deterministic content_id at first match

Untagged libraries get a path-derived local: content_id at scan time and
only learn their provider IDs later, when the match worker confirms a
result. Previously that id was never folded back in, so untagged-then-matched
items kept a per-server local: placeholder forever and never converged across
servers (re-ID was deferred to a migration rerun).

mergeAndPersist now promotes a local: skeleton to its deterministic
provider-anchored id at the moment of first confirmed match, via a single
new gate (canonicalizeLocalContentID):

  - target id already taken  -> merge onto it (existing rebind machinery)
  - target id free           -> rename in place

The rename is a single SQL function (silo_rename_content_id); FK children
follow via ON UPDATE CASCADE added to the content_id family, so a fresh
skeleton moves a handful of rows rather than the full-table remap the bulk
migration does. The guard is one IsLocal prefix check, so tagged content and
all refreshes pay nothing, and the move is self-healing under retry.

Verified: gofmt/vet/build clean; migrate-validate passes; migration applies
on the real schema (up/down/up), FKs gain ON UPDATE CASCADE while keeping
ON DELETE; functional test confirms series PK move + series_id cascade +
provider-id sweep, and movie rename.

Follow-ups (noted in docs): recomposeSeriesChildIDs for a series that
accumulated episodes before matching; a lockstep test for the soft-ref list.

* fix(catalog): harden content_id parsing and merge per review

Address review feedback on the deterministic content_id work:

- history_source.go: gate the anchored-episode display-id transform on the
  full five-part episode shape (split_part parts 2-5 non-empty), not just the
  'episode:' prefix, so a malformed id can't transform to 'series:broken:' and
  vanish at the media_items join. Shared anchoredEpisodePredicate drives both
  the null-poisoned join key and the series-recovery expression.
- contentid.go: unexport the provider-precedence slices so no package can
  mutate the frozen SchemeVersion ordering at runtime.
- contentid.go: add parseAnchored to validate the exact per-kind arity and
  numeric season/episode suffixes; SeriesIDFromContentID and IsProviderAnchored
  now fail closed on truncated/malformed ids (e.g. "episode:tvdb:296762").
- canonicalize.go: distinguish catalog.ErrItemNotFound from transient lookup
  errors (a real error no longer masquerades as "target free"), and allow a
  matched local source to be consolidated onto the canonical row instead of
  orphaning a duplicate.

* refactor(contentid): URL-safe "-" separator in content_id

Use "-" instead of ":" to join content_id components
(movie-tmdb-228064, episode-tvdb-296762-1-5, local-<hex>). "-" is an RFC 3986
unreserved character, so a content_id is URL-safe verbatim: encodeURIComponent
is a no-op and the id is its own tidy path segment (/item/series-tvdb-296762)
with no %3A escaping. The stored value equals the URL value, so there is no
encode/decode boundary and an operator can grep the id straight out of a URL or
log. Every component is [a-z0-9]+ (or "tt"+digits), so "-" is unambiguous.

Pre-release format finalization: this branch is unmerged, so no deployed data
carries ":" ids — the migration mints the "-" form fresh and no re-migration is
needed. Still SchemeVersion 1.

- contentid.go: single `sep` constant drives construction and parsing so the two
  can never drift; all constructors/parsers and doc examples updated.
- history_source.go: split_part transform and the anchored-episode predicate use
  '-'; kept in lockstep with the package via a code comment.
- 20260612130000_deterministic_content_id.sql: derivation and season/episode
  composition emit '-'; LIKE filters match 'series-%'.
- docs/architecture/deterministic-content-id.md: format spec + rationale for the
  separator choice; this is the design doc the change is derived from.

Client-side: the web frontend treats content_id as an opaque string (no
splitting/regex), so no client changes are required; existing
encodeURIComponent call sites simply stop emitting %3A.

* docs(contentid): show why hash/bigint rejected in probe-cost table

Add Cross-server deterministic / Zero-join show transform / Human-readable
columns to the index-probe-cost comparison so the trade-off is legible at a
glance: the 128-bit hash and bigint surrogate are faster but each give up a
load-bearing property, and the structured key is the only all-checkmark row.

* docs(contentid): order probe-cost table to end on the structured key

* docs(contentid): label fenced blocks and drop stray EOF tags

Per CodeRabbit review: add 'text' language to three fenced code blocks
(MD040) and remove accidental </content></invoke> artifacts at EOF.

* fix(catalog): remap array-valued content_id soft references in deterministic id migration

The value-remap migration (20260612130000) enumerates the reference graph by FK
plus a scalar name+type sweep (text/varchar/bpchar). That misses
trending_discover_snapshots.content_ids: it is text[] (excluded by the type
filter), named content_ids not content_id (excluded by the name list), and
cannot carry an FK — so the bulk remap left those arrays holding stale Sonyflake
ids that resolve to nothing until the snapshot regenerates. A counterexample to
the migration's "self-protecting, cannot orphan" invariant.

Remap the array element-wise in both directions (Up old->new, Down new->old),
preserving order and leaving collision/unmatched elements untouched; a WHERE
EXISTS guard skips empty/unaffected arrays so array_agg never collapses the NOT
NULL column to NULL. Mirror the gap in silo_rename_content_id (20260614120000)
with array_replace for the single-value runtime rename so the two stay in
lockstep.

Verified on PG18: mixed/collision/empty arrays remap correctly and round-trip
clean; runtime array_replace preserves order.

Surfaced reviewing #155. The jellycompat restart-decode regression and the
atomicity-wording nit are posted as review comments, not addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(jellycompat): pack content_id into compat UUID reversibly so item ids survive restarts

Addresses the restart-decode regression raised in review of #155. With
content_id now a structured string instead of a numeric Sonyflake,
EncodeStringID sent every item/season id down the one-way SHA1 path, making
decode depend on an in-memory reverse map. That map is cold after a process
restart (the codec is a process-lifetime singleton), so a client presenting a
previously-issued item UUID — resume-from-home, deep link, detail page, image,
userdata — got "unknown compat id" until the item was re-listed.

Make the encoding reversible instead of stateful:

- internal/contentid: add Pack/Unpack, a bit-packed, fixed-budget (<=15 byte)
  binary form of a structured or local content_id. digitCount preserves
  provider-id leading zeros (e.g. imdb tt0944947); structured forms are
  self-delimiting; the local form fills the budget exactly. Provider ids that
  overflow uint64 return ok=false.
- Shrink ForLocal to a 112-bit (sha256(path)[:14]) hash so a local id packs
  losslessly into the 15-byte UUID payload. 112 bits is far beyond any single
  server's local-item count. No other code assumed the old width.
- internal/jellycompat: EncodeStringID packs item/season content_ids into the
  UUID (byte 0 = kind, bytes 1..15 = packed, non-zero tag distinguishes it from
  the numeric encoding); DecodeStringID unpacks first and re-packs to confirm,
  so an opaque id whose bytes merely parse is rejected and falls through to the
  map. Numeric ids and arbitrary names (genres, studios) are unchanged.

Net: item/season ids decode by pure computation — stable across restarts and
across instances — with no lookup table. Only the rare unpackable content_id and
non-content names still use the in-memory map.

TDD: round-trip property tests in contentid (all kinds, leading zeros, reject
cases) and a cross-instance decode test in jellycompat that fails on the old
hash+map path. Full contentid + jellycompat suites green; production code
golangci-clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(migrate): make the migration run timeout configurable (SILO_MIGRATE_TIMEOUT)

The boot-path migration runner hardcoded a 5-minute context timeout. The
deterministic-content-id value-remap (20260612130000) does a full-table COLLATE
rewrite + 65-column remap that needs ~20 min on a real dataset (615k items /
2M episodes), so it was cancelled at 5 min. Worse, Postgres keeps the orphaned
backend running (holding AccessExclusive locks) until it notices the dead client
at a statement boundary, while the goose session advisory lock releases on
disconnect — so each 5-min boot retry piled a new attempt behind the previous
one's locks. The migration never applied; the server boot-looped.

Make the timeout configurable via SILO_MIGRATE_TIMEOUT (a Go duration like
"60m"); 0 or negative disables the deadline for a one-off heavy migration. Default
stays 5m. All three entry points (migrate-status, --migrate-only, boot) honor it.

Required for the deterministic-content-id migration to apply on any real-sized
database, not just dev — the 5m cap made the PR undeployable at scale.

Follow-up (not here): on cancellation the runner should actively terminate its
backend so a future timeout cannot orphan a lock-holding statement.

TDD: MigrationTimeout parsing (default/override/zero/invalid) + MigrationContext
deadline behavior.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(contentid): require exact length for local ids in Unpack

Tighten the tagLocal branch of Unpack from `len(body) < localHashLen` to an
equality check. The local form fills the compat-UUID payload exactly (no
padding), so a body of any other length is non-canonical; matching it exactly
keeps Unpack a strict fail-closed inverse of Pack for the fixed-length branch,
which decodes client-supplied UUIDs.

Not applied to the structured branch (a review suggestion proposed the same
change there): structured ids are self-delimiting and the compat layer pads them
with trailing zeros to fill the 15-byte UUID payload, so ignoring trailing bytes
is intentional and documented. Rejecting them would make every structured id
fail to decode — the jellycompat cross-instance test guards against that.

Not a live bug today (the only caller passes u[1:] from a 16-byte UUID, so body
is always exactly localHashLen, and idcodec re-packs to verify), but it is the
correct contract and zero-risk. Adds a regression test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:34:13 -04:00

107 lines
4.2 KiB
Go

package metadata
import (
"context"
"errors"
"fmt"
"github.com/Silo-Server/silo-server/internal/catalog"
"github.com/Silo-Server/silo-server/internal/contentid"
)
// providerIDsStruct adapts the denormalized provider-id map carried on a
// MetadataResult to the contentid.ProviderIDs shape used for id derivation.
func providerIDsStruct(m map[string]string) contentid.ProviderIDs {
return contentid.ProviderIDs{
Tmdb: m["tmdb"],
Imdb: m["imdb"],
Tvdb: m["tvdb"],
}
}
// canonicalizeLocalContentID promotes a local skeleton to its deterministic,
// provider-anchored content_id once a confirmed match has supplied provider IDs
// (the untagged-then-matched re-ID). It returns the canonical id, which equals
// from when there is nothing to do.
//
// Tagged content and every refresh hit only the IsLocal guard, so the common
// path is free. When there is work to do, the promotion is one of:
//
// - target already taken: the two are the same logical item, so merge this
// skeleton onto the existing row using the shared rebind machinery; or
// - target free: rename in place. FK children move via ON UPDATE CASCADE (see
// migration 20260614120000_content_id_online_reid), so for a fresh skeleton
// this is a handful of rows, not the full-table remap the bulk migration does.
//
// It must run with the provider-dedup lock held (mergeAndPersist holds it), so a
// concurrent match of the same title cannot claim the target underneath us. The
// rename is self-healing: if it loses a rare race with skeleton creation and the
// unique constraint fires, the match returns an error, retries, and takes the
// merge branch on the next pass.
//
// Note: a series that already had season/episode rows before it matched keeps
// those children on their Sonyflake ids (ForSeason/ForEpisode need a series
// anchor). At first match the children usually do not exist yet, so this is
// rare; re-deriving them is a deferred follow-up (recomposeSeriesChildIDs).
func (s *MetadataService) canonicalizeLocalContentID(
ctx context.Context,
from string,
ids contentid.ProviderIDs,
itemType string,
) (string, error) {
if s == nil || !contentid.IsLocal(from) {
return from, nil
}
// Derive with no path fallback: we only want a provider-anchored id here, not
// another local value.
target, err := deriveLogicalContentID(itemType, ids, "")
if err != nil {
return "", err
}
if !contentid.IsProviderAnchored(target) || target == from {
return from, nil
}
// Look up the target, distinguishing "free" (not-found) from a transient
// failure. Treating a real error as "target free" would fall through to the
// rename path and surface as a misleading unique-constraint conflict instead
// of a retryable error.
existing, err := s.itemRepo.GetByID(ctx, target)
switch {
case err == nil && existing != nil:
// A row already at the target id is the same logical item; merge onto it.
// allowMatchedSource: this also runs when refreshing an already-matched
// local item, so the source row may be 'matched'; we still consolidate it
// onto the canonical target rather than orphaning a duplicate. Safe under
// the provider-dedup lock the caller holds.
if err := s.rebindItemToExistingItem(ctx, from, target, true); err != nil {
return "", fmt.Errorf("merging local item %s into %s: %w", from, target, err)
}
return target, nil
case err != nil && !errors.Is(err, catalog.ErrItemNotFound):
return "", fmt.Errorf("looking up canonical target %s: %w", target, err)
}
// Target free: pure value-move.
if err := s.renameContentID(ctx, from, target); err != nil {
return "", err
}
return target, nil
}
// renameContentID moves a single content_id value (a media item, season or
// episode) to a new, currently-free id. FK children follow via ON UPDATE
// CASCADE; the silo_rename_content_id function also sweeps the unconstrained
// soft references. The function body runs as one statement, so the move is
// atomic.
func (s *MetadataService) renameContentID(ctx context.Context, from, to string) error {
if s.dbPool == nil {
return fmt.Errorf("rename content id requires database pool")
}
if _, err := s.dbPool.Exec(ctx, `SELECT silo_rename_content_id($1, $2)`, from, to); err != nil {
return fmt.Errorf("rename content_id %s -> %s: %w", from, to, err)
}
return nil
}