Commit Graph
6 Commits
Author SHA1 Message Date
1664c60425 fix(metadata): publish artwork revisions atomically (#399)
* fix(metadata): publish artwork revisions atomically

* fix(metadata): harden artwork revision cleanup

* fix(metadata): address artwork revision review findings

- restore image applies for all media_items types and reject unsupported
  target/image combinations with 400 before uploading; episodes coerce to
  stills and the web dialog no longer offers image tabs episodes can't use
- add WHEN clauses to displacement triggers and hoist to_jsonb so bulk
  catalog upserts that assign unchanged artwork columns skip the trigger
- make artworkkey the single variant-ladder owner: imagecache derives its
  widths from it and triggers store image_type instead of hardcoded
  variant arrays, expanded by the collector at deletion time
- sweep dormant registry rows periodically so references lost through
  untriggered surfaces degrade to slow cleanup instead of leaking
- park just-published revisions dormant, keep dormant rows dormant on
  re-cache, and batch the GC reference pre-check per run
- heal rows re-referencing a just-deleted revision via reconciler-style
  resets after the deletion commits
- share a per-URL image-loaded hook across DetailHero, ItemCard,
  SectionItemCard, GlobalSearch, and CollectionPosterCard
- deduplicate Cache/CacheBytes finalization and drop unused VariantPaths
  plumbing

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(catalog): cast reused timestamp parameter in revision upsert

Postgres cannot deduce one type for $3 used both as a plain value and
inside a CASE arm; the dev deploy surfaced it as SQLSTATE 42P08 on every
publication. Cast both uses and cover the arm/park/track upserts with
database-backed tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(metadata): address artwork revision review comments

- keep a durable heal path: deletion marks deleted_at instead of removing
  the registry row, so a failed post-delete heal retries with backoff and
  broken references never park; trackers clear the marker on re-upload
- never treat bare existence as an immutable-content match; backends
  without content verification rewrite the object
- exercise revisioned cover keys in scanner/enrichment fakes, compare the
  tracked manifest exactly, and honor cancellation in the blocking test
  deleter

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 17:33:21 -04:00
203a18ae83 feat(observability): OpenTelemetry logs+traces with secret redaction and slog standardization (#290)
* feat(observability): OpenTelemetry logs+traces with secret redaction

Part of #265. Adds opt-in OpenTelemetry (logs + traces) alongside the existing
stderr + opslog pipeline, plus secret redaction on all sinks. Default-off: with
no OTEL_* / SILO_OTEL_ENABLED config, behavior is unchanged.

Bootstrap (internal/telemetry):
- Setup() builds one shared resource, a TracerProvider (parent-based trace-id
  ratio sampler), a LoggerProvider, and the W3C TraceContext+Baggage propagator
  from env. It installs NO MeterProvider — metrics stay on Prometheus, and the
  built-in no-op global MeterProvider keeps the trace instrumentation libs from
  double-emitting. Shutdown is deferred with a flush timeout.
- Logs are bridged via otelslog fan-out (slog.MultiHandler), level-gated by the
  shared LevelVar and best-effort so a failing collector can't break the console
  or DB branches. stderr + opslog stay untouched.

Secret redaction (internal/logredact):
- A slog.Handler masks secret-keyed attributes (password, token, api_key,
  authorization, cookie, ...) — including .With-bound attrs, nested groups,
  secret-keyed group subtrees, and values behind a LogValuer — on the console
  and OTLP sinks, with a no-op fast path when a record has no secret keys.
  opslog.shouldRedact delegates to logredact.SecretKey so all sinks share one
  marker list.

Rotation is infra-managed (no custom file sink): container runtime for stderr,
collector/backend for OTLP, opslog partition-pruning for the DB. Documented in
docs/architecture/observability.md.

Verification: go build ./..., go vet, gofmt -l — clean; go test
./internal/telemetry/ ./internal/logredact/ -race pass.

AI-use disclosure: implemented with AI assistance (Claude Code), including
adversarial reviews that hardened the bootstrap and fixed two redaction leak
paths; reviewed by the author.

* refactor(observability): slog context+component sweep, sloglint gate (phase 3)

Part of #265. Builds on the OTel bootstrap + redaction commit.

Standardizes every log call site onto the context-carrying slog variants so
records correlate with the active OpenTelemetry trace, and locks the standard
in with a machine gate so future code (human- or AI-authored) can't drift back.

- Call-site sweep: converted the remaining slog.<Level>(...) calls to the
  slog.<Level>Context(ctx, ...) form wherever a context.Context is in scope
  (background/init calls with no ctx are left as-is), across 183 files. Applied
  via a type-aware AST codemod. Log levels and message strings are preserved
  verbatim; a component attr (canonical per-package name) is added to direct
  package-level slog calls. Bound-logger calls keep their existing .With
  bindings. The main.go and telemetry package conversions rode with their file
  in the previous commit to keep each file within a single commit.
- Enforcement (.golangci.yml): enable sloglint with context=scope, static-msg,
  key-naming-case=snake, no-mixed-args. After the sweep all four report zero
  violations repo-wide (tests included), so make lint / CI now blocks any
  regression to the non-context form. The gate ships with the sweep because it
  cannot be green until the legacy sites are converted.

Metrics remain on Prometheus; no behavior change to /metrics or Grafana.

Verification: go build ./..., go vet ./..., gofmt -l — clean; sloglint (all 4
rules) 0 violations repo-wide; log levels verified unchanged.

AI-use disclosure: implemented with AI assistance (Claude Code), including the
codemod; reviewed by the author.

* fix(observability): honor per-signal OTLP protocol and secret WithGroup names

Two Codex review findings on PR #290:

- telemetry: OTEL_EXPORTER_OTLP_{TRACES,LOGS}_PROTOCOL now override the
  generic OTEL_EXPORTER_OTLP_PROTOCOL per signal, so mixed collector
  setups (e.g. HTTP logs + gRPC traces) build the right exporter.
- logredact: entering a group whose name is secret-bearing (e.g.
  WithGroup("authorization")) now masks every leaf in that subtree,
  matching how slog.Group("authorization", ...) is masked as a whole.

* fix(observability): address review feedback on telemetry bootstrap

- Telemetry setup failure no longer kills boot: Setup returns usable
  no-op providers alongside the error and main logs and continues with
  telemetry disabled, honoring the best-effort contract.
- Honor OTEL_TRACES_SAMPLER (always_on/off, traceidratio, parentbased_*
  variants); unsupported values fall back to parentbased_traceidratio.
- Attach node identity as semconv service.instance.id instead of the
  non-semconv node.name.
- Rename opslog retention-scope log attrs to target_component/target_level
  so they no longer collide with the canonical component routing key, and
  tag those lines with component=opslog.
- Fix stale levelGated comment casing; use WarnContext in the telemetry
  shutdown defer; document the LogValuer double-resolve on the redaction
  slow path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 08:53:52 -04:00
430224a1b9 perf: cut home-screen, Continue Watching, and Latest latency; cache shared home rails (#292)
* perf(jellycompat,sections): bound resume scan, batch leaf detail progress, widen section concurrency

Three low-risk fixes from the section-fetch performance investigation
(docs/superpowers/plans/2026-07-03-section-fetch-performance.md):

- jellycompat: bound loadProgressPage at resumeScanMaxRows=300 so a single
  request never pages through more than that many in-progress rows. The cap is
  unconditional: it also covers the sparse-visible-set case (a heavy watcher
  whose recent rows are mostly dismissed/superseded, or a Series/Season-only
  request that matches no leaf in-progress row), where the page never fills and
  the loop would otherwise scan the entire history — previously an O(history)
  scan reaching tens of seconds. In the common case the loop exits far earlier,
  so the cap only bounds the pathological worst case; 300 leaves ample headroom
  to fill a ~20-item Continue Watching page. Beyond the cap the reported total
  is a clamped lower bound. Covered by TestLoadProgressPage_BoundsScanForSparseVisibleSet.
- jellycompat: batch the leaf-item (movie/episode) progress lookup in
  GetItemDetailsByIDs via ListProgressWithCompletedHistory instead of a
  per-item GetProgressWithCompletedHistory (~100 sequential queries for a
  50-item detail page). Series keep the per-item episode-rollup path (they own
  no progress row). Output is unchanged; a batch-lookup failure is now logged
  rather than silently dropping played state for the whole page.
- sections: raise fetchAllMaxConcurrency 4 -> 6 to cut FetchAll wave count for
  large home layouts, staying within the default 20-conn pool.

Part of the home/continue-watching latency work.

AI-use disclosure: implemented with AI (Claude) assistance.

* perf(jellycompat): keep Latest browse on the cross-library fast path under isPlayed

/Items/Latest with isPlayed=false is the highest-frequency compat browse
(~10.8k calls/day). The played overlay can't be pushed into SQL, so browse
over-fetches and filters locally. The cross-library recently_added fast path
(BrowseRecentlyAddedAcrossLibraries: one ~1ms index walk per library) was
gated on Offset==0, so a heavy watcher who had already seen the newest items
needed a 2nd chunk and fell through to BrowsePage — a whole-catalog
MIN(first_seen_at) + GROUP BY HashAggregate over ~147k movies measured at
~755ms per call (0.8-1.6s observed end-to-end).

Fetch the entire over-fetch budget (maxScannedRows) in a single merged
fast-path walk instead of paging into BrowsePage, so the loop fills from one
call. The clamp caveat (MaxLimit=1000 leaves a fall-through only for
requestedLimit>200, off the Latest hot path) is documented inline.

Part of the home/browse latency work.

AI-use disclosure: implemented with AI (Claude) assistance.

* fix(jellycompat): scope resume scan cap to resume path and bound the fast-path loop

Addresses PR #292 review feedback:

- Codex (P2): the resumeScanMaxRows cap was applied unconditionally in the
  general loop, which also paginates the completed (watched-items) list. Gate it
  on resumeFiltered so the completed path keeps exact TotalRecordCount and deep
  StartIndex pagination. Covered by TestLoadProgressPage_CompletedScanNotCapped.
- CodeRabbit (Critical): the earlier raw-offset fast-path loop — the default
  Continue Watching shape and the sections-fallback route — had the same
  unbounded-scan bug and was not covered by the cap (the existing test forces
  EnableTotalRecordCount=true, routing around it). Bound it with the same
  resumeScanMaxRows guard. Covered by
  TestLoadProgressPage_BoundsFastPathScanForSparseVisibleSet.
- CodeRabbit (Minor): tag the doc's fenced example blocks as text to satisfy
  markdownlint MD040.

AI-use disclosure: implemented with AI (Claude) assistance.

* perf(sections): cache shared user-agnostic home rails per access scope

Home-screen rails that are identical for everyone who can see the same
libraries (recently added, recently released, genre, trending on server,
most watched, new to library, critically acclaimed, award winners, format
showcase, seasonal, mood, trending discover, admin-curated lists, and
library collections) were rebuilt from Postgres once per request, per user.
Only the overlay on top of each row (watched flags, play position, presigned
poster URLs) is actually per-user.

Insert a process-global resolved-list cache at the FetchOne choke point in
internal/sections. Each cacheable row is built once per access scope, held
with a 15m TTL, and refreshed in the background 3m before expiry;
singleflight collapses cold-miss stampedes into a single build. The per-user
overlay still runs fresh in buildSectionsResponse, so no profile state is
ever shared. Random and per-user rows (continue watching, next up,
recommendations, hidden gems, forgotten favorites, activity feed, user
collections) bypass the cache.

The access-scope key captures every access boundary the fetch path enforces
-- section identity (type + id + config hash) + item limit + accessible and
disabled libraries + max content rating + excluded media types + name prefix
+ allowed-content-id allowlist -- and nothing per-user, so entries are safely
shared. Empty membership is never cached (avoids freezing a transiently empty
rail); background refreshes are bounded by a timeout.

Scale (analytical, derived from the cache behavior -- not a measured
latency): for the user-agnostic rows, Postgres section-query volume collapses
from O(rows x concurrent requests) to O(rows x distinct access scopes) per
15m refresh window, because most users share a handful of access scopes.
Illustrative -- 40 cacheable rows on a home screen, 1000 concurrent users
falling into ~5 distinct access scopes:

  - before: ~40 x 1000 = ~40,000 section queries per wave of home loads
  - after:  ~40 x 5    = ~200 builds per 15m window (plus one background
            refresh per row per scope), i.e. a warm home load runs zero
            section queries for these rows.

That is a ~99% reduction in shared section-query load at that concurrency;
the win grows with concurrency and shrinks as access-scope diversity rises.

Design/plan doc added under docs/superpowers/plans/.

* perf(jellycompat): serve per-library Latest via the cached recently-added section

A jellyfin-compat per-library /Items/Latest rail is the same user-agnostic
list as the native "recently added" library rail -- both order by
mil.first_seen_at DESC. It was rebuilt on every request through
directContentService.BrowseItems, missing the resolved-list cache entirely.

Route per-library Latest for movies and series libraries through the native
section fetch instead, so it reuses the shared cache. HandleLatest resolves
the library's type once, and for a movies/series library builds a synthetic
SectionRecentlyAdded with the same type + config + limit + access scope the
native rail uses and calls FetchOne; the per-user overlay (favorites,
progress, episode targets, presign) is extracted into buildLatestItemDTOs and
shared by both the native and BrowseItems paths, so no overlay logic is
duplicated. Cached *models.MediaItem values are read-only -- LocalizeItemModels
deep-copies before any presign mutation.

To let the two surfaces share one entry, resolvedListCacheKey no longer
includes the arbitrary section ID: every cacheable section type derives its
membership from type + config + limit + scope, never from its own ID (audited
all 14 cacheable types plus the library-collection path; the sole s.ID read
lives in the non-cacheable user-collection branch). A native recently-added
rail and the compat Latest for the same library + scope now collapse to ONE
cache entry, built once and reused. Access-scope isolation is unchanged --
the removed ID never carried access information, and every access boundary
(libraries, rating cap, excluded types, content allow-list, name prefix)
still keys the entry.

Guardrails: the native path is restricted to movies and series libraries;
every other library type (ebook, music, manga, mixed) is ignored and keeps
its exact BrowseItems behavior -- important because an unfiltered
recently-added fetch would otherwise surface non-video items to Jellyfin
clients that only expect video. Deeper pages, played-filter and
backdrop-required requests, a client asking for a type other than the
library's own, and any FetchOne error also fall back to BrowseItems.

Chosen over an alternative that gave the synthetic section a deterministic ID
(which kept two separate cache entries): both returned identical data with
similar complexity, so the shared-entry design won.

* fix(sections,jellycompat): post-review fixes for the shared-list cache and Latest path

Consolidates fixes from the branch's adversarial review and PR #292 review
comments into one commit:

- Latest fast path: fall back to BrowseItems when a request carries a genre,
  name-prefix, or person filter (the synthetic recently-added section cannot
  express these, so serving it unfiltered would return a wrong, broader set).
  Eligibility is decided by latestFastPathEligible and covered by a test.
- Clamp the /Items/Latest page size to compatBrowseMaxLimit before building the
  section, matching the BrowseItems fallback, so a large client Limit can't drive
  an oversized recently-added fetch or explode the shared cache key with unbounded
  ItemLimit values.
- Evict expired entries from the process-global resolvedListCache: resolvedListSet
  sweeps expired keys at most once per minute, bounding the map to scopes seen
  within one TTL window. Covered by TestResolvedListCacheEvictsExpiredEntries.
- Log a short digest of the cache key (resolvedListLogKey) instead of the raw key
  in the background-refresh panic/error paths, since the key embeds
  user-controlled access-scope fields such as NamePrefix.

Skipped review comments (verified already fixed or stale against current code):
the resume fast-path scan bound and watched-items cap (04d2e795) and the docs
fence-language tags (already addressed).

Build, vet, and go test -race pass for internal/sections and internal/jellycompat.

* perf(plugins): cache plugin installations in-memory, invalidated on lifecycle change

## Problem
Every poster/image on a warm home rail re-read plugin_installations from
Postgres to answer "is this plugin enabled?" and to acquire the plugin client
(Source A: metadata chain buildProviders enabled-check; Source B: ensureClient
-> loadInstallation). Plugin-resolved image URLs are never URL-cached, so the
plugin source and the DB read behind it fired again on every identical warm
request; 100% of images in the target library are plugin-backed.

## Solution
- Guarded in-memory installation cache (map[int]*Installation + RWMutex) in
  plugins.Service. loadInstallation reads through it; the requireEnabled gate
  stays after the cache read so ErrInstallationDisabled semantics are unchanged.
  invalidateInstallationCache clears it and is self-registered as a lifecycle
  hook, so Service.OnLifecycleChange wipes it on install/enable/disable/update/
  uninstall.
- A generation counter closes an invalidate-vs-repopulate race: captured before
  installations.GetByID and re-checked under the write lock, so a row fetched
  before a lifecycle mutation is never written into a freshly cleared cache
  (would otherwise resurrect a just-disabled plugin).
- Route the metadata chain enabled-check through the same cache via a structural
  InstallationEnabledChecker interface (nil-safe: falls back to the pool query
  when no checker is injected), wired in cmd/silo/main.go.

## Post-review fix (auto-update reliability blocker)
AutoUpdateService mutated installations (new InstallPath, old dir deleted) on
the default auto update policy without firing OnLifecycleChange, leaving the
cache stale and breaking plugins with "stored plugin manifest mismatch" until
restart. It now takes an onChange callback wired to Service.OnLifecycleChange
and fires it once per Check run that mutated a row.

## Verification
go build/vet, go test ./internal/plugins/... ./internal/metadata/... (-race).
Tests: cache hit/invalidation, racing-invalidation guard, IsInstallationEnabled,
auto-update fires onChange.

## AI-use disclosure
Implemented with AI assistance (Claude).

* perf(jellycompat): batch per-item presign, and enrich series on the cached Latest path

## Problem
List rails presigned each item's poster/backdrop/logo/still image individually
(~160 singular resolver calls for a 40-item page where 4 batched calls suffice),
and ItemsHandler carried a near-verbatim duplicate of the batch presigner.

## Solution (batching)
Promote the batch presigner to a shared package-level presignCompatListItems
(presign_list.go) with a generic collectImagePaths[T]; convert the per-item
loops (cached home/Latest rail, favorites, batch loaders, userdata favorites) to
one batched PresignImageURLsWithExpiry per image type per page; batch the
season/episode collections; delete the three duplicate presign helpers. URL
output is unchanged (verified byte-for-byte).

## Post-review fix (series Latest data-parity regression)
The native cached Latest fast path built items via compatListItemsFromModels +
buildLatestItemDTOs and never ran the series watch-state rollup, so a series
library's Latest lost Played / UnplayedItemCount and page 1 disagreed with the
BrowseItems fallback. enrichSeriesUserData is promoted to the ContentService
interface and called on the native path (reused, not duplicated).

## Verification
go build/vet, go test ./internal/jellycompat/... ./internal/catalog/...
Tests: bounded presign invocation counts + per-item URL mapping; series rollup
populated on the native Latest path.

## AI-use disclosure
Implemented with AI assistance (Claude).

* perf(sections): gate personalized rails out of the shared cache; widen refresh lead

## Problem
1. The shared home-rail cache whitelisted custom_filter/genre sections by TYPE
   alone, but those route through fetchFiltered -> ParseQueryDefinition and can
   carry personalized (per-profile) rules/sorts (watched, favorited,
   in_watchlist, in_progress, last_watched; sorts progress/date_viewed/plays).
   Their membership is per-profile yet the cache key excludes userID/profileID,
   so a personalized rail built for one profile was served to others in the same
   access scope for up to 15m -- a cross-profile watchlist/watch-state leak.
2. The background-refresh lead was tuned so steady traffic is served a warm
   entry from a longer soft window.

## Solution
- Add QueryDefinition.IsPersonalized() (reusing the existing
  QueryFieldRequiresProfile/QuerySortRequiresProfile helpers).
  isCacheableSectionType now parses the section QueryDefinition and refuses to
  cache custom_filter/genre when personalized; non-personalized definitions stay
  cacheable. Seasonal/mood/trending build their definitions server-side and stay
  unconditionally cacheable.
- resolvedListRefreshLead 3m -> 10m (soft threshold builtAt+5min instead of
  builtAt+12min).

## Verification
go build/vet, go test ./internal/sections/... ./internal/catalog/... (-race).
Test: personalized custom_filter/genre not cacheable; non-personalized are.

## AI-use disclosure
Implemented with AI assistance (Claude).

* fix(sections,metadata): post-review fixes for shared cache and plugin chain staleness

Addresses three review findings on PR #292:

- sections: canonicalize section config JSON before hashing so configs
  differing only in whitespace/field order share a cache entry (native +
  jellycompat rail sharing). Added TestHashSectionConfigCanonicalizes.
- metadata: invalidate the resolved-chain cache on plugin lifecycle
  changes; the installation-enabled check already reads the invalidated
  plugin cache, but resolveChainCached could serve a stale provider chain
  for up to chainCacheTTL after a provider's availability changed.
- jellycompat: move ctx to the first parameter of presignCompatListItems
  for consistency with the other presign helpers.

Skipped the episode-image presign batching nitpick: the resolver already
dedupes+singleflights, so it is a Minor perf-only item not worth the
two-pass refactor risk in this pass.

* fix(sections,jellycompat): harden shared rail cache and Latest fast path per review

Addresses the eight findings from the deep review of this PR:

- Detach the blocking cold-miss rebuild from the singleflight leader's
  request context (context.WithoutCancel + the shared 30s build timeout)
  so one client disconnect no longer fails every collapsed waiter and
  leaves the entry uncached.
- Stop client-controlled values minting unbounded cache entries: the
  compat Latest fast path now always fetches a fixed 100-row budget and
  slices to the requested limit (one entry per scope+library instead of
  one per Limit value), and an unrecognized MaxOfficialRating string
  disqualifies the fast path instead of entering the global cache key.
- Add release_date to the sections item projection/scan so movies served
  via the Latest fast path keep PremiereDate (Jellyfin default-set field)
  in parity with the BrowseItems fallback.
- Fall back to per-item progress lookups when the batched leaf progress
  query fails, restoring one-item-at-a-time degradation instead of
  blanking played state for the whole page.
- Derive cache eligibility from a single source of truth: fetchSection
  and isCacheableSectionType now share the userAgnosticSectionFetcher
  table, whose no-userID/profileID signature makes a fetcher drop out of
  the cacheable set at compile time if it ever gains per-profile inputs.
- Decide Latest fast-path eligibility off the actual browse params the
  fallback would receive, so any filter later added to buildBrowseParams
  automatically disqualifies the cached path; share one
  compatDefaultBrowseLimit constant between both paths.
- Extract AccessFilter.WriteAccessScopeCacheKey as the shared, security-
  critical serializer for all access-scoped caches (resolved-list,
  editorial candidates, audiobook groups); the editorial key now captures
  ExcludedMediaTypes, which its loaders already applied in SQL.
- Strip leaked agent-transcript markup from the section-fetch plan doc.

go build ./..., go vet, gofmt clean; go test -race on
internal/sections, internal/catalog, internal/jellycompat passes
(TestBeginWebOperation* failures are the known pre-existing flakes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 02:01:30 -04:00
CoffeeKnyteandGitHub 4fb8f6a711 fix(metadata): merge duplicate people on refresh instead of looping on 23505 (#250)
* fix(metadata): merge duplicate people on refresh instead of looping on 23505

The background person-refresh worker re-selected the same people every cycle.
When a refresh resolved an external id (tmdb_id/imdb_id) already held by another
people row, the UPDATE violated a partial-unique index and raised SQLSTATE
23505; the tx rolled back so updated_at never advanced and FindCandidates
re-qualified the row forever. The underlying cause is two people rows for the
same human created from credits ingested with disjoint id sets (one tmdb-only,
one imdb-only) that BatchFindOrCreate never reconciled.

PersonRepository.Update now reconciles the collision instead of failing:

- The common no-collision path stays a single plain transaction (no added cost).
- On a 23505, updateResolvingConflicts runs the whole reconciliation in ONE
  transaction, retrying the write via savepoints so it commits atomically: it
  locks both rows FOR UPDATE in id order, then either merges the partner into
  the survivor (repoint item_people skipping duplicate credits, fold the
  partner's ids/fields onto the survivor, delete the partner) or, when the rows
  are not confidently the same human, drops just the conflicting id.
- canMergePeople requires compatible ids AND matching names, so a provider that
  hands the same id to two genuinely different people cannot trigger a
  destructive delete; that case falls to the non-destructive drop.
- A row merged away concurrently surfaces as pgx.ErrNoRows, which the refresh
  service maps to ErrPersonNotFound.

Existing stuck rows self-heal: they are still re-selected each cycle and now
merge (or drop) instead of looping, draining the warning population to zero.

Adds unit tests for the merge-decision logic (guard, compatibility,
field-folding, constraint mapping).

AI-use disclosure: implemented and adversarially reviewed with AI assistance.

* fix(metadata): preserve survivor's existing id when declining a person merge

The non-mergeable branch of resolveExternalIDConflict blanked the
conflicting external-id field before retrying the write. Because
execPersonUpdate is a full-row UPDATE, the retry persisted an empty
string and returned success, silently dropping a previously-valid
provider id (e.g. the admin PATCH path that mutates an existing id into
a colliding value) instead of leaving the row unchanged as the 23505
did.

Restore the locked survivor's currently-persisted value for the field so
the retried write is a no-op on that column: it commits without looping,
without deleting a possibly-distinct person, and without blanking an id
the survivor already held. Writing a row's own current value back can
never violate the unique index, so the field will not re-trigger the
conflict. The refresh-worker path is unchanged (survivor value is empty).

Convert clearExternalIDField into a general setExternalIDField setter and
extend its unit test to cover set-to-value and set-to-empty.
2026-07-01 09:05:58 -04:00
14ffc91dfb [codex] Expand provider image cache queue (#176)
* feat(metadata): expand provider image cache queue

* fix(metadata): harden provider image cache queue

Addresses bug-review feedback from Codex/CodeRabbit on the metadata image
cache pipeline. All findings validated against the code before fixing;
false positives (rows/connection deadlock, PhotoSourcePath merge coupling)
were confirmed non-issues and left unchanged.

- Honor metadata.cache_images for the background processor. The
  cache_metadata_images task was registered whenever S3 was configured,
  so merely enabling object storage downloaded the entire provider-artwork
  catalog even with caching disabled. Add ImageCacheProcessor.SetEnabled,
  gate RunOnce/RunUntilIdle on it, and wire it (with hot reload) from
  cfg.Metadata.CacheImages in main.go.
- Guard terminal job updates with lease ownership. EnqueueBatch can
  repurpose a running row with a new source; MarkSucceeded/MarkFailed
  keyed on id alone let a stale worker finalize the replacement job and
  drop the new artwork. Thread locked_by through and add
  status='running' AND locked_by=$n guards.
- Avoid uploading stale jobs onto the live artwork key. Verify the
  target still references the job's source (CurrentTargetSourcePath)
  before CacheImage, so a job whose source an admin/refresh already
  replaced cannot overwrite the deterministic storage object.
- COALESCE nullable external IDs in EnqueueExistingProviderArtwork. A
  NULL tmdb_id/tvdb_id/imdb_id on any candidate failed the scan and
  aborted the whole cache run; matches the existing item_repo pattern.
- Stop re-downloading the catalog every 30 days. Discovery now skips
  targets whose *_path is already a cached relative path, making the
  cached row the durable dedup marker instead of the prunable job row.
- Decouple catalog sweeps from queue draining. RunOnce no longer runs
  discovery per batch; RunUntilIdle sweeps only when the queue drains and
  throttles full sweeps to every 15m, so idle installs stop full-scanning
  every entity table each minute.
- Requeue claimed-but-unstarted jobs on cancellation. Acquire the
  semaphore before spawning workers and RequeueClaimed any jobs not yet
  started, instead of leaving them locked until the 15m lease expires.
- Skip the backoff sleep after the final upload attempt in
  putObjectWithRetry (saves ~1.5s on permanent failures).
- Add the s3/file/local/upload/generated exclusion to the seasons and
  episodes backfill in migration 20260617184537 for consistency with the
  later migration (the bad backfill was inert downstream, but the
  asymmetry is removed).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 10:07:58 -04:00
Silo Server Migration c085b12fd1 Initial Silo migration 2026-05-22 23:26:56 -04:00