* docs: define ebook architecture matching audiobooks * docs: plan ebook audiobook-parity implementation * feat: add ebook scanner parser foundation * fix: harden ebook scanner foundation * fix: handle ebook isbn labels * fix: guard ebook subtree scans * feat: scan ebook libraries in core * fix: preserve ebook scan people credits * fix: refresh ebook scan metadata safely * feat: persist ebook series membership * test: cover ebook series persistence decisions * fix: address ebook scanner PR review * docs: clarify ebook foundation PR scope * feat: add ebook metadata enricher * fix: harden ebook poster cache * feat: wire ebook metadata sync task * feat: expose ebook library metadata setup * feat: add ebook catalog scope support * feat: add ebook detail view * feat: label ebook file versions by format * feat: use file-size copy for downloads * feat: use file language in download dialog * test: cover ebook detail authors and downloads * fix: drop narrator credits from ebook scanner merges * fix: align ebook collection filters with book media * fix: drop asin provider ids from ebook enrichment * fix: force ebook people refresh for stale narrators * chore: omit ebook planning docs from branch * feat: add ebook detail related content * feat: add ebook reader file entrypoint * feat: render ebooks with foliate reader * feat: persist ebook reader progress * feat: add ebook reader controls * feat: extract ebook pdf metadata * feat: favor scanner isbn during ebook enrichment * feat: extract fbz ebook metadata * feat: count cbz ebook pages * feat: show ebook file page counts * feat: show ebook download summaries * feat: switch ebook reader files * feat: prefer epub for ebook read action * feat: surface ebook reader progress * feat: sync ebook reader progress cache * feat: hide ebook read action for unsupported files * feat: filter ebook reader file selector * fix: serve fbz ebook archives with reader mime type * fix: detect fbz ebooks from compound filename * fix: authorize fbz ebooks from compound filename * fix: scope ebook catalog facets * fix: reject narrator queries for ebooks * fix: build ebook recommendation text from authors * fix: include ebooks in embedding eligibility * fix: include ebooks in recommendation media mix * fix: include ebooks in recently added recommendations * feat: include ebook progress in recommendation signals * feat: include ebooks in continue watching sections * feat: include ebooks in catalog progress metrics * fix: read ebook isbn from epub metadata * fix: filter ebook asin provider aliases * fix: fall back from unsupported ebook reader files * fix: sort ebook catalogs by reader progress * fix: filter ebook catalogs by reader progress * fix: include ebooks in last watched catalog filters * feat: reflect ebook reader progress in item user state * feat: share ebook progress state across item surfaces * feat: report ebook scan progress * fix: include ebook activity in recommendations * fix: expose ebook reader progress on item detail * fix: support ebook subtree scans * fix: honor profile header for ebook item progress * fix: add ebook library default sections * fix: route ebook continue cards to reader * fix: hide watched toggle for ebooks * fix: route ebook watch tonight cards to reader * fix: route ebook hero actions to reader * fix: detect archive ebook reader formats by filename * feat: cache embedded ebook covers during scan * fix: encode ebook hero reader links * fix: persist non-epub ebook reader progress * fix: scope narrator catalog badges to audiobooks * fix: merge ebook reader progress during item repair * fix: label ebook progress filters as read * fix: show ebook related rails as book covers * fix: remove txt ebook reader support * fix: reject txt ebook reader files * fix: label ebook advanced filters as read * fix: label ebook personalized sorts as read * fix: remove plain text reader loader path * test: cover ebook unread catalog rules * fix: preserve ebook reader library context * fix: link ebook genres with library scope * fix: encode related rail item links * fix: encode catalog card item links * fix: encode hero and continue item links * fix: encode watch tonight item links * fix: encode recommendation and search item links * test: cover ebook scan format set * fix: label ebook search results clearly * fix: make global search prompt media neutral * fix: encode catalog read API ids * fix: encode item API ids * fix: include ebook reader vendor in docker build * fix: make ebook reader build clean * fix: clean ebook embedded descriptions * docs: plan ebook reader shell parity * feat: add ebook reader shell controls * fix: widen ebook scrolled reader flow * fix: remove scrolled reader content width cap * docs: plan ebook reader full parity * feat: persist ebook reader config * feat: add ebook annotations and bookmarks * feat: add ebook reader tools and aids * feat: add ebook advanced reader settings * fix: keep ebook reader panel in viewport * fix: use foliate sizing units for ebook scroll flow * fix: keep ebook settings controls readable * fix: simplify ebook reader settings controls * feat(ebooks): extract local covers during scan (#98) * feat(ebooks): extract local covers during scan * fix(ebooks): read nullable poster paths during cover scan * fix(catalog): coalesce nullable media artwork fields * fix(ebooks): group sibling formats by book identity * fix(ebooks): tolerate legacy ebook metadata encodings * fix(ebooks): decode PDF hex metadata strings * fix(ebooks): harden local cover extraction and format grouping Address review findings on the local cover scan: - Restrict generic sidecar covers (cover.jpg, folder.png, ...) to single-book directories, always accept images named after the book file, and apply exactly one cover per reconcile with sidecar taking precedence over the embedded cover. - Replace the read-then-write poster update with an atomic conditional UPDATE (ItemRepository.SetLocalPoster) so provider/admin artwork is never clobbered by concurrent writers, and refresh locally owned posters when the extracted cover bytes change (thumbhash compare). - Preserve UTF-8 PDF Info strings (including a UTF-8 BOM) instead of forcing everything through Windows-1252; the cp1252 fallback now only applies to non-UTF-8 bytes. - Select EPUB covers by manifest media-type with properties="cover-image" outranking the EPUB2 meta name="cover" id, so XHTML cover pages no longer shadow the real image. - Order CBZ pages naturally (2.jpg before 10.jpg, ch2/ before ch10/) when picking the cover page, via a single O(n) min-scan. - Bump the ebook content group key scheme to version 2 and reprocess rows written under older versions so pre-existing libraries gain sibling-format grouping instead of accumulating duplicates. - Group different formats only (a same-format sibling with colliding sparse metadata stays a separate item) and stop a joining sibling's embedded metadata from overwriting a provider-matched item. - Decode any IANA-labelled OPF/FB2 XML charset (windows-1251, koi8-r, shift_jis, ...) via x/net/html/charset, and wire the charset reader into FB2 parsing which previously had none. - Strip the full .fb2.zip double extension from filename-derived titles and group keys. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * feat(ebooks): add reader profiles and ruler (#99) * feat(ebooks): extract local covers during scan * fix(ebooks): read nullable poster paths during cover scan * fix(catalog): coalesce nullable media artwork fields * fix(ebooks): group sibling formats by book identity * fix(ebooks): tolerate legacy ebook metadata encodings * fix(ebooks): decode PDF hex metadata strings * feat(ebooks): add reader profiles and ruler * fix(ebooks): address reader ruler and profile review findings - skip renderer setStyles/render when computed styles and attributes are unchanged, so ruler position updates no longer re-style the book view - drag the ruler via a local draft that commits on release, with the surface rect cached at pointer-down - migrate font values persisted before the generic stacks (Inter, Georgia, Merriweather, legacy serif) so the font select never renders blank, with a Custom fallback option for unknown values - make the ruler band click-through and move dragging to a dedicated keyboard-accessible slider handle so links and text selection keep working under the band - share font stacks between options and profiles via READER_FONT_STACKS - surface the active reading profile, move presets to the top of the settings panel, and drop the redundant profile button aria-labels Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ebooks): resolve prefer-const lint error in readest document lib `pnpm run lint` failed on the branch because `direction` is never reassigned in getDirection; split the destructure so only the reassigned `writingMode` stays mutable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Merge branch 'main' into work/ebooks-reader-base Brings the ebook integration branch up to date with main (audiobook library redesign, continue-watching rework and card affordances, quic-go bump, jellycompat fixes). Conflict resolutions favor main's generalized mechanisms and register ebooks with them: - media scope validation goes through IsValidMediaScope (now including "ebook" alongside main's "video" group scope), in Go and in the web filter/search types - continue-watching uses main's typed rails; reading-type sections pull resume points from ebook_reader_progress and the ebook library default section is wired to ContinueTypeConfig(ContinueTypeReading) - item_repo keeps main's derived select-list machinery (itemColumnExpr) and both poster accessors (GetPoster/SetLocalPoster for ebook covers, GetPosterPath for audiobook covers) - web cards/hero/watch-tonight adopt main's buildMediaPlayHref helpers, which now route ebooks to /reader/ebook and encode content ids; ebook affordances (BookOpen icon, Read verb, percent-read subtitle) carry over onto main's reworked components - LibraryForm ebook support ported into main's refactored useLibraryForm/libraryTypes modules Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(docker): copy foliate-js vendor into Dockerfile.dev frontend stage foliate-js is a file:vendor/foliate-js dependency, so pnpm install needs the vendor directory before the lockfile install layer. The production Dockerfile already copies it; the dev image was missed, breaking make dev-deploy with ENOENT on /app/web/vendor/foliate-js. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ebooks): render Continue Reading sections as upright poster cards All-ebook continue sections previously fell through to the horizontal 16:9 wide card; include ebooks in the poster-variant check so book covers render in their natural 2:3 framing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ui): stop related-rail highlight ring clipping on detail pages Move the current-item ring onto the cover artwork with a themed ring-offset color (matching the sidebar profile highlight) and give the scroll container top headroom so the ring is not cut off by overflow-x-auto. Applies to both ebook and audiobook detail rails. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scanner): harden ebook scanning against data loss and bad metadata - Reconcile missing ebook files like video/audio, with real per-root walk failure tracking (failed/unmounted roots are excluded from deletion), symlinked-root support via the shared logical walker, and the empty-root cleanup allowance before any destructive reconciliation. - Create ebook items as 'pending' so enrichment can promote them to 'matched' (backfill migration included), and protect matched items from re-scan clobbering: title/year skipped, people/series fill-empty only. - PDF metadata: scan head + tail windows (non-linearized PDFs keep the Info dict at the end), require proper key delimiters, head values win. - Cap plain .fb2 reads like .fbz entries; drop .md as an ebook format. - gofmt internal/scanner/audiobook.go (pre-existing drift). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ebooks): make enrichment failures non-terminal with dedicated backoff state - Provider errors now record a failure (capped retries) instead of stamping last_refreshed, which permanently excluded items after transient outages. - Unconfigured metadata chains and the scan-window membership race skip the item without stamping or burning a retry. - Failure tracking moves to a new ebook_enrichment_state table, decoupling it from media_items.refresh_failures (shared with metadata refresh debt). - Preserve non-author people credits when persisting enrichment results. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(catalog): gate ebook progress on hidden history and centralize threshold - Apply user_history_hidden_items gating (video semantics) to the ebook watched/in-progress filters, progress sort plan, and Continue Reading. - Continue Reading pages past dismissed items via the shared collector and dedupes items across pages (also fixes the video path's latent exposure). - Centralize the 0.9 finished threshold as models.EbookFinishedProgressThreshold with a single SQL-interpolated mirror in catalog. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(recommendations): correct watcher counting and wire ebook taste signals - itemWatchersQuery dedupes to distinct (watcher, item) rows so one binge-watcher can no longer satisfy minWatchers; the eligibility floor now counts distinct accounts rather than profiles. - Hidden-history gating on GetEbookReaderProgressForUser (signal reader). - Ebook reading produces canonical implicit taste signals (weighted like the equivalent movie progress ratio); ebooks join taste-seed candidates. - Stale GetRecentlyAddedItems doc comment corrected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(api): harden ebook reader endpoints and serve a Content-Security-Policy - Serve a CSP on all SPA HTML responses: blob/srcdoc book iframes inherit it, so script-src 'self' 'wasm-unsafe-eval' blocks script execution from malicious book content (sandbox alone is defeated by the WebKit allow-scripts requirement). Threat model documented on the constant. - X-Content-Type-Options: nosniff on frontend, jellycompat, and ebook file responses; MIME resolution can no longer fall through to octet-stream for an admitted ebook file. - Annotation PATCH: presence-aware field semantics (absent keeps, present sets/clears), invariant re-validation on the merged row, and an atomic SELECT ... FOR UPDATE read-merge-write. - Request size caps (413) on progress/config/annotation writes; Content-Disposition via mime.FormatMediaType; hidden-history gating in the shared ebook progress lister; FK-cascade indexes for reader tables. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(api): native read-state endpoints for ebooks - POST/DELETE /watched/{id} accepts ebook content IDs: mark read upserts progress 1.0 preserving the reader's file/location (or picks the preferred reader file for never-opened books); mark unread mirrors video unwatch semantics and deletes the progress row. - /history/remove accepts ebooks: hides via user_history_hidden_items without touching the reading position (hidden != unread; next reading activity resurfaces the book, mirroring video re-watch). - Access-filter checks match the video branch; shared logic lives in ebook_read_state.go. Sort metrics/user-state thresholds use the shared constant; profile-header fallback deduplicated. Clients: response is {type: "ebook", affected_count: 1, played: bool}; the existing watched SSE event fires. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(web): harden the ebook reader UI - Open-flow race: cancellation checked after every await with full stale-run teardown (no wrong-file progress saves, no leaked views/blob URLs); book.destroy() on cleanup. - Progress: monotonic stale-response guard; visibilitychange flush uses the refresh-capable client, pagehide uses keepalive; per-book cross-format progress documented as deliberate. - Settings: side effects out of the setState updater; local edits no longer clobbered by late server config; pending saves flushed on unmount/pagehide. - TTS: generation token so Stop actually stops (Chromium/Firefox synthetic events); Media Session uninstalled on unmount. - External book links: http(s) only, opened with noopener,noreferrer. - apiBlob 512 MiB guard with a user-facing error; fraction bookmarks navigable; search-result key collisions fixed; dead e-ink code removed; getLibrarySortRelevanceScope deduplicated; md format dropped. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(web): mark read/unread affordances for ebooks - Item detail gets a Mark Read/Unread button; card menus drop the ebook gate and share type-aware labels/toasts (also dedupes audiobook wording). - Watched-state invalidation includes the reader progress query key so the Continue button and percent refresh after toggling. - Continue Reading dismiss copy for ebooks; dismissal path now URL-encodes item IDs (ebook content IDs can contain reserved characters). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: record the PR #124 review and hardening pass Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
958 lines
40 KiB
Go
958 lines
40 KiB
Go
package metadata
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"fmt"
|
|
"log/slog"
|
|
"regexp"
|
|
"strconv"
|
|
"strings"
|
|
"time"
|
|
|
|
"github.com/jackc/pgx/v5"
|
|
"github.com/jackc/pgx/v5/pgxpool"
|
|
)
|
|
|
|
const providerIDRepairDefaultBatchSize = 250
|
|
|
|
type ProviderIDIntegrityStats struct {
|
|
Scanned int
|
|
CleanInserts int
|
|
ProvisionalConflictsRepaired int
|
|
MatchedCanonicalizations int
|
|
SkippedUnresolved int
|
|
Errors int
|
|
RemainingEstimate int
|
|
}
|
|
|
|
type ProviderIDIntegrityRepairer struct {
|
|
pool *pgxpool.Pool
|
|
}
|
|
|
|
func NewProviderIDIntegrityRepairer(pool *pgxpool.Pool) *ProviderIDIntegrityRepairer {
|
|
return &ProviderIDIntegrityRepairer{pool: pool}
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) Run(ctx context.Context, batchSize int) (ProviderIDIntegrityStats, error) {
|
|
var stats ProviderIDIntegrityStats
|
|
if r == nil || r.pool == nil {
|
|
return stats, nil
|
|
}
|
|
if batchSize <= 0 {
|
|
batchSize = providerIDRepairDefaultBatchSize
|
|
}
|
|
cursor := ""
|
|
for {
|
|
if err := ctx.Err(); err != nil {
|
|
return stats, err
|
|
}
|
|
batch, err := r.fetchDriftBatch(ctx, cursor, batchSize)
|
|
if err != nil {
|
|
return stats, err
|
|
}
|
|
if len(batch) == 0 {
|
|
break
|
|
}
|
|
for _, row := range batch {
|
|
stats.Scanned++
|
|
cursor = row.ContentID
|
|
outcome, err := r.repairDriftRow(ctx, row)
|
|
if err != nil {
|
|
stats.Errors++
|
|
slog.Warn("metadata: provider-id drift row repair failed",
|
|
"content_id", row.ContentID,
|
|
"item_type", row.ItemType,
|
|
"status", row.Status,
|
|
"error", err,
|
|
)
|
|
continue
|
|
}
|
|
switch outcome {
|
|
case "clean_insert":
|
|
stats.CleanInserts++
|
|
case "provisional_conflict":
|
|
stats.ProvisionalConflictsRepaired++
|
|
case "matched_canonicalization":
|
|
stats.MatchedCanonicalizations++
|
|
default:
|
|
stats.SkippedUnresolved++
|
|
}
|
|
}
|
|
if len(batch) < batchSize {
|
|
break
|
|
}
|
|
}
|
|
remaining, err := r.countDrift(ctx)
|
|
if err == nil {
|
|
stats.RemainingEstimate = remaining
|
|
}
|
|
return stats, nil
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) fetchDriftBatch(ctx context.Context, cursor string, batchSize int) ([]providerIDDriftRow, error) {
|
|
rows, err := r.pool.Query(ctx, `
|
|
SELECT mi.content_id,
|
|
mi.type,
|
|
mi.status,
|
|
COALESCE(mi.imdb_id, ''),
|
|
COALESCE(mi.tmdb_id, ''),
|
|
COALESCE(mi.tvdb_id, '')
|
|
FROM media_items mi
|
|
WHERE mi.type IN ('movie', 'series')
|
|
AND mi.content_id > $2
|
|
AND (COALESCE(mi.imdb_id, '') <> ''
|
|
OR COALESCE(mi.tmdb_id, '') <> ''
|
|
OR COALESCE(mi.tvdb_id, '') <> '')
|
|
AND EXISTS (
|
|
SELECT 1
|
|
FROM (
|
|
VALUES
|
|
('tmdb', COALESCE(mi.tmdb_id, '')),
|
|
('tvdb', COALESCE(mi.tvdb_id, '')),
|
|
('imdb', COALESCE(mi.imdb_id, ''))
|
|
) AS legacy(provider, provider_id)
|
|
WHERE legacy.provider_id <> ''
|
|
AND NOT EXISTS (
|
|
SELECT 1
|
|
FROM media_item_provider_ids pid
|
|
WHERE pid.content_id = mi.content_id
|
|
AND pid.item_type = mi.type
|
|
AND pid.provider = legacy.provider
|
|
AND pid.provider_id = legacy.provider_id
|
|
)
|
|
)
|
|
ORDER BY mi.content_id ASC
|
|
LIMIT $1
|
|
`, batchSize, cursor)
|
|
if err != nil {
|
|
return nil, fmt.Errorf("querying provider-id drift rows: %w", err)
|
|
}
|
|
defer rows.Close()
|
|
|
|
var batch []providerIDDriftRow
|
|
for rows.Next() {
|
|
var row providerIDDriftRow
|
|
if err := rows.Scan(&row.ContentID, &row.ItemType, &row.Status, &row.IMDbID, &row.TMDBID, &row.TVDBID); err != nil {
|
|
return nil, fmt.Errorf("scanning provider-id drift row: %w", err)
|
|
}
|
|
batch = append(batch, row)
|
|
}
|
|
if err := rows.Err(); err != nil {
|
|
return nil, fmt.Errorf("iterating provider-id drift rows: %w", err)
|
|
}
|
|
return batch, nil
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) countDrift(ctx context.Context) (int, error) {
|
|
var count int
|
|
err := r.pool.QueryRow(ctx, `
|
|
SELECT COUNT(*)
|
|
FROM media_items mi
|
|
WHERE mi.type IN ('movie', 'series')
|
|
AND (COALESCE(mi.imdb_id, '') <> ''
|
|
OR COALESCE(mi.tmdb_id, '') <> ''
|
|
OR COALESCE(mi.tvdb_id, '') <> '')
|
|
AND EXISTS (
|
|
SELECT 1
|
|
FROM (
|
|
VALUES
|
|
('tmdb', COALESCE(mi.tmdb_id, '')),
|
|
('tvdb', COALESCE(mi.tvdb_id, '')),
|
|
('imdb', COALESCE(mi.imdb_id, ''))
|
|
) AS legacy(provider, provider_id)
|
|
WHERE legacy.provider_id <> ''
|
|
AND NOT EXISTS (
|
|
SELECT 1
|
|
FROM media_item_provider_ids pid
|
|
WHERE pid.content_id = mi.content_id
|
|
AND pid.item_type = mi.type
|
|
AND pid.provider = legacy.provider
|
|
AND pid.provider_id = legacy.provider_id
|
|
)
|
|
)
|
|
`).Scan(&count)
|
|
if err != nil {
|
|
return 0, fmt.Errorf("counting remaining provider-id drift: %w", err)
|
|
}
|
|
return count, nil
|
|
}
|
|
|
|
type providerIDDriftRow struct {
|
|
ContentID string
|
|
ItemType string
|
|
Status string
|
|
IMDbID string
|
|
TMDBID string
|
|
TVDBID string
|
|
}
|
|
|
|
type canonicalCandidate struct {
|
|
ContentID string
|
|
Status string
|
|
CreatedAt time.Time
|
|
ActiveFiles int
|
|
LibraryCount int
|
|
}
|
|
|
|
type providerIDEntry struct {
|
|
Provider string
|
|
ProviderID string
|
|
}
|
|
|
|
type providerIDOwner struct {
|
|
ContentID string
|
|
Status string
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) repairDriftRow(ctx context.Context, row providerIDDriftRow) (string, error) {
|
|
entries := providerEntriesFromDrift(row)
|
|
if len(entries) == 0 {
|
|
return "unresolved", nil
|
|
}
|
|
owners, err := r.findProviderIDOwners(ctx, row.ContentID, row.ItemType, entries)
|
|
if err != nil {
|
|
return "", err
|
|
}
|
|
if len(owners) == 0 {
|
|
if err := r.insertProviderIDsForDriftRow(ctx, row, entries); err != nil {
|
|
return "", err
|
|
}
|
|
return "clean_insert", nil
|
|
}
|
|
if len(owners) != 1 {
|
|
return "unresolved", nil
|
|
}
|
|
var owner providerIDOwner
|
|
for _, candidate := range owners {
|
|
owner = candidate
|
|
}
|
|
switch {
|
|
case isProvisionalOwnershipStatus(owner.Status) && isConfirmedOwnershipStatus(row.Status):
|
|
if err := r.moveProvisionalProviderClaim(ctx, owner.ContentID, row, entries); err != nil {
|
|
return "", err
|
|
}
|
|
return "provisional_conflict", nil
|
|
case isConfirmedOwnershipStatus(owner.Status) && isProvisionalOwnershipStatus(row.Status):
|
|
if _, err := canonicalizeProviderIDDuplicateInto(ctx, r.pool, row.ContentID, owner.ContentID, false); err != nil {
|
|
return "", err
|
|
}
|
|
return "matched_canonicalization", nil
|
|
case isConfirmedOwnershipStatus(owner.Status) && isConfirmedOwnershipStatus(row.Status):
|
|
if _, err := canonicalizeProviderIDDuplicate(ctx, r.pool, row.ContentID, owner.ContentID, true); err != nil {
|
|
return "", err
|
|
}
|
|
return "matched_canonicalization", nil
|
|
default:
|
|
return "unresolved", nil
|
|
}
|
|
}
|
|
|
|
func providerEntriesFromDrift(row providerIDDriftRow) []providerIDEntry {
|
|
entries := make([]providerIDEntry, 0, 3)
|
|
if strings.TrimSpace(row.TMDBID) != "" {
|
|
entries = append(entries, providerIDEntry{Provider: "tmdb", ProviderID: strings.TrimSpace(row.TMDBID)})
|
|
}
|
|
if strings.TrimSpace(row.TVDBID) != "" {
|
|
entries = append(entries, providerIDEntry{Provider: "tvdb", ProviderID: strings.TrimSpace(row.TVDBID)})
|
|
}
|
|
if strings.TrimSpace(row.IMDbID) != "" {
|
|
entries = append(entries, providerIDEntry{Provider: "imdb", ProviderID: strings.TrimSpace(row.IMDbID)})
|
|
}
|
|
return entries
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) findProviderIDOwners(
|
|
ctx context.Context,
|
|
contentID string,
|
|
itemType string,
|
|
entries []providerIDEntry,
|
|
) (map[string]providerIDOwner, error) {
|
|
owners := make(map[string]providerIDOwner)
|
|
for _, entry := range entries {
|
|
var owner providerIDOwner
|
|
err := r.pool.QueryRow(ctx, `
|
|
SELECT pid.content_id, mi.status
|
|
FROM media_item_provider_ids pid
|
|
JOIN media_items mi ON mi.content_id = pid.content_id
|
|
WHERE pid.provider = $1
|
|
AND pid.provider_id = $2
|
|
AND pid.item_type = $3
|
|
AND pid.content_id <> $4
|
|
LIMIT 1
|
|
`, entry.Provider, entry.ProviderID, itemType, contentID).Scan(&owner.ContentID, &owner.Status)
|
|
if err != nil {
|
|
if errors.Is(err, pgx.ErrNoRows) {
|
|
continue
|
|
}
|
|
return nil, fmt.Errorf("finding owner for %s=%s: %w", entry.Provider, entry.ProviderID, err)
|
|
}
|
|
owners[owner.ContentID] = owner
|
|
}
|
|
return owners, nil
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) insertProviderIDsForDriftRow(ctx context.Context, row providerIDDriftRow, entries []providerIDEntry) error {
|
|
tx, err := r.pool.BeginTx(ctx, pgx.TxOptions{})
|
|
if err != nil {
|
|
return fmt.Errorf("begin clean provider-id repair: %w", err)
|
|
}
|
|
defer tx.Rollback(ctx) //nolint:errcheck
|
|
if _, err := tx.Exec(ctx, `DELETE FROM media_item_provider_ids WHERE content_id = $1`, row.ContentID); err != nil {
|
|
return fmt.Errorf("clear existing provider ids for %s: %w", row.ContentID, err)
|
|
}
|
|
for _, entry := range entries {
|
|
if _, err := tx.Exec(ctx, `
|
|
INSERT INTO media_item_provider_ids (content_id, item_type, provider, provider_id, created_at, updated_at)
|
|
VALUES ($1, $2, $3, $4, NOW(), NOW())
|
|
`, row.ContentID, row.ItemType, entry.Provider, entry.ProviderID); err != nil {
|
|
return fmt.Errorf("insert repaired provider id %s for %s: %w", entry.Provider, row.ContentID, err)
|
|
}
|
|
}
|
|
if err := tx.Commit(ctx); err != nil {
|
|
return fmt.Errorf("commit clean provider-id repair: %w", err)
|
|
}
|
|
return nil
|
|
}
|
|
|
|
func (r *ProviderIDIntegrityRepairer) moveProvisionalProviderClaim(ctx context.Context, ownerContentID string, row providerIDDriftRow, entries []providerIDEntry) error {
|
|
tx, err := r.pool.BeginTx(ctx, pgx.TxOptions{})
|
|
if err != nil {
|
|
return fmt.Errorf("begin provisional provider-id repair: %w", err)
|
|
}
|
|
defer tx.Rollback(ctx) //nolint:errcheck
|
|
var ownerStatus string
|
|
err = tx.QueryRow(ctx, `SELECT status FROM media_items WHERE content_id = $1 FOR UPDATE`, ownerContentID).Scan(&ownerStatus)
|
|
if err != nil {
|
|
if errors.Is(err, pgx.ErrNoRows) {
|
|
return nil
|
|
}
|
|
return fmt.Errorf("loading status for provisional owner %s: %w", ownerContentID, err)
|
|
}
|
|
if isConfirmedOwnershipStatus(ownerStatus) {
|
|
return fmt.Errorf("refusing to move provider IDs from %s: status is now %q", ownerContentID, ownerStatus)
|
|
}
|
|
if _, err := tx.Exec(ctx, `DELETE FROM media_item_provider_ids WHERE content_id = $1`, ownerContentID); err != nil {
|
|
return fmt.Errorf("clear provisional provider ids for %s: %w", ownerContentID, err)
|
|
}
|
|
if _, err := tx.Exec(ctx, `DELETE FROM media_item_provider_ids WHERE content_id = $1`, row.ContentID); err != nil {
|
|
return fmt.Errorf("clear existing provider ids for %s: %w", row.ContentID, err)
|
|
}
|
|
for _, entry := range entries {
|
|
if _, err := tx.Exec(ctx, `
|
|
INSERT INTO media_item_provider_ids (content_id, item_type, provider, provider_id, created_at, updated_at)
|
|
VALUES ($1, $2, $3, $4, NOW(), NOW())
|
|
`, row.ContentID, row.ItemType, entry.Provider, entry.ProviderID); err != nil {
|
|
return fmt.Errorf("insert provider id %s for %s after provisional repair: %w", entry.Provider, row.ContentID, err)
|
|
}
|
|
}
|
|
if err := tx.Commit(ctx); err != nil {
|
|
return fmt.Errorf("commit provisional provider-id repair: %w", err)
|
|
}
|
|
return nil
|
|
}
|
|
|
|
func (s *MetadataService) canonicalizeProviderIDDuplicate(
|
|
ctx context.Context,
|
|
leftContentID string,
|
|
rightContentID string,
|
|
allowMatchedSource bool,
|
|
) (string, error) {
|
|
if s == nil || s.dbPool == nil {
|
|
return "", nil
|
|
}
|
|
return canonicalizeProviderIDDuplicate(ctx, s.dbPool, leftContentID, rightContentID, allowMatchedSource)
|
|
}
|
|
|
|
// clearProvisionalProviderIDsLocked removes provider ID rows for an item only
|
|
// after re-confirming, under a row-level lock, that the item is still in a
|
|
// provisional ownership status. Without the lock, a parallel goroutine could
|
|
// promote the item to matched between the caller's status check and our
|
|
// delete, silently wiping a confirmed item's provider IDs.
|
|
func clearProvisionalProviderIDsLocked(ctx context.Context, pool *pgxpool.Pool, contentID string) error {
|
|
contentID = strings.TrimSpace(contentID)
|
|
if contentID == "" {
|
|
return fmt.Errorf("content_id is required")
|
|
}
|
|
tx, err := pool.BeginTx(ctx, pgx.TxOptions{})
|
|
if err != nil {
|
|
return fmt.Errorf("begin provisional provider-id clear: %w", err)
|
|
}
|
|
defer tx.Rollback(ctx) //nolint:errcheck
|
|
|
|
var status string
|
|
err = tx.QueryRow(ctx, `SELECT status FROM media_items WHERE content_id = $1 FOR UPDATE`, contentID).Scan(&status)
|
|
if err != nil {
|
|
if errors.Is(err, pgx.ErrNoRows) {
|
|
return nil
|
|
}
|
|
return fmt.Errorf("loading status for provisional provider-id clear of %s: %w", contentID, err)
|
|
}
|
|
if isConfirmedOwnershipStatus(status) {
|
|
return fmt.Errorf("refusing to clear provider IDs of %s: status is now %q", contentID, status)
|
|
}
|
|
if _, err := tx.Exec(ctx, `DELETE FROM media_item_provider_ids WHERE content_id = $1`, contentID); err != nil {
|
|
return fmt.Errorf("clearing provisional provider IDs from %s: %w", contentID, err)
|
|
}
|
|
if err := tx.Commit(ctx); err != nil {
|
|
return fmt.Errorf("commit provisional provider-id clear: %w", err)
|
|
}
|
|
return nil
|
|
}
|
|
|
|
func canonicalizeProviderIDDuplicate(
|
|
ctx context.Context,
|
|
pool *pgxpool.Pool,
|
|
leftContentID string,
|
|
rightContentID string,
|
|
allowMatchedSource bool,
|
|
) (string, error) {
|
|
leftContentID = strings.TrimSpace(leftContentID)
|
|
rightContentID = strings.TrimSpace(rightContentID)
|
|
if leftContentID == "" || rightContentID == "" || leftContentID == rightContentID {
|
|
return leftContentID, nil
|
|
}
|
|
tx, err := pool.BeginTx(ctx, pgx.TxOptions{})
|
|
if err != nil {
|
|
return "", fmt.Errorf("begin provider-id canonicalization: %w", err)
|
|
}
|
|
defer tx.Rollback(ctx) //nolint:errcheck
|
|
|
|
canonical, source, err := chooseCanonicalProviderOwner(ctx, tx, leftContentID, rightContentID)
|
|
if err != nil {
|
|
return "", err
|
|
}
|
|
if !allowMatchedSource && isConfirmedOwnershipStatus(source.Status) {
|
|
return "", fmt.Errorf("refusing to canonicalize matched source %s without allowMatchedSource", source.ContentID)
|
|
}
|
|
if isConfirmedOwnershipStatus(canonical.Status) && isConfirmedOwnershipStatus(source.Status) && bothCandidatesHaveContent(canonical, source) {
|
|
return "", fmt.Errorf("refusing to merge matched items %s and %s: both have files or library memberships (likely separate editions)", canonical.ContentID, source.ContentID)
|
|
}
|
|
if err := canonicalizeMediaItemReferencesTx(ctx, tx, source.ContentID, canonical.ContentID); err != nil {
|
|
return "", err
|
|
}
|
|
if _, err := tx.Exec(ctx, `DELETE FROM media_items WHERE content_id = $1`, source.ContentID); err != nil {
|
|
return "", fmt.Errorf("delete duplicate media item %s: %w", source.ContentID, err)
|
|
}
|
|
if err := tx.Commit(ctx); err != nil {
|
|
return "", fmt.Errorf("commit provider-id canonicalization: %w", err)
|
|
}
|
|
return canonical.ContentID, nil
|
|
}
|
|
|
|
func canonicalizeProviderIDDuplicateInto(
|
|
ctx context.Context,
|
|
pool *pgxpool.Pool,
|
|
sourceID string,
|
|
canonicalID string,
|
|
allowMatchedSource bool,
|
|
) (string, error) {
|
|
sourceID = strings.TrimSpace(sourceID)
|
|
canonicalID = strings.TrimSpace(canonicalID)
|
|
if sourceID == "" || canonicalID == "" {
|
|
return "", fmt.Errorf("source and canonical content IDs are required")
|
|
}
|
|
if sourceID == canonicalID {
|
|
return canonicalID, nil
|
|
}
|
|
tx, err := pool.BeginTx(ctx, pgx.TxOptions{})
|
|
if err != nil {
|
|
return "", fmt.Errorf("begin provider-id canonicalization: %w", err)
|
|
}
|
|
defer tx.Rollback(ctx) //nolint:errcheck
|
|
|
|
candidates := make(map[string]canonicalCandidate, 2)
|
|
rows, err := tx.Query(ctx, `
|
|
SELECT mi.content_id,
|
|
mi.status,
|
|
mi.created_at,
|
|
(SELECT COUNT(*) FROM media_files mf WHERE mf.content_id = mi.content_id AND mf.missing_since IS NULL) AS active_files,
|
|
(SELECT COUNT(*) FROM media_item_libraries mil WHERE mil.content_id = mi.content_id) AS library_count
|
|
FROM media_items mi
|
|
WHERE mi.content_id = ANY($1::text[])
|
|
ORDER BY mi.content_id
|
|
FOR UPDATE
|
|
`, []string{sourceID, canonicalID})
|
|
if err != nil {
|
|
return "", fmt.Errorf("locking duplicate provider owners: %w", err)
|
|
}
|
|
for rows.Next() {
|
|
var candidate canonicalCandidate
|
|
if err := rows.Scan(&candidate.ContentID, &candidate.Status, &candidate.CreatedAt, &candidate.ActiveFiles, &candidate.LibraryCount); err != nil {
|
|
rows.Close()
|
|
return "", fmt.Errorf("scanning duplicate provider owner lock: %w", err)
|
|
}
|
|
candidates[candidate.ContentID] = candidate
|
|
}
|
|
if err := rows.Err(); err != nil {
|
|
rows.Close()
|
|
return "", fmt.Errorf("iterating duplicate provider owner locks: %w", err)
|
|
}
|
|
rows.Close()
|
|
source, sourceOK := candidates[sourceID]
|
|
canonical, canonicalOK := candidates[canonicalID]
|
|
if !sourceOK || !canonicalOK {
|
|
return "", fmt.Errorf("expected two duplicate provider owners, got %d", len(candidates))
|
|
}
|
|
if !allowMatchedSource && isConfirmedOwnershipStatus(source.Status) {
|
|
return "", fmt.Errorf("refusing to canonicalize matched source %s without allowMatchedSource", sourceID)
|
|
}
|
|
// Scanner-created provisional rows can already own local files/library
|
|
// memberships. If the matched owner also has content, merging here would
|
|
// collapse separate editions before the source is confirmed.
|
|
if bothCandidatesHaveContent(canonical, source) {
|
|
return "", fmt.Errorf("refusing to merge items %s and %s: both have files or library memberships (likely separate editions)", canonical.ContentID, source.ContentID)
|
|
}
|
|
if err := canonicalizeMediaItemReferencesTx(ctx, tx, sourceID, canonicalID); err != nil {
|
|
return "", err
|
|
}
|
|
if _, err := tx.Exec(ctx, `DELETE FROM media_items WHERE content_id = $1`, sourceID); err != nil {
|
|
return "", fmt.Errorf("delete duplicate media item %s: %w", sourceID, err)
|
|
}
|
|
if err := tx.Commit(ctx); err != nil {
|
|
return "", fmt.Errorf("commit provider-id canonicalization: %w", err)
|
|
}
|
|
return canonicalID, nil
|
|
}
|
|
|
|
func chooseCanonicalProviderOwner(ctx context.Context, tx pgx.Tx, leftContentID, rightContentID string) (canonical canonicalCandidate, source canonicalCandidate, err error) {
|
|
rows, err := tx.Query(ctx, `
|
|
SELECT mi.content_id,
|
|
mi.status,
|
|
mi.created_at,
|
|
(SELECT COUNT(*) FROM media_files mf WHERE mf.content_id = mi.content_id AND mf.missing_since IS NULL) AS active_files,
|
|
(SELECT COUNT(*) FROM media_item_libraries mil WHERE mil.content_id = mi.content_id) AS library_count
|
|
FROM media_items mi
|
|
WHERE mi.content_id = ANY($1::text[])
|
|
ORDER BY mi.content_id
|
|
FOR UPDATE
|
|
`, []string{leftContentID, rightContentID})
|
|
if err != nil {
|
|
return canonicalCandidate{}, canonicalCandidate{}, fmt.Errorf("loading duplicate provider owners: %w", err)
|
|
}
|
|
defer rows.Close()
|
|
|
|
candidates := make([]canonicalCandidate, 0, 2)
|
|
for rows.Next() {
|
|
var c canonicalCandidate
|
|
if err := rows.Scan(&c.ContentID, &c.Status, &c.CreatedAt, &c.ActiveFiles, &c.LibraryCount); err != nil {
|
|
return canonicalCandidate{}, canonicalCandidate{}, fmt.Errorf("scanning duplicate provider owner: %w", err)
|
|
}
|
|
candidates = append(candidates, c)
|
|
}
|
|
if err := rows.Err(); err != nil {
|
|
return canonicalCandidate{}, canonicalCandidate{}, fmt.Errorf("iterating duplicate provider owners: %w", err)
|
|
}
|
|
if len(candidates) != 2 {
|
|
return canonicalCandidate{}, canonicalCandidate{}, fmt.Errorf("expected two duplicate provider owners, got %d", len(candidates))
|
|
}
|
|
|
|
best := candidates[0]
|
|
other := candidates[1]
|
|
if betterCanonicalCandidate(other, best) {
|
|
best, other = other, best
|
|
}
|
|
return best, other, nil
|
|
}
|
|
|
|
// bothCandidatesHaveContent reports whether both items appear to hold real
|
|
// user-curated content (active files or library memberships). When that is
|
|
// true for two matched items, they are likely intentional separate entries
|
|
// (e.g. director's vs. theatrical cut sharing an external ID) and we refuse
|
|
// to silently destroy one by merging.
|
|
func bothCandidatesHaveContent(left, right canonicalCandidate) bool {
|
|
leftHasContent := left.ActiveFiles > 0 || left.LibraryCount > 0
|
|
rightHasContent := right.ActiveFiles > 0 || right.LibraryCount > 0
|
|
return leftHasContent && rightHasContent
|
|
}
|
|
|
|
func betterCanonicalCandidate(left, right canonicalCandidate) bool {
|
|
if (left.ActiveFiles > 0) != (right.ActiveFiles > 0) {
|
|
return left.ActiveFiles > 0
|
|
}
|
|
if left.LibraryCount != right.LibraryCount {
|
|
return left.LibraryCount > right.LibraryCount
|
|
}
|
|
if !left.CreatedAt.Equal(right.CreatedAt) {
|
|
return left.CreatedAt.Before(right.CreatedAt)
|
|
}
|
|
return left.ContentID < right.ContentID
|
|
}
|
|
|
|
// mediaItemMergeStep is one statement in the ordered media-item merge sequence
|
|
// run by canonicalizeMediaItemReferencesTx. Every step references $1 (the
|
|
// source content ID); steps that move data onto the canonical row also
|
|
// reference $2 (the canonical content ID).
|
|
type mediaItemMergeStep struct {
|
|
name string
|
|
sql string
|
|
}
|
|
|
|
// mergeStepPlaceholderRe matches positional query placeholders ($1, $2, …).
|
|
var mergeStepPlaceholderRe = regexp.MustCompile(`\$(\d+)`)
|
|
|
|
// maxPlaceholder returns the highest positional placeholder number referenced
|
|
// in sql (0 if none). It reads the full placeholder number, so "$20" is 20 and
|
|
// is never confused with "$2" — unlike a substring check.
|
|
func maxPlaceholder(sql string) int {
|
|
highest := 0
|
|
for _, m := range mergeStepPlaceholderRe.FindAllStringSubmatch(sql, -1) {
|
|
if n, err := strconv.Atoi(m[1]); err == nil && n > highest {
|
|
highest = n
|
|
}
|
|
}
|
|
return highest
|
|
}
|
|
|
|
// mergeStepArgs returns the positional arguments for a merge step, sized to the
|
|
// placeholders the SQL actually binds. Passing too many args (e.g. an unused
|
|
// $2) makes pgx reject the Exec with "mismatched param and argument count"
|
|
// under the default QueryExecModeCacheStatement, aborting the whole merge
|
|
// transaction. Every merge step binds $1 (source) and optionally $2
|
|
// (canonical); TestMergeStepPlaceholdersAreBounded enforces that contract, so
|
|
// the default branch is unreachable for the defined steps and fails loud if a
|
|
// future step violates it.
|
|
func mergeStepArgs(sql, sourceID, canonicalID string) []any {
|
|
switch n := maxPlaceholder(sql); n {
|
|
case 1:
|
|
return []any{sourceID}
|
|
case 2:
|
|
return []any{sourceID, canonicalID}
|
|
default:
|
|
panic(fmt.Sprintf("merge step binds unsupported placeholder count %d: %q", n, sql))
|
|
}
|
|
}
|
|
|
|
var mediaItemMergeSteps = []mediaItemMergeStep{
|
|
{"merge provider ids", `
|
|
INSERT INTO media_item_provider_ids (content_id, item_type, provider, provider_id, created_at, updated_at)
|
|
SELECT $2, item_type, provider, provider_id, created_at, NOW()
|
|
FROM media_item_provider_ids
|
|
WHERE content_id = $1
|
|
ON CONFLICT DO NOTHING`},
|
|
{"drop source provider ids", `DELETE FROM media_item_provider_ids WHERE content_id = $1`},
|
|
{"merge legacy provider columns", `
|
|
UPDATE media_items dest
|
|
SET imdb_id = COALESCE(NULLIF(dest.imdb_id, ''), src.imdb_id),
|
|
tmdb_id = COALESCE(NULLIF(dest.tmdb_id, ''), src.tmdb_id),
|
|
tvdb_id = COALESCE(NULLIF(dest.tvdb_id, ''), src.tvdb_id),
|
|
updated_at = NOW()
|
|
FROM media_items src
|
|
WHERE dest.content_id = $2 AND src.content_id = $1`},
|
|
{"move media files", `UPDATE media_files SET content_id = $2, updated_at = NOW() WHERE content_id = $1`},
|
|
{"move seasons", `UPDATE seasons SET series_id = $2, updated_at = NOW() WHERE series_id = $1`},
|
|
{"move episodes", `UPDATE episodes SET series_id = $2, updated_at = NOW() WHERE series_id = $1`},
|
|
{"merge library memberships", `
|
|
INSERT INTO media_item_libraries (content_id, media_folder_id, first_seen_at)
|
|
SELECT $2, media_folder_id, MIN(first_seen_at)
|
|
FROM media_item_libraries
|
|
WHERE content_id = $1
|
|
GROUP BY media_folder_id
|
|
ON CONFLICT (content_id, media_folder_id) DO UPDATE
|
|
SET first_seen_at = LEAST(media_item_libraries.first_seen_at, EXCLUDED.first_seen_at)`},
|
|
{"delete source library memberships", `DELETE FROM media_item_libraries WHERE content_id = $1`},
|
|
{"move root claims", `UPDATE media_item_roots SET content_id = $2, last_seen_at = NOW() WHERE content_id = $1`},
|
|
{"move group claims", `UPDATE media_item_groups SET content_id = $2, last_seen_at = NOW() WHERE content_id = $1`},
|
|
{"merge collection items", `
|
|
INSERT INTO library_collection_items (collection_id, media_item_id, position, source_rank, created_at, updated_at)
|
|
SELECT collection_id, $2, MIN(position), MIN(source_rank), MIN(created_at), NOW()
|
|
FROM library_collection_items
|
|
WHERE media_item_id = $1
|
|
GROUP BY collection_id
|
|
ON CONFLICT (collection_id, media_item_id) DO UPDATE
|
|
SET position = LEAST(library_collection_items.position, EXCLUDED.position),
|
|
source_rank = LEAST(library_collection_items.source_rank, EXCLUDED.source_rank),
|
|
updated_at = NOW()`},
|
|
{"delete source collection items", `DELETE FROM library_collection_items WHERE media_item_id = $1`},
|
|
{"merge personal collection items", `
|
|
INSERT INTO user_personal_collection_items (user_id, collection_id, media_item_id, position, added_at)
|
|
SELECT user_id, collection_id, $2, MIN(position), MIN(added_at)
|
|
FROM user_personal_collection_items
|
|
WHERE media_item_id = $1
|
|
GROUP BY user_id, collection_id
|
|
ON CONFLICT (user_id, collection_id, media_item_id) DO UPDATE
|
|
SET position = LEAST(user_personal_collection_items.position, EXCLUDED.position),
|
|
added_at = LEAST(user_personal_collection_items.added_at, EXCLUDED.added_at)`},
|
|
{"delete source personal collection items", `DELETE FROM user_personal_collection_items WHERE media_item_id = $1`},
|
|
{"merge favorites", `
|
|
INSERT INTO user_favorites (user_id, profile_id, media_item_id, added_at)
|
|
SELECT user_id, profile_id, $2, MIN(added_at)
|
|
FROM user_favorites
|
|
WHERE media_item_id = $1
|
|
GROUP BY user_id, profile_id
|
|
ON CONFLICT (user_id, profile_id, media_item_id) DO UPDATE
|
|
SET added_at = LEAST(user_favorites.added_at, EXCLUDED.added_at)`},
|
|
{"delete source favorites", `DELETE FROM user_favorites WHERE media_item_id = $1`},
|
|
{"merge watchlist", `
|
|
INSERT INTO user_watchlist (user_id, profile_id, media_item_id, added_at)
|
|
SELECT user_id, profile_id, $2, MIN(added_at)
|
|
FROM user_watchlist
|
|
WHERE media_item_id = $1
|
|
GROUP BY user_id, profile_id
|
|
ON CONFLICT (user_id, profile_id, media_item_id) DO UPDATE
|
|
SET added_at = LEAST(user_watchlist.added_at, EXCLUDED.added_at)`},
|
|
{"delete source watchlist", `DELETE FROM user_watchlist WHERE media_item_id = $1`},
|
|
{"merge ratings", `
|
|
INSERT INTO user_ratings (user_id, profile_id, media_item_id, rating, rated_at)
|
|
SELECT DISTINCT ON (user_id, profile_id) user_id, profile_id, $2, rating, rated_at
|
|
FROM user_ratings
|
|
WHERE media_item_id = $1
|
|
ORDER BY user_id, profile_id, rated_at DESC
|
|
ON CONFLICT (user_id, profile_id, media_item_id) DO UPDATE
|
|
SET rating = CASE WHEN EXCLUDED.rated_at >= user_ratings.rated_at THEN EXCLUDED.rating ELSE user_ratings.rating END,
|
|
rated_at = GREATEST(user_ratings.rated_at, EXCLUDED.rated_at)`},
|
|
{"delete source ratings", `DELETE FROM user_ratings WHERE media_item_id = $1`},
|
|
{"merge progress", `
|
|
INSERT INTO user_watch_progress (
|
|
user_id, profile_id, media_item_id, position_seconds, duration_seconds, completed,
|
|
updated_at, last_file_id, last_resolution, last_hdr, last_codec_video, last_edition_key
|
|
)
|
|
SELECT user_id, profile_id, $2, position_seconds, duration_seconds, completed,
|
|
updated_at, last_file_id, last_resolution, last_hdr, last_codec_video, last_edition_key
|
|
FROM user_watch_progress
|
|
WHERE media_item_id = $1
|
|
ON CONFLICT (user_id, profile_id, media_item_id) DO UPDATE
|
|
SET position_seconds = CASE WHEN EXCLUDED.updated_at >= user_watch_progress.updated_at THEN EXCLUDED.position_seconds ELSE user_watch_progress.position_seconds END,
|
|
duration_seconds = GREATEST(user_watch_progress.duration_seconds, EXCLUDED.duration_seconds),
|
|
completed = user_watch_progress.completed OR EXCLUDED.completed,
|
|
updated_at = GREATEST(user_watch_progress.updated_at, EXCLUDED.updated_at),
|
|
last_file_id = CASE WHEN EXCLUDED.updated_at >= user_watch_progress.updated_at THEN EXCLUDED.last_file_id ELSE user_watch_progress.last_file_id END,
|
|
last_resolution = CASE WHEN EXCLUDED.updated_at >= user_watch_progress.updated_at THEN EXCLUDED.last_resolution ELSE user_watch_progress.last_resolution END,
|
|
last_hdr = CASE WHEN EXCLUDED.updated_at >= user_watch_progress.updated_at THEN EXCLUDED.last_hdr ELSE user_watch_progress.last_hdr END,
|
|
last_codec_video = CASE WHEN EXCLUDED.updated_at >= user_watch_progress.updated_at THEN EXCLUDED.last_codec_video ELSE user_watch_progress.last_codec_video END,
|
|
last_edition_key = CASE WHEN EXCLUDED.updated_at >= user_watch_progress.updated_at THEN EXCLUDED.last_edition_key ELSE user_watch_progress.last_edition_key END`},
|
|
{"delete source progress", `DELETE FROM user_watch_progress WHERE media_item_id = $1`},
|
|
{"merge ebook reader progress", `
|
|
INSERT INTO ebook_reader_progress (
|
|
user_id, profile_id, content_id, file_id, location, progress, updated_at
|
|
)
|
|
SELECT user_id, profile_id, $2, file_id, location, progress, updated_at
|
|
FROM ebook_reader_progress
|
|
WHERE content_id = $1
|
|
ON CONFLICT (user_id, profile_id, content_id) DO UPDATE
|
|
SET file_id = CASE WHEN EXCLUDED.updated_at >= ebook_reader_progress.updated_at THEN EXCLUDED.file_id ELSE ebook_reader_progress.file_id END,
|
|
location = CASE WHEN EXCLUDED.updated_at >= ebook_reader_progress.updated_at THEN EXCLUDED.location ELSE ebook_reader_progress.location END,
|
|
progress = CASE WHEN EXCLUDED.updated_at >= ebook_reader_progress.updated_at THEN EXCLUDED.progress ELSE ebook_reader_progress.progress END,
|
|
updated_at = GREATEST(ebook_reader_progress.updated_at, EXCLUDED.updated_at)`},
|
|
{"delete source ebook reader progress", `DELETE FROM ebook_reader_progress WHERE content_id = $1`},
|
|
{"move history", `UPDATE user_watch_history SET media_item_id = $2 WHERE media_item_id = $1`},
|
|
{"merge hidden history", `
|
|
INSERT INTO user_history_hidden_items (user_id, profile_id, media_item_id, hidden_before, updated_at)
|
|
SELECT user_id, profile_id, $2, MAX(hidden_before), MAX(updated_at)
|
|
FROM user_history_hidden_items
|
|
WHERE media_item_id = $1
|
|
GROUP BY user_id, profile_id
|
|
ON CONFLICT (user_id, profile_id, media_item_id) DO UPDATE
|
|
SET hidden_before = GREATEST(user_history_hidden_items.hidden_before, EXCLUDED.hidden_before),
|
|
updated_at = GREATEST(user_history_hidden_items.updated_at, EXCLUDED.updated_at)`},
|
|
{"delete source hidden history", `DELETE FROM user_history_hidden_items WHERE media_item_id = $1`},
|
|
{"delete duplicate home item dismissals", `
|
|
DELETE FROM user_home_item_dismissals src
|
|
USING user_home_item_dismissals dest
|
|
WHERE src.media_item_id = $1
|
|
AND dest.media_item_id = $2
|
|
AND dest.user_id = src.user_id
|
|
AND dest.profile_id = src.profile_id
|
|
AND dest.surface = src.surface`},
|
|
{"move home item dismissals", `UPDATE user_home_item_dismissals SET media_item_id = $2 WHERE media_item_id = $1`},
|
|
{"remap home item dismissals series", `UPDATE user_home_item_dismissals SET series_id = $2 WHERE series_id = $1`},
|
|
{"delete duplicate audio preferences", `
|
|
DELETE FROM user_audio_preferences src
|
|
USING user_audio_preferences dest
|
|
WHERE src.series_id = $1
|
|
AND dest.series_id = $2
|
|
AND dest.user_id = src.user_id
|
|
AND dest.profile_id = src.profile_id`},
|
|
{"move audio preferences", `UPDATE user_audio_preferences SET series_id = $2 WHERE series_id = $1`},
|
|
{"delete duplicate subtitle preferences", `
|
|
DELETE FROM user_subtitle_preferences src
|
|
USING user_subtitle_preferences dest
|
|
WHERE src.series_id = $1
|
|
AND dest.series_id = $2
|
|
AND dest.user_id = src.user_id
|
|
AND dest.profile_id = src.profile_id`},
|
|
{"move subtitle preferences", `UPDATE user_subtitle_preferences SET series_id = $2 WHERE series_id = $1`},
|
|
{"delete duplicate series playback preferences", `
|
|
DELETE FROM user_series_playback_preferences src
|
|
USING user_series_playback_preferences dest
|
|
WHERE src.series_id = $1
|
|
AND dest.series_id = $2
|
|
AND dest.user_id = src.user_id
|
|
AND dest.profile_id = src.profile_id`},
|
|
{"move series playback preferences", `UPDATE user_series_playback_preferences SET series_id = $2 WHERE series_id = $1`},
|
|
{"merge plex sync item bindings timestamps", `
|
|
UPDATE plex_sync_item_bindings dest
|
|
SET last_seen_at = GREATEST(dest.last_seen_at, src.last_seen_at),
|
|
updated_at = NOW()
|
|
FROM plex_sync_item_bindings src
|
|
WHERE src.media_item_id = $1
|
|
AND dest.connection_id = src.connection_id
|
|
AND (dest.media_item_id = $2 OR dest.plex_rating_key = src.plex_rating_key)
|
|
AND dest.media_item_id <> src.media_item_id`},
|
|
{"delete duplicate plex sync item bindings", `
|
|
DELETE FROM plex_sync_item_bindings src
|
|
USING plex_sync_item_bindings dest
|
|
WHERE src.media_item_id = $1
|
|
AND dest.connection_id = src.connection_id
|
|
AND (dest.media_item_id = $2 OR dest.plex_rating_key = src.plex_rating_key)
|
|
AND dest.media_item_id <> src.media_item_id`},
|
|
{"move remaining plex sync item bindings", `UPDATE plex_sync_item_bindings SET media_item_id = $2, updated_at = NOW() WHERE media_item_id = $1`},
|
|
{"delete duplicate plex sync item state", `
|
|
DELETE FROM plex_sync_item_state src
|
|
USING plex_sync_item_state dest
|
|
WHERE src.media_item_id = $1
|
|
AND dest.media_item_id = $2
|
|
AND dest.mapping_id = src.mapping_id`},
|
|
{"move plex sync item state", `UPDATE plex_sync_item_state SET media_item_id = $2, updated_at = NOW() WHERE media_item_id = $1`},
|
|
{"move webhook sync item state", `UPDATE webhook_sync_item_state SET media_item_id = $2, updated_at = NOW() WHERE media_item_id = $1`},
|
|
{"move watch together rooms selected", `UPDATE watch_together_rooms SET selected_content_id = $2 WHERE selected_content_id = $1`},
|
|
{"move watch together suggestions", `UPDATE watch_together_suggestions SET content_id = $2 WHERE content_id = $1`},
|
|
{"move admin playback history", `UPDATE playback_history_admin SET media_item_id = $2 WHERE media_item_id = $1`},
|
|
{"move user downloads", `UPDATE user_downloads SET media_item_id = $2 WHERE media_item_id = $1`},
|
|
{"move downloads", `UPDATE downloads SET content_id = $2, updated_at = NOW() WHERE content_id = $1`},
|
|
{"merge embeddings", `
|
|
INSERT INTO media_item_embeddings (media_item_id, embedding, model, canonical_text, created_at, updated_at)
|
|
SELECT $2, embedding, model, canonical_text, created_at, updated_at
|
|
FROM media_item_embeddings
|
|
WHERE media_item_id = $1
|
|
ON CONFLICT (media_item_id) DO NOTHING`},
|
|
{"delete source embeddings", `DELETE FROM media_item_embeddings WHERE media_item_id = $1`},
|
|
{"merge item localizations", `
|
|
INSERT INTO media_item_localizations (
|
|
content_id, language, title, sort_title, overview, tagline, poster_path, poster_thumbhash,
|
|
backdrop_path, backdrop_thumbhash, logo_path, created_at, updated_at
|
|
)
|
|
SELECT $2, language, title, sort_title, overview, tagline, poster_path, poster_thumbhash,
|
|
backdrop_path, backdrop_thumbhash, logo_path, created_at, NOW()
|
|
FROM media_item_localizations
|
|
WHERE content_id = $1
|
|
ON CONFLICT (content_id, language) DO NOTHING`},
|
|
{"delete source localizations", `DELETE FROM media_item_localizations WHERE content_id = $1`},
|
|
{"dedupe people", `
|
|
DELETE FROM item_people src
|
|
USING item_people dest
|
|
WHERE src.content_id = $1
|
|
AND dest.content_id = $2
|
|
AND dest.person_id = src.person_id
|
|
AND dest.kind = src.kind
|
|
AND dest.character = src.character`},
|
|
{"move people", `UPDATE item_people SET content_id = $2 WHERE content_id = $1`},
|
|
{"merge refresh debt", `
|
|
INSERT INTO metadata_refresh_debt (
|
|
target_type, content_id, priority, reason_mask, next_refresh_at,
|
|
claimed_at, lease_expires_at, last_attempt_at, last_success_at,
|
|
attempt_count, last_error, updated_at
|
|
)
|
|
SELECT target_type, $2, priority, reason_mask, next_refresh_at,
|
|
claimed_at, lease_expires_at, last_attempt_at, last_success_at,
|
|
attempt_count, last_error, updated_at
|
|
FROM metadata_refresh_debt
|
|
WHERE target_type = 'item' AND content_id = $1
|
|
ON CONFLICT (target_type, content_id) DO UPDATE
|
|
SET priority = GREATEST(metadata_refresh_debt.priority, EXCLUDED.priority),
|
|
reason_mask = metadata_refresh_debt.reason_mask | EXCLUDED.reason_mask,
|
|
next_refresh_at = LEAST(metadata_refresh_debt.next_refresh_at, EXCLUDED.next_refresh_at),
|
|
attempt_count = GREATEST(metadata_refresh_debt.attempt_count, EXCLUDED.attempt_count),
|
|
last_error = COALESCE(NULLIF(metadata_refresh_debt.last_error, ''), EXCLUDED.last_error),
|
|
updated_at = NOW()`},
|
|
{"delete source refresh debt", `DELETE FROM metadata_refresh_debt WHERE target_type = 'item' AND content_id = $1`},
|
|
{"merge stale media ids", `
|
|
INSERT INTO stale_media_ids (content_id, provider, provider_id, first_seen_at, last_seen_at)
|
|
SELECT $2, provider, provider_id, first_seen_at, last_seen_at
|
|
FROM stale_media_ids
|
|
WHERE content_id = $1
|
|
ON CONFLICT (content_id, provider) DO UPDATE
|
|
SET provider_id = COALESCE(NULLIF(stale_media_ids.provider_id, ''), EXCLUDED.provider_id),
|
|
first_seen_at = LEAST(stale_media_ids.first_seen_at, EXCLUDED.first_seen_at),
|
|
last_seen_at = GREATEST(stale_media_ids.last_seen_at, EXCLUDED.last_seen_at)`},
|
|
{"delete source stale media ids", `DELETE FROM stale_media_ids WHERE content_id = $1`},
|
|
{"move watch provider history exports", `UPDATE watch_provider_history_exports SET media_item_id = $2, updated_at = NOW() WHERE media_item_id = $1`},
|
|
{"move watch provider scrobble sessions", `UPDATE watch_provider_scrobble_sessions SET media_item_id = $2, updated_at = NOW() WHERE media_item_id = $1`},
|
|
{"merge watch provider favorites by provider key", `
|
|
UPDATE watch_provider_favorite_items dest
|
|
SET remote_present = dest.remote_present OR src.remote_present,
|
|
local_present = dest.local_present OR src.local_present,
|
|
last_seen_remote_at = GREATEST(dest.last_seen_remote_at, src.last_seen_remote_at),
|
|
last_seen_local_at = GREATEST(dest.last_seen_local_at, src.last_seen_local_at),
|
|
last_exported_at = GREATEST(dest.last_exported_at, src.last_exported_at),
|
|
last_error = COALESCE(NULLIF(dest.last_error, ''), src.last_error),
|
|
updated_at = NOW()
|
|
FROM watch_provider_favorite_items src
|
|
WHERE src.media_item_id = $1
|
|
AND dest.media_item_id = $2
|
|
AND src.connection_id = dest.connection_id
|
|
AND src.provider_item_key <> ''
|
|
AND src.provider_item_key = dest.provider_item_key`},
|
|
{"merge watch provider favorites by connection", `
|
|
UPDATE watch_provider_favorite_items dest
|
|
SET provider_item_key = CASE
|
|
WHEN dest.provider_item_key = '' THEN src.provider_item_key
|
|
ELSE dest.provider_item_key
|
|
END,
|
|
kind = COALESCE(NULLIF(dest.kind, ''), src.kind),
|
|
title = COALESCE(NULLIF(dest.title, ''), src.title),
|
|
year = CASE WHEN dest.year = 0 THEN src.year ELSE dest.year END,
|
|
remote_present = dest.remote_present OR src.remote_present,
|
|
local_present = dest.local_present OR src.local_present,
|
|
last_seen_remote_at = GREATEST(dest.last_seen_remote_at, src.last_seen_remote_at),
|
|
last_seen_local_at = GREATEST(dest.last_seen_local_at, src.last_seen_local_at),
|
|
last_exported_at = GREATEST(dest.last_exported_at, src.last_exported_at),
|
|
last_error = COALESCE(NULLIF(dest.last_error, ''), src.last_error),
|
|
updated_at = NOW()
|
|
FROM watch_provider_favorite_items src
|
|
WHERE src.media_item_id = $1
|
|
AND dest.media_item_id = $2
|
|
AND src.connection_id = dest.connection_id`},
|
|
{"delete duplicate watch provider favorites", `
|
|
DELETE FROM watch_provider_favorite_items src
|
|
USING watch_provider_favorite_items dest
|
|
WHERE src.media_item_id = $1
|
|
AND dest.media_item_id = $2
|
|
AND src.connection_id = dest.connection_id
|
|
AND (src.provider_item_key = dest.provider_item_key OR dest.media_item_id = $2)`},
|
|
{"move remaining watch provider favorites", `UPDATE watch_provider_favorite_items SET media_item_id = $2, updated_at = NOW() WHERE media_item_id = $1`},
|
|
}
|
|
|
|
func canonicalizeMediaItemReferencesTx(ctx context.Context, tx pgx.Tx, sourceID, canonicalID string) error {
|
|
if sourceID == "" || canonicalID == "" || sourceID == canonicalID {
|
|
return nil
|
|
}
|
|
if err := ensureSeriesCanMove(ctx, tx, sourceID, canonicalID); err != nil {
|
|
return err
|
|
}
|
|
for _, step := range mediaItemMergeSteps {
|
|
if _, err := tx.Exec(ctx, step.sql, mergeStepArgs(step.sql, sourceID, canonicalID)...); err != nil {
|
|
return fmt.Errorf("%s: %w", step.name, err)
|
|
}
|
|
}
|
|
return nil
|
|
}
|
|
|
|
func ensureSeriesCanMove(ctx context.Context, tx pgx.Tx, sourceID, canonicalID string) error {
|
|
var sourceType string
|
|
if err := tx.QueryRow(ctx, `SELECT type FROM media_items WHERE content_id = $1`, sourceID).Scan(&sourceType); err != nil {
|
|
if errors.Is(err, pgx.ErrNoRows) {
|
|
return fmt.Errorf("source media item %s not found", sourceID)
|
|
}
|
|
return fmt.Errorf("loading source item type: %w", err)
|
|
}
|
|
if sourceType != "series" {
|
|
return nil
|
|
}
|
|
var conflictCount int
|
|
if err := tx.QueryRow(ctx, `
|
|
SELECT COUNT(*)
|
|
FROM seasons src
|
|
JOIN seasons dest
|
|
ON dest.series_id = $2
|
|
AND dest.season_number = src.season_number
|
|
WHERE src.series_id = $1
|
|
`, sourceID, canonicalID).Scan(&conflictCount); err != nil {
|
|
return fmt.Errorf("checking duplicate season conflicts: %w", err)
|
|
}
|
|
if conflictCount > 0 {
|
|
return fmt.Errorf("series %s has %d season conflicts with canonical series %s", sourceID, conflictCount, canonicalID)
|
|
}
|
|
return nil
|
|
}
|