Files
silo-server/migrations/146_audiobook_title_cleanup.up.sql
T
RXWatcherandClaude Opus 4.7 14a05bba9f feat(audiobooks): comprehensive UX, scanner, and collections work
Detail page:
- collapse Chapters section behind header toggle (73-chapter pages no
  longer push content below the fold)
- square cover frames throughout (audiobook covers are Audible-style 1:1,
  not 2:3 book portrait)
- narrator picker dropdown when multiple narrations of the same book exist
- clickable genre badges (route to /audiobooks?genre=X)
- reordered so credits/rails sit above the chapter list
- Play-from-Start button forces remount via playToken counter (was a
  no-op when player was already at position 0)
- regression test for the Play-from-Start fix

Mini bar / Now Listening:
- mini bar respects --app-sidebar-offset so it stops getting covered by
  the desktop sidebar
- Now Listening adds overflow scroll + a labeled "Back to player" button
  so controls are never inaccessible on short viewports

Library page:
- infinite scroll (replaces Previous/Next pagination)
- genre filter chip with X-to-clear

Scanner:
- 8-worker parallel reconcile (env SILO_AUDIOBOOK_SCAN_WORKERS to override)
- file-path-first dedup so cleaned titles don't collapse separate
  narrations
- title cleanup at write time (strips "Read by X" / "(unabridged)"
  suffixes; original tag preserved in original_title)
- audiobook_series upsert from tag-derived series_name/series_position
- secondary dedup check (author + narrator + year + duration ±0.5% +
  title-prefix) so two folders of the same book attach to one row

Backend detail handler:
- new fetchAlsoByAuthor, fetchInSeries (with series row > 1 entry guard),
  fetchSimilar (embedding-first with shared-genre fallback),
  fetchOtherNarrations (regex-strips narrator suffix to group siblings)
- audiobookDetailResponse gained also_by_author, in_series,
  similar_audiobooks, other_narrations fields
- list endpoint accepts a genre query param

Embeddings:
- BuildEmbeddingText branches on item.Type == "audiobook" to use
  author/narrator credits instead of cast/director/writer
- mediaTypeLabel helper centralizes movie / "TV series" / audiobook
- ListEmbeddingTextCandidates SQL mirrors the Go branching exactly
- ItemsNeedingEmbedding and TotalMediaItemCount loosen status='matched'
  gate to also include audiobooks (which don't go through TMDB match)
- FindSimilar gains a mediaType filter so cross-type results never appear
- callers in similar.go and personal.go pass the source item's type

Collections:
- MediaAudiobook MediaKind + audiobook(s) case in templateEligibleForLibrary
  (stops offering broken movie/TV templates to audiobook libraries)
- useAddItemToCollection hook (user + admin/library endpoints)
- AddToCollectionDialog wired into audiobook detail and movie/series
  ActionBar overflow menu
- ManualCollectionItemsEditor gained a search-and-add panel with
  debounced live results
- QueryDefinition.media_scope, QuerySortRelevanceScope, ALL_MEDIA_SCOPES
  extended to include "audiobook"
- CatalogFilterBar gained an Audiobooks media scope option
- parseCatalogMediaScope (backend) accepts "audiobook"

Migrations:
- 145_audiobook_series: per-book series_name/series_index with a
  best-effort title-pattern backfill for the existing corpus
- 146_audiobook_title_cleanup: strips narrator suffix / (unabridged)
  noise from existing titles, preserving raw in original_title

Scripts:
- scripts/dedup_audiobooks.py: one-shot merge for "Title" vs
  "Title: Subtitle" duplicates, file-path-stable, dry-run by default

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 19:47:54 +02:00

45 lines
1.7 KiB
SQL

-- Clean narrator/edition suffixes out of audiobook titles. The narrator
-- is already captured separately in item_people (kind=8) from the file's
-- narrator tag, so leaving it in the title field is duplicate data that
-- visibly clutters the UI ("Cytonic (UK Version: Read by Sophie Aldred)"
-- → "Cytonic"). The raw title is preserved in original_title for
-- forensic reference.
--
-- The regex must stay in lockstep with the Go scanner.stripNarratorSuffix
-- helper. Both run case-insensitive and target a trailing block of the
-- form: optional separator + optional "(UK Version:|US Version:|...|)"
-- + "read by X" + optional close-paren.
WITH cleaned AS (
SELECT mi.content_id,
mi.title AS raw_title,
trim(regexp_replace(
regexp_replace(
regexp_replace(
mi.title,
E'\\s*\\(?\\s*[-:,]?\\s*(UK Version:?|US Version:?)?\\s*[Rr]ead [Bb]y [A-Za-z0-9., ''&]+\\)?\\s*$',
'',
'g'
),
E'\\s*\\(unabridged\\)\\s*', ' ', 'gi'
),
E'\\s+', ' ', 'g'
)) AS clean_title
FROM media_items mi
WHERE mi.type = 'audiobook'
AND (mi.title ~* '\s+read by ' OR mi.title ~* '\(unabridged\)')
)
UPDATE media_items mi
SET
title = c.clean_title,
sort_title = LOWER(c.clean_title),
original_title = CASE
WHEN COALESCE(mi.original_title, '') = '' THEN c.raw_title
ELSE mi.original_title
END,
updated_at = NOW()
FROM cleaned c
WHERE mi.content_id = c.content_id
AND c.clean_title <> ''
AND c.clean_title <> c.raw_title;