Files
silo-server/internal/scanner/audiobook.go
T
RXWatcherandClaude Opus 4.7 14a05bba9f feat(audiobooks): comprehensive UX, scanner, and collections work
Detail page:
- collapse Chapters section behind header toggle (73-chapter pages no
  longer push content below the fold)
- square cover frames throughout (audiobook covers are Audible-style 1:1,
  not 2:3 book portrait)
- narrator picker dropdown when multiple narrations of the same book exist
- clickable genre badges (route to /audiobooks?genre=X)
- reordered so credits/rails sit above the chapter list
- Play-from-Start button forces remount via playToken counter (was a
  no-op when player was already at position 0)
- regression test for the Play-from-Start fix

Mini bar / Now Listening:
- mini bar respects --app-sidebar-offset so it stops getting covered by
  the desktop sidebar
- Now Listening adds overflow scroll + a labeled "Back to player" button
  so controls are never inaccessible on short viewports

Library page:
- infinite scroll (replaces Previous/Next pagination)
- genre filter chip with X-to-clear

Scanner:
- 8-worker parallel reconcile (env SILO_AUDIOBOOK_SCAN_WORKERS to override)
- file-path-first dedup so cleaned titles don't collapse separate
  narrations
- title cleanup at write time (strips "Read by X" / "(unabridged)"
  suffixes; original tag preserved in original_title)
- audiobook_series upsert from tag-derived series_name/series_position
- secondary dedup check (author + narrator + year + duration ±0.5% +
  title-prefix) so two folders of the same book attach to one row

Backend detail handler:
- new fetchAlsoByAuthor, fetchInSeries (with series row > 1 entry guard),
  fetchSimilar (embedding-first with shared-genre fallback),
  fetchOtherNarrations (regex-strips narrator suffix to group siblings)
- audiobookDetailResponse gained also_by_author, in_series,
  similar_audiobooks, other_narrations fields
- list endpoint accepts a genre query param

Embeddings:
- BuildEmbeddingText branches on item.Type == "audiobook" to use
  author/narrator credits instead of cast/director/writer
- mediaTypeLabel helper centralizes movie / "TV series" / audiobook
- ListEmbeddingTextCandidates SQL mirrors the Go branching exactly
- ItemsNeedingEmbedding and TotalMediaItemCount loosen status='matched'
  gate to also include audiobooks (which don't go through TMDB match)
- FindSimilar gains a mediaType filter so cross-type results never appear
- callers in similar.go and personal.go pass the source item's type

Collections:
- MediaAudiobook MediaKind + audiobook(s) case in templateEligibleForLibrary
  (stops offering broken movie/TV templates to audiobook libraries)
- useAddItemToCollection hook (user + admin/library endpoints)
- AddToCollectionDialog wired into audiobook detail and movie/series
  ActionBar overflow menu
- ManualCollectionItemsEditor gained a search-and-add panel with
  debounced live results
- QueryDefinition.media_scope, QuerySortRelevanceScope, ALL_MEDIA_SCOPES
  extended to include "audiobook"
- CatalogFilterBar gained an Audiobooks media scope option
- parseCatalogMediaScope (backend) accepts "audiobook"

Migrations:
- 145_audiobook_series: per-book series_name/series_index with a
  best-effort title-pattern backfill for the existing corpus
- 146_audiobook_title_cleanup: strips narrator suffix / (unabridged)
  noise from existing titles, preserving raw in original_title

Scripts:
- scripts/dedup_audiobooks.py: one-shot merge for "Title" vs
  "Title: Subtitle" duplicates, file-path-stable, dry-run by default

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 19:47:54 +02:00

249 lines
8.3 KiB
Go

package scanner
import (
"context"
"fmt"
"os"
"path/filepath"
"regexp"
"sort"
"strings"
)
// narratorSuffixRE matches a trailing "read by X" / "(Read by X)" /
// "(UK Version: Read by X)" / "- read by X" pattern that some audiobook
// taggers put in the title field. The narrator is already captured as
// item_people kind=8 from the dedicated narrator tag, so removing this
// noise yields a clean human-readable title without losing any data.
//
// The narrator body deliberately excludes the dash so titles like
// "Series Name Read by Foo N - Book Title" don't get mistakenly
// truncated (the "Read by Foo" there is part of the series name, not a
// narrator credit). Real narrator suffixes never contain a `-` after the
// "read by".
var narratorSuffixRE = regexp.MustCompile(`(?i)\s*\(?\s*[-:,]?\s*(UK Version:?|US Version:?)?\s*read by [A-Za-z0-9., '&]+\)?\s*$`)
// unabridgedTokenRE matches a parenthesized "(unabridged)" anywhere in
// the title (sometimes mid-string between series and book). Stripped
// because it's a format marker, not part of the work's name.
var unabridgedTokenRE = regexp.MustCompile(`(?i)\s*\(unabridged\)\s*`)
// collapseSpacesRE squashes any runs of whitespace into a single space.
// Used after the strip passes since removing a mid-string token can
// leave double spaces behind.
var collapseSpacesRE = regexp.MustCompile(`\s+`)
// stripNarratorSuffix removes the narrator-suffix noise and "(unabridged)"
// markers from a title. Returns the input unchanged when no match.
// Kept in sync with the SQL `regexp_replace` used by migration 146 so
// the scanner write path and one-shot backfill produce identical output.
func stripNarratorSuffix(title string) string {
cleaned := narratorSuffixRE.ReplaceAllString(title, "")
cleaned = unabridgedTokenRE.ReplaceAllString(cleaned, " ")
cleaned = collapseSpacesRE.ReplaceAllString(cleaned, " ")
return strings.TrimSpace(cleaned)
}
// parsedAudiobook is the structured output of parseAudiobookFolder.
// The scanner write path (Task 8) converts this into media_items +
// media_files + item_people rows.
type parsedAudiobook struct {
Title string
Author string
Narrator string
Series string
SeriesPosition string
Year int
ASIN string
Overview string
Genres []string
Publisher string
ReleaseDate string
Language string
Files []parsedAudiobookFile
}
// parsedAudiobookFile is one audio file belonging to a parsed audiobook.
// For single-file .m4b audiobooks there is exactly one entry; for
// multi-file folders (Task 7) there is one per file.
type parsedAudiobookFile struct {
Path string
Chapters []ChapterInfo
Duration int // seconds
Bitrate int // kbps
CodecAudio string // aac, mp3, opus, flac
Container string // m4b, mp3, mka, ...
AudioChannels int
}
// parseAudiobookFolder reads a single audiobook folder and returns its
// structured representation. Recognized layouts:
// - one audio file in the folder, optionally with embedded chapters
// - multiple audio files in the folder; each becomes its own
// parsedAudiobookFile with a single synthesized chapter (title =
// filename stem); metadata comes from the first file's tags
//
// Returns an error wrapping os.ErrNotExist when the folder contains zero
// audio files, so the caller can skip it.
func parseAudiobookFolder(ctx context.Context, ffprobePath string, folderPath string) (*parsedAudiobook, error) {
entries, err := os.ReadDir(folderPath)
if err != nil {
return nil, fmt.Errorf("read audiobook folder %s: %w", folderPath, err)
}
var audioFiles []string
for _, entry := range entries {
if entry.IsDir() {
continue
}
if SupportsAudioFile(entry.Name()) {
audioFiles = append(audioFiles, filepath.Join(folderPath, entry.Name()))
}
}
if len(audioFiles) == 0 {
return nil, fmt.Errorf("audiobook folder %s: %w", folderPath, os.ErrNotExist)
}
sort.Strings(audioFiles)
book := &parsedAudiobook{}
if len(audioFiles) == 1 {
probed, err := ProbeFile(ctx, ffprobePath, audioFiles[0])
if err != nil {
return nil, fmt.Errorf("probe audiobook file %s: %w", audioFiles[0], err)
}
book.populateFromTags(probed.FormatTags)
book.Files = []parsedAudiobookFile{{
Path: audioFiles[0],
Chapters: probed.Chapters,
Duration: probed.Duration,
Bitrate: probed.Bitrate,
CodecAudio: probed.CodecAudio,
Container: probed.Container,
AudioChannels: probed.AudioChannels,
}}
return book, nil
}
// Multi-file case: read header from the first file, synthesize one
// chapter per file with title = filename stem.
probedFirst, err := ProbeFile(ctx, ffprobePath, audioFiles[0])
if err != nil {
return nil, fmt.Errorf("probe first audiobook file %s: %w", audioFiles[0], err)
}
book.populateFromTags(probedFirst.FormatTags)
book.Files = make([]parsedAudiobookFile, 0, len(audioFiles))
for i, path := range audioFiles {
stem := strings.TrimSuffix(filepath.Base(path), filepath.Ext(path))
probed, err := ProbeFile(ctx, ffprobePath, path)
if err != nil {
return nil, fmt.Errorf("probe audiobook file %s: %w", path, err)
}
book.Files = append(book.Files, parsedAudiobookFile{
Path: path,
Chapters: []ChapterInfo{{
Index: i,
Title: stem,
StartSeconds: 0,
EndSeconds: float64(probed.Duration),
}},
Duration: probed.Duration,
Bitrate: probed.Bitrate,
CodecAudio: probed.CodecAudio,
Container: probed.Container,
AudioChannels: probed.AudioChannels,
})
}
return book, nil
}
// populateFromTags fills the audiobook's header fields (Title, Author,
// Narrator, Series, Year) from the ffprobe format tags. Tags are
// lower-cased by normalizeFormatTags upstream.
func (b *parsedAudiobook) populateFromTags(tags map[string]string) {
b.Title = firstNonEmpty(tags["title"], tags["album"])
b.Author = firstNonEmpty(tags["album_artist"], tags["artist"], tags["composer"])
b.Narrator = firstNonEmpty(tags["narrator"], tags["performer"], tags["composer"])
b.Series = firstNonEmpty(tags["series"], tags["mvnm"], tags["album"])
b.SeriesPosition = firstNonEmpty(tags["series-part"], tags["mvin"], tags["movement"])
b.ASIN = firstNonEmpty(tags["asin"], tags["audible_asin"], tags["com.audible.asin"])
b.Overview = firstNonEmpty(tags["description"], tags["summary"], tags["comment"], tags["©cmt"])
b.Publisher = firstNonEmpty(tags["publisher"], tags["label"], tags["©pub"])
b.ReleaseDate = firstNonEmpty(tags["releasedate"], tags["release_date"], tags["date"], tags["year"], tags["releasetime"])
b.Language = firstNonEmpty(tags["language"], tags["lang"])
if year := firstNonEmpty(tags["date"], tags["year"], tags["releasetime"]); year != "" {
if y := parseTagYear(year); y > 0 {
b.Year = y
}
}
b.Genres = parseGenresFromTags(tags)
}
// parseGenresFromTags pulls genres from the `genre` tag (which is often
// slash-separated like "Mystery, Thriller & Suspense/Suspense/Fiction")
// plus tmp_Genre1..tmp_Genre5 sub-tags written by some audiobook taggers.
// Returns a deduplicated list in input order.
func parseGenresFromTags(tags map[string]string) []string {
seen := make(map[string]struct{})
var out []string
add := func(g string) {
g = strings.TrimSpace(g)
if g == "" {
return
}
key := strings.ToLower(g)
if _, ok := seen[key]; ok {
return
}
seen[key] = struct{}{}
out = append(out, g)
}
if raw := tags["genre"]; raw != "" {
for _, part := range strings.Split(raw, "/") {
for _, sub := range strings.Split(part, ",") {
add(sub)
}
}
}
for i := 1; i <= 5; i++ {
key := fmt.Sprintf("tmp_genre%d", i)
if v := tags[key]; v != "" {
add(v)
}
}
return out
}
// parseTagYear extracts a 4-digit year (e.g. 1900-9999) from a tag value
// that may be a bare year ("2024"), an ISO date ("2024-05-23"), or a
// padded form ("(2024)"). Returns 0 if no plausible year is found.
func parseTagYear(s string) int {
s = strings.TrimSpace(s)
for i := 0; i+4 <= len(s); i++ {
candidate := s[i : i+4]
if isAllDigits(candidate) {
year := 0
for _, c := range candidate {
year = year*10 + int(c-'0')
}
if year >= 1900 && year <= 9999 {
return year
}
}
}
return 0
}
func isAllDigits(s string) bool {
if s == "" {
return false
}
for _, c := range s {
if c < '0' || c > '9' {
return false
}
}
return true
}