Files
silo-server/internal/scanner/scan_state.go
T
CoffeeKnyteandGitHub 021e54a03c fix(scanner): fix slow-scan regressions from #319 and #322 (#341)
* fix(scanner): stop classifying "other" content folders as extras

Regression from #322 (trailers and extras for movies and series), which
introduced the extrasDirKinds map.

The extras directory classifier mapped the generic labels "other" and
"others" to ExtraKindOther. These are not part of the Jellyfin/Plex extras
folder convention the map claims to mirror, and they collide with real
content-scope folder names.

A library organized as "movies/other/<Title (year) {ids}>/<file>" tripped
the depth-2 ancestor lookup in classifyExtraPath: every title two levels
under the scope folder "other" was classified as an "other"-kind extra. Such
files are partitioned out of primary root/group inference and matching, then
deferred in processExtraFiles because their parent cannot resolve (they are
the primary titles, not children of one). The result on one deployment was
~10k movies under a folder named "other" funneled through the slow extras
path every scan (parent-unresolved deferrals at ~9.5/s), stalling the scan
and freezing that scope for new/changed primary content.

Remove "other"/"others" from extrasDirKinds. The ExtraKindOther kind stays
reachable through genuine convention labels (extra/extras/interviews/
scenes/shorts). Add regression coverage asserting titles under a scope folder
named other/others stay primary.

* perf(scanner): rewrite identity-only changes without re-probing

A pure identity/grouping change on an already-probed file — a
root_assignment_changed or group_assignment_changed reason with nothing
else — used to fall into the full update branch, which unconditionally ran
ffprobe (probeFile) and then upserted every column, including probe columns,
from the freshly built row. When a group-key or root scheme changes across
the library (see #319), this reprobed nearly every file on the next scan:
an incremental scan that normally takes ~1h ran 7h+ as a full-library
ffprobe storm, even though the media bytes were untouched.

Add a metadata-only update path in processFile: when identityOnlyUpdateReasons
reports every reason is a root/group reassignment, rewrite just the derived
identity columns via the new FileRepository.UpdateIdentity and skip ffprobe,
OSHash, and marker fetch entirely. UpdateIdentity issues a targeted UPDATE of
the root/group/identity and edition/presentation columns only, mirroring
Upsert's column handling, and leaves probe data, file bytes/mtime/hash,
subtitles, chapters, markers, and content/episode/extra linkage intact. The
stored group key converges to the recomputed value on the next scan, so the
file takes the unchanged fast-path thereafter — without a probe storm.

The shared identity-column population is extracted into populateScanIdentity
so the full path and the metadata-only path stay in lockstep.

Verification: unit test for the identityOnlyUpdateReasons classifier; a
DB-backed test (skipped without SILO_TEST_DATABASE_URL) asserting UpdateIdentity
rewrites grouping while preserving probe/linkage columns; the UPDATE statement
was also exercised against the live schema inside a rolled-back transaction.

* fix(scanner): harden identity fast path and extras scope classification

Review follow-ups for the two scan-regression fixes on this branch,
addressing both Codex review comments on PR #341 plus adversarial-review
findings.

Identity fast path (processFile/UpdateIdentity):

- Gate the metadata-only path on existing.ExtraID == "": a row still
  linked as an extra reaching processFile is being reclassified as
  primary, and only the full upsert clears extra linkage; UpdateIdentity
  would have frozen it out of matching forever (match backlog filters
  extra_id IS NULL).
- Gate on existing.FileHash != "": the full path backfills the OSHash
  and fetches hash-keyed S3 intro/credits markers, which no later scan
  reason would repair; hash-less legacy rows now take the full path once
  instead of silently losing that repair channel. file_hash is added to
  the scan-state row shape to support the gate.
- Clear match_suppressed_at like every other scan write, so files with
  fresh identity re-enter the match backlog (suppression is documented
  as lasting "until retried or seen by a new scan").
- Write media_folder_id, mirroring Upsert's ON CONFLICT reassignment.
- Return ErrFileNotFound when the row vanished mid-scan (concurrent
  delete) and fall through to the full upsert path instead of surfacing
  a per-file scan error.
- Return only the row id instead of RETURNING all ~75 columns: the fast
  path fires once per file during library-wide grouping migrations, and
  dragging the track/chapter JSONB payloads along for a million rows
  dominated the cost of the path built to be cheap.
- Extract identityColumnDefaults shared by Upsert and UpdateIdentity so
  the defaulting rules cannot drift, and drop the no-op editionConfidence
  indirection copied between them.
- Use populateScanIdentity in the new-file insert path too; it still
  carried a verbatim copy of the extracted block (with a provably dead
  existingByPath lookup).

Extras classification:

- Restore "other" to extrasDirKinds: it is part of both the documented
  Jellyfin and Plex extras-folder conventions (the removed-label fix
  overshot and broke "movies/<Title>/Other/<file>" libraries, ingesting
  their extras as bogus primary titles). "others" stays removed - it is
  in neither convention.
- Replace label removal with the structural guard the PR had deferred:
  classifyExtraPath now rejects a supplemental-named directory sitting
  at library-scope depth (the dir, any supplemental ancestor, or the
  first non-supplemental ancestor is a configured library root). This
  fixes the original "movies/other/<Title>" defer-storm generically,
  covering every convention label (shorts, scenes, extras, ...) used as
  a content-scope folder.
- Scope extras parent binding by folder.Paths instead of the walk roots,
  so a subtree scan targeting a single movie folder still binds that
  movie's own extras instead of deferring them.

Tests: eligibility-gate unit tests, scope-guard classifier cases
(convention Other/ inside a title binds; scope-level other/shorts stay
primary), and the DB-backed UpdateIdentity test now also covers folder
moves, suppression clearing, and ErrFileNotFound. Full scanner suite ran
green against a migrated scratch PostgreSQL 17 container.

* refactor(scanner): simplify extras scope guard to title-folder rule

Replace the ancestor-walking supplementalDirAtScopeDepth loop with the
plain rule it was approximating: a convention-named directory counts as
an extras dir only when it sits inside a title folder — it must not be a
configured library root or directly under one. Same outcome for the
layouts that matter (movies/other/<Title> stays primary, <Title>/Other
classifies), less machinery.

* test(scanner): assert all rewritten identity columns in UpdateIdentity test

* fix(scanner): make extras scope classification structure-aware

The title-folder rule from 53632022 anchored on library roots, so it
missed both directions: chained convention dirs at the root
("movies/extras/behind the scenes/clip.mkv") classified as extras with
an unresolvable parent (deferred forever), and category folders nested
below the root ("movies/4K/other/<Title>/") still misclassified their
titles.

Replace the root-distance heuristic with the structural property that
actually distinguishes the two cases: a convention-named directory only
counts as an extras dir when its owner (first non-supplemental
ancestor) is a title folder — a directory that holds media of its own.
The new extrasClassifier derives that from the scan's walked path list
(no extra I/O): movie folders must hold a file directly beside the
extras dir; series folders may hold episodes one level down in season
folders (media hiding inside a folder's own extras dirs doesn't count).
Library roots never qualify. Watch-event scans, which have no walked
list, probe ownership with bounded os.ReadDir instead.

This handles title folders at any depth below the root and keeps
scope/category folders primary at any depth, with two known edges: a
title folder holding only extras (its media file missing) stays primary
until the file appears, and a mixed dir holding both loose media and a
category folder degrades to deferral, never wrong linkage.

resolveExtraParent's inline supplemental-chain walk is extracted into
the shared firstNonSupplementalAncestor.
2026-07-08 11:18:07 -04:00

322 lines
9.4 KiB
Go

package scanner
import (
"context"
"encoding/json"
"fmt"
"time"
"github.com/Silo-Server/silo-server/internal/models"
"github.com/jackc/pgx/v5"
)
// scanStateFile is the lightweight media_files row shape used by library scans.
// It intentionally excludes large JSON payloads like tracks and chapters.
type scanStateFile struct {
ID int
ContentID string
ExtraID string
CanonicalRootPath string
ObservedRootPath string
ContentGroupKey string
GroupKeyVersion int
BaseTitle string
BaseYear int
BaseType string
IdentityConfidence string
IdentityJSON []byte
FilePath string
FileSize int64
FileModifiedAt *time.Time
FileHash string
CodecVideo string
CodecAudio string
Resolution string
Container string
Duration int
EditionRaw string
EditionKey string
EditionConfidence *float64
EditionSource string
PresentationKind string
PresentationGroupKey string
PresentationPartIndex int
MultiEpisodeStart int
MultiEpisodeEnd int
ProbeSource string
ProbeUpdatedAt *time.Time
MissingSince *time.Time
HasVideoTracks bool
HasAudioTracks bool
HasChapters bool
ExternalSubtitlePaths []string
}
const scanStateColumns = `id, content_id, extra_id,
canonical_root_path, observed_root_path, content_group_key, group_key_version,
base_title, base_year, base_type, identity_confidence, identity_json,
file_path, file_size, file_modified_at, file_hash,
codec_video, codec_audio, resolution, container, duration,
edition_raw, edition_key, edition_confidence, edition_source,
presentation_kind, presentation_group_key, presentation_part_index,
multi_episode_start, multi_episode_end,
probe_source, probe_updated_at, missing_since,
COALESCE(jsonb_typeof(video_tracks) = 'array' AND jsonb_array_length(video_tracks) > 0, FALSE) AS has_video_tracks,
COALESCE(jsonb_typeof(audio_tracks) = 'array' AND jsonb_array_length(audio_tracks) > 0, FALSE) AS has_audio_tracks,
chapters IS NOT NULL AS has_chapters,
COALESCE((
SELECT jsonb_agg(path)
FROM (
SELECT elem->>'path' AS path
FROM jsonb_array_elements(COALESCE(external_subtitles, '[]'::jsonb)) AS t(elem)
WHERE btrim(elem->>'path') <> ''
) subtitle_paths
), '[]'::jsonb) AS external_subtitle_paths`
func scanScanStateRow(row pgx.Row) (*scanStateFile, error) {
var state scanStateFile
var contentID *string
var extraID *string
var canonicalRootPath, observedRootPath, contentGroupKey, baseTitle, baseType *string
var groupKeyVersion, baseYear *int
var identityConfidence *string
var identityJSON []byte
var fileModifiedAt *time.Time
var fileHash *string
var codecVideo, codecAudio, resolution, container *string
var duration *int
var editionRaw, editionKey, editionSource *string
var presentationKind, presentationGroupKey *string
var presentationPartIndex, multiEpisodeStart, multiEpisodeEnd *int
var probeSource *string
var externalSubtitlePathsJSON []byte
if err := row.Scan(
&state.ID,
&contentID,
&extraID,
&canonicalRootPath,
&observedRootPath,
&contentGroupKey,
&groupKeyVersion,
&baseTitle,
&baseYear,
&baseType,
&identityConfidence,
&identityJSON,
&state.FilePath,
&state.FileSize,
&fileModifiedAt,
&fileHash,
&codecVideo,
&codecAudio,
&resolution,
&container,
&duration,
&editionRaw,
&editionKey,
&state.EditionConfidence,
&editionSource,
&presentationKind,
&presentationGroupKey,
&presentationPartIndex,
&multiEpisodeStart,
&multiEpisodeEnd,
&probeSource,
&state.ProbeUpdatedAt,
&state.MissingSince,
&state.HasVideoTracks,
&state.HasAudioTracks,
&state.HasChapters,
&externalSubtitlePathsJSON,
); err != nil {
return nil, fmt.Errorf("scanning scan state row: %w", err)
}
if contentID != nil {
state.ContentID = *contentID
}
if extraID != nil {
state.ExtraID = *extraID
}
if canonicalRootPath != nil {
state.CanonicalRootPath = *canonicalRootPath
}
if observedRootPath != nil {
state.ObservedRootPath = *observedRootPath
}
if contentGroupKey != nil {
state.ContentGroupKey = *contentGroupKey
}
if groupKeyVersion != nil {
state.GroupKeyVersion = *groupKeyVersion
}
if baseTitle != nil {
state.BaseTitle = *baseTitle
}
if baseYear != nil {
state.BaseYear = *baseYear
}
if baseType != nil {
state.BaseType = *baseType
}
if identityConfidence != nil {
state.IdentityConfidence = *identityConfidence
}
if len(identityJSON) > 0 {
state.IdentityJSON = append([]byte(nil), identityJSON...)
}
state.FileModifiedAt = fileModifiedAt
if fileHash != nil {
state.FileHash = *fileHash
}
if codecVideo != nil {
state.CodecVideo = *codecVideo
}
if codecAudio != nil {
state.CodecAudio = *codecAudio
}
if resolution != nil {
state.Resolution = *resolution
}
if container != nil {
state.Container = *container
}
if duration != nil {
state.Duration = *duration
}
if editionRaw != nil {
state.EditionRaw = *editionRaw
}
if editionKey != nil {
state.EditionKey = *editionKey
}
if editionSource != nil {
state.EditionSource = *editionSource
}
if presentationKind != nil {
state.PresentationKind = *presentationKind
}
if presentationGroupKey != nil {
state.PresentationGroupKey = *presentationGroupKey
}
if presentationPartIndex != nil {
state.PresentationPartIndex = *presentationPartIndex
}
if multiEpisodeStart != nil {
state.MultiEpisodeStart = *multiEpisodeStart
}
if multiEpisodeEnd != nil {
state.MultiEpisodeEnd = *multiEpisodeEnd
}
if probeSource != nil {
state.ProbeSource = *probeSource
}
if len(externalSubtitlePathsJSON) > 0 {
if err := json.Unmarshal(externalSubtitlePathsJSON, &state.ExternalSubtitlePaths); err != nil {
return nil, fmt.Errorf("unmarshaling external subtitle paths: %w", err)
}
}
return &state, nil
}
func scanScanStateRows(rows pgx.Rows) ([]*scanStateFile, error) {
defer rows.Close()
files := make([]*scanStateFile, 0)
for rows.Next() {
file, err := scanScanStateRow(rows)
if err != nil {
return nil, err
}
files = append(files, file)
}
if err := rows.Err(); err != nil {
return nil, fmt.Errorf("iterating scan state rows: %w", err)
}
return files, nil
}
// GetScanStateByFolder returns the lightweight scan-state rows for a folder.
func (r *FileRepository) GetScanStateByFolder(ctx context.Context, folderID int) ([]*scanStateFile, error) {
query := `SELECT ` + scanStateColumns + ` FROM media_files WHERE media_folder_id = $1 ORDER BY file_path ASC`
rows, err := r.pool.Query(ctx, query, folderID)
if err != nil {
return nil, fmt.Errorf("querying scan state by folder: %w", err)
}
return scanScanStateRows(rows)
}
// GetScanStateByFolderAndPathPrefix returns lightweight scan-state rows for a
// folder subtree.
func (r *FileRepository) GetScanStateByFolderAndPathPrefix(ctx context.Context, folderID int, pathPrefix string) ([]*scanStateFile, error) {
query := `SELECT ` + scanStateColumns + ` FROM media_files
WHERE media_folder_id = $1
AND (file_path = $2 OR file_path LIKE $3 ESCAPE '\')
ORDER BY file_path ASC`
rows, err := r.pool.Query(ctx, query, folderID, pathPrefix, pathPrefixLike(pathPrefix))
if err != nil {
return nil, fmt.Errorf("querying scan state by folder and path prefix: %w", err)
}
return scanScanStateRows(rows)
}
func scanStateFromMediaFile(file *models.MediaFile) *scanStateFile {
if file == nil {
return nil
}
return &scanStateFile{
ID: file.ID,
ContentID: file.ContentID,
ExtraID: file.ExtraID,
CanonicalRootPath: file.CanonicalRootPath,
ObservedRootPath: file.ObservedRootPath,
ContentGroupKey: file.ContentGroupKey,
GroupKeyVersion: file.GroupKeyVersion,
BaseTitle: file.BaseTitle,
BaseYear: file.BaseYear,
BaseType: file.BaseType,
IdentityConfidence: file.IdentityConfidence,
IdentityJSON: append([]byte(nil), file.IdentityJSON...),
FilePath: file.FilePath,
FileSize: file.FileSize,
FileModifiedAt: file.FileModifiedAt,
FileHash: file.FileHash,
CodecVideo: file.CodecVideo,
CodecAudio: file.CodecAudio,
Resolution: file.Resolution,
Container: file.Container,
Duration: file.Duration,
EditionRaw: file.EditionRaw,
EditionKey: file.EditionKey,
EditionConfidence: file.EditionConfidence,
EditionSource: file.EditionSource,
PresentationKind: file.PresentationKind,
PresentationGroupKey: file.PresentationGroupKey,
PresentationPartIndex: file.PresentationPartIndex,
MultiEpisodeStart: file.MultiEpisodeStart,
MultiEpisodeEnd: file.MultiEpisodeEnd,
ProbeSource: file.ProbeSource,
ProbeUpdatedAt: file.ProbeUpdatedAt,
MissingSince: file.MissingSince,
HasVideoTracks: len(file.VideoTracks) > 0,
HasAudioTracks: len(file.AudioTracks) > 0,
HasChapters: file.Chapters != nil,
ExternalSubtitlePaths: externalSubtitlePaths(file.ExternalSubtitles),
}
}
func externalSubtitlePaths(subs []models.ExternalSubtitle) []string {
if len(subs) == 0 {
return nil
}
paths := make([]string, 0, len(subs))
for _, sub := range subs {
if sub.Path != "" {
paths = append(paths, sub.Path)
}
}
return paths
}