Files
silo-server/internal/rootcheck/rootcheck.go
T
8fc054c15d fix(scanner): never purge files under unreachable library roots (#372)
* fix(scanner): never purge files under unreachable library roots

An unreachable root is not a removed root. When one root of a multi-root
library dies (unmounted share, dead drive) while another root still has
files, the whole-library empty-root guard does not fire — the surviving
root produced files — so the scan marks everything under the dead root
missing_since (desired: hides it from browse/playback) and then, with the
default scanner.empty_trash_after_scan=true + 24h file_removal_grace, the
next scan after the grace hard-deletes every row under the dead root. A
week-long drive outage silently destroys the root's entire catalog state:
probe data, intro/credits markers, file hashes. Worse, membership
reconciliation immediately purges media_items whose only files lived on
the dead root, cascading user collections (library_collection_items has
ON DELETE CASCADE) and deleting cached artwork.

This change makes "temporarily offline" survivable:

- Probe each configured root at scan start (os.Stat + IsDir + ReadDir,
  factored into the new internal/rootcheck package and shared with the
  admin mount-check endpoint). Unreachable roots are skipped by the walk
  but their scopes still reconcile, so files are still marked missing.
- The trash sweep (DeleteMissingByFolder) now excludes rows whose path
  sits under an unreachable root, using the same exact-path + escaped
  prefix-LIKE matching as ListIDsOutsideRoots (a sibling root that merely
  shares a string prefix is never protected). With all roots reachable
  the emitted SQL is unchanged.
- Membership removal still happens — browse/home hide items via
  media_item_libraries, so removal is what keeps a dead-root-only title
  out of the catalog — but the orphan media_items purge exempts items
  whose files sit under an unreachable root. Their metadata, artwork,
  and collection links survive; when the root returns, the upsert clears
  missing_since and syncPresentLibraryState re-inserts the membership,
  restoring the item with zero re-probing or re-matching.
- The folder surfaces scan_warning_code='dead_root' with a message naming
  the unreachable roots; a fully healthy scan or a successful mount check
  clears it, mirroring empty_root. The admin UI shows a badge and banner.
- Deliberate deletion is untouched: removing a path from the library
  config still purges via ListIDsOutsideRoots, files under reachable
  roots keep the exact 24h-grace purge, the empty-root guard and the
  autoscan dead-mount guard are unchanged.

The audiobook/podcast/ebook reconcile paths share the same folder-wide
sweep and orphan purge, so they get the same guard.

Covered by tests: an end-to-end two-root scan (root dies -> rows survive
a zero-grace sweep and warning is set; root returns -> rows resurrect
with their original ids and the warning clears; deleting a file under a
reachable root still purges), repo-level sweep-protection and
sibling-prefix tests, orphan-purge exemption, and rootcheck unit tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(scanner): probe uncompacted roots and take dead-root path on full outage

Review follow-ups: (1) probe every configured path instead of the compacted
traversal roots, so a nested child mount that dies under a reachable parent
is still protected from the sweep; (2) when every configured root is
unreachable, bypass the empty-root confirm flow (without consuming the
one-time cleanup allowance), mark files missing, and raise dead_root instead
of empty_root; (3) dead_root warning banner no longer shows empty-root
confirm-deletion guidance as its fallback hint.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(scanner): simplify dead-root protection plumbing

- extract pathscope.CoverageClauses as the single builder for the
  exact-path + escaped prefix-LIKE root predicate; scanner's
  rootCoverageClauses delegates to it and catalog's
  excludeOrphansUnderProtectedPrefixes reuses it instead of hand-rolling
  the same clause loop
- extract Scanner.sweepMissingAndReconcile to replace the identical
  trash-sweep + membership-reconcile + S3-image-cleanup block that was
  triplicated across the audiobook, ebook, and podcast scans (callers
  keep their flavor-specific log lines so messages stay constant)
- add unreachableConfiguredRoots helper for the repeated
  probeUnreachableRoots(ctx, folder.ID, cleanScanRoots(folder.Paths))
  expression in scanPaths and ScanFile
- drop the unread Path field from rootcheck.Result
- move the dead/empty-root warning text constants in AdminLibraries.tsx
  out of the middle of the import block

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scanner): close dead-root protection gaps found in review

Remediates the confirmed findings from the deep review of this PR:

- Scoped audiobook scans (autoscan file events, subtree scans) ran the
  folder-wide sweep while probing only the scoped clone's Paths, so a
  healthy-subtree event could hard-delete a dead sibling root's rows.
  sweepMissingAndReconcile now reloads the folder's configured roots
  from the DB and probes them uncompacted, which also protects nested
  child mounts in the audiobook/ebook/podcast reconcilers.

- A lost mount that leaves an empty, stat-able mountpoint probed as
  reachable and kept the historical purge timeline. A reachable root
  that is a literally empty directory while cataloged rows remain under
  it is now treated as suspect: rows are only marked missing, the sweep
  and orphan purge exempt it, dead_root is raised, and the mount-check
  endpoint reports it (additive suspect_empty field) instead of
  clearing the warning. Arming the one-time empty-cleanup allowance
  completes the deletion, including in the mixed case where other
  roots are healthy. Roots that still have directory entries keep the
  historical grace-then-purge path.

- Confirmed empty cleanup (allow_empty_cleanup_once) no longer
  force-deletes rows under probe-dead roots: an outage is not a
  confirmation, so a dead sibling root's catalog survives a confirmed
  cleanout of a reachable empty root.

- Root probes are now bounded (rootcheck.ProbeWithTimeout, 5s): a hung
  network mount degrades into the protected unreachable path with a
  probe_timeout error code instead of stalling every scan of the
  folder indefinitely.

- Documented the cross-library limitation of the orphan-purge
  exemption next to the query it applies to.

All behavior is pinned by new DB-backed tests (suspect-empty
protection + confirmed completion, confirmed-cleanup dead-root
survival, scoped/nested-root sweep protection, suspect-root query,
probe timeout).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scanner): address dead-root review findings

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-16 13:58:44 -04:00

224 lines
5.9 KiB
Go

// Package rootcheck probes library root paths for reachability: a root is
// reachable when it exists, is a directory, and can be listed. It is shared
// by the scanner's dead-root protection and the admin mount-check endpoint so
// both agree on what "unreachable" means.
package rootcheck
import (
"context"
"errors"
"io"
"os"
"sync"
"time"
)
// Error codes reported by Probe. They are part of the admin mount-check API
// response contract.
const (
ErrCodeNotFound = "not_found"
ErrCodePermissionDenied = "permission_denied"
ErrCodeNotDirectory = "not_directory"
ErrCodeReadFailed = "read_failed"
ErrCodeStatFailed = "stat_failed"
ErrCodeTimeout = "probe_timeout"
)
// DefaultProbeTimeout bounds how long a single probe may block. A dead mount
// usually errors within milliseconds, but a hung network filesystem
// (hard-mounted NFS, wedged SMB/FUSE) blocks stat/readdir indefinitely —
// probes run on scan and request hot paths, so a hung mount must degrade
// into "unreachable" rather than stall the caller.
const DefaultProbeTimeout = 5 * time.Second
// DefaultProbeConcurrency bounds the number of distinct roots started by one
// batch. Repeated probes for the same path are also coalesced process-wide.
const DefaultProbeConcurrency = 8
// Result describes the outcome of probing a single root path.
type Result struct {
Reachable bool
// Empty is set for a reachable directory with zero entries. A completely
// empty root is the on-disk signature of a lost mount (the mountpoint
// directory remains, its contents vanished with the mount), which a
// reachability check alone cannot detect.
Empty bool
ErrorCode string // empty when Reachable
ErrorMessage string // empty when Reachable
}
// Probe checks that path exists, is a directory, and can be listed.
func Probe(path string) Result {
res := Result{Reachable: true}
info, err := os.Stat(path)
switch {
case err != nil:
res.Reachable = false
res.ErrorCode, res.ErrorMessage = classify(err, false)
case !info.IsDir():
res.Reachable = false
res.ErrorCode, res.ErrorMessage = ErrCodeNotDirectory, "Path is not a directory"
default:
dir, err := os.Open(path)
if err == nil {
_, err = dir.Readdirnames(1)
_ = dir.Close()
}
if errors.Is(err, io.EOF) {
res.Empty = true
} else if err != nil {
res.Reachable = false
res.ErrorCode, res.ErrorMessage = classify(err, true)
}
}
return res
}
type probeCall struct {
done chan struct{}
result Result
}
type probeCoordinator struct {
mu sync.Mutex
inFlight map[string]*probeCall
}
func newProbeCoordinator() *probeCoordinator {
return &probeCoordinator{inFlight: make(map[string]*probeCall)}
}
var sharedProbes = newProbeCoordinator()
// ProbeWithTimeout runs Probe but gives up once timeout elapses or ctx is
// done, reporting the root unreachable with ErrCodeTimeout. Concurrent calls
// for the same path share one underlying syscall so a wedged mount cannot
// accumulate one blocked goroutine per scan or mount check.
func ProbeWithTimeout(ctx context.Context, path string, timeout time.Duration) Result {
return sharedProbes.probe(ctx, path, timeout, func(path string) Result { return Probe(path) })
}
func (c *probeCoordinator) probe(
ctx context.Context,
path string,
timeout time.Duration,
probe func(string) Result,
) Result {
if timeout <= 0 {
timeout = DefaultProbeTimeout
}
c.mu.Lock()
call := c.inFlight[path]
if call == nil {
call = &probeCall{done: make(chan struct{})}
c.inFlight[path] = call
go func() {
call.result = probe(path)
close(call.done)
c.mu.Lock()
delete(c.inFlight, path)
c.mu.Unlock()
}()
}
c.mu.Unlock()
return awaitProbe(ctx, timeout, call.done, func() Result { return call.result })
}
// ProbeManyWithTimeout probes paths concurrently while preserving input order.
func ProbeManyWithTimeout(ctx context.Context, paths []string, timeout time.Duration) []Result {
return probeMany(ctx, paths, timeout, DefaultProbeConcurrency, ProbeWithTimeout)
}
func probeMany(
ctx context.Context,
paths []string,
timeout time.Duration,
limit int,
probe func(context.Context, string, time.Duration) Result,
) []Result {
results := make([]Result, len(paths))
if len(paths) == 0 {
return results
}
if limit <= 0 || limit > len(paths) {
limit = len(paths)
}
jobs := make(chan int)
var workers sync.WaitGroup
workers.Add(limit)
for range limit {
go func() {
defer workers.Done()
for i := range jobs {
results[i] = probe(ctx, paths[i], timeout)
}
}()
}
for i := range paths {
jobs <- i
}
close(jobs)
workers.Wait()
return results
}
func probeBounded(ctx context.Context, timeout time.Duration, probe func() Result) Result {
if timeout <= 0 {
timeout = DefaultProbeTimeout
}
done := make(chan Result, 1)
go func() { done <- probe() }()
var ctxDone <-chan struct{}
if ctx != nil {
ctxDone = ctx.Done()
}
timer := time.NewTimer(timeout)
defer timer.Stop()
select {
case res := <-done:
return res
case <-ctxDone:
case <-timer.C:
}
return timeoutResult()
}
func awaitProbe(ctx context.Context, timeout time.Duration, done <-chan struct{}, result func() Result) Result {
var ctxDone <-chan struct{}
if ctx != nil {
ctxDone = ctx.Done()
}
timer := time.NewTimer(timeout)
defer timer.Stop()
select {
case <-done:
return result()
case <-ctxDone:
case <-timer.C:
}
return timeoutResult()
}
func timeoutResult() Result {
return Result{
Reachable: false,
ErrorCode: ErrCodeTimeout,
ErrorMessage: "Probe timed out; filesystem is not responding",
}
}
func classify(err error, isRead bool) (string, string) {
switch {
case errors.Is(err, os.ErrNotExist):
return ErrCodeNotFound, "Path does not exist"
case errors.Is(err, os.ErrPermission):
return ErrCodePermissionDenied, "Permission denied"
case isRead:
return ErrCodeReadFailed, "Failed to read directory"
default:
return ErrCodeStatFailed, "Failed to stat path"
}
}