- service.go: reject trailing data after the decoded manifest object.
Decoder.More() only reports array/object iteration, so a stray closing
delimiter (e.g. {...}}) slipped through where json.Unmarshal used to
reject it. Require the stream to reach io.EOF after decoding on both
the received and embedded sides; add a regression test.
- repo.go: split the list projection from cleanup. reportListSelectSQL
keeps the app_build JSONB extraction for the admin list; new
reportCleanupSelectSQL omits it so retention/stale batches don't touch
each candidate's manifest JSONB just to delete a row.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
- AdminDiagnostics list: fix regression where rows dereferenced the
now-omitted manifest for app_build. Project app_build server-side out
of manifest JSONB into both list and detail responses (cheap
COALESCE(manifest->'report'->>'app_build','')), split the TS type into
DiagnosticReportSummary (list, no manifest) and DiagnosticReport
(detail, with manifest), and read report.app_build in the row/detail.
- embeddedManifestMatches: decode with json.Decoder + UseNumber so large
integers above 2^53 (e.g. log_summary.lines) can't collapse to the same
float and falsely match; re-assert no-trailing-data strictness.
- Quota reservation (SKIP): reserving the client-claimed archive.bytes is
sound because archiveMatches requires claimed==actual before MarkReady,
so no stored report exceeds its reservation; documented in a code comment.
- Multipart parts: reject a wrong-name/wrong-content-type part without
calling part.Close(), which would drain up to the bundle limit while
holding the in-flight slot; abandon it so malformed uploads fail promptly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
- service: reject supplied child-profile attribution with a distinct
ErrChildProfileForbidden (403 child_profile_forbidden) instead of
silently dropping it as if the profile were not found; a profile that
is simply not the user's still drops attribution unchanged
- repo: add a manifest-free list projection (reportListSelectSQL /
scanReportSummary) for admin list and retention/stale cleanup queries
so they no longer drag the full manifest JSONB per row; keep the full
projection for GetByID/DeleteByID and mark Manifest omitempty
- cleanup: delete/mark the DB row before the blob in retention and stale
loops so a mid-run DB failure can't leave a ready report pointing at a
missing bundle; blob-delete failures are logged with bucket/keys for
orphan cleanup to reap rather than aborting the run (shared helper with
the admin DeleteReport path)
- admin: reject diagnostics settings where max_bytes_per_user would fall
below max_bundle_bytes (and the reciprocal), which would make every
max-size upload fail quota
- router/demo: route POST /diagnostics/reports through DemoGuard and block
the reports prefix in demo mode while keeping GET /diagnostics/status
available
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
- schema: add crash/report.type conditionals (allOf if/then) so a
crash/anr/native_crash/hang/abnormal_exit manifest requires `crash`
and a `manual` manifest forbids it, matching ValidateManifest.
- service: reject uploads where X-Profile-Id and manifest.report.profile_id
are both present but differ (new ErrProfileMismatch, mapped to 400
profile_mismatch) instead of silently preferring the header; single-source
and matching cases unchanged. Adds service tests for mismatch, match, and
header-only attribution.
- schema: require manifest.json as the first archive.entries element via
prefixItems (contains retained for validators without prefixItems support).
- schema: document that maxLength is a character-count bound while the server
enforces UTF-8 byte length, via a top-level note and per-field notes on the
free-text device_summary and crash fields.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
- Extend the upload write deadline alongside the read deadline so a slow
upload finishing after the integrated server's 120s WriteTimeout can still
return its success response instead of timing out a report that succeeded.
- Reject child-profile attribution for diagnostics: wire the attribution
validator through a shared profile lookup that reports IsChild and drop
attribution for child profiles, which must not perform diagnostics actions.
- Assert the download test captures the clicked anchor and checks its blob:
href and silo-diagnostics-<short_id>.tar.gz filename, not just cleanup.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
- settings.go: cap the parsed cleanup interval at 7 days before converting to
time.Duration so a huge configured value can't overflow int64 nanoseconds and
wrap into a tiny/negative interval; add boundary tests.
- settings.go: propagate genuine settings read failures from LoadSettings
(missing/empty -> default, error -> fail) so a transient DB error surfaces
retryably instead of silently reporting uploads disabled or wrong quotas.
- bundle.go: validate non-manifest bundle entries while streaming with bounded
memory -- device.json and crash/*.json must be a single JSON object,
logs.jsonl/breadcrumbs.jsonl must be newline-delimited JSON objects with a
per-line byte cap (new contract.MaxLogLineBytes); binary members stay opaque.
- diagnostics upload handler: extend the read deadline per-route via
http.ResponseController.SetReadDeadline (10m) so slow mobile uploads of large
bundles aren't cut off by the shared 30s server ReadTimeout.
- web admin download: request the ?proxy=1 streaming path directly so downloads
work when S3Private is only server-reachable and errors can surface in-page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
- bundle: reject tar entry names that differ from their trimmed form instead
of normalizing padded names into the allowlist
- repo: reserve expected bytes on receiving rows and count receiving+ready in
the per-user byte quota so concurrent/multi-node uploads can't overshoot
- contract: require the crash object for event report types and keep it absent
for manual; add contract tests
- settings/service: seed diagnostics.server_instance_id atomically via
insert-if-absent and adopt the winning value across nodes
- bundle/service: capture the embedded manifest.json during ValidateBundle and
reject reports whose embedded manifest disagrees with the part-1 manifest
(minus archive); add tests
- admin: delete the DB row before the blob on DeleteReport; log bucket/key when
the blob delete fails instead of leaving a visible report with a missing bundle
- bundle: reject PAX/GNU tar formats and extension records that smuggle bytes
past validation; add a PAX-archive rejection test
- migration: add CHECK constraints for state, report_type, and platform
- docs: add text/jsonc language identifiers to the two unfenced code blocks
- cleanup: log-and-continue per report and aggregate errors so one poisoned
report no longer blocks the whole run; update tests
- tasks: give diagnostics its own cleanup interval key instead of reusing the
opslog key, and bound the startup settings lookup
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
The report detail already rendered app_version (app_build); the list
rows showed only v<app_version>. TestFlight triage needs the build
number at a glance, so list rows now render it from the stored
manifest, e.g. "v1.4.2 (20841)".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
Two validator behaviors made the contract unimplementable for clients
using standard tar libraries:
- Any byte after the tar end-of-archive marker was rejected, but GNU
tar, Python tarfile, and Apache Commons Compress all pad the archive
with zero blocks to a record boundary. Accept up to 64 KiB of zero
padding; any non-zero trailing data is still rejected.
- uncompressed_bytes was computed as the sum of entry payloads, which
no tar-producing client observes. Define it as the total decompressed
tar stream (headers, end-of-archive marker, and padding included) —
the byte count between a client's tar writer and gzip writer, and
what gzip -l reports. Documented in the design doc and contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
pgx binds a nil Go slice as SQL NULL, which bypasses the column's '{}'
default and violates its NOT NULL constraint, so reports without
playback session ids failed to insert. Bind an empty slice instead.
The ingest path also swallowed the underlying insert and bundle
validation errors, logging only a generic rejection reason; both sites
now log the real error so failures are diagnosable from server logs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
PutObjectStream omitted Content-Length because client-reported sizes are
untrusted, but Cloudflare R2 rejects unsized PutObject bodies with
411 MissingContentLength. Route streaming uploads through the SDK's
multipart manager, which buffers fixed-size parts (8 MiB, sequential)
and sends each with a known length, preserving bounded memory for
untrusted stream sizes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh
* ci(docker): publish multi-arch images (linux/amd64 + linux/arm64)
The Dockerfile was already arch-aware (TARGETARCH in the Jellyfin apt
repo); this adds linux/arm64 to the buildx platform list so the pushed
manifest list serves both x86 servers and arm64 hosts (Apple Silicon,
Graviton, Pi 4/5). Verified locally on the arm64 runner host: image
builds and both silo and jellyfin-ffmpeg run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci(docker): move image builds to GitHub-hosted runners
Replaces the single self-hosted job with a per-arch matrix (ubuntu-latest
for amd64, ubuntu-24.04-arm for arm64) so each platform builds natively
with no emulation, pushing by digest, plus a merge job that stitches the
digests into one tagged manifest list. Layer caching moves from the
persistent local builder to type=gha per-platform scopes.
The private-SDK constraint that originally forced self-hosted no longer
applies: silo-plugin-sdk is public and resolves from proxy.golang.org.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* ci(docker): use normalized IMAGE_LC in merge-job metadata step
Not a functional fix — metadata-action lowercases the images input
itself (proven by run 29768620803) — but keeps all image references in
the merge job on the explicit IMAGE_LC form.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(web): collapse per-scan rows on the Libraries page
Bulk single-file scans (e.g. autoscan picking up a season drop) rendered
one full-width row per scan under the library, each with its own stop
button that actually cancelled every scan for the library. With 70+
queued files the table became an unusable wall of paths.
Each library's scans now collapse into one summary line (N running ·
N queued) with a single cancel control, compact rows for the first few
scans (file basename, full path on hover, progress for running ones),
and a Show all expander capped to a scrollable height. The Scan Queue
popover gets the same running-first collapse per library group, shows
basenames instead of full paths, and caps its staggered fade-in delay
so deep rows don't appear seconds late.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): address scan-collapse review feedback
- Keep the 'Entire library' target visible on pathless (full-library)
scan rows in the inline active-work list.
- Drop the redundant '+ N more' button in LibraryScanTasks; the header
'Show all N' toggle is the single expand control there.
- Extract useCollapsedScans so the table rows and Scan Queue popover
share one collapse implementation, and always derive the popover
target from formatActiveScanTarget instead of a hardcoded fallback.
- Update the active-scan progress test for the split mode/target markup.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
List/detail/download/delete for uploaded diagnostic reports at
/admin/diagnostics, beside the Logs viewer: server-filter parity with the
admin API, cursor pagination, manifest + device summary detail sheet,
playback-session links into the filtered logs view, presign/stream
download handling, and feature-state banner. Extracts a shared
authenticated apiResponse() helper used by api() and apiDownload().
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XppCCycoaskCsW7ja1fZct
* fix(userstore): batch allowed-library and profile lookups when listing
Listing profiles cost 1 + P queries: one allowed-libraries lookup fired
per profile row while the cursor was still open, which on Postgres also
checks out a second pooled connection mid-scan. The admin sessions
dashboard makes it worse, calling ListProfiles once per streaming user
just to resolve names. The SQLite store had the same pattern for both
profiles and collections, even though the Postgres collections path was
already written with array_agg to avoid exactly this.
Collect the rows first, then fetch the child lists in one batched query
and stitch them together in Go. No behaviour change, just fewer round
trips.
* fix(userstore): avoid sqlite batch variable limits
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
* fix(api): throttle api_keys last_used_at writes in auth middleware
Every API key request spawned a goroutine that ran an UPDATE on
api_keys, so a key driving HLS segments or a polling integration hit the
table with one write per request, and a stalled database could pile
those goroutines up without bound. The jellycompat authenticator already
guards this same write with a once-per-minute throttle per key; the main
middleware was missing it.
Bring the two in line. Track the last write per key ID and only launch
the update once a minute has passed, with a timeout on the background
write. The map is keyed by key ID so it stays bounded.
* fix(auth): bound API key last-used throttling
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Implements slice 1 of docs/design/2026-07-19-client-diagnostics.md: the
versioned contract (schemas, fixtures, Go validator), storage-validated
diagnostics.uploads_enabled gate, account-scoped status endpoint, hardened
streaming multipart ingest with quota reservation and a receiving/ready/
failed report state machine, S3 streaming puts, acting-admin report API
(list/detail/download/delete with audit events), and the retention +
orphan-reconciliation cleanup task.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XppCCycoaskCsW7ja1fZct
Cross-repo spec and rollout plan for opt-in client crash reporting and
debug-log upload to the user's own Silo server: silo-server ingest, storage,
admin API, and retention; silo-android and silo-apple capture, consent, and
upload. Disabled by default on both server and client.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XppCCycoaskCsW7ja1fZct
* feat(activity): refine play-method tags and add a Jellyfin-client pill
Two related tagging improvements to the admin activity views, squashed:
Split audio transcodes into their own tag. The Play Method summary and
Server Activity popover bucketed every session by its raw play_method,
lumping real video transcodes together with video-copy HLS repackages and
having no separate tag for audio-only transcodes. Classify each session by
the per-stream decisions the backend already reports:
- video re-encoded -> "transcode" (yellow)
- only audio re-encoded -> "audio" (red)
- streams only repackaged -> "remux" (blue, incl. video-copy HLS)
- nothing touched -> "direct" (green)
ordered direct -> remux -> transcode -> audio across the distribution bar,
legend, method filter/sort, the per-row badge, and the Server Activity
stream counts.
Add a Jellyfin-client "JF" pill. Sessions from a Jellyfin-ecosystem client
(Jellyfin Web, Findroid, Swiftfin, Infuse, etc.) get a purple "JF" pill
next to the play-method tag. Detection is UI-only: isJellyfinSession()
positively matches client_name (set from the Jellyfin MediaBrowser auth
header) and then the raw user agent against the known Jellyfin client
tokens, mirroring the server's client-labeling list. The pill is orthogonal
to the method classification — a session can be both "transcode" and JF.
Pure UI/presentation change; no backend behavior changes.
AI-use disclosure: implemented with AI assistance (Claude Code).
* fix(web): cache-control on SPA shell so deploys bust stale UI
The frontend handler served index.html with no cache directives, leaving
freshness to browser/CDN heuristics. A stale index.html at a CDN edge kept
serving old content-hashed bundles, so a client-side hard refresh couldn't
recover — one browser would show the new UI while another showed the old.
Apply the standard SPA cache policy:
- index.html (and SPA-route fallbacks): no-cache + a truncated-SHA-256
ETag, so the shell is cached but revalidated on every load and answers
an unchanged request with a cheap 304.
- /assets/* (Vite content-hashed bundles): public, max-age=31536000,
immutable — cached indefinitely; a new build changes the filename hash,
which busts them automatically.
- other stable-named bundled files (sw.js, icons, fonts): no-cache, so a
changed service worker or icon can't stay stuck in a cache.
Caching is preserved (no no-store anywhere); only the tiny HTML shell is
revalidated, which is what busts a stale UI on deploy.
* fix(activity): compute the method bucket server-side and unify every session surface
Review follow-ups for the play-method tags (PR #387):
- The server now emits effective_play_method (additive field) from the same
per-stream decisions that drive the badges, so all consumers — web, realtime
popover, and the Android/Apple admin views later — agree on the bucket
instead of each client re-reducing raw play_method. Rows with an unknown
play_method (stale rows from older nodes) stay unbucketed rather than being
misreported as audio transcodes off the bare transcode_audio flag; the web
fallback classifier mirrors that and reports "unknown".
- Jellyfin-ecosystem detection moved server-side as is_jellyfin_client, owned
next to the client-labeling rules so the two lists cannot drift; the web
token list is gone. Adds kodi/mpv/delfin/finamp, which reach Silo only
through the Jellyfin compat surface.
- The dashboard stream cards, stats session table, and household streams panel
now use the same classification as the activity page and popover — they
previously showed contradictory tags for the same live session.
- One shared method->label/color table in adminActivityPresentation.ts
replaces the four independent copies (METHOD_META + three switches); the
method column sort now uses the shared cost-order comparator instead of
alphabetical; dead "copy"/"hls" order entries removed and the reachable
"unknown" bucket is styled.
* fix(server): make SPA revalidation RFC-compliant and stop rebuilding the shell per request
Review follow-ups for the SPA cache policy (PR #387):
- Stable-URL bundled files (sw.js, icons, vendor bundles) now carry a content
ETag. The embedded FS has no modtimes, so http.FileServer emits no validator
of its own — no-cache alone forced a full re-download of multi-megabyte
vendor trees on every use because there was nothing to revalidate against.
- Shell and favicon conditional requests go through http.ServeContent, which
implements RFC 9110 If-None-Match semantics (weak comparison, ETag lists).
The previous exact string compare never matched once a fronting proxy
compressed the response and weakened the ETag to W/"...", silently killing
the 304 path in the most common deployment topology.
- The rendered shell (index read + branding render + SHA-256) is cached per
branding snapshot via the new Snapshot.RenderKey instead of being rebuilt on
every request — the 304 revalidation that no-cache makes the common case now
costs two header writes. The misnamed weakContentETag (it emits a strong
validator) is renamed contentETag.
* fix(activity): show the JF pill on every session surface, not just the mobile row
Review comments on PR #387: the JF pill only rendered inside Admin
Activity's sm:hidden mobile row, so the desktop table — and the other
session surfaces that now share the method classification — never
identified Jellyfin-compat sessions.
Extract the pill into a shared JellyfinSessionPill component (renders
nothing for native sessions) and drop it into the Admin Activity desktop
client line, the dashboard stream cards, the household streams panel,
and the stats active-session table.
* fix(playback): sync real encode decisions and client identity for compat transcodes
Review comments on PR #387:
- Jellyfin HLS sessions that copy video and re-encode only audio synced as
full video transcodes: ensureUpstreamPlayback resets transcodeAudio for the
transcode transport method, and the TargetCodecVideo "copy" decision lived
only in TranscodeOpts. A new SessionManager.SetTranscodeStreamDetails
mirrors the actual decisions onto the upstream session when the transcode
starts (local and remote-node paths, via an optional interface so test
fakes are unaffected), so these sessions now bucket as "audio"/"remux".
- Transcode recipe cards now record TranscodeAudio derived from the opts
(only an explicit "copy" leaves audio untouched — empty runs ffmpeg's aac
default), so a session rebuilt after a restart keeps the same bucket.
- Recipe cards carry client name/version/user-agent, and reconstruction
restores them, so the admin client label and the JF pill survive server
restarts; the compat fallback card populates them from the live
MediaBrowser request. Deliberately not projected into stream-token claims,
where a user agent would bloat every stream URL.
* feat(api): capability endpoint for the live-session activity fields
Review comment on PR #387: effective_play_method and is_jellyfin_client are
omitempty, so an independently deployed client cannot distinguish an older
server from a supported one reporting an unknown method or a non-Jellyfin
session. GET /admin/sessions/capabilities advertises both fields plus the
closed bucket vocabulary, following the additive capability-endpoint rule
(same pattern as /collections/capabilities).
* fix(playback): treat empty target audio codec as an AAC re-encode in live state
ffmpeg defaults an empty target audio codec to AAC (appendAudioArgs), and the
new recipe logic already records that as an audio transcode — but the live
native path computed transcodeAudio=false for an empty codec, so the running
stream reported remux until a restart flipped it to audio. Extract the
predicate into playback.TranscodesAudio, share it across the live path, the
recipe card, and the compat mirror, and make appendAudioArgs case-insensitive
so the ffmpeg switch agrees with the predicate for any spelling.
Part of #387 review follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(jellycompat): re-sync sessions after recording compat encode decisions
ensureUpstreamPlayback flushes the session (compat_start) before
ensureTranscodeSession / startRemoteTranscode record the actual codec
decisions, and that later mutation triggered no sync — so the admin view
showed a video-copy stream as a full video transcode until the periodic
reconciler ran. Trigger syncSessionsNow after the details are recorded
successfully; the helper is shared, so both the local and remote-node
paths are covered.
Part of #387 review follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(catalog): typo-tolerant search for the Postgres (non-Meilisearch) path
## What this does (plain language)
When someone searches the library and misspells a title — "intersteller",
"godfathr", "jurasic" — the Postgres-backed search used to return nothing,
because it only did exact full-text matching. This adds a "did you mean"
fallback: when the normal search finds little or nothing, we run a second,
typo-tolerant lookup and surface the closest titles.
This only affects deployments that search via Postgres (the fallback path).
Meilisearch already does its own typo tolerance and is left untouched.
## Why not just make the main query fuzzy
The obvious approach — OR a trigram similarity match into the main search — is a
performance trap. The trigram operator is "lossy", so Postgres re-checks every
near-miss candidate by rebuilding three title search-vectors per row. On a real
library that turned routine searches into multi-second queries.
Measured on a 175k-title dev database:
- exact full-text only: ~60 ms
- fuzzy OR'd into the main query: ~217 ms (and far worse on prod-sized data)
## How it works
The fuzzy arm is a completely separate query (buildFuzzySearchSQL). It matches
only on the trigram-indexed title_normalized column and ranks only by
similarity() on that same column — it never touches the title search-vectors, so
it pays no per-row rebuild. It runs only when the exact search is "sparse" (fewer
than 5 hits) and the query is long enough for the trigram index to help (>= 4
characters), so the common case stays on the fast exact path. It is wired into
SearchPage (not just the thin Search wrapper) so the catalog search provider
benefits too.
## Measured on the live 183k-title catalog (read-only EXPLAIN ANALYZE)
- exact query for a typo: ~0.8 ms (0 hits -> triggers the fallback)
- fuzzy fallback query: ~5-27 ms, always via the trigram index, with no
search-vector rebuild
- "intersteller" -> Interstellar (similarity 0.63)
- "godfathr" -> GodFather (0.58), The Godfather
- "breakin" (134 exact hits) -> fuzzy correctly does NOT fire
Shared scope predicates (type / library / access / manga-exclusion) are extracted
into appendSearchScopeFilters so the exact and fuzzy queries filter identically.
Adapted from the earlier feat/search-fuzzy-fallback prototype onto main's current
SearchPage / includeTotal architecture.
* refactor(catalog): correct fuzzy-search pagination and parse the query once
Follow-up to the fuzzy fallback, from an adversarial code review. Two things: a
pagination correctness fix and a small performance/readability cleanup. Both were
validated against the live 183k-title catalog.
## The pagination bug (plain language)
Fuzzy results are shown after the exact results, as one combined list. The first
version stitched that list together with page-offset math, and got the math wrong
past the first page:
- the reported result count grew as you paged (page 1 said "31 results",
page 2 said "33");
- titles shown on page 1 could reappear on page 2;
- paging far past the end still ran the (pointless) fuzzy query every time;
- a tiny page size (e.g. an autocomplete asking for 3) could hide the fuzzy
results behind a page the client was told did not exist.
## The fix
Because the fuzzy fallback only runs when exact results are sparse (< 5) and the
fuzzy part is capped at 50, the whole combined list is tiny. So instead of
fragile per-page offset math, we now fetch that small combined list once and take
the requested slice in memory. Every page is then correct by construction: stable
total, no repeats, no wasted work past the end.
Before -> after, typo search "intersteller" (21 results, page size 5):
- total reported on page 2: 31 then 33 (drifting) -> 21 (stable)
- repeated titles across pages: yes -> none
- request past the end (offset 500): 2-3 DB queries -> 0 extra queries
- autocomplete (page size 1): fuzzy hidden -> paginates correctly
Cursor-style callers (that don't ask for a total) can't locate the boundary
between the two blocks on a later page, so they now get the fuzzy results as a
single terminal first page — no misleading "more results" flag.
## The cleanup
The raw query string was being parsed three times per search (once for the
eligibility check, once in each SQL builder). It is now parsed once in SearchPage
and passed down; the shared search-text derivation is extracted so the two
builders can't disagree; and the normalized form the eligibility gate needs is
precomputed at parse time. ("Performance first", per the repo guidelines.)
Also considered and rejected: excluding exact hits from the fuzzy query with a
NOT(full-text) clause instead of by id. It reintroduced the search-vector rebuild
the whole design avoids — measured ~51 ms vs ~20 ms on the worst case — so
id-based exclusion stayed.
Known limitation: the fuzzy path re-reads the small exact block in a second
query, so a title written in the sub-millisecond gap between the two reads could
be missed until the next search. Harmless and inherent to a multi-query design.
* fix(catalog): close fuzzy-search library-scope leak and restore small-limit cursor recall
Addresses two findings from the PR #386 review bots.
## Library-scope leak (Codex P1)
The search scope helper shared by the FTS query and the trigram fuzzy fallback
filtered libraries with `JOIN media_item_libraries mil` +
`NOT (mil.media_folder_id = ANY($disabled))`. An item linked to BOTH a disabled
and a non-disabled library fans out to two joined rows; the non-disabled row
satisfies the deny check, GROUP BY collapses the item back, and it surfaces in
search results despite the disabled library. Because the new fuzzy fallback
reuses this helper, typo searches could leak disabled-library items too.
appendSearchScopeFilters now delegates to the leak-safe
appendLibraryAccessConditions (access_filter.go), which emits item-scoped
EXISTS/NOT EXISTS subqueries — the same form GetByIDs/EnsureAccessible already
use — and needs no membership JOIN. The disabled-only path keeps its
argument-free positive-membership EXISTS so orphan items don't slip through a
vacuous NOT EXISTS. The scored CTEs keep GROUP BY (now required only for the
MAX() ranking aggregates). New regression test pins the EXISTS/NOT EXISTS shape
and the absence of a JOIN for both the FTS and fuzzy builders.
## Small-limit cursor recall (Codex P2)
In cursor mode (include_total=false) the FTS probe fetched only limit+1 rows.
For a tiny caller limit (e.g. an autocomplete asking for 2) with a few incidental
exact hits, that made ftsHasMore true, so the block never looked "sparse" and the
typo fallback never fired — and subsequent offsets are barred from triggering it,
so the fuzzy results were unreachable entirely.
SearchPage now floors the cursor-mode probe at fuzzyFallbackThreshold rows, and
execSearchBlock returns the pre-trim row count so sparsity is judged as
`fetched < threshold` independent of the caller's page size. The returned page is
still trimmed to limit with correct hasMore. Exact mode is unchanged (it judges
sparsity by the page-independent window count).
* fix(catalog): harden fuzzy-search fallback per adversarial review
Addresses the confirmed findings from a deep review of the fuzzy-search
fallback:
- Cursor mode now enters the fallback only when the whole sparse FTS
block fits the caller's page, so the terminal fuzzy page can never
hide exact matches the plain hasMore path would have surfaced
(jellycompat clients with EnableTotalRecordCount=false lost matches).
- execSearchBlock takes a querier and returns its untrimmed rows;
SearchPage hands the already-fetched block to the fallback instead of
re-running an identical FTS query on every sparse search.
- Fuzzy truncation is detected with LIMIT cap+1 instead of a
COUNT(*) OVER () window count that only fed a debug log; truncated
exact-mode responses now report total_exact=false rather than
presenting the cap as an exact count.
- The fuzzy query runs in a transaction pinning
pg_trgm.similarity_threshold via SET LOCAL, so match quality cannot
drift with cluster configuration.
- When the FTS block has real hits, fuzzy augmentation demands
similarity >= 0.45 so correctly-spelled sparse queries only gain
near-identical titles instead of base-threshold trigram noise.
- filterCatalogSearchItems no longer erases fuzzy matches on the
filtered/sorted/prefix resolver path: a typo token is never a
substring of the titles it matched, which left typo search returning
zero results there while the plain search box showed matches.
- The cursor probe floor applies only when the fallback can fire;
cursor fuzzy fetches no more rows than the terminal page can serve.
- slog.Debug -> slog.DebugContext (sloglint); reuse
contentIDsFromMediaItems instead of a duplicate helper; document the
title-only fuzzy scope.
Verified against the dev deployment: stable totals across pages, no
duplicates, small-limit cursor recall restored, filtered-path typo
search working, ~160ms typo-path latency.
Part of PR #386.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(catalog): reach long titles via strict word similarity in fuzzy search
Full-string trigram similarity is diluted by every extra trigram a long
title contributes, so a typo of one word could never reach titles like
"Avengers: Endgame" ("avegners" scores ~0.38 against "avengers" but far
below threshold against the full title). Swap the fuzzy predicate from %
to <<% (strict_word_similarity), which scores the query against the best
word-boundary extent of the title. At equal thresholds <<% is a strict
superset of %, and the existing gin_trgm_ops index serves both — no
migration needed.
The SET LOCAL pin moves to pg_trgm.strict_word_similarity_threshold and
is load-bearing: the 0.6 server default would reject ordinary one-edit
typos outright.
Ranking is strict word similarity first with whole-title similarity()
as tie-break, so near-identical short titles ("The Avengers") sort above
long titles that merely contain the matched word.
The 0.45 augmentation floor deliberately stays on whole-title
similarity(): word similarity rates embedded prefix words far too high
("coral" scores 0.5 against "coraline"), which dev testing showed would
flood a correctly-spelled sparse query with 27 noise rows. Zero-hit
(true typo) queries skip the floor, so the new long-title recall applies
where it matters.
Dev-verified: "avegners" now returns The Avengers first, then Avengers
Grimm / Avengers: Endgame; "coraline" still returns exactly its 4 real
titles; cursor small-limit recall, filtered-path typo search, pagination
stability, and ~160ms typo-path latency all unchanged.
Part of PR #386.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(scanner): recover malformed video durations
* fix(scanner): preserve probe failure semantics
* fix(scanner): harden malformed-duration recovery after review
Remediates the deep-review findings on the duration recovery machinery:
- Ignore attached_pic cover-art streams when deciding whether a file is
video: audiobooks/music with embedded artwork no longer fail import or
persist 1-second durations, and never trigger the packet scan.
- Cap the audio-only duration path (raised ceiling, not removed) so
malformed audio containers cannot persist multi-year durations.
- Keep the parsed probe when the packet fallback fails instead of
discarding codecs/tracks/resolution with a hard ProbeFile error.
- Apply the implausibly-short guard to the end-minus-start fallbacks so
collapsed absolute-timestamp spans reach the packet scan too.
- Make the legacy-duration repair one-shot: rows reprobed by the fixed
parser are authoritative, ending the infinite reprobe loop for
genuinely short large clips.
- Teach the library-scan repair predicate (needsCriticalProbeRepairScanState)
the same legacy signature so scans repair collapsed durations instead
of deferring to request-time repair.
- Grant the packet-scan timeout to any reprobe likely to hit the
fallback (Duration<=0 video files), not just the legacy shape.
- Share the implausibly-short thresholds between probe and repair layers
(closes the 100-500MiB repair gap) and deduplicate frame-rate parsing
within the scanner package.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* feat(metadata): register builtin NFO provider and broaden parsing
Phases A and B of the #216 local-NFO work, implemented test-first.
Registration & hint-first identity (Phase A):
- Migration seeds a reserved kind='builtin' silo.builtin installation
and an 'nfo' metadata capability (default_enabled=false, priority 1
for movie/series) with a partial unique index and documented Down.
- In-process builtin provider registry (internal/metadata/builtin.go);
buildProviders returns the registered provider for builtin rows.
- Guard rails keep the reserved row out of every plugin surface (user
plugin-settings, installations list, image resolvers, preload,
auto-update, store Delete, mutation handlers -> 409); silo.builtin is
a reserved manifest id.
- Startup sync materializes legacy content_level='' chains per level,
then appends builtin capabilities disabled via
AppendProviderToAllChains (idempotent); resolveEnabledProvidersBy
priority now respects default_enabled=false.
- NFO uniqueids seed the trusted-hint machinery via IdentityHintProvider
with per-mode conflict policy (stored IDs win on scheduled refresh,
NFO wins on manual refresh, Identify skips NFO); ID-less candidates
are excluded from provider-priority tie-breaks and nfo never counts
as corroboration.
- Web chain-editor empty-state gate is now server-derived so builtin
providers are reachable on plugin-less servers.
Parser breadth & sidecar hardening (Phase B):
- Parser covers the practical Kodi/Jellyfin field set for <movie> and
<tvshow>: original title, tagline, runtime, dates, content rating,
genres/studios/countries/tags, multi-source ratings with scale
normalization, cast with roles/order, director/credits. Empty
collections stay nil so merge early-returns apply.
- findNFO parses candidates and falls through on read/parse failure or
root-type mismatch, so a stray movie.nfo cannot shadow tvshow.nfo;
GetMetadata gains the same ContentType guard Search has.
- New FieldReleaseDates lock gates Year/ReleaseDate/First+LastAirDate
in merge (Go) and the edit-metadata dialog (web), closing the gap
where a manual refresh re-applied NFO dates over admin corrections.
- Merge-contract tests pin NFO fill semantics, genres whole-list
first-provider-wins, and NFO edits propagating on manual refresh only.
- Docs: new admin wiki page (supported fields, merge semantics,
naming-supplies-structure contract), index bullet, sidecar wording
revision, v1-scope feature-detection note.
Zero behavior change while the provider is disabled (default); pinned
by CI-mode and DB-gated test suites.
Part of #216
AI-use disclosure: implemented with Claude Code (Fable 5) via
spec-driven TDD and agent-assisted implementation.
* feat(metadata): ingest local sidecar artwork and read series-depth NFO
Phases C and D of the #216 local-NFO work, implemented test-first, plus
the mixed-library use-case pins. Together these deliver the headline
case: a series absent from every remote database (e.g. a fitness
library) scans into a fully presented show -> named seasons -> titled
episodes tree from NFO files and sidecar art alone.
Local sidecar artwork through the S3 image cache (Phase C):
- The NFO provider implements ImageProvider: poster/backdrop/logo
sidecar discovery with a fixed precedence map, symlink/non-regular
rejection, an 8 MiB cap, and file:// source URLs at rating 0. Generic
filenames apply only via the sidecar search paths, so a shared
folder.jpg in a flat multi-movie directory applies to none.
- file:// becomes a live local source scheme: routed into *_source_path
(never *_path), accepted by every image enqueue gate, attributed as
provider "local", excluded from cached-path detection.
- The image-cache processor caches local files with lexical-on-logical
confinement to the library roots, open-handle reads with re-checks,
the same variant widths as remote art, and stable (7-day) failure
classification. Keys land under
local/{contentType}/{contentID}/{hash8}/{imageType}; superseded
prefixes are cleaned on re-cache and item deletion.
- applyIfBetter gains a local exemption so rating-0 local art can fill
matched items without being stickily displaced; ImageRequest carries
additive sidecar path context.
Series depth (Phase D):
- SeasonsRequest/EpisodesRequest carry additive local path context
(series roots, per-season directories, per-episode file paths),
derived from naming at match time and reconstructed on refresh.
- season.nfo supplies season name/plot; NFO season numbers are advisory
(directory-derived number wins with a Warn - naming owns structure).
<episodedetails> gains aired/runtime/ratings; <basename>.nfo titles
episodes and <basename>-thumb.ext supplies thumbs; filename SxxEyy
wins over NFO numbers.
- Episode NFOs work without a season.nfo (provider seasons unioned with
on-disk seasons); SynthesizeFallbackEpisodes always runs after persist
so NFO-less episodes keep synthesized rows. Season/episode file:// art
rides the Phase C pipeline unchanged.
- Migration adds season:1/episode:1 to the builtin NFO capability's
default_priority (still default_enabled=false).
Mixed sports-library use case (tests only, no product change):
- Pins the classification contract for one library holding movie-shaped
and show-shaped content (WWE PPV events as movies next to a "WWE
SmackDown" show, NASCAR/F1/FIFA with partial TVDB/TMDB data): naming
decides movie-vs-series per file before any provider runs; the NFO
supplies metadata/identity but never flips type (ContentType guard);
the per-root Type override is the correction path.
- NFO-driven type classification at scan time is recorded as an explicit
deferred open question.
Part of #216
AI-use disclosure: implemented with Claude Code (Fable 5) via
spec-driven TDD and agent-assisted implementation.
* docs(metadata): document local NFO metadata architecture
Add a single as-built architecture page
(docs/architecture/local-nfo-metadata.md) for the #216 local-NFO
feature: the builtin registration model, hint-first identity semantics,
the file:// -> S3 artwork pipeline and its deployment constraint, series
depth, the mixed-library classification contract, and known limitations.
This replaces the working implementation plan, the per-phase specs, and
the narrow sidecar-artwork note, which were planning drafts and are left
untracked; admin-facing behavior remains in the wiki.
Part of #216
AI-use disclosure: planned, drafted, and consolidated with Claude Code
(Fable 5) using multi-agent exploration and adversarial review.
* fix(metadata): address PR review findings on NFO builtin provider
Fold in the valid, low-risk fixes surfaced by automated review on #390:
- imagecache: extract validateCacheRequest so CacheBytes (the local
sidecar season/episode path) enforces the same episode-requires-season
guard as Cache, preventing distinct episodes' art from colliding under
one S3 key.
- image_cache_processor: close the sidecar symlink-swap window by
rejecting the opened handle unless os.SameFile matches the Lstat'd
file, so a leaf swapped to a symlink can't pull an out-of-root target
into the public cache.
- plugins: guard the reserved builtin installation row in the store's
Update, matching Delete, so its version/enabled/capabilities can never
be rewritten even if a mutation slips past the HTTP layer.
- cmd/silo: bound SyncBuiltinProviderChains with a 30s timeout so a stuck
DB round-trip fails fast at startup instead of hanging.
- metadata: panic instead of silently no-op'ing on an invalid
RegisterBuiltinProvider call (init-time programmer error).
- docs: correct the media-folder-and-naming NFO paragraph to state
season/episode NFOs and sidecar artwork are actively read.
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
* fix(metadata): publish artwork revisions atomically
* fix(metadata): harden artwork revision cleanup
* fix(metadata): address artwork revision review findings
- restore image applies for all media_items types and reject unsupported
target/image combinations with 400 before uploading; episodes coerce to
stills and the web dialog no longer offers image tabs episodes can't use
- add WHEN clauses to displacement triggers and hoist to_jsonb so bulk
catalog upserts that assign unchanged artwork columns skip the trigger
- make artworkkey the single variant-ladder owner: imagecache derives its
widths from it and triggers store image_type instead of hardcoded
variant arrays, expanded by the collector at deletion time
- sweep dormant registry rows periodically so references lost through
untriggered surfaces degrade to slow cleanup instead of leaking
- park just-published revisions dormant, keep dormant rows dormant on
re-cache, and batch the GC reference pre-check per run
- heal rows re-referencing a just-deleted revision via reconciler-style
resets after the deletion commits
- share a per-URL image-loaded hook across DetailHero, ItemCard,
SectionItemCard, GlobalSearch, and CollectionPosterCard
- deduplicate Cache/CacheBytes finalization and drop unused VariantPaths
plumbing
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(catalog): cast reused timestamp parameter in revision upsert
Postgres cannot deduce one type for $3 used both as a plain value and
inside a CASE arm; the dev deploy surfaced it as SQLSTATE 42P08 on every
publication. Cast both uses and cover the arm/park/track upserts with
database-backed tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(metadata): address artwork revision review comments
- keep a durable heal path: deletion marks deleted_at instead of removing
the registry row, so a failed post-delete heal retries with backoff and
broken references never park; trackers clear the marker on re-upload
- never treat bare existence as an immutable-content match; backends
without content verification rewrite the object
- exercise revisioned cover keys in scanner/enrichment fakes, compare the
tracked manifest exactly, and honor cancellation in the blocking test
deleter
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(web): update metadata for selected version
* refactor(web): simplify selected duration lookup
* refactor(web): share the hero runtime formatter via mediaFormat
MovieContent and EpisodeContent carried identical local formatDuration
helpers; move the minutes-based badge formatter into lib/mediaFormat as
formatRuntimeMinutes alongside the other canonical media-spec helpers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(playback): run orphaned-transcode cleanup in the background at startup
The native and Jellyfin-compat routers swept stale per-session transcode
dirs synchronously during NewRouter, before the listener bound. On a slow
network filesystem this blocked startup for 80+s (64 leftover dirs on the
last deploy), so restart-reconnect clients were turned away and the health
check reported the server unhealthy the whole time.
Move both sweeps into a background goroutine (StartBackgroundOrphanCleanup)
so the listener comes up immediately and the cleanup runs concurrently. The
delete logic is unchanged: same active-session snapshot and MaxTokenTTL
age-sparing, only later. A package-level mutex serializes concurrent sweeps
of the shared transcode root so the two background sweeps can't race on
os.RemoveAll.
Part of #412
* fix(transcode): background the node boot-time transcode-dir sweep
A dedicated transcode node swept leftover transcode dirs synchronously in
NewServer, before startStandaloneServer bound its listener. On a slow
network filesystem that delete blocked the node from coming online at boot,
the same startup-stall class as the main server.
Move the sweep into the shared StartBackgroundOrphanCleanup goroutine so the
node's listener binds immediately. Backgrounding required an age guard: the
sweep previously ran as a full wipe (minAge=0) with an empty active-set,
which was only safe because it completed before any request could arrive.
Run concurrently that would race a token-carried reconstruct writing into
TranscodeDir/<sessionID>, deleting segments a fresh ffmpeg is producing.
Passing MaxTokenTTL spares any dir younger than the max token lifetime —
exactly the ones a still-valid reconnect could reconstruct — while dirs
older than any surviving token (never reconstructable) are still reclaimed.
Part of #412
* feat(playback): reclaim orphaned transcode dirs periodically, not just at boot
The orphaned-transcode sweep only ran at startup on both the central server
and transcode nodes, so it only ever reclaimed dirs left by an ungraceful
prior shutdown. During a long uptime the in-memory session reapers delete the
dirs of sessions they still track, but a dir whose owning session was dropped
without its RemoveAll succeeding becomes an "untracked orphan" with no runtime
GC — on a box that runs for weeks these accumulate until the next restart.
Add StartPeriodicOrphanCleanup: an immediate background sweep followed by an
hourly re-run bound to a lifecycle context. Wire it on all three surfaces —
native API and Jellyfin-compat (via deps.AppContext) and the transcode node
(via a new Server.StartOrphanSweeper(appCtx), replacing its boot-only sweep).
When no context is supplied (tests) it degrades to a single boot-time sweep so
no ticker goroutine outlives the caller. The sweep stays age-guarded at
MaxTokenTTL, so nothing reconstructable is ever reaped.
Because the node sweep now runs during live traffic, it snapshots the live
job set (Server.activeSessionIDs) and spares those dirs by id rather than by
age alone — a long-lived session that only re-serves already-written segments
stops advancing its dir mtime, which age could otherwise misclassify. In
integrated mode the native and compat sweeps share one TranscodeDir but each
snapshots only its own manager's live set; the resulting cross-manager reap of
a >24h idle dir is bounded (rebuilds from token/recipe) and documented at both
call sites.
Part of #412
The orphaned session sweep in CleanupOrphanedTranscodeDirs treats every
subdirectory of the transcode dir as a dead session, but the subtitle
cache lives in there too. Cache hits only bump the entry file mtimes,
never the directory's, so an active cache that had not seen a new
extraction in 24 hours looked stale and got wiped in one go. That forces
a full ffmpeg demux on the next PGS subtitle selection, which is the
exact cost the cache exists to avoid, and the transcode node's boot wipe
removed it unconditionally.
Skip the cache directory in the sweep. It validates entries on lookup
and runs its own LRU eviction, so leaving it alone is safe.
* fix(scanner): never purge files under unreachable library roots
An unreachable root is not a removed root. When one root of a multi-root
library dies (unmounted share, dead drive) while another root still has
files, the whole-library empty-root guard does not fire — the surviving
root produced files — so the scan marks everything under the dead root
missing_since (desired: hides it from browse/playback) and then, with the
default scanner.empty_trash_after_scan=true + 24h file_removal_grace, the
next scan after the grace hard-deletes every row under the dead root. A
week-long drive outage silently destroys the root's entire catalog state:
probe data, intro/credits markers, file hashes. Worse, membership
reconciliation immediately purges media_items whose only files lived on
the dead root, cascading user collections (library_collection_items has
ON DELETE CASCADE) and deleting cached artwork.
This change makes "temporarily offline" survivable:
- Probe each configured root at scan start (os.Stat + IsDir + ReadDir,
factored into the new internal/rootcheck package and shared with the
admin mount-check endpoint). Unreachable roots are skipped by the walk
but their scopes still reconcile, so files are still marked missing.
- The trash sweep (DeleteMissingByFolder) now excludes rows whose path
sits under an unreachable root, using the same exact-path + escaped
prefix-LIKE matching as ListIDsOutsideRoots (a sibling root that merely
shares a string prefix is never protected). With all roots reachable
the emitted SQL is unchanged.
- Membership removal still happens — browse/home hide items via
media_item_libraries, so removal is what keeps a dead-root-only title
out of the catalog — but the orphan media_items purge exempts items
whose files sit under an unreachable root. Their metadata, artwork,
and collection links survive; when the root returns, the upsert clears
missing_since and syncPresentLibraryState re-inserts the membership,
restoring the item with zero re-probing or re-matching.
- The folder surfaces scan_warning_code='dead_root' with a message naming
the unreachable roots; a fully healthy scan or a successful mount check
clears it, mirroring empty_root. The admin UI shows a badge and banner.
- Deliberate deletion is untouched: removing a path from the library
config still purges via ListIDsOutsideRoots, files under reachable
roots keep the exact 24h-grace purge, the empty-root guard and the
autoscan dead-mount guard are unchanged.
The audiobook/podcast/ebook reconcile paths share the same folder-wide
sweep and orphan purge, so they get the same guard.
Covered by tests: an end-to-end two-root scan (root dies -> rows survive
a zero-grace sweep and warning is set; root returns -> rows resurrect
with their original ids and the warning clears; deleting a file under a
reachable root still purges), repo-level sweep-protection and
sibling-prefix tests, orphan-purge exemption, and rootcheck unit tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(scanner): probe uncompacted roots and take dead-root path on full outage
Review follow-ups: (1) probe every configured path instead of the compacted
traversal roots, so a nested child mount that dies under a reachable parent
is still protected from the sweep; (2) when every configured root is
unreachable, bypass the empty-root confirm flow (without consuming the
one-time cleanup allowance), mark files missing, and raise dead_root instead
of empty_root; (3) dead_root warning banner no longer shows empty-root
confirm-deletion guidance as its fallback hint.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(scanner): simplify dead-root protection plumbing
- extract pathscope.CoverageClauses as the single builder for the
exact-path + escaped prefix-LIKE root predicate; scanner's
rootCoverageClauses delegates to it and catalog's
excludeOrphansUnderProtectedPrefixes reuses it instead of hand-rolling
the same clause loop
- extract Scanner.sweepMissingAndReconcile to replace the identical
trash-sweep + membership-reconcile + S3-image-cleanup block that was
triplicated across the audiobook, ebook, and podcast scans (callers
keep their flavor-specific log lines so messages stay constant)
- add unreachableConfiguredRoots helper for the repeated
probeUnreachableRoots(ctx, folder.ID, cleanScanRoots(folder.Paths))
expression in scanPaths and ScanFile
- drop the unread Path field from rootcheck.Result
- move the dead/empty-root warning text constants in AdminLibraries.tsx
out of the middle of the import block
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scanner): close dead-root protection gaps found in review
Remediates the confirmed findings from the deep review of this PR:
- Scoped audiobook scans (autoscan file events, subtree scans) ran the
folder-wide sweep while probing only the scoped clone's Paths, so a
healthy-subtree event could hard-delete a dead sibling root's rows.
sweepMissingAndReconcile now reloads the folder's configured roots
from the DB and probes them uncompacted, which also protects nested
child mounts in the audiobook/ebook/podcast reconcilers.
- A lost mount that leaves an empty, stat-able mountpoint probed as
reachable and kept the historical purge timeline. A reachable root
that is a literally empty directory while cataloged rows remain under
it is now treated as suspect: rows are only marked missing, the sweep
and orphan purge exempt it, dead_root is raised, and the mount-check
endpoint reports it (additive suspect_empty field) instead of
clearing the warning. Arming the one-time empty-cleanup allowance
completes the deletion, including in the mixed case where other
roots are healthy. Roots that still have directory entries keep the
historical grace-then-purge path.
- Confirmed empty cleanup (allow_empty_cleanup_once) no longer
force-deletes rows under probe-dead roots: an outage is not a
confirmation, so a dead sibling root's catalog survives a confirmed
cleanout of a reachable empty root.
- Root probes are now bounded (rootcheck.ProbeWithTimeout, 5s): a hung
network mount degrades into the protected unreachable path with a
probe_timeout error code instead of stalling every scan of the
folder indefinitely.
- Documented the cross-library limitation of the orphan-purge
exemption next to the query it applies to.
All behavior is pinned by new DB-backed tests (suspect-empty
protection + confirmed completion, confirmed-cleanup dead-root
survival, scoped/nested-root sweep protection, suspect-root query,
probe timeout).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(scanner): address dead-root review findings
---------
Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
* feat(plugins): group Apps sidebar by plugin manifest category
Implements the plugin SDK's documented PluginManifest.category semantics
(silo-plugin-sdk proto/silo/plugin/v1/common.proto): a slash-delimited
path that groups plugins in the user-facing Apps section, e.g.
"Books/Audiobooks" lands under Apps -> Books. The field existed in the
manifest proto but silo-server never surfaced it.
Server: the user plugin-settings list/detail responses now include an
additive-only `category,omitempty` string sourced from the already-loaded
manifest via GetCategory(); no new parsing paths.
Web: PluginSettingsSummary gains `category?: string`, and AppSidebar
groups Apps entries by the FIRST segment of the category path (one level
of grouping for now; deeper segments intentionally ignored, documented
against the SDK contract). When fewer than 2 distinct categories exist
among the visible app links, today's flat list under the single "Apps"
header is kept; with 2+ categories, per-category sub-headers render via
the existing SidebarSectionHeader (labels hide in the collapsed sidebar
the same way other section headers do). Uncategorized plugins fall under
"Other", which always sorts last.
Tests: Go unit tests for the summary converter (category passthrough and
JSON omission when empty) and vitest coverage for the pure
groupAppNavLinks helper plus grouped/flat/collapsed sidebar rendering.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(plugins): use generic category examples in comments and tests
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* refactor(web): simplify Apps sidebar link list rendering
- fold the duplicated <ul> list markup in the grouped and flat Apps
branches into a single renderAppNavList helper so the list styling
cannot drift between the two render paths
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>