Commit Graph
846 Commits
Author SHA1 Message Date
8dea4b9056 fix(playback): stop a broken Dolby Vision RPU from hanging playback (#485)
* fix(playback): stop a broken Dolby Vision RPU from hanging playback

A Profile 7 source whose RPU ffmpeg cannot parse took the whole session
down. The dovi_rpu bitstream filter does not fail cleanly — it rejects
every packet while ffmpeg runs on, so one observed session emitted
376,316 stderr lines before the process was killed, no manifest was ever
built, and the client got a 503 after ~10 seconds that it showed as an
endless spinner:

  [dovi_rpu] Failed to read unit 1 (type 39).
  [vost#0:0/copy] Error applying bitstream filters to a packet:
      Invalid data found ... Invalid SEI message: payload_size too large

Whether the strip works is a property of the file, not of ffmpeg, so
SupportsDoviRPUFilter cannot answer it — but asking ffmpeg to strip two
seconds to the null muxer can, in about a second. Profile 7 sources are
probed once each on the start path and the result is cached; a source
that fails is copied without the filter, which leaves the base layer and
plays. Refusing to play at all does not.

The probe reads stderr rather than trusting the exit code: ffmpeg treats
a per-packet bitstream-filter error as non-fatal and exits 0, which is
exactly how a stream that could never start reached a live session.

Cache is keyed on path plus size so a replaced file is re-probed, and
bounded like the letterbox cache. A nil probe keeps stripping, since most
Profile 7 sources need it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* chore(playback): use InfoContext in the rpu probe

* fix(playback): decide the DV RPU strip in the plan, not at the transport

The probe answered the right question in the wrong place. Suppressing the
bitstream filter as the transport was being built left the plan still
promising DynamicRange "hdr10", Claims.Video.HDR10 and the
dolby_vision_metadata_removed / hdr10_base_layer_preserved validated
claims, and left session.RemuxDVMode persisted as strip_to_hdr10 — so the
client was told it was getting clean HDR10 while receiving Profile 7 with
dangling RPUs, and every path that re-derives the filter from the durable
session put the hanging filter straight back:

  - HandleStartTranscode (quality change, seek, burn-in restart), which
    also re-derives it from file.PrimaryDVProfile() == 7 with no plan at all
  - the remote audio-switch restart, from updatedSession.RemuxDVMode
  - the progressive-remux transport, which plan_v3 evaluates *before* the
    HLS branches, so for many clients the broken file never reached the
    probe in the first place

The verdict is now a planner input alongside the transformation
registries: the registries answer whether the executor carries the
transformation, this answers whether the file does. A source that fails it
is never planned onto a strip route, so the plan's claims, RemuxDVMode and
every restart derived from them agree with what the pipeline can produce.
With no tone-map recipe in this tree an HDR10-only client has no route
left, so it gets a dv_conversion_unsupported terminal naming the real
cause rather than a generic HDR message; a client that can run its own DV
transformation still gets that route, with a degradation warning
explaining why the server route was dropped.

The two paths that bypass the plan entirely are gated at the executor:
legacy/auto remux neutralizes the profile exactly as it already does for a
missing dovi_rpu filter, and the explicit v3 strip recipe fails loudly
rather than emit dangling RPUs under an HDR10 claim.

The probe itself:

  - Tri-state verdict. Only a stderr-confirmed rejection is cached. A
    timeout, a cancelled request or an ffmpeg that will not start is
    inconclusive: the strip is kept and nothing is written, so one client
    disconnecting can no longer disable the strip for a file permanently.
  - Singleflight, so concurrent or retried starts of a title share one run.
  - Timeout cut to 6s, inside the budget a client waits on the manifest,
    and off the session lifecycle lock now that it runs at planning time.
  - Keyed on size and mtime as well as path and binary, so a file replaced
    in place with the same length is re-probed.
  - stderr capture bounded at 64 KiB; the markers are in the first lines
    and a rejecting filter emits a pair per frame.
  - The head-only coverage is stated rather than asserted: this catches a
    source that rejects from the first access unit, not one that breaks an
    hour in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(playback): keep the RPU probe alive when its caller leaves

The probe rode the caller's context, so a leader whose client disconnected
returned dvRPUUnknown — and CanStrip then handed that fail-open to every
follower already blocked on the shared call, even though their own requests
were still alive. One client walking away was enough to give a live session
the strip the source cannot survive, which is the hang this whole change
exists to prevent. The work was also thrown away, so the next start paid for
the probe again.

Run it under context.WithoutCancel instead. dvRPUProbeTimeout still bounds
it, so nothing is left running; a verdict reached after the leader has gone
is still correct and still worth caching. Follower behaviour is unchanged:
a follower whose own request is cancelled still leaves immediately rather
than waiting on someone else's probe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-27 08:15:25 -04:00
baa33768d6 fix(scanner): stop reporting an unusable ffprobe as an empty folder (#468)
* fix(scanner): stop reporting an unusable ffprobe as an empty folder

parseAudiobookFolder and parsePodcastShow signalled "this folder holds no
audio files" by wrapping os.ErrNotExist, and their reconcile callers skipped
on that. exec also wraps fs.ErrNotExist when the configured ffprobe binary
cannot be run, so a wrong playback.ffmpeg_path made every candidate folder
look empty: the scan logged processed=N failed=0, indexed nothing, and gave
the operator no clue why the library stayed empty.

Introduce an errFolderHasNoMedia sentinel that deliberately does not wrap
os.ErrNotExist, and skip on that instead. A folder that disappears between
the scan walk and the parse still maps to the sentinel, so a mid-scan rename
or delete stays a quiet skip rather than a scan failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scanner): bound the all-failed scan summary instead of joining every failure

The sentinel change in this PR makes a previously-unreachable path reachable.
Before it, a misconfigured ffprobe made every candidate folder look empty, so
the scan skipped everything and `failed` stayed 0 — the `failedCount ==
processedCount` branch never fired. Now that an unusable ffprobe propagates as
a real failure, that branch is the expected outcome of a first scan with a bad
`playback.ffmpeg_path`, and it joins one wrapped error per failed folder.

On the 240k-folder library the scan code is written for, `errors.Join` over
that slice produces a ~64 MB error string (measured) that is written verbatim
into `scan_runs.error_message` and republished over the Redis events channel
and the admin SSE stream. The `failures` slice itself also grew unbounded for
the whole scan even when the all-failed guard could not fire (any rescan with
`skipped > 0`), holding hundreds of megabytes across a multi-hour scan before
discarding it.

Add a `scanFailures` collector that retains the first 20 failures and counts
the rest, joining them with a trailing "and N more failures (elided)". The
retained sample still names the cause, which is the entire purpose of the
summary. The same 64 MB case now produces 5.4 KB.

The ebook and manga scans have the identical shape and the same exposure via
their own probe failures, so all four call sites share the collector rather
than fixing the two audio paths alone. `failMu` now guards only `cancelErr`
and is renamed `cancelMu` to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-26 23:35:36 -04:00
172beb99ef fix(watch-together): stop a dropped socket reading as the host leaving (#487)
* feat(watch-together): make vote rooms actually vote

selection_mode has been stored, normalized and published since the
feature landed, and nothing has ever read it. A "vote" room behaved
exactly like a host_pick one: members could suggest and vote, the tally
was recorded and broadcast, and then the host promoted whatever they
liked regardless of it.

In a vote room the host now starts the winner rather than choosing it.
Promoting anything other than the leading suggestion is refused, because
being able to overrule the tally makes the mode host_pick with extra
steps and turns the vote counts on everyone else's screen into
decoration.

The winner is the head of the repository's existing ordering
(vote_count DESC, created_at ASC): most votes, ties to whoever suggested
first — deterministic, and re-suggesting a title cannot jump the queue.

A room where nobody has voted has no winner and says so, rather than
quietly promoting the oldest suggestion as though a vote had happened.

host_pick rooms are untouched: the host still promotes freely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* fix(watch-together): close the second door into a vote room's selection

Gating PromoteSuggestion left SelectItem wide open: it is host-only but
was not gated by selection mode, so the host of a vote room could set any
title directly and bypass the vote entirely. Enforcing the tally on one
path and not the other makes the vote counts on everyone else's screen
decoration.

A vote room now refuses a direct selection outright. The winner is the
only way in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* fix(watch-together): stop a dropped socket reading as the host leaving

hostDisconnectTTL was 15 seconds, which treated any transient drop as a
departure. An explicit leave and an explicit close already tear the room
down immediately, so this timer only ever covers a host who has NOT said
they are going — and at 15s a host who backgrounded the app, moved
between screens, or hit a brief network blip lost the room for everyone
with a "host_left" nobody could explain.

Two minutes survives a reconnect or an app switch, and is short enough
that a genuinely departed host does not leave a room open all evening.
The janitor still reaps idle rooms independently.

This matters for what the clients are growing into: a room you stay in
while you browse for something to suggest. A client that drops its socket
when the lobby leaves composition should cost you a reconnect, not the
room.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* fix(watch-together): let a vote room actually start its winner

The vote gate landed on both doors into a room's selection, but promoting
the winner walks through SelectItem to commit — so the gate meant to stop
the host bypassing the vote also stopped the vote itself. Vote rooms could
not start playback by any route.

Split the commit path: SelectItem keeps the gate for direct requests, and
PromoteSuggestion goes through the internal path once it has confirmed the
suggestion is the winner. Map ErrVoteRoomSelection in the promote handler
too, so a future regression there reads as a conflict rather than a 500.

Add service-level tests for both gates — the previous tests only covered
the pure winnerFrom helper, which is why the suite stayed green while vote
rooms were non-functional.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-26 22:01:31 -04:00
148c9291c5 feat(web): add Connect Apps settings page for compat sign-in (#488)
* feat(web): add Connect Apps settings page for compat sign-in

Jellyfin-protocol clients offer one username box and one password box and
never prompt for a profile, so signing in requires `account#Profile` and
`password#PIN`. Nothing in the product taught that syntax, and every way of
getting it wrong surfaces as the same "invalid username or password", so it
became a recurring support burden.

Add Settings -> Connect Apps, which states the credentials for the signed-in
account rather than describing them in the abstract. The page is segmented by
app type: the two formats are never shown at once, each side names the apps it
covers, and the compat side is visually distinct so it cannot be mistaken for
the normal Silo login. The compat listener's separate address is shown too,
since pointing a client at the Silo address fails identically to a bad
password.

Backend adds GET /api/v1/compat/connect-info, an account-scoped read of the
compat listener's enabled flag, public URL, and server name. The admin status
endpoint already covers this ground for operators, but it also reports install
paths and version provenance, so it stays admin-only; this returns only what a
client learns by connecting anyway. It is auth-only and deliberately not
profile-scoped, since it describes how to sign in.

ConnectInfoForConfig shares the enabled-flag precedence with
WebComponentStatusForConfig via compatEnabled, so the two endpoints cannot
disagree about whether compat is on.

The page declines to display a username it knows cannot work: profile names
permit `#` but the resolver splits at the last one, so `alice#Movie #2` parses
as account "alice#Movie". Such profiles get an explanation and no copy button
instead of a string that fails to authenticate.

Part of #432

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web): harden Connect Apps against misleading sign-in states

Review of the first pass found six ways the page could state something
untrue. Each one matters more than usual here, because the page exists
specifically to stop people guessing at credentials.

Report the running listener, not the stored setting. jellyfin_compat.enabled
is restart-required and cmd/silo builds the compat server from the boot
config alone, so the stored value describes intent. Reporting it promised
credentials for a listener that does not exist yet, or claimed the API was
off while the running one kept serving. ConnectInfo now returns the boot-time
state plus a pending_restart flag, and the page distinguishes "not running
yet" from "turned off". The admin status endpoint keeps reporting configured
intent, so the two intentionally diverge until a restart.

Stop presenting fetch failures as a disabled compat API. React Query clears
isLoading on error, so a failed request rendered "the compatibility API is
turned off" and sent users to an admin about a setting that was fine. A
failed profile list was worse: it fell through to an empty list and offered
the bare account name, which silently drops the profile suffix. Both now
withhold credentials and say the load failed.

Detect accounts that cannot use password login. Compat login is hardwired to
the local provider, which rejects accounts with local_password_login_enabled
false before checking any password, so SSO and plugin-provisioned accounts
can never authenticate. The page told them to type a password anyway; it now
says the compat API cannot accept the account.

Flag loopback compat addresses. jellyfin_compat.public_url defaults to
http://127.0.0.1:8096, which resolves to the client device on the phones and
TVs this page names. An untouched default was offered as the exact address to
copy; it is now explained instead of presented as usable.

Apply the #-in-profile-name guard to the summary list too. The selected-profile
field withheld an unusable username while "Every profile at a glance"
reintroduced it two sections below.

Read only the three settings this endpoint consumes. GetAll on the encrypted
repository decrypted every stored secret on each authenticated page view, and
an unrelated decryption failure would have silently dropped valid compat
overrides.

Invalidate the connect-info cache when an admin saves jellyfin_compat.*
settings. The address applies without a restart, but the page cached it for
five minutes and kept offering the old value for copying.

Part of #432

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 21:07:48 -04:00
99d205676f fix(metadata): prevent stale cross-provider IDs (#480)
* fix(metadata): prevent stale cross-provider IDs

* fix(metadata): address stale ID review findings

* fix(migrations): build the stale-ID primary key concurrently

ALTER TABLE ... ADD PRIMARY KEY builds the index under ACCESS EXCLUSIVE,
blocking reads and writes on stale_media_ids for the whole build. Create the
wider unique index with CREATE UNIQUE INDEX CONCURRENTLY and attach it with
ADD CONSTRAINT ... PRIMARY KEY USING INDEX instead; all three key columns are
already NOT NULL, so the attach is metadata-only. Same treatment on the
rollback path, plus the repo's INVALID-remnant cleanup so a failed concurrent
build is not silently accepted by IF NOT EXISTS.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 11:19:33 -04:00
02203d9e40 fix(playback): trust the server's media runtime end to end (#482)
* fix(scanner): reject durations that imply an impossible bitrate

The duration-plausibility rule only rejected videos of 10 seconds or less,
so a feature film that probed as 61 seconds passed untouched and persisted.
Clients then had nothing trustworthy to anchor on: Android's grow-only
duration ratchet has no floor to hold when the catalog value is wrong, so
the playback engine's growing-HLS-window duration won and a 90-minute movie
displayed as ~1 minute.

Size and duration together pin an implied bitrate, which separates the two
cases the absolute floor conflates. A genuine short clip has an ordinary
bitrate; a 100 GB file claiming 61 seconds implies ~13 Gbps. The ceiling
sits far above any real medium, so legitimate content cannot trip it — and
unlike the absolute floor, it does not false-positive on a genuine
high-bitrate short.

Also bump the repair-rule revision marker so rows judged by the previous,
weaker rule are re-checked once under this one. Without that bump an
improved rule never reaches the rows it was written for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(playback): publish source runtime in v3 plans and stop faking the copy seek window

Two defects with one root: a v3 plan described where playback sits without
ever stating how long the media is.

Add source.duration_seconds. It is the file's full runtime, never
`total - source_start` and never adjusted by timeline_offset_seconds, and it
is omitted rather than null when unknown — clients that coerce null to a
numeric default would read it as zero, the exact value this field exists to
stop them inventing. It is set in SourceDescriptorFromFileV3, the single
place every delivery already flows through, so direct play, progressive
remux, HLS remux and HLS transcode all carry it.

Until now the v3 plan omitted duration entirely, so clients fell back to the
playback engine. On an HLS copy remux the server intentionally serves
FFmpeg's still-growing playlist, so the engine reports the length produced
so far. With no server-supplied runtime to anchor on, a feature film played
back as a couple of minutes. The legacy protocol already answered this
correctly via fileDurationSeconds; this restores parity.

Separately, the copy branch published seek_window_end_seconds as the media
runtime. That made the window look *complete*, which clients read as proof
that any target inside it is locally seekable, so they native-seek past the
produced head of a growing playlist instead of asking for a reanchor. Leave
the end open: an incomplete window plus can_seek_anywhere=false routes every
seek through the server, which is what legacy did before v3 added the bound.

Advertise plan_source_duration_v1 so a client can distinguish "this server
does not populate the field" from "this server knows the runtime is
genuinely unknown" — without it, both look like an absent field and a client
cannot tell whether its own catalog fallback is still required.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web): pair the exit position with the media runtime, not the element duration

The player's exit state converts its position to media time but took the
duration from the video element, which is player-local. On a remux or
transcode stream the element only covers the window produced so far, so the
two values live in different coordinate systems.

Resuming a movie 50 minutes in makes that concrete: the exit position is
~3060s of media time while the element reports ~120s. The progress cache
then evaluates `position >= duration`, marks the item completed, latches the
watched badge, and — because completion clears the resume point — resets
position to 0. Exiting a resumed movie destroyed the resume point and
claimed it had been watched.

The server's runtime is authoritative and already expressed in media time,
so prefer it and fall back to the element only when no server value exists.
The rule moves into mediaTimeline.ts next to the coordinate conversions it
depends on, which is also what makes it testable — VideoPlayer itself has no
test harness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 00:12:29 -04:00
QuickandGitHub 10394b0a05 fix(scanner): stop hiding media when a library root is offline (#472)
* fix(scanner): stop hiding media when a library root is offline

A scan that cannot read a library root found no files there, so every
cataloged file under it was marked missing. Catalog reads all filter on
missing_since IS NULL, so marking is equivalent to deletion from a user's
point of view: the title leaves browse, search and next-up, and playback
answers "Source media file is missing" for media that is intact on disk.

The dead-root protection already existed but only guarded the destructive
operations. protectedConfiguredRoots was computed *after* the marking loop
in scanPaths, and applyScopedScan received the protected set but applied it
only to its force-delete branch. So an unreachable root could not lose its
rows, but could still have its entire catalog hidden until the next
successful scan.

On a CephFS deployment whose per-library subvolume mounts flap, this marked
190 present files missing in a single day — 15% of all missing-flagged rows
were files sitting untouched on disk, some flagged more than 20 hours after
their last write.

Hoist the probe above the marking loop and skip files under an unreachable
or suspect-empty root in both the folder and scoped paths. Pass the
unreachable set to the walked-scope call too: a nested child mount can die
under a healthy parent, and its rows are inside the parent's scope.

An offline root tells us nothing about whether its files exist. The only
safe reading is to leave them alone and let the next good scan decide.

Genuine deletions under a reachable root are unaffected and still marked
and swept on the same schedule as before.

Report the count as ScanResult.MissingSkippedProtected and log it, so an
operator can tell "my library shrank" from "my mount dropped".

Known gap: a suspect-empty *nested child* root under a healthy parent is
still marked missing, because the suspect set is not resolved until after
the walk loop. Unreachable roots — the case observed in production — are
covered.

* fix(scanner): protect suspect-empty and partially-walked roots too

Addresses review findings on #472. The original change guarded missing-marking
against probe-unreachable roots, but left three ways for a storage fault to
still hide a healthy library.

Suspect-empty detection was reactive. suspectEmptyRoots asked
ListRootsWithOnlyMissingFiles, which returns a root only once it has NO live
rows left. On the first scan after a mount drops — the moment that matters —
the rows are still live, so the root was not classified suspect and the scan
marked everything missing. The protection then engaged on the next scan, in
time to protect the wreckage. Ask ListRootsWithCatalogedFiles instead: any
cataloged row under an empty-but-reachable root is the lost-mount signature.
Intentional emptying is still reachable through the operator's one-time
cleanup allowance, which is the deliberate path for it.

Nested suspect-empty children were unprotected. Root compaction sends only the
populated parent through the walked-scope branch, which received only
unreachableRoots, so an empty child mountpoint had its rows marked missing on
its parent scanning cleanly. Pass the suspect set as well.

Partial walks were treated as authoritative. walkLogicalTree deliberately
swallows per-entry Lstat/ReadDir failures so one bad file cannot abort a scan
of a million, and collectLogicalFilePaths passed nil for the failure counter —
so the video path had no signal at all. A mount dying partway through
traversal produced a short file list indistinguishable from a large deletion.
Thread the counter through, and exclude a scope whose walk came back
incomplete from missing reconciliation, mirroring what the ebook scanner
already does via ebookRootScan.failed.

Also extract the duplicated mark-missing loop into markMissingExcludingProtected
so the folder and scoped paths cannot drift, and correct two comments that
still described the pre-fix "files are marked missing" behaviour — the exact
text a future reader would have trusted when reintroducing this bug.

TestScanFolderNestedSuspectEmptyChildRootProtection asserted the old
behaviour and is updated accordingly.

* fix(scanner): scope walk-failure protection and stop pruning on partial walks

Addresses the second Codex review round on #472. The previous commit's
incomplete-walk protection was too blunt in one direction and applied too late
in another.

Walk failures were counted, not located, and any non-zero count protected the
whole library root. A dangling symlink is both common and permanent, so that
would have suppressed missing-file reconciliation for its entire root on every
future scan — genuinely deleted titles would stay live indefinitely. That is
the same class of bug as the one this PR fixes, pointing the other way.
recordWalkFailure now records the logical path of each unreadable entry, and
only those paths are protected. Per-entry failures record the child path, so a
dangling symlink protects itself and nothing else, while a directory that
cannot be read protects its subtree.

Snapshot and group pruning ran before the protection. reconcileScannedRoots
and reconcileScannedGroups delete whatever the walk did not see, and both run
ahead of the missing-file guard, so a partial walk still dropped root
snapshots, observed locations and group locations for the unread portion —
corrupting later metadata matching even though the media_files rows survived.
Upserting what was seen is always safe; pruning now waits for a scan that read
the whole tree.

The confirmed-cleanup allowance was consumed to no effect for nested suspect
children. The walked-parent branch protected them unconditionally and runs
before the allowance is consumed, and an already-reconciled scope cannot be
revisited — so arming the allowance burned the confirmation while the child's
rows stayed live forever. Read the allowance without consuming it before the
walk loop, and honour it there. Unreachable roots stay protected either way:
an outage is never a confirmation to erase a catalog.

Two new regression tests, plus signature updates in the ebook pipeline, which
already tracked walk failures and now shares the path-based representation.

* fix(scanner): re-probe nested roots and gate group pruning on walk completeness

Third Codex review round on #472; both findings confirmed.

Group pruning ignored walk completeness in the subtree path. scanPaths passed
the completeness decision to reconcileScannedRoots but left
reconcileScannedGroups on !allowEmptyRootGuard, which is always true for
ScanSubtree — so a subtree scan that hit an unreadable directory still replaced
group snapshots and locations from a partial inventory. Same rule now applies
to both.

Nested roots were not re-probed before their parent was reconciled. Root
compaction folds a child mount into its parent for traversal, so a child that
is healthy at the initial probe but drops before the parent is walked leaves no
scope of its own, and the post-walk re-probe only revisits scopes that walked
empty. The parent walks files, looks healthy, and the child's rows are marked
missing on its success. reprobeNestedRoots re-checks this root's configured
children immediately before reconciling, protecting any that have since become
unreachable — or suspect-empty, unless the operator has confirmed cleanup.

Also guard suspectEmptyRoots against a nil file repository, matching
emptyCleanupArmed: without a catalog there is nothing to protect.

* fix(scanner): keep re-probed outages protected through folder-wide cleanup

Fourth Codex review round on #472; both findings confirmed. The first could
destroy data.

reprobeNestedRoots protected a root it found offline only for the scope being
reconciled, then discarded the result. The folder-wide membership reconcile and
the trash sweep afterwards rebuilt their protected set from the initial probe
alone, so rows under a child that dropped mid-scan — already marked missing and
past the removal grace — were hard-deleted by the very scan that noticed the
outage. Accumulate those roots in reprobedRoots, fold them into
protectedScanRoots, and reuse that set for the membership reconcile and sweep
instead of rebuilding. They now also land in ScanResult.UnreachableRoots so the
folder warning reflects the outage rather than presenting a partial scan as
clean.

Snapshot and group pruning was enabled for scopes that were never walked. The
gate was len(walkFailures) == 0, but an unreachable root gets nil walkRoots, so
it has no walk and therefore no failures — and pruning then deleted its
snapshots, observed locations and group locations even though its media rows
were protected. The same held for a suspect-empty child compacted into a
populated parent. Pruning now additionally requires that the scope was actually
walked and contains no protected path.

The new test pins that the sweep honours the protected set it is given. It does
not reproduce the mid-scan race itself: staging that needs the drop to land
between the probe and the walk, which a test cannot reach without hooks. That
path is covered by inspection, and the test comment says so rather than
implying coverage it does not have.

* fix(scanner): route every protection source through one folder-wide set

Fifth Codex review round on #472. Two P1s, one of them the second data-loss
path in this area — and the direct sibling of the one fixed in 35326adc, which
is the reason this commit changes the structure rather than patching another
edge.

Rows beneath a directory the walk could not read were protected only inside
applyScopedScan. The folder-wide protected set was rebuilt from the probe
results alone, so DeleteMissingByFolder could permanently delete rows past the
removal grace under a subtree this scan never managed to read — deleting on the
strength of an observation that was never made.

The recurring defect is structural: protection is discovered in several places
(initial probe, mid-loop re-probe, per-scope walk failures) and consumed in
several more (scoped reconcile, membership reconcile, trash sweep), and each
fix so far has wired up one edge and missed another. Every source now
accumulates folder-wide and every consumer reads the combined set, so a new
source has one place to register instead of several to remember.

reprobeNestedRoots classified from two probe batches. It called
probeUnreachableRoots, then suspectEmptyRoots probed the same paths again; a
child dropping between the samples was reachable to the first and discarded by
the second, which only returns reachable-and-empty roots. It now classifies
both states from one batch, so the disconnect it exists to catch cannot fall
between its own probes.

Re-probed roots kept their classification instead of being collapsed into
unreachableRoots, which had been reporting a suspect-empty child as
unreachable and giving operators contradictory failure information.

The new regression test is verified to fail with the propagation disabled and
pass with it, rather than assumed to cover the path.

Not addressed: the cleanup allowance is read without being reserved, so two
overlapping full scans of one folder can both observe it armed. Narrow, needs
a transactional reserve in the scan-claim query, and is left for follow-up
rather than bundled here.

* fix(scanner): resolve root protection before scoped metadata pruning

Sixth Codex review round on #472.

scanPaths pruned before it knew what was protected. reconcileScannedRoots and
reconcileScannedGroups ran roughly 160 lines ahead of protectedConfiguredRoots,
so a ScanSubtree of a mount that dropped but left a reachable empty mountpoint
walked clean, reported no failures, and pruned root snapshots and observed and
group locations against that empty inventory — preserving the media rows while
deleting the metadata describing them. Protection is now resolved before any
reconciliation, and both prunes share one decision, matching applyScopedScan.

Pending empty scopes never re-probed their nested children. A parent whose only
media lives in a child walks empty when that child drops, so it lands in
pendingEmptyScopes rather than the populated-scope branch where
reprobeNestedRoots ran. Probing the parent alone proves nothing: it still holds
the child's bare mountpoint directory, so it reads present and non-empty. With
a healthy sibling keeping the folder-wide empty guard quiet, nothing protected
the child. Both branches now re-probe.

MissingSkippedProtected never left the scanner. Both ingest-to-result
conversions copied every other cleanup count but not this one, and
events.ScanRunResult had no field, so scan history, completion events and API
responses reported an all-zero no-op for a scan that skipped files because
storage was offline. Added as a new field, which is additive under the v1 API
rules.

Test honesty: the new test does NOT exercise the pending-scope re-probe. It
empties the child before the scan, so the initial probe classifies it and
protection arrives by that path — verified by confirming the test still passes
with the re-probe disabled. It is named and commented for what it does cover.
The mid-scan race behind both re-probe fixes needs the drop to land between the
probe and the walk, which is not reachable from a test without hooks; those
fixes rest on inspection.
2026-07-25 14:47:12 -04:00
QuickandClaude Opus 5 482efdf140 chore: add T3 Code project icon
T3 Code resolves a repository icon from t3.json's iconPath, falling back to
well-known locations such as assets/icon.png. Neither resolved in this repo,
so it showed the default folder icon in the T3 Code sidebar.

The icon reuses the canonical Silo mark unchanged over a navy backdrop,
so the Silo repos stay distinguishable at the ~14px the sidebar renders.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 18:37:51 +00:00
QuickandClaude Opus 5 9e7fe79590 docs: record live TV, IPTV, and .strm as permanent non-goals
Live TV, OTA/DVB tuners, IPTV, EPG/XMLTV guide sync, DVR, and .strm
remote-URL shortcuts are permanently out of scope for Silo. The
deciding factor is app store distribution: the first-party iOS, tvOS,
macOS, and Android clients ship through Apple and Google, and a server
that plays arbitrary remote stream URLs puts the entire client suite at
risk of rejection or takedown, not just the feature. Secondarily, live
TV is a separate product surface whose reliability burden competes with
the core playback path.

This was an undocumented boundary until now, and contributors spent real
effort against it (#419, #420, #474). Write it down in docs/non-goals.md
and summarize it in AGENTS.md so both humans and agents see it before
proposing or implementing in this area.

Refs #474, #419, #295

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 15:07:42 +00:00
355508f6e7 fix(requests): retire stalled targets when presence confirms the media (#470)
reconcileRequest completes a presence-confirmed request only when it has no live targets, so a quality-agnostic TMDB hit cannot orphan an in-flight download. That gate assumes the router eventually moves every target to a terminal state; when it does not, the request is pinned open forever even though the media is in the library.

Retire targets stuck in queued past a 24h horizon with no status transition when presence confirms the media, and let the existing target aggregate drive request status. Targets actively downloading are never retired. Each retirement logs at WARN, since reaching this path means a router is misbehaving.

Part of #469

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 02:54:30 -04:00
Quick a102630365 feat(web): allow tailnet and hostname access to dev server
- Extend Vite `allowedHosts` with `.ts.net`, the machine hostname, and a `VITE_ALLOWED_HOSTS` override
- Document reaching the dev UI over Tailscale in the web-ui-testing skill
2026-07-25 05:27:52 +00:00
Quick 51ef906025 docs: add agent skills and silo-dev helper script
- Add skills for dev-environment debugging, jellycompat diagnosis, Discord triage, and web UI testing under .claude/skills (symlinked as .agents/skills)
- Add scripts/silo-dev plus .silo-dev.env.example for driving local or remote Silo deployments
- Ignore .silo-dev.env and reference the shared skill config in AGENTS.md
2026-07-25 05:16:19 +00:00
ece3578e91 docs: rightsize repo context for Claude 5 context engineering (#467)
Applies the practices from Anthropic's "new rules of context engineering
for Claude 5 generation models" to this repo's always-on context.

AGENTS.md (symlinked as CLAUDE.md) drops content derivable from the
filesystem — the module-by-module structure tour, the Makefile target
list, and most of the style section, all of which Claude reads directly
from internal/, the Makefile, .golangci and web/.prettierrc. What stays
is the part that isn't derivable: the Goose migration rules, the v1
additive-only API contract, multi-repo boundaries, and the workspace
gotchas. Also fixes a dead pointer to .claude/skills/deployment-debugging,
which does not exist; the runbook is the dev-environment-debugging skill.

The external-contributor AI disclosure block moves to
docs/ai-contributions.md, reached by a one-line pointer, so it costs
nothing in the common case where no external PR is being prepared.

issue-to-pr sheds the generic agent hygiene now covered by the harness
system prompt and keeps the Silo-specific gates. Its commit trailer no
longer pins a stale model name.

scripts/jellycompat-diff.sh replaces the hand-typed curl/python
one-liners the jellycompat-diagnosis skill used to carry as prose. It
unions item keys across all returned items rather than reading item[0],
which was hiding fields present on only some items.

Verified: make verify-local-paths, bash -n, shellcheck, and a mock
two-server run of the diff script.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 15:15:19 -04:00
ee31a1f0e2 feat(playback): formalize resumable direct streams and stall observability (#464)
* feat(playback): formalize resumable direct streams and stall observability

Implements #443: strong stat-based ETag + If-Range on original-file direct
play (via http.ServeContent), stream-end outcome classification in
RollingDeadlineWriter (stalled_reap vs client_gone vs completed) with a
structured log event and Prometheus counters, the direct_stream_resume_v1
protocol-v3 capability, and a contract doc. Progressive remux is explicitly
excluded from the resume contract.

Code written by OpenAI Codex CLI (gpt-5.6-sol) from a Claude-authored spec;
reviewed and verified by Claude.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): harden direct stream resume contract

* test(playback): cover resume platform contracts

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 14:22:12 -04:00
22fec4ed2d feat(metadata): add resilient match queue diagnostics (#463)
* feat(metadata): add resilient match queue diagnostics

* fix(metadata): harden match queue lifecycle

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-24 12:18:49 -04:00
383973ec22 feat(metadata): improve match accuracy and localized titles (#461)
* feat(metadata): improve match accuracy and localized titles

* fix(metadata): address matching review findings

* test(catalog): align empty alias snapshot scope

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-24 12:02:52 -04:00
QuickandGitHub 2bc6ebb013 Merge pull request #460 from Silo-Server/agent/repair-anchored-identity-links
fix(scanner): repair historical provider-anchored merges
2026-07-23 17:38:19 -04:00
QuickandGitHub 5e2dd91d6b Merge pull request #456 from Silo-Server/feat/admin-settings-contract
fix(admin): enforce settings contracts end to end
2026-07-23 16:00:55 -04:00
Quick104 920629e0ad fix(admin): address settings contract review findings 2026-07-23 15:14:52 -04:00
Quick104 e625d574e3 fix(scanner): repair historical provider-anchored merges 2026-07-23 15:06:39 -04:00
Quick104 eee63772a7 feat: add AI disclosure requirements to contributing guidelines and issue templates 2026-07-23 14:37:38 -04:00
QuickandGitHub 4d1c97655d Merge pull request #397 from Rhainland/fix/episode-search
feat(search): add TV episode search support
2026-07-23 14:29:57 -04:00
QuickandGitHub 7cfdabb096 Merge pull request #458 from Silo-Server/fix/section-add-insecure-context
fix(web): make Add Section work on insecure (plain-HTTP) origins
2026-07-23 12:06:10 -04:00
Quick104andClaude Fable 5 0834edcfc0 style(web): keep uuid formatting within 100-char line width
Addresses CodeRabbit review on PR #458.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:58:36 -04:00
Quick104andClaude Fable 5 82379fa3e2 fix(web): make Add Section work on insecure (plain-HTTP) origins
Both add-section paths on Settings > Home Screen called crypto.randomUUID()
unguarded. Browsers only expose randomUUID in secure contexts, so on
self-hosted servers accessed over plain HTTP the click handler threw
synchronously and the Add section button appeared dead while Cancel still
worked. api/client.ts and plexAuth.ts already carried ad-hoc fallbacks for
the same problem; extract a shared lib/uuid helper (UUIDv4 via
crypto.getRandomValues, available in insecure contexts) and use it at all
four call sites.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:51:43 -04:00
QuickandGitHub 5d7eec7a62 Merge pull request #457 from Silo-Server/t3code/review-discord-issue
fix(autoscan): match rewrite rules against UNC webhook paths
2026-07-23 11:42:33 -04:00
Quick104andClaude Fable 5 65d94459f4 fix(autoscan): preserve trailing separator through path rewrites
Review follow-up: normalizePath strips a trailing slash, but the slash is
semantic for legacy-scope changes — filepath.Dir("/x/Show/") is the
directory itself while filepath.Dir("/x/Show") is its parent, so dropping
it widened targeted directory notifications into parent/library scans.
applyRewrites now records whether the incoming path ended with a separator
(either form) and restores it on the returned path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:40:38 -04:00
Quick104andClaude Fable 5 6da6cf7a71 fix(search): opt episode documents out of embedder vectors explicitly
A userProvided Meilisearch embedder rejects any document that omits
_vectors entirely, failing the whole indexing task once episode documents
enter a rebuild batch. Episode docs now carry the documented
`_vectors.<embedder>: null` opt-out, and attachDocumentVectors runs the
per-document pass even for episode-only batches.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:40:20 -04:00
Quick104 3c56ed606c fix(admin): enforce settings contracts end to end 2026-07-23 11:24:55 -04:00
QuickandGitHub c0b2e18173 Merge pull request #433 from RXWatcher/fix/ebook-enrichment-architecture
feat(ebooks): decouple metadata enrichment from scans via a durable queue
2026-07-23 11:18:42 -04:00
Quick104andClaude Fable 5 89b7d25084 fix(autoscan): match rewrite rules against UNC webhook paths
Incoming webhook paths were only separator-swapped (normalizeSeparators),
while rewrite From values went through normalizePath, which also collapses
duplicate slashes. A Windows UNC root from a Windows-hosted arr
(\\NAS\Media\TV -> //NAS/Media/TV) therefore never prefix-matched its
rewrite rule (/NAS/Media/TV), so every import logged "webhook paths
matched no library folder" and nothing scanned.

applyRewrites now normalizes the incoming path with the same normalizePath
used for the stored From, and normalizes the joined result so a
trailing-slash To cannot produce a doubled separator.

Reported via internal Discord thread (Sonarr on Windows with a UNC TV
root posting to the autoscan webhook).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:14:45 -04:00
Quick104andClaude Fable 5 84cd0acb0c fix(ebooks): keep legacy-lane rows in their lane until terminal outcomes
Scanner re-enqueues (priority 100) could silently promote pending legacy
backlog rows (priority -100) into the incremental lane via the enqueue
upsert's GREATEST, and the fail/release requeue branches hardcoded 100
regardless of the row's lane. A mass mtime shift or group-key-version
bump would have moved the entire legacy backlog out from under the
backfill task's pacing controls into the scheduled sync task.

Lane changes now happen only through terminal outcomes (complete or
discard); enqueue, fail, and release preserve a negative priority.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 11:10:31 -04:00
Quick104 2e45a9c015 Merge remote-tracking branch 'origin/main' into pr-397-devmerge
# Conflicts:
#	internal/catalog/item_repo.go
2026-07-23 11:01:21 -04:00
QuickandGitHub 0d3b94d2dd Merge pull request #452 from Silo-Server/dependabot/go_modules/google.golang.org/grpc-1.82.1
build(deps): bump google.golang.org/grpc from 1.81.1 to 1.82.1
2026-07-23 10:53:12 -04:00
QuickandGitHub 3f951c8454 Merge pull request #453 from Silo-Server/fix/jellycompat-watch-scrobbling
fix(jellycompat): make watch scrobbling reliable
2026-07-23 09:20:02 -04:00
Quick104 f6615ef8d7 fix(jellycompat): confirm authoritative watch stops 2026-07-23 09:15:19 -04:00
Quick104 ac1dd99c72 fix(jellycompat): address scrobble lifecycle review 2026-07-22 23:22:09 -04:00
Quick104 487dc84829 fix(jellycompat): deduplicate playback negotiations
- Replace unstarted negotiations for the same device and item
- Apply deduplication atomically across durable store instances
2026-07-22 22:01:02 -04:00
Quick104 166c5ef32f Add reliable Jellycompat watch scrobbling
- Forward start, pause, resume, and stop events with stable media identities
- Persist and retry terminal scrobbles across teardown and restart paths
- Reject ambiguous playback-report route matches
2026-07-22 21:41:05 -04:00
dependabot[bot]andGitHub 5ddec8fa4e build(deps): bump google.golang.org/grpc from 1.81.1 to 1.82.1
Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.81.1 to 1.82.1.
- [Release notes](https://github.com/grpc/grpc-go/releases)
- [Commits](https://github.com/grpc/grpc-go/compare/v1.81.1...v1.82.1)

---
updated-dependencies:
- dependency-name: google.golang.org/grpc
  dependency-version: 1.82.1
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-22 22:26:46 +00:00
QuickandGitHub a0507c78eb Merge pull request #447 from Silo-Server/codex/jellycompat-color-range
fix(playback): preserve and expose video color range
2026-07-22 13:38:16 -04:00
Quick104 4e42cb3053 Improve admin diagnostics controls
- Add a client uploads toggle with status refresh and feedback
- Replace native date filters with calendar and time pickers
- Improve responsive filter layout and test upload settings
2026-07-22 13:13:52 -04:00
rxwatcherandClaude Fable 5 bd97f3c123 fix(web): return a cleanup from the preload-error installer
Review feedback: window is shared across vitest tests, so repeated
installs accumulated listeners. The installer now returns a remover and
the tests detach in afterEach.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 8dc2dcbc13 fix(web): reload once when a deployed chunk fails to import
After a deploy, an open tab still references the previous build's
content-hashed chunks and the first lazy navigation dies with "Failed to
fetch dynamically imported module". Handle Vite's vite:preloadError by
reloading onto the current build, guarded to at most one reload per
minute so a persistently missing chunk cannot reload-loop.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 bcc78e7332 fix(server): return 404 for missing content-hashed assets
A /assets/ chunk from a previous build no longer exists after a deploy;
serving the SPA shell at that URL makes the browser fail dynamic imports
on a text/html module. Exclude /assets/ from the SPA fallback so the
miss surfaces as a 404 the client can react to.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 845d969af1 feat(ebooks): run legacy backfill automatically
Give the backfill task a default 15-minute interval trigger. With the
rate-limit cooldown floor each run meets a fresh ready-set, a saturated
batch trips the zero-progress breaker, and an empty lane exits in
milliseconds, so the backlog drains at provider speed unattended. The
canary claim cap and batch delay keep their semantics, and operators can
retune or disable the trigger through the admin task UI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 dfce4972e4 fix(ebooks): floor rate-limited requeue delays
The ebook-metadata plugin attaches ~1s RetryInfo to ResourceExhausted
errors as request-pacing advice for its internal token bucket. Adopting
that hint verbatim as the queue horizon made rate-limited rows claimable
again immediately, so every backfill run re-claimed the same saturated
tail. Clamp rate-limited requeues to a 15m floor (SILO_EBOOK_RATE_LIMIT_COOLDOWN
to tune); hints above the floor are honored up to the existing 24h cap.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 a679b5e1a1 docs(ebooks): backfill automation + rate-limit cooldown design
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 e18bc3c568 docs(ebooks): add enrichment architecture plan
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00
rxwatcherandClaude Fable 5 a2e6cd8d5f docs(scan): document vanished-path rejection branch
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 13:43:57 +02:00