Commit Graph
880 Commits
Author SHA1 Message Date
CoffeeKnyte fc65dcaf16 fix(playback): make session liveness server-observed
Client progress reports could keep a zero-byte phantom session alive and,
worse, make it outrank a genuinely-serving stream when the enforcer picked
over-cap victims.

`streammonitor.LiveLocalSessions` substituted `Session.LastActivityAt` for a
zero `LastServedAt`, and `LastActivityAt` is advanced by `UpdateProgress` and
by the realtime WebSocket hello/ack/result handlers. Since
`streamenforcer.selectVictims` keeps the `limit` most-recently-served streams,
a progress-only phantom sorted ahead of a real stream and the real one was
trimmed instead. Reaping had the same root cause: `sessionIsInactiveLocked`
keyed idleness on `LastActivityAt`, so a client that kept pinging held a
session open forever.

Per decision A5 (Option C), client progress is now UI metadata only and never
feeds enforcement or reaping:

- `LiveLocalSessions` projects `LastServedAt` verbatim, emitting an empty
  timestamp when the session has never served, so it sorts as the stalest
  over-cap victim.
- `sessionIsInactiveLocked` measures idleness from `LastServedAt`, falling back
  only to `StartedAt`.
- A configurable never-served window (`DefaultUnservedSessionGrace`, 2m, via
  `SetUnservedSessionGrace`) keeps a legitimately slow start from being reaped
  before its first byte, without granting a phantom unbounded life. It is a
  separate knob rather than a hardcoded floor so it cannot silently override
  `SetLivenessGracePeriods`.

In-flight transports remain exempt, so direct-play and remux long pours and
per-segment HLS serves are unaffected.

Paused sessions with an open realtime/WebSocket connection are exempt from
reaping. That preserves the issue #243 fix (reaping a paused transcode froze
clients) while staying within Option C: an open, ping-checked connection is
server-observed, unlike a client's reported progress, and the session still
consumes one of the user's cap slots.

Part 1 of 3 for the Batch 4 liveness/replica work.

Part of #305
2026-07-31 00:03:44 +00:00
CoffeeKnyte 4288c7c5c6 docs(playback): score the revocation batch and record its accepted limits
Marks #13, #7/M1, A2/#6, A1/#3 and A3/#5 resolved, and records the limits this
batch accepts rather than leaving them for the next reviewer to rediscover.

Retracts the last of the two false claims flagged in the serve-path docs pass:
the kill switch now does "keep a stream dead" for the token's reconstructable
life. The other -- monitoring does not "never trust client progress" -- is still
false and stays flagged until decision A5 lands.

Newly documented accepted limits:

- A user cutoff cannot cut an API-key-owned pour. API-key credentials carry no
  issue time, so they take the zero credential time and IsRevoked's documented
  fail-open contract applies. Deliberate, logged, and better than substituting
  time.Now(), which actively defeats the cutoff.
- Two central replicas revoking the same key can still race the Redis mirror,
  because mirrorToRedis is an unconditional SET rather than an atomic merge.
  Same-process writes are serialized by opMu; cross-replica convergence needs
  A6's shared picture.
- The over-cap TTL setting affects future revocations only. Monotonic expiry
  means it cannot shorten an existing kill; only an explicit unrevoke clears one,
  and the enforcer may recreate it while the over-count persists.
- Per-login logout cuts are still unavailable: cutting one device's live streams
  needs per-login identity in the stream credential, which is the
  authorization-generation model rather than the iat model chosen here.

Records that the enforcer's repeat kill is non-extending and why: with monotonic
expiry and a 30s evaluation loop, a plain long TTL would renew a wrong kill
indefinitely for as long as any stale record survived.

Part of #305.
2026-07-30 12:27:19 +00:00
CoffeeKnyte e5bf0155ad fix(playback): make revocation state converge and cutoffs credential-accurate
Five defects in revocation state and credential semantics. Lands after the
tracker-lifecycle batch on purpose: raising the over-cap TTL is only safe once
the count feeding it is trustworthy.

#13 -- a longer old revocation suppressed a newer cutoff. applyLocal kept or
replaced the WHOLE record by expiry, so when the existing revocation expired
later the new one was dropped entirely, including its newer RevokedAt. The
durable upsert did the same, with a comment documenting it as intentional.
RevokedAt is the user-kill CUTOFF, so this left a credential issued between the
two cutoffs valid -- a second admin kill after a user re-authenticates silently
failed to cut them. The two fields now merge independently: ExpiresAt stays
monotonic, RevokedAt advances to the later value, and reason follows the newer
cutoff. Both superseded comments are replaced rather than left contradicting the
code. Session-kind revocation still ignores RevokedAt, so the enforcer's
re-revoke cannot weaken a session kill.

Also fixed while here: Redis received the merged record but pub/sub published the
raw input one, so under pub/sub-only delivery (Redis down) an edge got the newer
short record without the older long expiry and lost monotonicity. Both now carry
the merged record.

#7 + M1 -- Postgres could indefinitely block the urgent Redis kill.
RevokeWithWarnings held the global opMu across all propagation, stripped the
caller's deadline with WithoutCancel, and did the durable Postgres upsert BEFORE
Redis, on a pool with no statement timeout. The local kill still applied, so
playback on that process was fine -- but edge propagation, pub/sub, the admin
response and every later revoke/unrevoke stalled behind the lock. Redis and
pub/sub now go first, and the detached context is bounded. WithoutCancel is kept
deliberately: propagation must outlive an aborted admin request.

opMu scope is deliberately NOT narrowed. mirrorToRedis is an unconditional SET
with no atomic merge, so same-process serialization is what stops an older value
overwriting a newer one; narrowing the lock would also let an unrevoke interleave
with a revoke's propagation. Bounding the context caps how long the lock can be
held, which is the actual reported harm. The remaining cross-replica race -- two
central replicas racing the same SET -- is documented, not half-fixed; it needs
A6's shared picture.

A2 / #6 -- a missed unrevoke got resurrected. In-memory tombstones already
existed, but being process-local they did not survive a restart or reach a
replica that missed the pub/sub event, so maintain's durable self-heal
re-Upserted the surviving entry and the ban returned. Tombstones are now durable,
via two nullable columns on stream_revocations rather than a second table: a
tombstone is a state of the same key, and it needs its own expiry horizon
separate from the revocation's. The upsert rejects a stale replica's write while
a tombstone is live but lets a genuinely newer revocation clear it, and warm
paths apply tombstones BEFORE revocations so an un-banned key cannot be restored
as a live kill. Tombstones are pruned on the same sweep, so the table cannot grow
without bound.

A1 / #3 -- over-cap kills reopened after 5 minutes while the token stayed
reconstructable for 24h. The TTL now derives from playback.MaxTokenTTL rather
than duplicating 24h, behind a validated setting.

Critically, the enforcer uses a revoke-if-absent path rather than re-revoking.
Expiry is monotonic and the enforcer re-evaluates every 30s, so a plain long TTL
would slide expiry forward by another full lifetime on every pass -- making a
wrong kill effectively permanent for as long as any stale record persisted, with
only an explicit unrevoke to recover it. Admin Revoke keeps its monotonic
behaviour; only the enforcer's own repeat kill is non-extending. The setting is
documented as affecting future revocations only, since monotonic expiry means it
cannot shorten one already issued.

A3 / #5 -- the user cutoff compared against a fresh time.Now() taken at request
entry, so a request from a pre-cutoff login could look post-cutoff and escape the
kill. The credential time is now the access token's iat.

Two deliberate choices worth stating. API-key credentials carry no issue time, so
they pass the zero time and, per IsRevoked's documented contract, are never
matched by a user cutoff: a user kill provably cannot cut an API-key-owned pour.
That is an accepted, logged, documented hole -- and strictly better than
time.Now(), which actively defeats the cutoff. And jellycompat uses the compat
session's CreatedAt rather than the bridged Silo token's iat, because that token
refreshes without a new Jellyfin login, so its iat would advance on refresh and
let a refreshed credential slip past a cutoff.

Stream tokens are now bound to their route: a token whose SessionID does not match
the URL's session_id is rejected with 403 instead of being silently ignored,
matching the reconstruction helper that already refused a different session.

Per-login logout cuts remain out of scope -- they need per-login identity in the
stream credential. S5 (the (sessionID, userID, startedAt) clump) is rejected as
ceremony now that A3 is the iat option rather than the generation model.

#12 (closing an RSS feed does not cut its current pour) is deferred: it needs a
namespaced revocation id that cannot collide with real session ids, that id
threaded onto public feed requests, and protection against a new feed inheriting
an old tombstone.

Part of #305.
2026-07-30 12:27:05 +00:00
CoffeeKnyte b1502d70f5 docs(playback): score the tracker-lifecycle batch and correct the async claim
Marks GAP-15 resolved and records the three wrong-over-cap-count defects the
tracker-lifecycle batch closed, with why they had to precede the revocation
batch: decision A1 removes the self-healing that limits the damage of a miscount.

Notes why the v3 identity split had not yet caused visible harm -- fresh v3
starts sent no owner attribution at all, so the transport record landed under
user 0, which the enforcer skips, silently exempting the stream from the cap
rather than double-counting it. That is a worse failure than the double count it
masked, and worth recording so the next reader does not "fix" only the visible
half.

Corrects the "fully async monitoring" claim instead of rushing the queue: the
first Redis projection write per session is synchronous so the record is visible
before the request returns, and later liveness and byte updates ride the refresh
tick. The consequence -- a slow Redis adds latency to the first request of a
stream -- is now stated. The ordered, lifecycle-aware projection queue is
deliberately deferred until its startup, drain, cleanup ordering, backpressure
and refresh interaction can be designed together; a naive fire-and-forget
projection is exactly what caused the ghost-session defect this batch fixed.

Follow-up list renumbered accordingly; the GAP-14 (A7) and opMu items are
retained, not dropped.

Part of #305.
2026-07-30 09:10:22 +00:00
CoffeeKnyte ecb4555eec fix(playback): make edge tracking lifecycle-safe and session identity canonical
Five defects that all produced a WRONG over-cap count, which is why they land
before the revocation batch: decision A1 raises the over-cap revocation TTL from
5m to ~24h, removing the self-healing that currently limits the damage of a
miscount. A false positive after A1 blocks a legitimate stream for a day, so the
count has to be trustworthy first.

#1 -- overlapping edge requests deleted a live stream. Tracker.sessions was a
set and Remove tore down all state plus the Redis key, while both proxy pour
handlers deferred removal unconditionally. Two overlapping Range GETs on one
session id -- ordinary seek behaviour -- meant the first to finish deleted the
record while the second was still pouring, and later AddBytes calls were then
dropped because AddBytes ignores bytes for a session with no live record. The
stream went invisible to authoritative monitoring while still serving.

Track now returns a Lease that the request-scoped caller releases exactly once;
teardown happens when the last live lease is released. A plain refcount would
have been wrong: Track(A) -> Remove -> Track(B) -> Release(A) decrements B, and
clamping at zero does not help because the count legitimately belongs to B. That
is not hypothetical -- the transcode node deliberately replaces sessions under
the same id so a quality switch does not orphan ffmpeg, and it calls
unconditional Remove from its reaper and stop paths. So each generation carries
an epoch, Remove and Cleanup bump it, and a release from a superseded generation
is a logged no-op. Lease identity is a set rather than a counter, which makes a
duplicate release detectable instead of silently destructive.

The transcode node keeps using Remove: its Track calls are not request-scoped and
are correctly owned by session lifecycle. "Every Track needs a paired Release" is
true only of the request-scoped callers.

#8 -- async transcode tracking could leave a permanent ghost. The tracking write
ran as a bare goroutine with a WithoutCancel context, so if stop won the race the
delayed Track recreated the record after cleanup -- and because it landed in
sessions, Snapshot treated it as live until Remove and it NEVER idle-expired. A
permanent phantom inflating its owner's count, able to trigger false over-cap
kills of that user's real streams. The write now takes the per-session lifecycle
lock that stop and reap already hold, and re-checks session pointer identity
before writing, so a stopped or replaced generation cannot resurrect a record.
Pointer identity rather than id equality is what makes same-id replacement safe.
The write stays off the request path -- the API server and the playback client
are blocked on the 202.

#9 + M3 -- protocol-v3 counted one stream twice. The stream token carries a
transport id distinct from the logical session id, and the node tracked under the
transport id while the API/proxy record used the logical one, so mergeStreams saw
two streams. M3 was the reason this had not yet bitten: the v3 fresh-start caller
sent no owner attribution at all, so the transport record landed under user 0,
which the enforcer skips -- silently exempting the stream from the cap entirely.
Fresh v3 starts now carry the logical session id and full owner attribution
(both were already in scope at the call site), and merging is keyed on logical
identity where present via one shared helper used by both merge functions, which
had already drifted apart once.

The enforcer view resolves SessionID to the logical id so a kill targets the real
session rather than a replaceable transport generation. The raw admin view keeps
the transport id and exposes logical_session_id as an additive omitempty field,
advertised on the node-sessions capability endpoint, so the v1 response shape is
unchanged.

GAP-15 -- edge transcode liveness was request-observed. touchTranscodeSession
fired before proxying, so hammering dead segment URLs advanced LastServedAt with
zero bytes served. Visibility and liveness are now separate operations:
EnsureEphemeral makes a session visible without claiming bytes were served, and
served-byte liveness advances only from a 2xx/206 upstream response. Previously
the proxy metered every upstream body regardless of status, so a node 404's error
body counted as served bytes -- moving the touch later would not have fixed it.

S4 -- LiveLocalSessions moved from the HTTP handlers package to streammonitor,
which owns monitoring. A background enforcer importing api/handlers was
backwards. Pure move; its existing mapping assertions moved with it. The
LastActivityAt fallback inside it is left as-is -- decision A5 removes it in the
liveness batch.

Verified with go test -race across nodesessions, proxy and transcodenode; the
overlap regression test was confirmed to fail under the old unconditional
teardown.

Part of #305.
2026-07-30 09:09:35 +00:00
CoffeeKnyte 8a9cf04a41 docs(playback): score the serve-path fixes and record the settled decisions
The two coverage matrices and the plan's as-built deltas described GAP-10..GAP-15
as open and asserted things the code did not do. Rescored against the code as it
now stands, and made the remaining overclaims explicit rather than leaving them
to be discovered by the next reviewer.

Marked resolved with the reasoning that produced each fix: GAP-10 (ABS Unwrap),
GAP-11 (ebook observability), GAP-12 (rolling-deadline cut latch, pulled forward
from the revocation batch because it makes GAP-10's fix inert), GAP-13
(BytesServed merged as a max). GAP-14 and GAP-15 stay open with their batch and,
for GAP-14, decision A7 attached.

Two claims are called out as STILL FALSE wherever they appear, so the blanket
phrasing does not creep back: the kill switch does not "keep a stream dead" (an
over-cap kill reopens after 5m until A1 lands) and monitoring does not "never
trust client progress" (the LastActivityAt fallback survives until A5 lands).

Documents a bound the plan never stated: the in-flight watcher polls, so a cut
lands within one interval -- 5s in production -- not instantly. Every "cut" claim
in these docs now says so.

Corrects two pieces of stale guidance that would have misled the next
implementer. Follow-up 0b said to reuse guardRevocationCut for the ebook routes:
that wrapper keys on a session_id URL param and an ?st= token those routes do not
carry, so it compiles and silently guards nothing. Follow-up 0c suggested a
sticky flag on RollingDeadlineWriter: unreachable as written, because the rolling
writer is constructed inside the serve helpers and wraps the writer the watcher
holds, so Unwrap() -- which walks toward the socket -- can never reach it.

Records decisions A1-A8 in one table so the follow-up issues inherit them, plus
one new finding for the A3 batch: streamRequestIdentity verifies a stream token's
signature but never checks its SessionID against the URL's session_id.

Expands the AI-use disclosure to the standard docs/ai-contributions.md requires:
tool, exact model IDs, involvement classification, and the adversarial findings
and resolution from both directions of the cross-model review -- including the
five defects found in the AI implementation and the one found in the AI review of
it. Also states plainly that pnpm is absent on the host used for this round, so
frontend evidence must come from CI, and that the jellycompat test package does
not compile on main, so the compat changes here are verified by reading only.

Part of #305.
2026-07-30 08:37:22 +00:00
CoffeeKnyte 70064823a6 feat(api): advertise stream monitoring and revocation capabilities
The admin kill-list endpoints and the transfers field on the live node-sessions
payload shipped with no capability advertisement, contrary to the repo's
additive-v1 rule that new features expose capability endpoints for feature
detection rather than relying on version sniffing. (S3)

Adds two endpoints, each mounted beside the surface it describes so an
advertisement cannot outlive the route it advertises:

  GET /admin/node-sessions/capabilities
  GET /admin/streams/revocations/capabilities

The node-sessions capability keeps schema support and runtime availability as
separate booleans. The transfers key is always present in the response shape
once that endpoint exists, but the process-local registry behind it is optional
wiring -- an edge deployment can legitimately serve it as an empty list forever.
Collapsing the two into one flag would advertise download monitoring that is not
actually running, so `transfers` reports the schema and `transfers_active`
reports the wiring.

The revocation capability advertises the closed {kind} vocabulary accepted by
DELETE /admin/streams/revocations/{kind}/{id}. To make that advertisement
impossible to desync, the accepted kinds move into a single map that both the
wire parser and the capability handler read -- previously the parser duplicated
the list in a switch, so a newly-accepted kind could go unadvertised and clients
would feature-detect an incomplete vocabulary. The kinds are returned sorted,
because map iteration order is randomised and this is a wire response that must
be stable across calls.

Tests cover the drift in both directions over the whole vocabulary rather than
sampling rejected strings, the sort stability, and that the handler returns a
copy so a caller cannot corrupt the package-level vocabulary.

Note the previous plan for this work put both capabilities on
/admin/sessions/capabilities. That was wrong: that endpoint documents the
Postgres-backed /admin/sessions payload, whereas transfers belongs to
/admin/node-sessions, which is gated on NodeRepo and may not be mounted at all.

No frontend change: web/src does not consume either endpoint.

Part of #305.
2026-07-30 08:31:48 +00:00
CoffeeKnyte 7adedf6429 fix(playback): close the serve-path monitoring and kill-switch gaps
Six defects on byte-serving paths, all of which made the PR's monitoring and
kill-switch claims narrower than documented.

#2/GAP-10 -- the ABS in-flight kill switch was a production no-op. accessLog
wraps every ABS route, and its statusRecorder implemented Write, WriteHeader,
Hijack and Flush but not Unwrap, so http.NewResponseController dead-ended and
SetWriteDeadline returned ErrNotSupported. A multi-GB audiobook pour survived a
RevokeUser. The existing test passed throughout because it called handlers
directly and never saw the middleware; the new test drives the mounted router
over a real socket, and both new assertions fail if Unwrap is removed again.

GAP-11 -- ebook, comic and PDF serving was invisible and un-killable: no meter,
no transfer record, no Refuse, no watcher, on a route that serves cbz/cbr/pdf
files routinely 100 MB-1 GB+. It now follows the ABS file-handler idiom. Note
guardRevocationCut is deliberately *not* reused here: it keys on a session_id
URL param and an ?st= token this route does not carry, so it would have
compiled and silently guarded nothing. Cap-exempt per decision A4 -- admission
is untouched and neither route consumes a stream slot.

#10 -- the no-proxy remote transcode hop forwarded segments through a bare
RollingDeadlineWriter, so bytes on the API hop went unaccounted for a supported
topology. Metering is scoped to media bodies; manifests are excluded so a
rewritten playlist is not counted as media, and a mid-copy failure is no longer
silently discarded.

#16 -- native, proxy and compat subtitle pours were entry-gated only. They now
carry a transport span, a meter and an in-flight watcher. Proxy subtitle bytes
are attributed only when a tracker record already exists: taking Track/Remove
lifecycle ownership per subtitle request would walk straight into the
overlapping-request defect (#1) that Batch 2 addresses. Compat subtitle
extraction is buffered and rejects bitmap formats, so a cut stops delivery but
not extraction already in progress.

M2 -- the proxy deferred tracker Remove with the request context, which is
already canceled on client disconnect, so the Redis DEL never happened and the
key lingered until TTL -- a false over-cap window that could get a legitimate
stream killed. Cleanup now uses a short bounded context.

GAP-13 -- mergeStreams took the freshest record wholesale and never merged
BytesServed, so a stream that poured 8 GiB at an edge could report 0. Merged as
a max, not a sum: the records are two observers of one pour. Fixed in
DedupeSessionInfos too, which had the same hole and feeds the admin view.

#11 needed no behaviour change -- that route was already metered, registered and
watched, and commit c24d8396 plus this Unwrap fix are what make its cut work.
Its comment claimed download-class exemption while the comment above it said the
?token= form is for iOS streaming; both facts and decision A4 are now stated.

Part of #305.
2026-07-30 06:11:43 +00:00
CoffeeKnyte c24d839681 fix(playback): latch the revocation cut against rolling write deadlines
WatchAndCut set the socket write deadline to now once and returned. On every
pour wrapped in httpstream.RollingDeadlineWriter, bump() pushes that deadline
back out to now+180s before the next write once the 15s bumpStep has elapsed,
and the constructor bumps immediately. Nothing re-armed the watcher, so a
revocation cut was reliable against a stalled pour and unreliable against a
fast-draining one -- weakest against exactly the ripping case it exists to stop.
(GAP-12)

The obvious fix does not work: the rolling writer is constructed *inside*
ServeDirectPlay/ServeRemux and wraps the writer the watcher holds, so it sits
*above* the watcher. Unwrap() walks toward the socket, so the watcher can never
reach it by writer introspection. The cut therefore has to travel by a side
channel.

Adds httpstream.CutLatch, carried on the request context, which RollingDeadline-
Writer consults in bump(). Once latched, the writer never extends the deadline
again -- a cut is a deliberate hang-up, not a stall. bump() re-checks the latch
after setting a future deadline so a concurrent cut cannot be lost to the
check/set race, and WatchAndCut now keeps re-applying the deadline on each tick
instead of returning after the first cut, as belt-and-braces for any writer
topology the latch does not reach.

A failed SetWriteDeadline is now logged instead of silently discarded, so the
next wrapper that breaks the Unwrap chain is loud rather than invisible. It is
logged once per watcher, since the re-applying tick would otherwise repeat it
every 5s for the life of the pour.

WatchAndCutContext and NewRollingDeadlineWriterCtx are added alongside the
existing signatures rather than replacing them, so this commit changes no
caller behaviour on its own. Options.WatchInterval makes the 5s poll injectable
for real-socket tests; the production default is unchanged.

Note the polling bound this leaves: a revoked pour keeps delivering for up to
one watch interval (5s in production) before the cut lands.

Part of #305.
2026-07-30 06:10:33 +00:00
CoffeeKnyte 09de4f08ec feat(downloads): monitor in-flight download pours in memory
Six routes poured full media with no monitor record and no byte measurement:
the two native download routes, the Jellyfin-compat download, both ABS file
variants, and the public ABS RSS feed file. They were the last invisible bytes
on the server.

- internal/transfers: a process-local, in-memory registry of active pours.
  Bounded (10k entries) with rate-limited "full" warnings, all request-derived
  strings normalized and length-clamped (ABS and jellycompat carry no bounded
  client name, so those fields are header-derived and untrusted), overflow-safe
  byte accumulation, deterministic snapshots, nil-safe throughout. No I/O on any
  path, and no persistence: a pour dies with the process, so durable rows would
  only need reaping after a crash.
- Reuses the existing playback.SessionMeteredWriter via ServedBytesRecorder
  rather than adding a second writer — that writer is where the sendfile
  (ReadFrom) and kill-cut (Unwrap) hazards live and both have regressed before.
- Deliberately NOT plumbed through streammonitor/streamenforcer, which are
  untouched. Downloads stay off the live-stream path by construction: there is
  no type, field or collection through which one can reach the enforcer, so a
  download can never be counted against max_streams or trimmed as an over-cap
  stream.
- Admin visibility is a sibling `transfers` array on the existing
  /admin/nodes/sessions response; the `sessions` array is byte-for-byte
  unchanged. Transfers appear only in the unfiltered listing, since a node_id
  filter targets an edge and these are process-local.
- Per-pour ids are unique, never the download id: concurrent and repeated
  GET/Range requests against one download row are legitimate and would collide.
  The id is minted in the handler (where the revocation watcher is armed) and
  the registry entry is opened in the service only after file/artifact
  resolution succeeds, so a failed auth or lookup never registers a transfer.
- Defer order is an invariant at every call site: End is registered before the
  meter's Close so LIFO flushes the tail first. Reversed, a final sub-1MiB flush
  lands on an unknown id and is silently lost. Pinned by a test.

No schema change and no migration: downloads.bytes_sent keeps its documented
meaning (a lifecycle marker set to file_size on completion, not a live counter)
and is untouched.

Phase 2 — killing a single download, and a standing per-user download block —
is deliberately deferred. A user revocation already cuts in-flight download
pours; it is a cutoff, so it does not refuse new ones. The unique per-pour id
exists so phase 2 only changes "" to that id at each WatchAndCut site.

Part of the stream monitoring & kill-switch epic.
2026-07-29 16:42:37 +00:00
CoffeeKnyte 561e00e4ca docs(playback): record the post-plan hardening pass in the plan's as-built deltas
The plan doc records post-plan changes in its As-built deltas section; the
follow-up audit's four findings (integrated byte accounting, the VERIFY-4
transcode grace, the operator kill-list API with unrevoke, and the previously
unguarded Audiobookshelf byte routes) belong there so the intent doc and the
as-built matrix agree. Also restates what stays deferred.

Part of the stream monitoring & kill-switch epic.
2026-07-29 15:22:34 +00:00
CoffeeKnyte f3e1fd3b38 docs(playback): refresh the monitoring and abuse coverage matrices
Both matrices are the stated review checklist for this area, so a stale claim
in them is a real defect.

- stream-abuse-matrix: "correction #1", row A8 and follow-up #1 claimed the
  async enforcer reads the raw users.max_streams column and is "dormant on a
  default install". That was fixed earlier on this branch by resolving
  SessionManager.EffectiveLimits (the same group-merged policy synchronous
  admission uses); all three now describe the shipped behaviour. Added rows for
  the ABS authenticated file route, the ABS public track and the public RSS
  feed file route.
- playback-paths matrix: added ABS to the serving/monitoring/kill tables, marked
  VERIFY-4 resolved with the served-at-driven transcode grace, recorded
  integrated BytesServed and the deliberately distinct server-observed
  LastServedAt, and documented the kill-list endpoints, Unrevoke, the tombstone
  and the publish-failure fail-safe.
- Restated as explicitly deferred rather than implied: multi-replica
  enforcement (VERIFY-3) and the transcode-node ghost-record sweep. Both need
  the monitoring write side restructured and deserve their own change.

Part of the stream monitoring & kill-switch epic.
2026-07-29 15:17:46 +00:00
CoffeeKnyte e35efddac7 fix(audiobooks): enforce the stream kill switch on ABS byte-serving routes
The Audiobookshelf-compat surface sits outside the monitoring and kill design
entirely — neither architecture matrix mentioned it. Three routes pour full
media and none consulted the kill switch:

- /(abs/)api/items/{id}/file/{ino} and /download — bearer auth, no revocation
  check, no in-flight cut.
- /(abs/)public/session/{sid}/track/{idx} — mounted outside bearerAuth (the
  session id is the capability); it held a transport marker but was unkillable.
- /feed/{slug}/file/{ino} — no auth at all, the slug is the capability;
  invisible and unkillable, and closing a feed only blocked the next request
  rather than cutting a pour already in flight.

Each surface now passes its real credential-issue time, because a user
revocation is a cutoff, not a ban: it matches only streams whose credential
predates it. ABS bearer tokens are stateless JWTs that OnUserSessionsRevoked
does not delete, so passing request-entry time (as the jellycompat login path
safely does) would have meant a user kill never refused a later ABS request.

- bearerAuth carries the JWT's iat into ctxAuth; the authenticated file route
  uses it.
- The public track uses the persisted playback session's StartedAt and passes
  the native session id so session-level kills land too, plus the shared
  metered writer for byte accounting.
- The feed file uses the feed's CreatedAt for an owner cutoff — a feed opened
  before the revocation dies, one opened after re-authenticating serves — and
  arms the in-flight cut.

The authenticated file route stays download-class: like the native and
jellycompat download routes it is exempt from the live-stream cap and from
streammonitor, and is covered by no download quota. It is now killable and
explicitly documented rather than quietly invisible. Bringing all three
download-class routes under one quota and one monitor record is tracked as a
follow-up rather than adding a fourth per-route model here.

Part of the stream monitoring & kill-switch epic.
2026-07-29 15:17:36 +00:00
CoffeeKnyte fca9f38a33 feat(api): operator-visible stream kill list with unrevoke
The kill switch had no operator surface. streamrevoke.Store.List() existed and
was called from nowhere, so an admin could only terminate one session by id or
revoke a user's streams as a side effect of editing their account — with no way
to see what was revoked, choose a TTL, or undo a mistake. Because expiry is
deliberately monotonic, a wrong 24h kill was irreversible.

- GET/POST/DELETE /api/v1/admin/streams/revocations, admin-only, additive.
- Explicit wire-to-internal kind mapping: the wire accepts "session" (and
  "sess"), the store key is "sess". Passing the wire string straight into a Key
  would create a revocation IsRevoked never consults.
- Validation: non-empty bounded session ids; canonical positive user ids
  (strconv.Itoa round-trip, so "01" is rejected — the cache key is the
  canonical form); ttl_seconds bounded to 30d so the duration cannot overflow;
  bounded reason and request body. DELETE of an absent key is idempotent.
- Store.Unrevoke, guarded by a bounded in-memory tombstone: the tombstone is
  installed and the local entry dropped BEFORE the slow durable/Redis deletes,
  so a concurrent poll reconcile cannot re-apply the row it just read. A newer
  Revoke on the same key clears the tombstone, so an unrevoke never suppresses
  a later legitimate kill. Tombstones age out with the kill they replaced.
- maintain() takes the same operation lock as Revoke/Unrevoke around its
  durable block, closing the window where a poll tick could re-Upsert a row
  Unrevoke had just deleted — invisible until a restart resurrected the kill.
- Durable self-heal now compares expiry, not mere presence, so a failed Revoke
  mirror leaves a stale shorter row that the next tick repairs.
- A failed unrevoke publish fails safe: other processes keep the kill until it
  expires. Propagation failures surface as warnings rather than weakening the
  local result.

Unrevoking an over-cap victim is legal but the async enforcer will re-revoke it
on its next pass while the user is still over cap; that is documented at the
endpoint.

Part of the stream monitoring & kill-switch epic.
2026-07-29 15:17:15 +00:00
CoffeeKnyte 4b2031fdf8 feat(playback): account served bytes and hold transcode liveness server-side
Integrated deployments recorded no served bytes at all: BytesServed was only
ever advanced by the edge writer, so a single-node install showed
bytes_served: 0 for every stream. Worse, integrated LastServedAt was mapped
from Session.LastActivityAt, which UpdateProgress advances — letting a client
influence which of its own streams the over-cap enforcer trims first.

- internal/playback: Session gains BytesServed and a distinct LastServedAt.
  LastServedAt advances only from server-observed events (AddServedBytes,
  BeginTransport, EndTransport) and never from a client progress report.
  LastActivityAt keeps its exact prior meaning and reaping semantics.
- internal/playback/metered_writer: one shared SessionMeteredWriter for every
  integrated pour. It forwards io.ReaderFrom so the kernel sendfile path
  survives, guards the fallback copy against re-entering ReadFrom, and
  implements Unwrap() so the revocation cut's SetWriteDeadline still reaches
  the socket. Both properties have regressed on this branch before (GAP-3,
  GAP-9) and are now pinned by a chain test.
- Wired at every integrated pour: native direct-play/remux and transcode
  segment, jellycompat direct/remux and HLS segment. Each site defers the tail
  flush; the wrapper alone loses the final partial chunk.
- Close VERIFY-4 (buffer-ahead evasion): an unpaused transcode gets a bounded
  10m grace measured from the server-observed clock. This covers local,
  cleanly-completed and offloaded transcodes uniformly, unlike an ffmpeg
  liveness probe, which sees only local processes and reports false once a
  copy-mode encode finishes ahead of playback. Paused sessions keep their
  load-bearing 30m grace.
- internal/nodesessions: the same idle window one layer out — transcode
  records idle out at 180s instead of 60s, applied through a single helper
  used by ActiveCount, Snapshot and refreshAll so the node's count and status
  view cannot disagree.

Part of the stream monitoring & kill-switch epic.
2026-07-29 15:16:59 +00:00
CoffeeKnyte 6ae1ea30c8 docs: adversarial abuse matrix + enforcement architecture decision
- docs/architecture/stream-abuse-matrix.md: red-team of this branch against
  ~29 abuse stories, each scored on detection vs enforcement, including the
  three places the branch's own docs oversold the code.
- docs/superpowers/plans/2026-07-07-abuse-cold-enforcer-architecture.md:
  the deny-list vs lease analysis and phased plan, with the final decision
  note: keep the single revocation pipeline with reason-scoped TTLs; the
  two-tier lease is dropped.

Part of #306.
2026-07-29 12:28:06 +00:00
CoffeeKnyte 0cc706e066 fix(playback): resolve group-merged cap in async stream enforcer
The enforcer's LimitFunc read the raw users.max_streams column, which is 0
("inherit from group") for every standard account since migrations
20260702180000/20260702190000 moved the real cap into the Default Group.
The enforcer treats limit <= 0 as unlimited, so the async over-cap brain
never trimmed anyone on a default install — only per-process synchronous
admission held, leaving the cross-node backstop it was built for a no-op.

Extract admission's limit lookup into a shared SessionLimitProvider
(GetByID + access.EffectivePolicyForUser) and feed the enforcer through
it, so admission and the enforcer can never disagree about a user's
effective cap again.

Part of #306.
2026-07-29 12:28:06 +00:00
CoffeeKnyte 53c9313b57 docs(playback): monitoring & kill-switch plan + as-built coverage matrix
Capture the design intent and the shipped coverage for the stream monitoring +
revocation kill-switch work, refreshed to match the final implementation and the
monitoring → kill-switch → docs commit layout.

- docs/superpowers/plans/2026-07-04-stream-monitoring-and-kill-switch.md: the
  implementation plan of record (detection / enforcement / async brain split),
  with a Status banner and as-built deltas noting what shipped differently
  (durable Postgres mirror wired non-optional, shared Refuse/WatchAndCut
  helpers, monotonic expiry, ownership carry-forward, first-class monitoring
  fields).
- docs/architecture/playback-paths-monitoring-kill-matrix.md: the as-built
  coverage matrix across server layout × playback type × route, GAP-1..GAP-9
  resolution notes, restart-durability axis, and open verification items
  (VERIFY-3/4, operator config keys, compat download quota).

Part of the stream monitoring & kill-switch epic.
2026-07-29 12:26:03 +00:00
CoffeeKnyte 0217cce2df feat(playback): stream kill switch + async over-cap enforcer
Add the enforcement layer on top of server-observed monitoring: a revocation
kill switch that stops any stream within ~120s and keeps it dead, plus an async
over-cap enforcer that drives kills off the live monitoring picture — entirely
off the per-segment hot path and with no client-protocol change.

- internal/streamrevoke: the central kill list. IsRevoked is a pure in-memory
  lookup safe on the request hot path; a Redis pub/sub + poll mirror keeps edge
  caches current, and a Postgres durable mirror lets kills survive a server
  restart AND a Redis flush so a restart-resilient stream cannot be reconstructed
  and re-served after being killed. A user revocation is a cutoff (kills tokens
  minted before it, spares post-reauth tokens), not a 24h ban.
- internal/streamenforcer: async over-cap brain — reads the monitoring snapshot
  and per-user limits, selects victims, and collapses every reason (exceeded
  limit, admin terminate, abuse) to the same action: write a revocation.
- Edge + native + jellycompat enforcement: proxy refuses revoked sessions on
  every request and cuts long direct-play/remux pours mid-stream; the transcode
  node guards both serve and the reconstruct path so a killed session is never
  re-spawned after a node restart; jellycompat serve surfaces close their
  kill-switch coverage holes.
- streamtoken.IssuedTime exposes the token iat the user-kill cutoff compares
  against; token IssuedTime + revocation guards wire through router, downloads,
  and admin terminate-by-id (with admin-list dedupe).
- Restore sendfile zero-copy on direct-play/remux byte counting so the monitor's
  served-byte accounting does not cost the sendfile fast path.
- migrations/sql: stream_revocations durable table.

Part of the stream monitoring & kill-switch epic.
2026-07-29 12:26:03 +00:00
CoffeeKnyte 22ffaad911 feat(playback): server-observed stream monitoring (async, no client trust)
Introduce a first-class, authoritative view of what is actually streaming,
observed server-side and never trusting client progress reports. This is the
base observation layer the kill switch and async over-cap enforcer build on.

- internal/streammonitor: live-stream snapshot model plus pluggable Sources
  (local func source, Redis source, multi-source fan-in) so a single node and a
  multi-node deployment expose the same picture.
- internal/nodesessions/tracker: serve-activity attribution — LastServedAt and
  served-byte counters advance from real serving, not client pings, giving an
  authoritative liveness signal.
- Client identity as monitoring attribution: Origin ("native" | "jellycompat")
  and ClientName ride the server-signed stream token (streamtoken.Claims) and
  the transcode-start request so an edge/transcode node — which never sees the
  originating API path — can stamp them onto its live-session record. These are
  attribution only: not byte-affecting and not a trust assertion.
- Serve-activity marks on the transcode node (MarkServed) so a node's own record
  reflects real serving instead of a LastServedAt frozen at start time.
- Admin observation surfaces: node/session listing carries owner + client
  identity and dedupes multi-record sessions.

Part of the stream monitoring & kill-switch epic.
2026-07-29 12:23:48 +00:00
rxwatcherandQuick 08035f3806 fix(admin): validate AI settings drafts 2026-07-28 21:49:41 -04:00
rxwatcherandQuick f0bc170113 feat(admin): clarify AI service configuration 2026-07-28 21:49:41 -04:00
7ab393fc3a fix(jellycompat): resolve item duration probed-first with runtime fallback (#493)
* fix(jellycompat): resolve item duration probed-first with runtime fallback

Jellyfin-protocol clients received no runtime at all for items whose catalog
runtime is 0. RunTimeTicks is omitempty, so a zero value is dropped from the
JSON entirely rather than sent as 0, and strict clients (Infuse) abandon
playback on those items. On the production deployment 5,245 movies have
media_items.runtime = 0 while 5,239 of them have a correct probed
media_files.duration.

Resolve duration at read time the way /api/v1 already does: probed file
duration first, catalog runtime as the fallback. The item row is deliberately
not backfilled — one item can have several versions of different lengths, so
per-file data does not belong there.

- scanner: FirstDurationsByContentIDs / FirstDurationsByEpisodeIDs, batched
  lookups using the same "first live file with duration > 0, ordered by id"
  rule as the v1 API's contentDurationSeconds. The episode_id IS NULL guard on
  the content-id query is load-bearing: every episode file carries its series'
  content_id, so without it a series row would report an episode's duration.
- catalog: optional batchDurationFetcher extension on DetailService, following
  the existing extraFileFetcher pattern so test fakes need no changes.
  Nil-receiver safe and fail-soft — a failed lookup logs and degrades to the
  catalog runtime rather than failing the page.
- jellycompat: DurationSeconds on upstreamListItem/upstreamEpisode, a shared
  runtimeTicks resolver, and fillListItemDurations wired into the nine page
  producers. Fixes the three sites that had no fallback (itemFromList,
  episodeFromUpstream, HandleSearchHints); the detail and PlaybackInfo paths
  were already correct.

This is additive within the v1 rules — it populates a field that was
previously omitted. No field is renamed, removed, retyped, or repurposed.

* fix(jellycompat): avoid duplicate duration lookups

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-28 21:33:30 -04:00
c52ca7dd7a feat(admin): identify compat sessions and Android devices in the live session view (#495)
* feat(admin): identify Android devices by model in live session view

Android clients that send a bare default User-Agent (e.g.
"Dalvik/2.1.0 (Linux; U; Android 11; AFTKRT Build/RS8180.3729N)")
showed up as "Dalvik" in the admin live-session view, which tells an
operator nothing about the device.

Parse the model code out of the UA (the token between the last ';' and
"Build/") and map the Amazon Fire TV family and NVIDIA Shield to product
names. Unknown but parseable models fall back to "Android · <MODEL>"
instead of "Dalvik"; multi-word models like "Pixel 7" are preserved
whole. This is display-only: the session still stores the raw model code
in its user agent, and no response field or contract changes.

* feat(admin): mark Jellyfin-compat sessions with the JF pill by origin

The admin "JF" pill was derived at read time by substring-matching a
token list against the client name / user agent. A real Jellyfin
client that authenticates through the compat surface but sends a bare
User-Agent and no MediaBrowser client name (e.g. a Fire TV app) got no
pill, even though it plainly came through the Jellyfin API.

Stamp compat origin as immutable identity at session creation and carry
it through to the admin view:

- ClientInfo.IsCompat is set true in the jellycompat auth path; newSession
  copies it onto Session.IsJellyfinCompat.
- The flag rides the durable RecipeCard (next to the client metadata that
  already exists so the pill survives reconstruction) and is restored in
  ReconstructSession, so a server restart keeps the pill.
- buildLiveSessionSync -> worker.SessionSync -> a new compat_origin column
  on playback_sessions_sync (added migration); the reconciler upserts,
  reloads, and compares it so origin changes still publish and unchanged
  rows do not churn.
- The handler ORs the stored origin with the existing name/UA heuristic,
  which stays as a fallback for rows written before this column existed.

is_jellyfin_client keeps the same name and type on the wire; it is only
sourced more accurately.

* fix(admin): correct Android device labels

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-28 21:27:58 -04:00
54cfc48ce3 fix(metadata): improve localized series matching (#490)
* fix(metadata): improve localized series matching

* test(metadata): strengthen consensus regressions

* fix(metadata): harden localized matching

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-28 21:02:51 -04:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
4eab955ec1 build(deps): bump github.com/quic-go/webtransport-go (#508)
Bumps [github.com/quic-go/webtransport-go](https://github.com/quic-go/webtransport-go) from 0.10.0 to 0.11.1.
- [Release notes](https://github.com/quic-go/webtransport-go/releases)
- [Commits](https://github.com/quic-go/webtransport-go/compare/v0.10.0...v0.11.1)

---
updated-dependencies:
- dependency-name: github.com/quic-go/webtransport-go
  dependency-version: 0.11.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-28 18:14:25 -04:00
95edf19389 feat(invitations): shareable claim links and open-in-app on the claim page (#509)
* feat(invitations): always return the claim link so admins can share it directly

The claim URL was only surfaced when email sending failed. Admins who want
to hand the link over another channel (chat, SMS) had no way to get it —
and the raw token exists only in the send/resend response, since the server
stores just its hash.

The create and resend flows now always include claim_url (additive on
/api/v1), and the admin UI keeps the dialog open after either action with
the link and a labeled Copy button. Truncation and stacked buttons keep the
unbreakable URL from forcing horizontal scroll on phone widths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): offer to open invite claims in the Android app

The Android app already registers silo://invite?server=...&token=... with
a full native claim flow, but nothing ever emitted that link — an https
invite always ended in the browser.

On Android user agents the claim page now leads with a prominent 'Open in
the Silo app' button carrying that deep link, with the web form kept below
as the fallback ('or set up in the browser'). The button is a plain anchor:
a user-tapped custom-scheme link is the one reliable path, and we never
fire it automatically since there is no installed-check and a miss surfaces
an OS error. The password field's autofocus is suppressed alongside it so
the keyboard doesn't push the button off screen. iOS is excluded until the
Apple app registers the scheme.

The server origin travels in the server param verbatim, so non-443 ports
and plain-http LAN servers need no extra convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 15:01:23 -04:00
63e18cf37f fix(playback): prevent transcode resolution upscaling (#503)
* docs(playback): design transcode resolution clamp

* docs(playback): plan transcode resolution clamp

* fix(playback): prevent transcode resolution upscaling

* refactor(playback): share transcode resolution tiers

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-28 13:07:05 -04:00
SuspenseandGitHub 5d6f6323d5 perf(web): smooth item sidebar collapse (#504) 2026-07-28 09:50:30 -04:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
315f843fa5 build(deps): bump react-router from 7.15.1 to 8.3.0 in /web (#507)
Bumps [react-router](https://github.com/remix-run/react-router/tree/HEAD/packages/react-router) from 7.15.1 to 8.3.0.
- [Release notes](https://github.com/remix-run/react-router/releases)
- [Changelog](https://github.com/remix-run/react-router/blob/main/packages/react-router/CHANGELOG.md)
- [Commits](https://github.com/remix-run/react-router/commits/react-router@8.3.0/packages/react-router)

---
updated-dependencies:
- dependency-name: react-router
  dependency-version: 8.3.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-28 09:17:14 -04:00
271a2e1741 feat: emailed invitations, claim + household setup, and server-driven onboarding tour (#501)
* feat(invitations): add emailed pre-provisioned invitations

Admins can invite a specific person by email: the invitation pre-binds
role, access group, and library access, and the invitee only chooses a
password. Their email address becomes their username, so login gains an
email fallback (username lookup first, email column only on miss for
inputs that parse as a bare address).

- invitations table: single-use token (SHA-256 at rest) bound to one
  address; a partial unique index makes resend-supersedes atomic; no
  users row exists until accept, so a typo'd address can't squat a
  username. Status is derived from timestamps, not stored.
- internal/invitations: repository, service, and branded email through
  the shared internal/mail sender. When SMTP is off the claim URL is
  returned for manual delivery instead of failing.
- Admin endpoints /admin/invitations (list/create/resend/revoke) beside
  the existing invite-codes routes; public claim endpoints
  /invitations/{token} (+/accept) rate-limited with the other auth
  endpoints. Unknown/expired/revoked/used tokens are indistinguishable.
- Accept returns the same login response shape as signup, so clients
  reuse their session plumbing.

Spec: docs/superpowers/specs/2026-07-27-invitations-and-onboarding-design.md
Plan: docs/superpowers/plans/2026-07-27-invitations-and-onboarding.md
Part of #215

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add invitation admin tab, claim page, and household setup

- Admin → Users gains an Invitations tab: compose (email, role, access
  group, libraries, note, first-profile and tour toggles), list with
  derived status, resend, revoke. When the server has no SMTP the create
  response's claim URL is surfaced for copy-paste instead of a fake
  success.
- /invite/:token claim page: everything but the password was decided at
  send time, so it asks for exactly one thing and lands the user signed
  in. Expired/used links get an explanatory card, not a 404.
- /household-setup ("Who's watching?"): profile tiles plus the existing
  ProfileEditorDialog, all through the existing /profiles endpoint —
  no new backend. "Just me for now" is a first-class exit.

Part of #215

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(onboarding): add server-driven onboarding manifest and state

GET /onboarding/flow returns the ordered first-run tour for this server
and profile: steps for disabled features (requests, watch together,
recommendations, notifications) are filtered out server-side, surface=tv
drops steps needing text entry, and child profiles never see stops they
can't act on. Copy lives in Go, so a wording fix is a deploy — clients
render step kinds they know and skip unknown ones by contract.

setting_choice steps name an explicit write target (profile_field /
setting / device_setting) because playback quality is a profile column,
not a settings key — the tour writes through the same APIs the settings
screens use.

Per-profile completion state lives in the user store (SQLite schema v14
+ a Postgres twin table), keyed by (profile_id, tour_id) with monotonic
completed/skipped timestamps: finishing on one device silences every
other; a later progress write can never un-complete.

Part of #215

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add the first-run feature tour

TourHost renders the server manifest as a modal overlay on Home: unknown
step kinds are skipped silently (the forward-compat contract), progress
posts per step, and setting_choice steps write real values through the
existing profile/settings mutations — by the last step the account is
genuinely configured. Skip is always one click and recorded server-side,
so no other device re-prompts. The tour ends by handing off to the
existing taste-seed picker, which now waits for the tour to finish
before its own redirect. Settings → Personalize gains a replay entry.

An invitation sent with show_tour=false plants a local hint that the
gate converts into a server-side skip for the first profile.

Part of #215

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): satisfy noUncheckedIndexedAccess in the tour's advance step

The Docker web build runs `tsc -b`, which applies the project's
noUncheckedIndexedAccess; the bounds check didn't narrow steps[next].
Look the step up once and branch on its presence instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): blur the whole app behind the tour, sidebar included

The tour overlay rendered inside the app layout, where an ancestor
creates a fixed-position containing block — inset-0 pinned to the
content pane, leaving the sidebar completely un-scrimmed. Portal the
dialog to <body> so the scrim truly covers the viewport, and raise the
backdrop blur from sm (4px) to xl (24px) so card titles and nav labels
aren't legible through it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(onboarding): name features by their UI labels in the tour copy

"Same movie, different couches" never said what the feature is called.
Every feature card now leads with the name the sidebar actually uses —
Watch Party, Requests, Watchlist, Calendar, Notifications — and says
where to find it, so the tour teaches vocabulary, not just concepts.
Server-side copy, so all three clients pick this up with no release.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(onboarding): add apps and Jellyfin-compat steps to the tour

Two new web-only feature cards near the end of the tour:

- "Take Silo with you" — native apps for iPhone/iPad/Apple TV and
  Android/Android TV, with outbound TestFlight and Play Store links.
  Steps gain an additive links field (label + url) that older clients
  ignore; the web TourHost renders them as external-link buttons.
- "Already use a Jellyfin app? It works here" — Infuse/VidHub/Findroid/
  Swiftfin connect via the Jellyfin API. Gated on
  jellyfin_compat.enabled (default-on, so unset counts as enabled;
  only an explicit "false" hides it).

Both steps are web-only: the apps card is pointless inside the apps it
advertises, and TV can't open store links. surface=phone/tv manifests
skip them, covered by tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): keep the tour card responsive on phone widths

Verified every step at 1600px, 390px, and 320px with an automated
overflow check. Fixes it found:

- Link buttons (apps step) now wrap and truncate instead of extending
  past the card edge.
- The footer wraps at very narrow widths, so the handoff step's wide
  primary button drops to its own line rather than overflowing.
- Progress pips hide on phones — decorative, and they crowded the
  Back/Next buttons.
- The card scrolls within 85dvh so a tall step never pins its buttons
  off-screen on landscape phones.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): render store links as branded badges in the tour

The apps step's plain outline buttons now render as store badges: the
Apple or Google Play mark with a store eyebrow (TestFlight beta /
Google Play) over the platform label — the familiar app-store badge
idiom. The brand is inferred from the link's host on the client, so
the server contract stays icon-free and non-store links keep the plain
external-link button. Labels drop the parenthesized store name the
eyebrow now carries.

Verified at 1600px and 390px with the overflow sweep: none.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 22:58:18 -04:00
54e184df85 feat(requests): enforce per-profile rating limits in discovery (#505)
* feat(requests): enforce per-profile rating limits in discovery

- Resolve each profile's max content rating and filter discovery, detail, and browse results against it, failing closed on missing ratings
- Reject request submissions for titles above the viewer's ceiling
- Add TMDB GetCertification backed by release_dates/content_ratings with a long-lived cache and singleflight
- Push certification.lte to TMDB for studio/network/genre browse as a cost pre-filter
- Backfill restricted section pages from a fixed window of TMDB pages to keep carousels populated and pagination stable

* fix(requests): address discovery rating review findings

- Preserve backfill overflow: sections use plain TMDB cursor semantics
  plus an additive next_page field instead of fixed windows, so an early
  stop never drops allowed titles from unconsumed pages (bit hardest at
  permissive R/TV-MA ceilings).
- Bound cold-path cost: DiscoverAll backfills at most 2 TMDB pages per
  section (vs 5 for a direct section request), capping worst-case cold
  certification hydration at 240 lookups instead of 600.
- Keep the TMDB prefilter a superset: rank-3 ceilings now push down
  certification.lte=NC-17/TV-MA rather than R, so titles the local
  ladder allows can't vanish upstream unrecoverably.
- Fail closed on foreign certifications: enforcement-path lookups use
  new US-only pickers (a Canadian PG no longer reads as US PG), while
  the display path keeps its any-country fallback. US multi-entry
  disagreements prefer the theatrical/real rating over festival NR.
- Detach shared certification fetches from the first caller's context
  (WithoutCancel + 30s bound) so one disconnecting client can't fail
  the singleflight result for concurrent waiters.
- Advertise enforcement via rating_restrictions_enforced on
  /requests/status so clients can feature-detect instead of
  version-sniffing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(requests): harden rating enforcement per second review pass

- GetDetail gates on the US-only enforcement certification (cached
  GetCertification) instead of the display rating, whose any-country
  fallback let a foreign "PG" pass the US ladder.
- pickUSMovieCertification takes the strictest recognized US rating when
  multiple release entries disagree ([PG, R] -> R); entry order is not
  meaningful and enforcement must not admit a title on its most lenient
  certificate.
- Certification singleflight uses DoChan so a canceled caller returns
  ctx.Err() immediately instead of blocking up to 30s on the detached
  shared fetch (which still completes for surviving waiters).
- Viewer rating ceiling resolves once per request and threads through
  discover/browse/detail enrichment (enrichPageWithCeiling); DiscoverAll
  drops from 12 scope resolutions per load to 1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 22:34:30 -04:00
1a784a1cdf fix(player): restart backward HLS seeks (#499)
Preserve negative player-time offsets for targets before a copy-mode HLS window so the existing transcode restart path handles them.

Co-authored-by: GPT-5.6 Sol <noreply@openai.com>
2026-07-27 16:23:45 -04:00
f235524365 feat(diagnostics): chunked report upload fallback for proxy body caps (#494)
* feat(diagnostics): chunked report upload fallback for proxy body caps

Diagnostics bundles can be up to max_bundle_bytes (10 MiB default), but a
reverse proxy in front of Silo commonly caps request bodies at nginx's
default client_max_body_size of 1 MiB. Such a proxy answers the single-shot
multipart upload with its own 413 before Silo ever sees the request, so any
report over the cap could never be delivered.

Add a chunked upload fallback under /api/v1/diagnostics/reports/uploads:

- POST   /                      {manifest, bundle_bytes} opens a session
- PUT    /{id}/chunks/{index}   streams one ≤768 KiB chunk (proxy-safe)
- POST   /{id}/complete         ingests the assembled bundle
- DELETE /{id}                  best-effort abandon

The assembled bundle goes through the exact same Ingest path as the
single-shot endpoint, so every content check (manifest contract, archive
sha/bytes/entries, quotas, profile attribution) applies identically.
Sessions reuse internal/uploads (the plugin chunked-upload spool manager)
plus a small owner map for per-user isolation; they spool to disk, expire
after 15 minutes, cap at one per user / 16 global, and complete shares the
existing per-user + global in-flight ingest limiter.

/diagnostics/status now advertises upload_chunk_bytes so clients can detect
support; older servers omit the field and clients treat that as
unsupported. The demo guard's diagnostics prefix gains PUT to cover the
chunk route.

Verified end to end against an OpenResty proxy with a 1m body cap: the
single-shot upload 413s, the same 1.6 MiB bundle uploads in three chunks
and lands as an accepted report; also exercised from the tvOS client's
fallback path in the simulator.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(diagnostics): harden chunked upload sessions per review

- Reserve the per-user slot and global cap atomically in init (a
  reservation map counted with live sessions), so concurrent inits by one
  account can no longer fan out past one session or transiently exceed the
  cap. Creation failures roll the reservation back.
- Move chunk body I/O outside the uploads.Manager mutex: a slow client
  streaming one chunk no longer serializes every other session's chunk
  writes, completes, and cancels. A per-chunk in-flight flag rejects
  duplicate concurrent writes to the same offset (ErrChunkBusy → 409), and
  cancel/expiry defer spool-directory removal to the last finishing
  writer.
- Chunk arrivals refresh the session expiry, making the TTL an idle
  timeout instead of an absolute deadline so a slow-but-progressing upload
  cannot expire mid-transfer.
- Extend the request read deadline on chunk PUTs and both deadlines on
  complete, matching the single-shot handler's slow-uplink handling.
- Keep the session when complete's availability re-check fails
  transiently (status load error → 500): only definitive
  disabled/storage-unavailable answers discard the spool, so a retried
  complete succeeds without re-uploading every chunk.
- Reclaim orphaned spool directories at startup (a restart previously
  stranded the old process's partial uploads forever) and sweep expired
  sessions on a timer instead of only from later init traffic.
- Document that session state is process-local and what that means for
  multi-replica deployments.

Adds concurrency/race tests (go test -race) for atomic admission,
same-chunk write exclusion, expiry refresh, transient-status retry, and
startup reclaim.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(diagnostics): count detached chunk writers and lift chunk PUT write deadline

Second review round:

- A canceled session whose slow chunk writer was still draining held a
  connection and spool disk but vanished from every count, so a
  cancel-and-reinit loop could stack unbounded live writers behind the
  16-session cap. The uploads manager now parks such sessions in a
  detached set (exposed as DetachedWriterSessions) until their last
  writer returns, and diagnostics init counts them in its admission gate.
- Chunk PUTs now extend the write deadline as well as the read deadline:
  on an uplink slow enough to eat the server's 120s WriteTimeout, the
  stored chunk's JSON acknowledgement would otherwise be lost and the
  client would retry an already-accepted chunk.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-27 12:07:37 -04:00
8dea4b9056 fix(playback): stop a broken Dolby Vision RPU from hanging playback (#485)
* fix(playback): stop a broken Dolby Vision RPU from hanging playback

A Profile 7 source whose RPU ffmpeg cannot parse took the whole session
down. The dovi_rpu bitstream filter does not fail cleanly — it rejects
every packet while ffmpeg runs on, so one observed session emitted
376,316 stderr lines before the process was killed, no manifest was ever
built, and the client got a 503 after ~10 seconds that it showed as an
endless spinner:

  [dovi_rpu] Failed to read unit 1 (type 39).
  [vost#0:0/copy] Error applying bitstream filters to a packet:
      Invalid data found ... Invalid SEI message: payload_size too large

Whether the strip works is a property of the file, not of ffmpeg, so
SupportsDoviRPUFilter cannot answer it — but asking ffmpeg to strip two
seconds to the null muxer can, in about a second. Profile 7 sources are
probed once each on the start path and the result is cached; a source
that fails is copied without the filter, which leaves the base layer and
plays. Refusing to play at all does not.

The probe reads stderr rather than trusting the exit code: ffmpeg treats
a per-packet bitstream-filter error as non-fatal and exits 0, which is
exactly how a stream that could never start reached a live session.

Cache is keyed on path plus size so a replaced file is re-probed, and
bounded like the letterbox cache. A nil probe keeps stripping, since most
Profile 7 sources need it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* chore(playback): use InfoContext in the rpu probe

* fix(playback): decide the DV RPU strip in the plan, not at the transport

The probe answered the right question in the wrong place. Suppressing the
bitstream filter as the transport was being built left the plan still
promising DynamicRange "hdr10", Claims.Video.HDR10 and the
dolby_vision_metadata_removed / hdr10_base_layer_preserved validated
claims, and left session.RemuxDVMode persisted as strip_to_hdr10 — so the
client was told it was getting clean HDR10 while receiving Profile 7 with
dangling RPUs, and every path that re-derives the filter from the durable
session put the hanging filter straight back:

  - HandleStartTranscode (quality change, seek, burn-in restart), which
    also re-derives it from file.PrimaryDVProfile() == 7 with no plan at all
  - the remote audio-switch restart, from updatedSession.RemuxDVMode
  - the progressive-remux transport, which plan_v3 evaluates *before* the
    HLS branches, so for many clients the broken file never reached the
    probe in the first place

The verdict is now a planner input alongside the transformation
registries: the registries answer whether the executor carries the
transformation, this answers whether the file does. A source that fails it
is never planned onto a strip route, so the plan's claims, RemuxDVMode and
every restart derived from them agree with what the pipeline can produce.
With no tone-map recipe in this tree an HDR10-only client has no route
left, so it gets a dv_conversion_unsupported terminal naming the real
cause rather than a generic HDR message; a client that can run its own DV
transformation still gets that route, with a degradation warning
explaining why the server route was dropped.

The two paths that bypass the plan entirely are gated at the executor:
legacy/auto remux neutralizes the profile exactly as it already does for a
missing dovi_rpu filter, and the explicit v3 strip recipe fails loudly
rather than emit dangling RPUs under an HDR10 claim.

The probe itself:

  - Tri-state verdict. Only a stderr-confirmed rejection is cached. A
    timeout, a cancelled request or an ffmpeg that will not start is
    inconclusive: the strip is kept and nothing is written, so one client
    disconnecting can no longer disable the strip for a file permanently.
  - Singleflight, so concurrent or retried starts of a title share one run.
  - Timeout cut to 6s, inside the budget a client waits on the manifest,
    and off the session lifecycle lock now that it runs at planning time.
  - Keyed on size and mtime as well as path and binary, so a file replaced
    in place with the same length is re-probed.
  - stderr capture bounded at 64 KiB; the markers are in the first lines
    and a rejecting filter emits a pair per frame.
  - The head-only coverage is stated rather than asserted: this catches a
    source that rejects from the first access unit, not one that breaks an
    hour in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(playback): keep the RPU probe alive when its caller leaves

The probe rode the caller's context, so a leader whose client disconnected
returned dvRPUUnknown — and CanStrip then handed that fail-open to every
follower already blocked on the shared call, even though their own requests
were still alive. One client walking away was enough to give a live session
the strip the source cannot survive, which is the hang this whole change
exists to prevent. The work was also thrown away, so the next start paid for
the probe again.

Run it under context.WithoutCancel instead. dvRPUProbeTimeout still bounds
it, so nothing is left running; a verdict reached after the leader has gone
is still correct and still worth caching. Follower behaviour is unchanged:
a follower whose own request is cancelled still leaves immediately rather
than waiting on someone else's probe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-27 08:15:25 -04:00
baa33768d6 fix(scanner): stop reporting an unusable ffprobe as an empty folder (#468)
* fix(scanner): stop reporting an unusable ffprobe as an empty folder

parseAudiobookFolder and parsePodcastShow signalled "this folder holds no
audio files" by wrapping os.ErrNotExist, and their reconcile callers skipped
on that. exec also wraps fs.ErrNotExist when the configured ffprobe binary
cannot be run, so a wrong playback.ffmpeg_path made every candidate folder
look empty: the scan logged processed=N failed=0, indexed nothing, and gave
the operator no clue why the library stayed empty.

Introduce an errFolderHasNoMedia sentinel that deliberately does not wrap
os.ErrNotExist, and skip on that instead. A folder that disappears between
the scan walk and the parse still maps to the sentinel, so a mid-scan rename
or delete stays a quiet skip rather than a scan failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(scanner): bound the all-failed scan summary instead of joining every failure

The sentinel change in this PR makes a previously-unreachable path reachable.
Before it, a misconfigured ffprobe made every candidate folder look empty, so
the scan skipped everything and `failed` stayed 0 — the `failedCount ==
processedCount` branch never fired. Now that an unusable ffprobe propagates as
a real failure, that branch is the expected outcome of a first scan with a bad
`playback.ffmpeg_path`, and it joins one wrapped error per failed folder.

On the 240k-folder library the scan code is written for, `errors.Join` over
that slice produces a ~64 MB error string (measured) that is written verbatim
into `scan_runs.error_message` and republished over the Redis events channel
and the admin SSE stream. The `failures` slice itself also grew unbounded for
the whole scan even when the all-failed guard could not fire (any rescan with
`skipped > 0`), holding hundreds of megabytes across a multi-hour scan before
discarding it.

Add a `scanFailures` collector that retains the first 20 failures and counts
the rest, joining them with a trailing "and N more failures (elided)". The
retained sample still names the cause, which is the entire purpose of the
summary. The same 64 MB case now produces 5.4 KB.

The ebook and manga scans have the identical shape and the same exposure via
their own probe failures, so all four call sites share the collector rather
than fixing the two audio paths alone. `failMu` now guards only `cancelErr`
and is renamed `cancelMu` to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-26 23:35:36 -04:00
172beb99ef fix(watch-together): stop a dropped socket reading as the host leaving (#487)
* feat(watch-together): make vote rooms actually vote

selection_mode has been stored, normalized and published since the
feature landed, and nothing has ever read it. A "vote" room behaved
exactly like a host_pick one: members could suggest and vote, the tally
was recorded and broadcast, and then the host promoted whatever they
liked regardless of it.

In a vote room the host now starts the winner rather than choosing it.
Promoting anything other than the leading suggestion is refused, because
being able to overrule the tally makes the mode host_pick with extra
steps and turns the vote counts on everyone else's screen into
decoration.

The winner is the head of the repository's existing ordering
(vote_count DESC, created_at ASC): most votes, ties to whoever suggested
first — deterministic, and re-suggesting a title cannot jump the queue.

A room where nobody has voted has no winner and says so, rather than
quietly promoting the oldest suggestion as though a vote had happened.

host_pick rooms are untouched: the host still promotes freely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* fix(watch-together): close the second door into a vote room's selection

Gating PromoteSuggestion left SelectItem wide open: it is host-only but
was not gated by selection mode, so the host of a vote room could set any
title directly and bypass the vote entirely. Enforcing the tally on one
path and not the other makes the vote counts on everyone else's screen
decoration.

A vote room now refuses a direct selection outright. The winner is the
only way in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* fix(watch-together): stop a dropped socket reading as the host leaving

hostDisconnectTTL was 15 seconds, which treated any transient drop as a
departure. An explicit leave and an explicit close already tear the room
down immediately, so this timer only ever covers a host who has NOT said
they are going — and at 15s a host who backgrounded the app, moved
between screens, or hit a brief network blip lost the room for everyone
with a "host_left" nobody could explain.

Two minutes survives a reconnect or an app switch, and is short enough
that a genuinely departed host does not leave a room open all evening.
The janitor still reaps idle rooms independently.

This matters for what the clients are growing into: a room you stay in
while you browse for something to suggest. A client that drops its socket
when the lobby leaves composition should cost you a reconnect, not the
room.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK

* fix(watch-together): let a vote room actually start its winner

The vote gate landed on both doors into a room's selection, but promoting
the winner walks through SelectItem to commit — so the gate meant to stop
the host bypassing the vote also stopped the vote itself. Vote rooms could
not start playback by any route.

Split the commit path: SelectItem keeps the gate for direct requests, and
PromoteSuggestion goes through the internal path once it has confirmed the
suggestion is the winner. Map ErrVoteRoomSelection in the promote handler
too, so a future regression there reads as a conflict rather than a 500.

Add service-level tests for both gates — the previous tests only covered
the pure winnerFrom helper, which is why the suite stayed green while vote
rooms were non-functional.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-26 22:01:31 -04:00
148c9291c5 feat(web): add Connect Apps settings page for compat sign-in (#488)
* feat(web): add Connect Apps settings page for compat sign-in

Jellyfin-protocol clients offer one username box and one password box and
never prompt for a profile, so signing in requires `account#Profile` and
`password#PIN`. Nothing in the product taught that syntax, and every way of
getting it wrong surfaces as the same "invalid username or password", so it
became a recurring support burden.

Add Settings -> Connect Apps, which states the credentials for the signed-in
account rather than describing them in the abstract. The page is segmented by
app type: the two formats are never shown at once, each side names the apps it
covers, and the compat side is visually distinct so it cannot be mistaken for
the normal Silo login. The compat listener's separate address is shown too,
since pointing a client at the Silo address fails identically to a bad
password.

Backend adds GET /api/v1/compat/connect-info, an account-scoped read of the
compat listener's enabled flag, public URL, and server name. The admin status
endpoint already covers this ground for operators, but it also reports install
paths and version provenance, so it stays admin-only; this returns only what a
client learns by connecting anyway. It is auth-only and deliberately not
profile-scoped, since it describes how to sign in.

ConnectInfoForConfig shares the enabled-flag precedence with
WebComponentStatusForConfig via compatEnabled, so the two endpoints cannot
disagree about whether compat is on.

The page declines to display a username it knows cannot work: profile names
permit `#` but the resolver splits at the last one, so `alice#Movie #2` parses
as account "alice#Movie". Such profiles get an explanation and no copy button
instead of a string that fails to authenticate.

Part of #432

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web): harden Connect Apps against misleading sign-in states

Review of the first pass found six ways the page could state something
untrue. Each one matters more than usual here, because the page exists
specifically to stop people guessing at credentials.

Report the running listener, not the stored setting. jellyfin_compat.enabled
is restart-required and cmd/silo builds the compat server from the boot
config alone, so the stored value describes intent. Reporting it promised
credentials for a listener that does not exist yet, or claimed the API was
off while the running one kept serving. ConnectInfo now returns the boot-time
state plus a pending_restart flag, and the page distinguishes "not running
yet" from "turned off". The admin status endpoint keeps reporting configured
intent, so the two intentionally diverge until a restart.

Stop presenting fetch failures as a disabled compat API. React Query clears
isLoading on error, so a failed request rendered "the compatibility API is
turned off" and sent users to an admin about a setting that was fine. A
failed profile list was worse: it fell through to an empty list and offered
the bare account name, which silently drops the profile suffix. Both now
withhold credentials and say the load failed.

Detect accounts that cannot use password login. Compat login is hardwired to
the local provider, which rejects accounts with local_password_login_enabled
false before checking any password, so SSO and plugin-provisioned accounts
can never authenticate. The page told them to type a password anyway; it now
says the compat API cannot accept the account.

Flag loopback compat addresses. jellyfin_compat.public_url defaults to
http://127.0.0.1:8096, which resolves to the client device on the phones and
TVs this page names. An untouched default was offered as the exact address to
copy; it is now explained instead of presented as usable.

Apply the #-in-profile-name guard to the summary list too. The selected-profile
field withheld an unusable username while "Every profile at a glance"
reintroduced it two sections below.

Read only the three settings this endpoint consumes. GetAll on the encrypted
repository decrypted every stored secret on each authenticated page view, and
an unrelated decryption failure would have silently dropped valid compat
overrides.

Invalidate the connect-info cache when an admin saves jellyfin_compat.*
settings. The address applies without a restart, but the page cached it for
five minutes and kept offering the old value for copying.

Part of #432

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 21:07:48 -04:00
99d205676f fix(metadata): prevent stale cross-provider IDs (#480)
* fix(metadata): prevent stale cross-provider IDs

* fix(metadata): address stale ID review findings

* fix(migrations): build the stale-ID primary key concurrently

ALTER TABLE ... ADD PRIMARY KEY builds the index under ACCESS EXCLUSIVE,
blocking reads and writes on stale_media_ids for the whole build. Create the
wider unique index with CREATE UNIQUE INDEX CONCURRENTLY and attach it with
ADD CONSTRAINT ... PRIMARY KEY USING INDEX instead; all three key columns are
already NOT NULL, so the attach is metadata-only. Same treatment on the
rollback path, plus the repo's INVALID-remnant cleanup so a failed concurrent
build is not silently accepted by IF NOT EXISTS.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 11:19:33 -04:00
02203d9e40 fix(playback): trust the server's media runtime end to end (#482)
* fix(scanner): reject durations that imply an impossible bitrate

The duration-plausibility rule only rejected videos of 10 seconds or less,
so a feature film that probed as 61 seconds passed untouched and persisted.
Clients then had nothing trustworthy to anchor on: Android's grow-only
duration ratchet has no floor to hold when the catalog value is wrong, so
the playback engine's growing-HLS-window duration won and a 90-minute movie
displayed as ~1 minute.

Size and duration together pin an implied bitrate, which separates the two
cases the absolute floor conflates. A genuine short clip has an ordinary
bitrate; a 100 GB file claiming 61 seconds implies ~13 Gbps. The ceiling
sits far above any real medium, so legitimate content cannot trip it — and
unlike the absolute floor, it does not false-positive on a genuine
high-bitrate short.

Also bump the repair-rule revision marker so rows judged by the previous,
weaker rule are re-checked once under this one. Without that bump an
improved rule never reaches the rows it was written for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(playback): publish source runtime in v3 plans and stop faking the copy seek window

Two defects with one root: a v3 plan described where playback sits without
ever stating how long the media is.

Add source.duration_seconds. It is the file's full runtime, never
`total - source_start` and never adjusted by timeline_offset_seconds, and it
is omitted rather than null when unknown — clients that coerce null to a
numeric default would read it as zero, the exact value this field exists to
stop them inventing. It is set in SourceDescriptorFromFileV3, the single
place every delivery already flows through, so direct play, progressive
remux, HLS remux and HLS transcode all carry it.

Until now the v3 plan omitted duration entirely, so clients fell back to the
playback engine. On an HLS copy remux the server intentionally serves
FFmpeg's still-growing playlist, so the engine reports the length produced
so far. With no server-supplied runtime to anchor on, a feature film played
back as a couple of minutes. The legacy protocol already answered this
correctly via fileDurationSeconds; this restores parity.

Separately, the copy branch published seek_window_end_seconds as the media
runtime. That made the window look *complete*, which clients read as proof
that any target inside it is locally seekable, so they native-seek past the
produced head of a growing playlist instead of asking for a reanchor. Leave
the end open: an incomplete window plus can_seek_anywhere=false routes every
seek through the server, which is what legacy did before v3 added the bound.

Advertise plan_source_duration_v1 so a client can distinguish "this server
does not populate the field" from "this server knows the runtime is
genuinely unknown" — without it, both look like an absent field and a client
cannot tell whether its own catalog fallback is still required.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web): pair the exit position with the media runtime, not the element duration

The player's exit state converts its position to media time but took the
duration from the video element, which is player-local. On a remux or
transcode stream the element only covers the window produced so far, so the
two values live in different coordinate systems.

Resuming a movie 50 minutes in makes that concrete: the exit position is
~3060s of media time while the element reports ~120s. The progress cache
then evaluates `position >= duration`, marks the item completed, latches the
watched badge, and — because completion clears the resume point — resets
position to 0. Exiting a resumed movie destroyed the resume point and
claimed it had been watched.

The server's runtime is authoritative and already expressed in media time,
so prefer it and fall back to the element only when no server value exists.
The rule moves into mediaTimeline.ts next to the coordinate conversions it
depends on, which is also what makes it testable — VideoPlayer itself has no
test harness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 00:12:29 -04:00
QuickandGitHub 10394b0a05 fix(scanner): stop hiding media when a library root is offline (#472)
* fix(scanner): stop hiding media when a library root is offline

A scan that cannot read a library root found no files there, so every
cataloged file under it was marked missing. Catalog reads all filter on
missing_since IS NULL, so marking is equivalent to deletion from a user's
point of view: the title leaves browse, search and next-up, and playback
answers "Source media file is missing" for media that is intact on disk.

The dead-root protection already existed but only guarded the destructive
operations. protectedConfiguredRoots was computed *after* the marking loop
in scanPaths, and applyScopedScan received the protected set but applied it
only to its force-delete branch. So an unreachable root could not lose its
rows, but could still have its entire catalog hidden until the next
successful scan.

On a CephFS deployment whose per-library subvolume mounts flap, this marked
190 present files missing in a single day — 15% of all missing-flagged rows
were files sitting untouched on disk, some flagged more than 20 hours after
their last write.

Hoist the probe above the marking loop and skip files under an unreachable
or suspect-empty root in both the folder and scoped paths. Pass the
unreachable set to the walked-scope call too: a nested child mount can die
under a healthy parent, and its rows are inside the parent's scope.

An offline root tells us nothing about whether its files exist. The only
safe reading is to leave them alone and let the next good scan decide.

Genuine deletions under a reachable root are unaffected and still marked
and swept on the same schedule as before.

Report the count as ScanResult.MissingSkippedProtected and log it, so an
operator can tell "my library shrank" from "my mount dropped".

Known gap: a suspect-empty *nested child* root under a healthy parent is
still marked missing, because the suspect set is not resolved until after
the walk loop. Unreachable roots — the case observed in production — are
covered.

* fix(scanner): protect suspect-empty and partially-walked roots too

Addresses review findings on #472. The original change guarded missing-marking
against probe-unreachable roots, but left three ways for a storage fault to
still hide a healthy library.

Suspect-empty detection was reactive. suspectEmptyRoots asked
ListRootsWithOnlyMissingFiles, which returns a root only once it has NO live
rows left. On the first scan after a mount drops — the moment that matters —
the rows are still live, so the root was not classified suspect and the scan
marked everything missing. The protection then engaged on the next scan, in
time to protect the wreckage. Ask ListRootsWithCatalogedFiles instead: any
cataloged row under an empty-but-reachable root is the lost-mount signature.
Intentional emptying is still reachable through the operator's one-time
cleanup allowance, which is the deliberate path for it.

Nested suspect-empty children were unprotected. Root compaction sends only the
populated parent through the walked-scope branch, which received only
unreachableRoots, so an empty child mountpoint had its rows marked missing on
its parent scanning cleanly. Pass the suspect set as well.

Partial walks were treated as authoritative. walkLogicalTree deliberately
swallows per-entry Lstat/ReadDir failures so one bad file cannot abort a scan
of a million, and collectLogicalFilePaths passed nil for the failure counter —
so the video path had no signal at all. A mount dying partway through
traversal produced a short file list indistinguishable from a large deletion.
Thread the counter through, and exclude a scope whose walk came back
incomplete from missing reconciliation, mirroring what the ebook scanner
already does via ebookRootScan.failed.

Also extract the duplicated mark-missing loop into markMissingExcludingProtected
so the folder and scoped paths cannot drift, and correct two comments that
still described the pre-fix "files are marked missing" behaviour — the exact
text a future reader would have trusted when reintroducing this bug.

TestScanFolderNestedSuspectEmptyChildRootProtection asserted the old
behaviour and is updated accordingly.

* fix(scanner): scope walk-failure protection and stop pruning on partial walks

Addresses the second Codex review round on #472. The previous commit's
incomplete-walk protection was too blunt in one direction and applied too late
in another.

Walk failures were counted, not located, and any non-zero count protected the
whole library root. A dangling symlink is both common and permanent, so that
would have suppressed missing-file reconciliation for its entire root on every
future scan — genuinely deleted titles would stay live indefinitely. That is
the same class of bug as the one this PR fixes, pointing the other way.
recordWalkFailure now records the logical path of each unreadable entry, and
only those paths are protected. Per-entry failures record the child path, so a
dangling symlink protects itself and nothing else, while a directory that
cannot be read protects its subtree.

Snapshot and group pruning ran before the protection. reconcileScannedRoots
and reconcileScannedGroups delete whatever the walk did not see, and both run
ahead of the missing-file guard, so a partial walk still dropped root
snapshots, observed locations and group locations for the unread portion —
corrupting later metadata matching even though the media_files rows survived.
Upserting what was seen is always safe; pruning now waits for a scan that read
the whole tree.

The confirmed-cleanup allowance was consumed to no effect for nested suspect
children. The walked-parent branch protected them unconditionally and runs
before the allowance is consumed, and an already-reconciled scope cannot be
revisited — so arming the allowance burned the confirmation while the child's
rows stayed live forever. Read the allowance without consuming it before the
walk loop, and honour it there. Unreachable roots stay protected either way:
an outage is never a confirmation to erase a catalog.

Two new regression tests, plus signature updates in the ebook pipeline, which
already tracked walk failures and now shares the path-based representation.

* fix(scanner): re-probe nested roots and gate group pruning on walk completeness

Third Codex review round on #472; both findings confirmed.

Group pruning ignored walk completeness in the subtree path. scanPaths passed
the completeness decision to reconcileScannedRoots but left
reconcileScannedGroups on !allowEmptyRootGuard, which is always true for
ScanSubtree — so a subtree scan that hit an unreadable directory still replaced
group snapshots and locations from a partial inventory. Same rule now applies
to both.

Nested roots were not re-probed before their parent was reconciled. Root
compaction folds a child mount into its parent for traversal, so a child that
is healthy at the initial probe but drops before the parent is walked leaves no
scope of its own, and the post-walk re-probe only revisits scopes that walked
empty. The parent walks files, looks healthy, and the child's rows are marked
missing on its success. reprobeNestedRoots re-checks this root's configured
children immediately before reconciling, protecting any that have since become
unreachable — or suspect-empty, unless the operator has confirmed cleanup.

Also guard suspectEmptyRoots against a nil file repository, matching
emptyCleanupArmed: without a catalog there is nothing to protect.

* fix(scanner): keep re-probed outages protected through folder-wide cleanup

Fourth Codex review round on #472; both findings confirmed. The first could
destroy data.

reprobeNestedRoots protected a root it found offline only for the scope being
reconciled, then discarded the result. The folder-wide membership reconcile and
the trash sweep afterwards rebuilt their protected set from the initial probe
alone, so rows under a child that dropped mid-scan — already marked missing and
past the removal grace — were hard-deleted by the very scan that noticed the
outage. Accumulate those roots in reprobedRoots, fold them into
protectedScanRoots, and reuse that set for the membership reconcile and sweep
instead of rebuilding. They now also land in ScanResult.UnreachableRoots so the
folder warning reflects the outage rather than presenting a partial scan as
clean.

Snapshot and group pruning was enabled for scopes that were never walked. The
gate was len(walkFailures) == 0, but an unreachable root gets nil walkRoots, so
it has no walk and therefore no failures — and pruning then deleted its
snapshots, observed locations and group locations even though its media rows
were protected. The same held for a suspect-empty child compacted into a
populated parent. Pruning now additionally requires that the scope was actually
walked and contains no protected path.

The new test pins that the sweep honours the protected set it is given. It does
not reproduce the mid-scan race itself: staging that needs the drop to land
between the probe and the walk, which a test cannot reach without hooks. That
path is covered by inspection, and the test comment says so rather than
implying coverage it does not have.

* fix(scanner): route every protection source through one folder-wide set

Fifth Codex review round on #472. Two P1s, one of them the second data-loss
path in this area — and the direct sibling of the one fixed in 35326adc, which
is the reason this commit changes the structure rather than patching another
edge.

Rows beneath a directory the walk could not read were protected only inside
applyScopedScan. The folder-wide protected set was rebuilt from the probe
results alone, so DeleteMissingByFolder could permanently delete rows past the
removal grace under a subtree this scan never managed to read — deleting on the
strength of an observation that was never made.

The recurring defect is structural: protection is discovered in several places
(initial probe, mid-loop re-probe, per-scope walk failures) and consumed in
several more (scoped reconcile, membership reconcile, trash sweep), and each
fix so far has wired up one edge and missed another. Every source now
accumulates folder-wide and every consumer reads the combined set, so a new
source has one place to register instead of several to remember.

reprobeNestedRoots classified from two probe batches. It called
probeUnreachableRoots, then suspectEmptyRoots probed the same paths again; a
child dropping between the samples was reachable to the first and discarded by
the second, which only returns reachable-and-empty roots. It now classifies
both states from one batch, so the disconnect it exists to catch cannot fall
between its own probes.

Re-probed roots kept their classification instead of being collapsed into
unreachableRoots, which had been reporting a suspect-empty child as
unreachable and giving operators contradictory failure information.

The new regression test is verified to fail with the propagation disabled and
pass with it, rather than assumed to cover the path.

Not addressed: the cleanup allowance is read without being reserved, so two
overlapping full scans of one folder can both observe it armed. Narrow, needs
a transactional reserve in the scan-claim query, and is left for follow-up
rather than bundled here.

* fix(scanner): resolve root protection before scoped metadata pruning

Sixth Codex review round on #472.

scanPaths pruned before it knew what was protected. reconcileScannedRoots and
reconcileScannedGroups ran roughly 160 lines ahead of protectedConfiguredRoots,
so a ScanSubtree of a mount that dropped but left a reachable empty mountpoint
walked clean, reported no failures, and pruned root snapshots and observed and
group locations against that empty inventory — preserving the media rows while
deleting the metadata describing them. Protection is now resolved before any
reconciliation, and both prunes share one decision, matching applyScopedScan.

Pending empty scopes never re-probed their nested children. A parent whose only
media lives in a child walks empty when that child drops, so it lands in
pendingEmptyScopes rather than the populated-scope branch where
reprobeNestedRoots ran. Probing the parent alone proves nothing: it still holds
the child's bare mountpoint directory, so it reads present and non-empty. With
a healthy sibling keeping the folder-wide empty guard quiet, nothing protected
the child. Both branches now re-probe.

MissingSkippedProtected never left the scanner. Both ingest-to-result
conversions copied every other cleanup count but not this one, and
events.ScanRunResult had no field, so scan history, completion events and API
responses reported an all-zero no-op for a scan that skipped files because
storage was offline. Added as a new field, which is additive under the v1 API
rules.

Test honesty: the new test does NOT exercise the pending-scope re-probe. It
empties the child before the scan, so the initial probe classifies it and
protection arrives by that path — verified by confirming the test still passes
with the re-probe disabled. It is named and commented for what it does cover.
The mid-scan race behind both re-probe fixes needs the drop to land between the
probe and the walk, which is not reachable from a test without hooks; those
fixes rest on inspection.
2026-07-25 14:47:12 -04:00
QuickandClaude Opus 5 482efdf140 chore: add T3 Code project icon
T3 Code resolves a repository icon from t3.json's iconPath, falling back to
well-known locations such as assets/icon.png. Neither resolved in this repo,
so it showed the default folder icon in the T3 Code sidebar.

The icon reuses the canonical Silo mark unchanged over a navy backdrop,
so the Silo repos stay distinguishable at the ~14px the sidebar renders.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 18:37:51 +00:00
QuickandClaude Opus 5 9e7fe79590 docs: record live TV, IPTV, and .strm as permanent non-goals
Live TV, OTA/DVB tuners, IPTV, EPG/XMLTV guide sync, DVR, and .strm
remote-URL shortcuts are permanently out of scope for Silo. The
deciding factor is app store distribution: the first-party iOS, tvOS,
macOS, and Android clients ship through Apple and Google, and a server
that plays arbitrary remote stream URLs puts the entire client suite at
risk of rejection or takedown, not just the feature. Secondarily, live
TV is a separate product surface whose reliability burden competes with
the core playback path.

This was an undocumented boundary until now, and contributors spent real
effort against it (#419, #420, #474). Write it down in docs/non-goals.md
and summarize it in AGENTS.md so both humans and agents see it before
proposing or implementing in this area.

Refs #474, #419, #295

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 15:07:42 +00:00
355508f6e7 fix(requests): retire stalled targets when presence confirms the media (#470)
reconcileRequest completes a presence-confirmed request only when it has no live targets, so a quality-agnostic TMDB hit cannot orphan an in-flight download. That gate assumes the router eventually moves every target to a terminal state; when it does not, the request is pinned open forever even though the media is in the library.

Retire targets stuck in queued past a 24h horizon with no status transition when presence confirms the media, and let the existing target aggregate drive request status. Targets actively downloading are never retired. Each retirement logs at WARN, since reaching this path means a router is misbehaving.

Part of #469

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 02:54:30 -04:00
Quick a102630365 feat(web): allow tailnet and hostname access to dev server
- Extend Vite `allowedHosts` with `.ts.net`, the machine hostname, and a `VITE_ALLOWED_HOSTS` override
- Document reaching the dev UI over Tailscale in the web-ui-testing skill
2026-07-25 05:27:52 +00:00
Quick 51ef906025 docs: add agent skills and silo-dev helper script
- Add skills for dev-environment debugging, jellycompat diagnosis, Discord triage, and web UI testing under .claude/skills (symlinked as .agents/skills)
- Add scripts/silo-dev plus .silo-dev.env.example for driving local or remote Silo deployments
- Ignore .silo-dev.env and reference the shared skill config in AGENTS.md
2026-07-25 05:16:19 +00:00
ece3578e91 docs: rightsize repo context for Claude 5 context engineering (#467)
Applies the practices from Anthropic's "new rules of context engineering
for Claude 5 generation models" to this repo's always-on context.

AGENTS.md (symlinked as CLAUDE.md) drops content derivable from the
filesystem — the module-by-module structure tour, the Makefile target
list, and most of the style section, all of which Claude reads directly
from internal/, the Makefile, .golangci and web/.prettierrc. What stays
is the part that isn't derivable: the Goose migration rules, the v1
additive-only API contract, multi-repo boundaries, and the workspace
gotchas. Also fixes a dead pointer to .claude/skills/deployment-debugging,
which does not exist; the runbook is the dev-environment-debugging skill.

The external-contributor AI disclosure block moves to
docs/ai-contributions.md, reached by a one-line pointer, so it costs
nothing in the common case where no external PR is being prepared.

issue-to-pr sheds the generic agent hygiene now covered by the harness
system prompt and keeps the Silo-specific gates. Its commit trailer no
longer pins a stale model name.

scripts/jellycompat-diff.sh replaces the hand-typed curl/python
one-liners the jellycompat-diagnosis skill used to carry as prose. It
unions item keys across all returned items rather than reading item[0],
which was hiding fields present on only some items.

Verified: make verify-local-paths, bash -n, shellcheck, and a mock
two-server run of the diff script.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-24 15:15:19 -04:00
ee31a1f0e2 feat(playback): formalize resumable direct streams and stall observability (#464)
* feat(playback): formalize resumable direct streams and stall observability

Implements #443: strong stat-based ETag + If-Range on original-file direct
play (via http.ServeContent), stream-end outcome classification in
RollingDeadlineWriter (stalled_reap vs client_gone vs completed) with a
structured log event and Prometheus counters, the direct_stream_resume_v1
protocol-v3 capability, and a contract doc. Progressive remux is explicitly
excluded from the resume contract.

Code written by OpenAI Codex CLI (gpt-5.6-sol) from a Claude-authored spec;
reviewed and verified by Claude.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): harden direct stream resume contract

* test(playback): cover resume platform contracts

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 14:22:12 -04:00
22fec4ed2d feat(metadata): add resilient match queue diagnostics (#463)
* feat(metadata): add resilient match queue diagnostics

* fix(metadata): harden match queue lifecycle

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-24 12:18:49 -04:00
383973ec22 feat(metadata): improve match accuracy and localized titles (#461)
* feat(metadata): improve match accuracy and localized titles

* fix(metadata): address matching review findings

* test(catalog): align empty alias snapshot scope

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-24 12:02:52 -04:00