d01b5de3bd7c45fb574d6d8fb5448b694fc940af
230
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cd9e358bbc |
fix(playback): fail closed on transfer-registry saturation
`transfers.Registry` caps at 10,000 entries. Past that `Begin` returned `ErrRegistryFull`, and every call site logged at Debug and served the file anyway with its byte updates discarded. With no connection cap anywhere, one actor could pin the registry and blind download-class monitoring for everyone — which is the monitoring the abuse controls depend on. Per decision A7, saturation now fails closed and is unreachable by a single actor in the first place: - A per-user concurrent-transfer cap (`playback.max_user_concurrent_transfers`, default 24) checked before the global limit, so one actor exhausts its own budget rather than the shared registry. 24 leaves headroom for multi-connection downloaders, which open 4-8 sockets per file and take a transfer id per concurrent Range request. `MaxPerUser` is a *pointer* option because a plain int cannot distinguish "unset, use the default" from "explicitly 0, unlimited". - All five download-class call sites refuse rather than serve unmonitored: 429 `transfer_limit_exceeded` for the per-user cap, 503 `monitoring_unavailable` for a full registry, both with `Retry-After`. The handoff listed four sites; `internal/api/handlers/ebook_reader.go` was the missing fifth. - The ABS file handler now admits the transfer *before* setting Content-Disposition and the audio Content-Type, so a rejection is not mislabeled as an audio attachment. Two correctness fixes fall out of doing this properly: - `End` now decrements the per-user count only when it actually removed an entry, and drops the map entry at zero so neither counts nor warning timestamps grow unbounded. - The two download handlers registered `defer transfers.End(transfer.ID)` *before* the service called `Begin`, so a duplicate id would have removed another request's live record. `Begin` moves into the handler beside its `End`, and the service keeps its late DownloadID/MediaFileID enrichment through a new `Registry.Annotate`. Saturation is logged at Warn inside the registry, rate-limited per user and carrying the route and user id, rather than once per rejected request at each call site. A7 asked for Warn at the call sites, but an actor parked at its cap would then generate unbounded warning traffic — a log-amplification vector of its own. No information is lost: route and user id are already on the Transfer. 429/503 are new failure statuses on existing endpoints, not repurposed ones. silo-android and silo-apple need follow-up to retry with backoff on `Retry-After`; the cap and the fail-closed behavior are advertised on `GET /downloads/capability` so clients can feature-detect instead of probing. The jellycompat call site is verified by reading only: that test package does not compile on origin/main (pre-existing, unrelated). Part 3 of 3 for the Batch 4 liveness/replica work. Part of #305 |
||
|
|
40f117bd2d |
test(playback): assert the admin session view omits an unserved last_served_at
`TestHandleListSessionsAddsSiblingTransfersWithoutChangingSessionShape` expected `last_served_at` to equal `Session.LastActivityAt`, encoding the client-progress fallback that |
||
|
|
e5bf0155ad |
fix(playback): make revocation state converge and cutoffs credential-accurate
Five defects in revocation state and credential semantics. Lands after the tracker-lifecycle batch on purpose: raising the over-cap TTL is only safe once the count feeding it is trustworthy. #13 -- a longer old revocation suppressed a newer cutoff. applyLocal kept or replaced the WHOLE record by expiry, so when the existing revocation expired later the new one was dropped entirely, including its newer RevokedAt. The durable upsert did the same, with a comment documenting it as intentional. RevokedAt is the user-kill CUTOFF, so this left a credential issued between the two cutoffs valid -- a second admin kill after a user re-authenticates silently failed to cut them. The two fields now merge independently: ExpiresAt stays monotonic, RevokedAt advances to the later value, and reason follows the newer cutoff. Both superseded comments are replaced rather than left contradicting the code. Session-kind revocation still ignores RevokedAt, so the enforcer's re-revoke cannot weaken a session kill. Also fixed while here: Redis received the merged record but pub/sub published the raw input one, so under pub/sub-only delivery (Redis down) an edge got the newer short record without the older long expiry and lost monotonicity. Both now carry the merged record. #7 + M1 -- Postgres could indefinitely block the urgent Redis kill. RevokeWithWarnings held the global opMu across all propagation, stripped the caller's deadline with WithoutCancel, and did the durable Postgres upsert BEFORE Redis, on a pool with no statement timeout. The local kill still applied, so playback on that process was fine -- but edge propagation, pub/sub, the admin response and every later revoke/unrevoke stalled behind the lock. Redis and pub/sub now go first, and the detached context is bounded. WithoutCancel is kept deliberately: propagation must outlive an aborted admin request. opMu scope is deliberately NOT narrowed. mirrorToRedis is an unconditional SET with no atomic merge, so same-process serialization is what stops an older value overwriting a newer one; narrowing the lock would also let an unrevoke interleave with a revoke's propagation. Bounding the context caps how long the lock can be held, which is the actual reported harm. The remaining cross-replica race -- two central replicas racing the same SET -- is documented, not half-fixed; it needs A6's shared picture. A2 / #6 -- a missed unrevoke got resurrected. In-memory tombstones already existed, but being process-local they did not survive a restart or reach a replica that missed the pub/sub event, so maintain's durable self-heal re-Upserted the surviving entry and the ban returned. Tombstones are now durable, via two nullable columns on stream_revocations rather than a second table: a tombstone is a state of the same key, and it needs its own expiry horizon separate from the revocation's. The upsert rejects a stale replica's write while a tombstone is live but lets a genuinely newer revocation clear it, and warm paths apply tombstones BEFORE revocations so an un-banned key cannot be restored as a live kill. Tombstones are pruned on the same sweep, so the table cannot grow without bound. A1 / #3 -- over-cap kills reopened after 5 minutes while the token stayed reconstructable for 24h. The TTL now derives from playback.MaxTokenTTL rather than duplicating 24h, behind a validated setting. Critically, the enforcer uses a revoke-if-absent path rather than re-revoking. Expiry is monotonic and the enforcer re-evaluates every 30s, so a plain long TTL would slide expiry forward by another full lifetime on every pass -- making a wrong kill effectively permanent for as long as any stale record persisted, with only an explicit unrevoke to recover it. Admin Revoke keeps its monotonic behaviour; only the enforcer's own repeat kill is non-extending. The setting is documented as affecting future revocations only, since monotonic expiry means it cannot shorten one already issued. A3 / #5 -- the user cutoff compared against a fresh time.Now() taken at request entry, so a request from a pre-cutoff login could look post-cutoff and escape the kill. The credential time is now the access token's iat. Two deliberate choices worth stating. API-key credentials carry no issue time, so they pass the zero time and, per IsRevoked's documented contract, are never matched by a user cutoff: a user kill provably cannot cut an API-key-owned pour. That is an accepted, logged, documented hole -- and strictly better than time.Now(), which actively defeats the cutoff. And jellycompat uses the compat session's CreatedAt rather than the bridged Silo token's iat, because that token refreshes without a new Jellyfin login, so its iat would advance on refresh and let a refreshed credential slip past a cutoff. Stream tokens are now bound to their route: a token whose SessionID does not match the URL's session_id is rejected with 403 instead of being silently ignored, matching the reconstruction helper that already refused a different session. Per-login logout cuts remain out of scope -- they need per-login identity in the stream credential. S5 (the (sessionID, userID, startedAt) clump) is rejected as ceremony now that A3 is the iat option rather than the generation model. #12 (closing an RSS feed does not cut its current pour) is deferred: it needs a namespaced revocation id that cannot collide with real session ids, that id threaded onto public feed requests, and protection against a new feed inheriting an old tombstone. Part of #305. |
||
|
|
ecb4555eec |
fix(playback): make edge tracking lifecycle-safe and session identity canonical
Five defects that all produced a WRONG over-cap count, which is why they land before the revocation batch: decision A1 raises the over-cap revocation TTL from 5m to ~24h, removing the self-healing that currently limits the damage of a miscount. A false positive after A1 blocks a legitimate stream for a day, so the count has to be trustworthy first. #1 -- overlapping edge requests deleted a live stream. Tracker.sessions was a set and Remove tore down all state plus the Redis key, while both proxy pour handlers deferred removal unconditionally. Two overlapping Range GETs on one session id -- ordinary seek behaviour -- meant the first to finish deleted the record while the second was still pouring, and later AddBytes calls were then dropped because AddBytes ignores bytes for a session with no live record. The stream went invisible to authoritative monitoring while still serving. Track now returns a Lease that the request-scoped caller releases exactly once; teardown happens when the last live lease is released. A plain refcount would have been wrong: Track(A) -> Remove -> Track(B) -> Release(A) decrements B, and clamping at zero does not help because the count legitimately belongs to B. That is not hypothetical -- the transcode node deliberately replaces sessions under the same id so a quality switch does not orphan ffmpeg, and it calls unconditional Remove from its reaper and stop paths. So each generation carries an epoch, Remove and Cleanup bump it, and a release from a superseded generation is a logged no-op. Lease identity is a set rather than a counter, which makes a duplicate release detectable instead of silently destructive. The transcode node keeps using Remove: its Track calls are not request-scoped and are correctly owned by session lifecycle. "Every Track needs a paired Release" is true only of the request-scoped callers. #8 -- async transcode tracking could leave a permanent ghost. The tracking write ran as a bare goroutine with a WithoutCancel context, so if stop won the race the delayed Track recreated the record after cleanup -- and because it landed in sessions, Snapshot treated it as live until Remove and it NEVER idle-expired. A permanent phantom inflating its owner's count, able to trigger false over-cap kills of that user's real streams. The write now takes the per-session lifecycle lock that stop and reap already hold, and re-checks session pointer identity before writing, so a stopped or replaced generation cannot resurrect a record. Pointer identity rather than id equality is what makes same-id replacement safe. The write stays off the request path -- the API server and the playback client are blocked on the 202. #9 + M3 -- protocol-v3 counted one stream twice. The stream token carries a transport id distinct from the logical session id, and the node tracked under the transport id while the API/proxy record used the logical one, so mergeStreams saw two streams. M3 was the reason this had not yet bitten: the v3 fresh-start caller sent no owner attribution at all, so the transport record landed under user 0, which the enforcer skips -- silently exempting the stream from the cap entirely. Fresh v3 starts now carry the logical session id and full owner attribution (both were already in scope at the call site), and merging is keyed on logical identity where present via one shared helper used by both merge functions, which had already drifted apart once. The enforcer view resolves SessionID to the logical id so a kill targets the real session rather than a replaceable transport generation. The raw admin view keeps the transport id and exposes logical_session_id as an additive omitempty field, advertised on the node-sessions capability endpoint, so the v1 response shape is unchanged. GAP-15 -- edge transcode liveness was request-observed. touchTranscodeSession fired before proxying, so hammering dead segment URLs advanced LastServedAt with zero bytes served. Visibility and liveness are now separate operations: EnsureEphemeral makes a session visible without claiming bytes were served, and served-byte liveness advances only from a 2xx/206 upstream response. Previously the proxy metered every upstream body regardless of status, so a node 404's error body counted as served bytes -- moving the touch later would not have fixed it. S4 -- LiveLocalSessions moved from the HTTP handlers package to streammonitor, which owns monitoring. A background enforcer importing api/handlers was backwards. Pure move; its existing mapping assertions moved with it. The LastActivityAt fallback inside it is left as-is -- decision A5 removes it in the liveness batch. Verified with go test -race across nodesessions, proxy and transcodenode; the overlap regression test was confirmed to fail under the old unconditional teardown. Part of #305. |
||
|
|
70064823a6 |
feat(api): advertise stream monitoring and revocation capabilities
The admin kill-list endpoints and the transfers field on the live node-sessions
payload shipped with no capability advertisement, contrary to the repo's
additive-v1 rule that new features expose capability endpoints for feature
detection rather than relying on version sniffing. (S3)
Adds two endpoints, each mounted beside the surface it describes so an
advertisement cannot outlive the route it advertises:
GET /admin/node-sessions/capabilities
GET /admin/streams/revocations/capabilities
The node-sessions capability keeps schema support and runtime availability as
separate booleans. The transfers key is always present in the response shape
once that endpoint exists, but the process-local registry behind it is optional
wiring -- an edge deployment can legitimately serve it as an empty list forever.
Collapsing the two into one flag would advertise download monitoring that is not
actually running, so `transfers` reports the schema and `transfers_active`
reports the wiring.
The revocation capability advertises the closed {kind} vocabulary accepted by
DELETE /admin/streams/revocations/{kind}/{id}. To make that advertisement
impossible to desync, the accepted kinds move into a single map that both the
wire parser and the capability handler read -- previously the parser duplicated
the list in a switch, so a newly-accepted kind could go unadvertised and clients
would feature-detect an incomplete vocabulary. The kinds are returned sorted,
because map iteration order is randomised and this is a wire response that must
be stable across calls.
Tests cover the drift in both directions over the whole vocabulary rather than
sampling rejected strings, the sort stability, and that the handler returns a
copy so a caller cannot corrupt the package-level vocabulary.
Note the previous plan for this work put both capabilities on
/admin/sessions/capabilities. That was wrong: that endpoint documents the
Postgres-backed /admin/sessions payload, whereas transfers belongs to
/admin/node-sessions, which is gated on NodeRepo and may not be mounted at all.
No frontend change: web/src does not consume either endpoint.
Part of #305.
|
||
|
|
7adedf6429 |
fix(playback): close the serve-path monitoring and kill-switch gaps
Six defects on byte-serving paths, all of which made the PR's monitoring and
kill-switch claims narrower than documented.
#2/GAP-10 -- the ABS in-flight kill switch was a production no-op. accessLog
wraps every ABS route, and its statusRecorder implemented Write, WriteHeader,
Hijack and Flush but not Unwrap, so http.NewResponseController dead-ended and
SetWriteDeadline returned ErrNotSupported. A multi-GB audiobook pour survived a
RevokeUser. The existing test passed throughout because it called handlers
directly and never saw the middleware; the new test drives the mounted router
over a real socket, and both new assertions fail if Unwrap is removed again.
GAP-11 -- ebook, comic and PDF serving was invisible and un-killable: no meter,
no transfer record, no Refuse, no watcher, on a route that serves cbz/cbr/pdf
files routinely 100 MB-1 GB+. It now follows the ABS file-handler idiom. Note
guardRevocationCut is deliberately *not* reused here: it keys on a session_id
URL param and an ?st= token this route does not carry, so it would have
compiled and silently guarded nothing. Cap-exempt per decision A4 -- admission
is untouched and neither route consumes a stream slot.
#10 -- the no-proxy remote transcode hop forwarded segments through a bare
RollingDeadlineWriter, so bytes on the API hop went unaccounted for a supported
topology. Metering is scoped to media bodies; manifests are excluded so a
rewritten playlist is not counted as media, and a mid-copy failure is no longer
silently discarded.
#16 -- native, proxy and compat subtitle pours were entry-gated only. They now
carry a transport span, a meter and an in-flight watcher. Proxy subtitle bytes
are attributed only when a tracker record already exists: taking Track/Remove
lifecycle ownership per subtitle request would walk straight into the
overlapping-request defect (#1) that Batch 2 addresses. Compat subtitle
extraction is buffered and rejects bitmap formats, so a cut stops delivery but
not extraction already in progress.
M2 -- the proxy deferred tracker Remove with the request context, which is
already canceled on client disconnect, so the Redis DEL never happened and the
key lingered until TTL -- a false over-cap window that could get a legitimate
stream killed. Cleanup now uses a short bounded context.
GAP-13 -- mergeStreams took the freshest record wholesale and never merged
BytesServed, so a stream that poured 8 GiB at an edge could report 0. Merged as
a max, not a sum: the records are two observers of one pour. Fixed in
DedupeSessionInfos too, which had the same hole and feeds the admin view.
#11 needed no behaviour change -- that route was already metered, registered and
watched, and commit
|
||
|
|
09de4f08ec |
feat(downloads): monitor in-flight download pours in memory
Six routes poured full media with no monitor record and no byte measurement: the two native download routes, the Jellyfin-compat download, both ABS file variants, and the public ABS RSS feed file. They were the last invisible bytes on the server. - internal/transfers: a process-local, in-memory registry of active pours. Bounded (10k entries) with rate-limited "full" warnings, all request-derived strings normalized and length-clamped (ABS and jellycompat carry no bounded client name, so those fields are header-derived and untrusted), overflow-safe byte accumulation, deterministic snapshots, nil-safe throughout. No I/O on any path, and no persistence: a pour dies with the process, so durable rows would only need reaping after a crash. - Reuses the existing playback.SessionMeteredWriter via ServedBytesRecorder rather than adding a second writer — that writer is where the sendfile (ReadFrom) and kill-cut (Unwrap) hazards live and both have regressed before. - Deliberately NOT plumbed through streammonitor/streamenforcer, which are untouched. Downloads stay off the live-stream path by construction: there is no type, field or collection through which one can reach the enforcer, so a download can never be counted against max_streams or trimmed as an over-cap stream. - Admin visibility is a sibling `transfers` array on the existing /admin/nodes/sessions response; the `sessions` array is byte-for-byte unchanged. Transfers appear only in the unfiltered listing, since a node_id filter targets an edge and these are process-local. - Per-pour ids are unique, never the download id: concurrent and repeated GET/Range requests against one download row are legitimate and would collide. The id is minted in the handler (where the revocation watcher is armed) and the registry entry is opened in the service only after file/artifact resolution succeeds, so a failed auth or lookup never registers a transfer. - Defer order is an invariant at every call site: End is registered before the meter's Close so LIFO flushes the tail first. Reversed, a final sub-1MiB flush lands on an unknown id and is silently lost. Pinned by a test. No schema change and no migration: downloads.bytes_sent keeps its documented meaning (a lifecycle marker set to file_size on completion, not a live counter) and is untouched. Phase 2 — killing a single download, and a standing per-user download block — is deliberately deferred. A user revocation already cuts in-flight download pours; it is a cutoff, so it does not refuse new ones. The unique per-pour id exists so phase 2 only changes "" to that id at each WatchAndCut site. Part of the stream monitoring & kill-switch epic. |
||
|
|
fca9f38a33 |
feat(api): operator-visible stream kill list with unrevoke
The kill switch had no operator surface. streamrevoke.Store.List() existed and was called from nowhere, so an admin could only terminate one session by id or revoke a user's streams as a side effect of editing their account — with no way to see what was revoked, choose a TTL, or undo a mistake. Because expiry is deliberately monotonic, a wrong 24h kill was irreversible. - GET/POST/DELETE /api/v1/admin/streams/revocations, admin-only, additive. - Explicit wire-to-internal kind mapping: the wire accepts "session" (and "sess"), the store key is "sess". Passing the wire string straight into a Key would create a revocation IsRevoked never consults. - Validation: non-empty bounded session ids; canonical positive user ids (strconv.Itoa round-trip, so "01" is rejected — the cache key is the canonical form); ttl_seconds bounded to 30d so the duration cannot overflow; bounded reason and request body. DELETE of an absent key is idempotent. - Store.Unrevoke, guarded by a bounded in-memory tombstone: the tombstone is installed and the local entry dropped BEFORE the slow durable/Redis deletes, so a concurrent poll reconcile cannot re-apply the row it just read. A newer Revoke on the same key clears the tombstone, so an unrevoke never suppresses a later legitimate kill. Tombstones age out with the kill they replaced. - maintain() takes the same operation lock as Revoke/Unrevoke around its durable block, closing the window where a poll tick could re-Upsert a row Unrevoke had just deleted — invisible until a restart resurrected the kill. - Durable self-heal now compares expiry, not mere presence, so a failed Revoke mirror leaves a stale shorter row that the next tick repairs. - A failed unrevoke publish fails safe: other processes keep the kill until it expires. Propagation failures surface as warnings rather than weakening the local result. Unrevoking an over-cap victim is legal but the async enforcer will re-revoke it on its next pass while the user is still over cap; that is documented at the endpoint. Part of the stream monitoring & kill-switch epic. |
||
|
|
4b2031fdf8 |
feat(playback): account served bytes and hold transcode liveness server-side
Integrated deployments recorded no served bytes at all: BytesServed was only ever advanced by the edge writer, so a single-node install showed bytes_served: 0 for every stream. Worse, integrated LastServedAt was mapped from Session.LastActivityAt, which UpdateProgress advances — letting a client influence which of its own streams the over-cap enforcer trims first. - internal/playback: Session gains BytesServed and a distinct LastServedAt. LastServedAt advances only from server-observed events (AddServedBytes, BeginTransport, EndTransport) and never from a client progress report. LastActivityAt keeps its exact prior meaning and reaping semantics. - internal/playback/metered_writer: one shared SessionMeteredWriter for every integrated pour. It forwards io.ReaderFrom so the kernel sendfile path survives, guards the fallback copy against re-entering ReadFrom, and implements Unwrap() so the revocation cut's SetWriteDeadline still reaches the socket. Both properties have regressed on this branch before (GAP-3, GAP-9) and are now pinned by a chain test. - Wired at every integrated pour: native direct-play/remux and transcode segment, jellycompat direct/remux and HLS segment. Each site defers the tail flush; the wrapper alone loses the final partial chunk. - Close VERIFY-4 (buffer-ahead evasion): an unpaused transcode gets a bounded 10m grace measured from the server-observed clock. This covers local, cleanly-completed and offloaded transcodes uniformly, unlike an ffmpeg liveness probe, which sees only local processes and reports false once a copy-mode encode finishes ahead of playback. Paused sessions keep their load-bearing 30m grace. - internal/nodesessions: the same idle window one layer out — transcode records idle out at 180s instead of 60s, applied through a single helper used by ActiveCount, Snapshot and refreshAll so the node's count and status view cannot disagree. Part of the stream monitoring & kill-switch epic. |
||
|
|
0217cce2df |
feat(playback): stream kill switch + async over-cap enforcer
Add the enforcement layer on top of server-observed monitoring: a revocation kill switch that stops any stream within ~120s and keeps it dead, plus an async over-cap enforcer that drives kills off the live monitoring picture — entirely off the per-segment hot path and with no client-protocol change. - internal/streamrevoke: the central kill list. IsRevoked is a pure in-memory lookup safe on the request hot path; a Redis pub/sub + poll mirror keeps edge caches current, and a Postgres durable mirror lets kills survive a server restart AND a Redis flush so a restart-resilient stream cannot be reconstructed and re-served after being killed. A user revocation is a cutoff (kills tokens minted before it, spares post-reauth tokens), not a 24h ban. - internal/streamenforcer: async over-cap brain — reads the monitoring snapshot and per-user limits, selects victims, and collapses every reason (exceeded limit, admin terminate, abuse) to the same action: write a revocation. - Edge + native + jellycompat enforcement: proxy refuses revoked sessions on every request and cuts long direct-play/remux pours mid-stream; the transcode node guards both serve and the reconstruct path so a killed session is never re-spawned after a node restart; jellycompat serve surfaces close their kill-switch coverage holes. - streamtoken.IssuedTime exposes the token iat the user-kill cutoff compares against; token IssuedTime + revocation guards wire through router, downloads, and admin terminate-by-id (with admin-list dedupe). - Restore sendfile zero-copy on direct-play/remux byte counting so the monitor's served-byte accounting does not cost the sendfile fast path. - migrations/sql: stream_revocations durable table. Part of the stream monitoring & kill-switch epic. |
||
|
|
22ffaad911 |
feat(playback): server-observed stream monitoring (async, no client trust)
Introduce a first-class, authoritative view of what is actually streaming,
observed server-side and never trusting client progress reports. This is the
base observation layer the kill switch and async over-cap enforcer build on.
- internal/streammonitor: live-stream snapshot model plus pluggable Sources
(local func source, Redis source, multi-source fan-in) so a single node and a
multi-node deployment expose the same picture.
- internal/nodesessions/tracker: serve-activity attribution — LastServedAt and
served-byte counters advance from real serving, not client pings, giving an
authoritative liveness signal.
- Client identity as monitoring attribution: Origin ("native" | "jellycompat")
and ClientName ride the server-signed stream token (streamtoken.Claims) and
the transcode-start request so an edge/transcode node — which never sees the
originating API path — can stamp them onto its live-session record. These are
attribution only: not byte-affecting and not a trust assertion.
- Serve-activity marks on the transcode node (MarkServed) so a node's own record
reflects real serving instead of a LastServedAt frozen at start time.
- Admin observation surfaces: node/session listing carries owner + client
identity and dedupes multi-record sessions.
Part of the stream monitoring & kill-switch epic.
|
||
|
|
f0bc170113 | feat(admin): clarify AI service configuration | ||
|
|
c52ca7dd7a |
feat(admin): identify compat sessions and Android devices in the live session view (#495)
* feat(admin): identify Android devices by model in live session view Android clients that send a bare default User-Agent (e.g. "Dalvik/2.1.0 (Linux; U; Android 11; AFTKRT Build/RS8180.3729N)") showed up as "Dalvik" in the admin live-session view, which tells an operator nothing about the device. Parse the model code out of the UA (the token between the last ';' and "Build/") and map the Amazon Fire TV family and NVIDIA Shield to product names. Unknown but parseable models fall back to "Android · <MODEL>" instead of "Dalvik"; multi-word models like "Pixel 7" are preserved whole. This is display-only: the session still stores the raw model code in its user agent, and no response field or contract changes. * feat(admin): mark Jellyfin-compat sessions with the JF pill by origin The admin "JF" pill was derived at read time by substring-matching a token list against the client name / user agent. A real Jellyfin client that authenticates through the compat surface but sends a bare User-Agent and no MediaBrowser client name (e.g. a Fire TV app) got no pill, even though it plainly came through the Jellyfin API. Stamp compat origin as immutable identity at session creation and carry it through to the admin view: - ClientInfo.IsCompat is set true in the jellycompat auth path; newSession copies it onto Session.IsJellyfinCompat. - The flag rides the durable RecipeCard (next to the client metadata that already exists so the pill survives reconstruction) and is restored in ReconstructSession, so a server restart keeps the pill. - buildLiveSessionSync -> worker.SessionSync -> a new compat_origin column on playback_sessions_sync (added migration); the reconciler upserts, reloads, and compares it so origin changes still publish and unchanged rows do not churn. - The handler ORs the stored origin with the existing name/UA heuristic, which stays as a fallback for rows written before this column existed. is_jellyfin_client keeps the same name and type on the wire; it is only sourced more accurately. * fix(admin): correct Android device labels --------- Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> |
||
|
|
95edf19389 |
feat(invitations): shareable claim links and open-in-app on the claim page (#509)
* feat(invitations): always return the claim link so admins can share it directly The claim URL was only surfaced when email sending failed. Admins who want to hand the link over another channel (chat, SMS) had no way to get it — and the raw token exists only in the send/resend response, since the server stores just its hash. The create and resend flows now always include claim_url (additive on /api/v1), and the admin UI keeps the dialog open after either action with the link and a labeled Copy button. Truncation and stacked buttons keep the unbreakable URL from forcing horizontal scroll on phone widths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(web): offer to open invite claims in the Android app The Android app already registers silo://invite?server=...&token=... with a full native claim flow, but nothing ever emitted that link — an https invite always ended in the browser. On Android user agents the claim page now leads with a prominent 'Open in the Silo app' button carrying that deep link, with the web form kept below as the fallback ('or set up in the browser'). The button is a plain anchor: a user-tapped custom-scheme link is the one reliable path, and we never fire it automatically since there is no installed-check and a miss surfaces an OS error. The password field's autofocus is suppressed alongside it so the keyboard doesn't push the button off screen. iOS is excluded until the Apple app registers the scheme. The server origin travels in the server param verbatim, so non-443 ports and plain-http LAN servers need no extra convention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
63e18cf37f |
fix(playback): prevent transcode resolution upscaling (#503)
* docs(playback): design transcode resolution clamp * docs(playback): plan transcode resolution clamp * fix(playback): prevent transcode resolution upscaling * refactor(playback): share transcode resolution tiers --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> |
||
|
|
271a2e1741 |
feat: emailed invitations, claim + household setup, and server-driven onboarding tour (#501)
* feat(invitations): add emailed pre-provisioned invitations
Admins can invite a specific person by email: the invitation pre-binds
role, access group, and library access, and the invitee only chooses a
password. Their email address becomes their username, so login gains an
email fallback (username lookup first, email column only on miss for
inputs that parse as a bare address).
- invitations table: single-use token (SHA-256 at rest) bound to one
address; a partial unique index makes resend-supersedes atomic; no
users row exists until accept, so a typo'd address can't squat a
username. Status is derived from timestamps, not stored.
- internal/invitations: repository, service, and branded email through
the shared internal/mail sender. When SMTP is off the claim URL is
returned for manual delivery instead of failing.
- Admin endpoints /admin/invitations (list/create/resend/revoke) beside
the existing invite-codes routes; public claim endpoints
/invitations/{token} (+/accept) rate-limited with the other auth
endpoints. Unknown/expired/revoked/used tokens are indistinguishable.
- Accept returns the same login response shape as signup, so clients
reuse their session plumbing.
Spec: docs/superpowers/specs/2026-07-27-invitations-and-onboarding-design.md
Plan: docs/superpowers/plans/2026-07-27-invitations-and-onboarding.md
Part of #215
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(web): add invitation admin tab, claim page, and household setup
- Admin → Users gains an Invitations tab: compose (email, role, access
group, libraries, note, first-profile and tour toggles), list with
derived status, resend, revoke. When the server has no SMTP the create
response's claim URL is surfaced for copy-paste instead of a fake
success.
- /invite/:token claim page: everything but the password was decided at
send time, so it asks for exactly one thing and lands the user signed
in. Expired/used links get an explanatory card, not a 404.
- /household-setup ("Who's watching?"): profile tiles plus the existing
ProfileEditorDialog, all through the existing /profiles endpoint —
no new backend. "Just me for now" is a first-class exit.
Part of #215
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(onboarding): add server-driven onboarding manifest and state
GET /onboarding/flow returns the ordered first-run tour for this server
and profile: steps for disabled features (requests, watch together,
recommendations, notifications) are filtered out server-side, surface=tv
drops steps needing text entry, and child profiles never see stops they
can't act on. Copy lives in Go, so a wording fix is a deploy — clients
render step kinds they know and skip unknown ones by contract.
setting_choice steps name an explicit write target (profile_field /
setting / device_setting) because playback quality is a profile column,
not a settings key — the tour writes through the same APIs the settings
screens use.
Per-profile completion state lives in the user store (SQLite schema v14
+ a Postgres twin table), keyed by (profile_id, tour_id) with monotonic
completed/skipped timestamps: finishing on one device silences every
other; a later progress write can never un-complete.
Part of #215
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(web): add the first-run feature tour
TourHost renders the server manifest as a modal overlay on Home: unknown
step kinds are skipped silently (the forward-compat contract), progress
posts per step, and setting_choice steps write real values through the
existing profile/settings mutations — by the last step the account is
genuinely configured. Skip is always one click and recorded server-side,
so no other device re-prompts. The tour ends by handing off to the
existing taste-seed picker, which now waits for the tour to finish
before its own redirect. Settings → Personalize gains a replay entry.
An invitation sent with show_tour=false plants a local hint that the
gate converts into a server-side skip for the first profile.
Part of #215
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): satisfy noUncheckedIndexedAccess in the tour's advance step
The Docker web build runs `tsc -b`, which applies the project's
noUncheckedIndexedAccess; the bounds check didn't narrow steps[next].
Look the step up once and branch on its presence instead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): blur the whole app behind the tour, sidebar included
The tour overlay rendered inside the app layout, where an ancestor
creates a fixed-position containing block — inset-0 pinned to the
content pane, leaving the sidebar completely un-scrimmed. Portal the
dialog to <body> so the scrim truly covers the viewport, and raise the
backdrop blur from sm (4px) to xl (24px) so card titles and nav labels
aren't legible through it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(onboarding): name features by their UI labels in the tour copy
"Same movie, different couches" never said what the feature is called.
Every feature card now leads with the name the sidebar actually uses —
Watch Party, Requests, Watchlist, Calendar, Notifications — and says
where to find it, so the tour teaches vocabulary, not just concepts.
Server-side copy, so all three clients pick this up with no release.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(onboarding): add apps and Jellyfin-compat steps to the tour
Two new web-only feature cards near the end of the tour:
- "Take Silo with you" — native apps for iPhone/iPad/Apple TV and
Android/Android TV, with outbound TestFlight and Play Store links.
Steps gain an additive links field (label + url) that older clients
ignore; the web TourHost renders them as external-link buttons.
- "Already use a Jellyfin app? It works here" — Infuse/VidHub/Findroid/
Swiftfin connect via the Jellyfin API. Gated on
jellyfin_compat.enabled (default-on, so unset counts as enabled;
only an explicit "false" hides it).
Both steps are web-only: the apps card is pointless inside the apps it
advertises, and TV can't open store links. surface=phone/tv manifests
skip them, covered by tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(web): keep the tour card responsive on phone widths
Verified every step at 1600px, 390px, and 320px with an automated
overflow check. Fixes it found:
- Link buttons (apps step) now wrap and truncate instead of extending
past the card edge.
- The footer wraps at very narrow widths, so the handoff step's wide
primary button drops to its own line rather than overflowing.
- Progress pips hide on phones — decorative, and they crowded the
Back/Next buttons.
- The card scrolls within 85dvh so a tall step never pins its buttons
off-screen on landscape phones.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat(web): render store links as branded badges in the tour
The apps step's plain outline buttons now render as store badges: the
Apple or Google Play mark with a store eyebrow (TestFlight beta /
Google Play) over the platform label — the familiar app-store badge
idiom. The brand is inferred from the link's host on the client, so
the server contract stays icon-free and non-store links keep the plain
external-link button. Labels drop the parenthesized store name the
eyebrow now carries.
Verified at 1600px and 390px with the overflow sweep: none.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
54e184df85 |
feat(requests): enforce per-profile rating limits in discovery (#505)
* feat(requests): enforce per-profile rating limits in discovery - Resolve each profile's max content rating and filter discovery, detail, and browse results against it, failing closed on missing ratings - Reject request submissions for titles above the viewer's ceiling - Add TMDB GetCertification backed by release_dates/content_ratings with a long-lived cache and singleflight - Push certification.lte to TMDB for studio/network/genre browse as a cost pre-filter - Backfill restricted section pages from a fixed window of TMDB pages to keep carousels populated and pagination stable * fix(requests): address discovery rating review findings - Preserve backfill overflow: sections use plain TMDB cursor semantics plus an additive next_page field instead of fixed windows, so an early stop never drops allowed titles from unconsumed pages (bit hardest at permissive R/TV-MA ceilings). - Bound cold-path cost: DiscoverAll backfills at most 2 TMDB pages per section (vs 5 for a direct section request), capping worst-case cold certification hydration at 240 lookups instead of 600. - Keep the TMDB prefilter a superset: rank-3 ceilings now push down certification.lte=NC-17/TV-MA rather than R, so titles the local ladder allows can't vanish upstream unrecoverably. - Fail closed on foreign certifications: enforcement-path lookups use new US-only pickers (a Canadian PG no longer reads as US PG), while the display path keeps its any-country fallback. US multi-entry disagreements prefer the theatrical/real rating over festival NR. - Detach shared certification fetches from the first caller's context (WithoutCancel + 30s bound) so one disconnecting client can't fail the singleflight result for concurrent waiters. - Advertise enforcement via rating_restrictions_enforced on /requests/status so clients can feature-detect instead of version-sniffing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(requests): harden rating enforcement per second review pass - GetDetail gates on the US-only enforcement certification (cached GetCertification) instead of the display rating, whose any-country fallback let a foreign "PG" pass the US ladder. - pickUSMovieCertification takes the strictest recognized US rating when multiple release entries disagree ([PG, R] -> R); entry order is not meaningful and enforcement must not admit a title on its most lenient certificate. - Certification singleflight uses DoChan so a canceled caller returns ctx.Err() immediately instead of blocking up to 30s on the detached shared fetch (which still completes for surviving waiters). - Viewer rating ceiling resolves once per request and threads through discover/browse/detail enrichment (enrichPageWithCeiling); DiscoverAll drops from 12 scope resolutions per load to 1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f235524365 |
feat(diagnostics): chunked report upload fallback for proxy body caps (#494)
* feat(diagnostics): chunked report upload fallback for proxy body caps
Diagnostics bundles can be up to max_bundle_bytes (10 MiB default), but a
reverse proxy in front of Silo commonly caps request bodies at nginx's
default client_max_body_size of 1 MiB. Such a proxy answers the single-shot
multipart upload with its own 413 before Silo ever sees the request, so any
report over the cap could never be delivered.
Add a chunked upload fallback under /api/v1/diagnostics/reports/uploads:
- POST / {manifest, bundle_bytes} opens a session
- PUT /{id}/chunks/{index} streams one ≤768 KiB chunk (proxy-safe)
- POST /{id}/complete ingests the assembled bundle
- DELETE /{id} best-effort abandon
The assembled bundle goes through the exact same Ingest path as the
single-shot endpoint, so every content check (manifest contract, archive
sha/bytes/entries, quotas, profile attribution) applies identically.
Sessions reuse internal/uploads (the plugin chunked-upload spool manager)
plus a small owner map for per-user isolation; they spool to disk, expire
after 15 minutes, cap at one per user / 16 global, and complete shares the
existing per-user + global in-flight ingest limiter.
/diagnostics/status now advertises upload_chunk_bytes so clients can detect
support; older servers omit the field and clients treat that as
unsupported. The demo guard's diagnostics prefix gains PUT to cover the
chunk route.
Verified end to end against an OpenResty proxy with a 1m body cap: the
single-shot upload 413s, the same 1.6 MiB bundle uploads in three chunks
and lands as an accepted report; also exercised from the tvOS client's
fallback path in the simulator.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(diagnostics): harden chunked upload sessions per review
- Reserve the per-user slot and global cap atomically in init (a
reservation map counted with live sessions), so concurrent inits by one
account can no longer fan out past one session or transiently exceed the
cap. Creation failures roll the reservation back.
- Move chunk body I/O outside the uploads.Manager mutex: a slow client
streaming one chunk no longer serializes every other session's chunk
writes, completes, and cancels. A per-chunk in-flight flag rejects
duplicate concurrent writes to the same offset (ErrChunkBusy → 409), and
cancel/expiry defer spool-directory removal to the last finishing
writer.
- Chunk arrivals refresh the session expiry, making the TTL an idle
timeout instead of an absolute deadline so a slow-but-progressing upload
cannot expire mid-transfer.
- Extend the request read deadline on chunk PUTs and both deadlines on
complete, matching the single-shot handler's slow-uplink handling.
- Keep the session when complete's availability re-check fails
transiently (status load error → 500): only definitive
disabled/storage-unavailable answers discard the spool, so a retried
complete succeeds without re-uploading every chunk.
- Reclaim orphaned spool directories at startup (a restart previously
stranded the old process's partial uploads forever) and sweep expired
sessions on a timer instead of only from later init traffic.
- Document that session state is process-local and what that means for
multi-replica deployments.
Adds concurrency/race tests (go test -race) for atomic admission,
same-chunk write exclusion, expiry refresh, transient-status retry, and
startup reclaim.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(diagnostics): count detached chunk writers and lift chunk PUT write deadline
Second review round:
- A canceled session whose slow chunk writer was still draining held a
connection and spool disk but vanished from every count, so a
cancel-and-reinit loop could stack unbounded live writers behind the
16-session cap. The uploads manager now parks such sessions in a
detached set (exposed as DetachedWriterSessions) until their last
writer returns, and diagnostics init counts them in its admission gate.
- Chunk PUTs now extend the write deadline as well as the read deadline:
on an uplink slow enough to eat the server's 120s WriteTimeout, the
stored chunk's JSON acknowledgement would otherwise be lost and the
client would retry an already-accepted chunk.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
8dea4b9056 |
fix(playback): stop a broken Dolby Vision RPU from hanging playback (#485)
* fix(playback): stop a broken Dolby Vision RPU from hanging playback
A Profile 7 source whose RPU ffmpeg cannot parse took the whole session
down. The dovi_rpu bitstream filter does not fail cleanly — it rejects
every packet while ffmpeg runs on, so one observed session emitted
376,316 stderr lines before the process was killed, no manifest was ever
built, and the client got a 503 after ~10 seconds that it showed as an
endless spinner:
[dovi_rpu] Failed to read unit 1 (type 39).
[vost#0:0/copy] Error applying bitstream filters to a packet:
Invalid data found ... Invalid SEI message: payload_size too large
Whether the strip works is a property of the file, not of ffmpeg, so
SupportsDoviRPUFilter cannot answer it — but asking ffmpeg to strip two
seconds to the null muxer can, in about a second. Profile 7 sources are
probed once each on the start path and the result is cached; a source
that fails is copied without the filter, which leaves the base layer and
plays. Refusing to play at all does not.
The probe reads stderr rather than trusting the exit code: ffmpeg treats
a per-packet bitstream-filter error as non-fatal and exits 0, which is
exactly how a stream that could never start reached a live session.
Cache is keyed on path plus size so a replaced file is re-probed, and
bounded like the letterbox cache. A nil probe keeps stripping, since most
Profile 7 sources need it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK
* chore(playback): use InfoContext in the rpu probe
* fix(playback): decide the DV RPU strip in the plan, not at the transport
The probe answered the right question in the wrong place. Suppressing the
bitstream filter as the transport was being built left the plan still
promising DynamicRange "hdr10", Claims.Video.HDR10 and the
dolby_vision_metadata_removed / hdr10_base_layer_preserved validated
claims, and left session.RemuxDVMode persisted as strip_to_hdr10 — so the
client was told it was getting clean HDR10 while receiving Profile 7 with
dangling RPUs, and every path that re-derives the filter from the durable
session put the hanging filter straight back:
- HandleStartTranscode (quality change, seek, burn-in restart), which
also re-derives it from file.PrimaryDVProfile() == 7 with no plan at all
- the remote audio-switch restart, from updatedSession.RemuxDVMode
- the progressive-remux transport, which plan_v3 evaluates *before* the
HLS branches, so for many clients the broken file never reached the
probe in the first place
The verdict is now a planner input alongside the transformation
registries: the registries answer whether the executor carries the
transformation, this answers whether the file does. A source that fails it
is never planned onto a strip route, so the plan's claims, RemuxDVMode and
every restart derived from them agree with what the pipeline can produce.
With no tone-map recipe in this tree an HDR10-only client has no route
left, so it gets a dv_conversion_unsupported terminal naming the real
cause rather than a generic HDR message; a client that can run its own DV
transformation still gets that route, with a degradation warning
explaining why the server route was dropped.
The two paths that bypass the plan entirely are gated at the executor:
legacy/auto remux neutralizes the profile exactly as it already does for a
missing dovi_rpu filter, and the explicit v3 strip recipe fails loudly
rather than emit dangling RPUs under an HDR10 claim.
The probe itself:
- Tri-state verdict. Only a stderr-confirmed rejection is cached. A
timeout, a cancelled request or an ffmpeg that will not start is
inconclusive: the strip is kept and nothing is written, so one client
disconnecting can no longer disable the strip for a file permanently.
- Singleflight, so concurrent or retried starts of a title share one run.
- Timeout cut to 6s, inside the budget a client waits on the manifest,
and off the session lifecycle lock now that it runs at planning time.
- Keyed on size and mtime as well as path and binary, so a file replaced
in place with the same length is re-probed.
- stderr capture bounded at 64 KiB; the markers are in the first lines
and a rejecting filter emits a pair per frame.
- The head-only coverage is stated rather than asserted: this catches a
source that rejects from the first access unit, not one that breaks an
hour in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(playback): keep the RPU probe alive when its caller leaves
The probe rode the caller's context, so a leader whose client disconnected
returned dvRPUUnknown — and CanStrip then handed that fail-open to every
follower already blocked on the shared call, even though their own requests
were still alive. One client walking away was enough to give a live session
the strip the source cannot survive, which is the hang this whole change
exists to prevent. The work was also thrown away, so the next start paid for
the probe again.
Run it under context.WithoutCancel instead. dvRPUProbeTimeout still bounds
it, so nothing is left running; a verdict reached after the leader has gone
is still correct and still worth caching. Follower behaviour is unchanged:
a follower whose own request is cancelled still leaves immediately rather
than waiting on someone else's probe.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
|
||
|
|
172beb99ef |
fix(watch-together): stop a dropped socket reading as the host leaving (#487)
* feat(watch-together): make vote rooms actually vote selection_mode has been stored, normalized and published since the feature landed, and nothing has ever read it. A "vote" room behaved exactly like a host_pick one: members could suggest and vote, the tally was recorded and broadcast, and then the host promoted whatever they liked regardless of it. In a vote room the host now starts the winner rather than choosing it. Promoting anything other than the leading suggestion is refused, because being able to overrule the tally makes the mode host_pick with extra steps and turns the vote counts on everyone else's screen into decoration. The winner is the head of the repository's existing ordering (vote_count DESC, created_at ASC): most votes, ties to whoever suggested first — deterministic, and re-suggesting a title cannot jump the queue. A room where nobody has voted has no winner and says so, rather than quietly promoting the oldest suggestion as though a vote had happened. host_pick rooms are untouched: the host still promotes freely. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK * fix(watch-together): close the second door into a vote room's selection Gating PromoteSuggestion left SelectItem wide open: it is host-only but was not gated by selection mode, so the host of a vote room could set any title directly and bypass the vote entirely. Enforcing the tally on one path and not the other makes the vote counts on everyone else's screen decoration. A vote room now refuses a direct selection outright. The winner is the only way in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK * fix(watch-together): stop a dropped socket reading as the host leaving hostDisconnectTTL was 15 seconds, which treated any transient drop as a departure. An explicit leave and an explicit close already tear the room down immediately, so this timer only ever covers a host who has NOT said they are going — and at 15s a host who backgrounded the app, moved between screens, or hit a brief network blip lost the room for everyone with a "host_left" nobody could explain. Two minutes survives a reconnect or an app switch, and is short enough that a genuinely departed host does not leave a room open all evening. The janitor still reaps idle rooms independently. This matters for what the clients are growing into: a room you stay in while you browse for something to suggest. A client that drops its socket when the lobby leaves composition should cost you a reconnect, not the room. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Db4dSxN9tH8yN7uUP549tK * fix(watch-together): let a vote room actually start its winner The vote gate landed on both doors into a room's selection, but promoting the winner walks through SelectItem to commit — so the gate meant to stop the host bypassing the vote also stopped the vote itself. Vote rooms could not start playback by any route. Split the commit path: SelectItem keeps the gate for direct requests, and PromoteSuggestion goes through the internal path once it has confirmed the suggestion is the winner. Map ErrVoteRoomSelection in the promote handler too, so a future regression there reads as a conflict rather than a 500. Add service-level tests for both gates — the previous tests only covered the pure winnerFrom helper, which is why the suite stayed green while vote rooms were non-functional. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> |
||
|
|
148c9291c5 |
feat(web): add Connect Apps settings page for compat sign-in (#488)
* feat(web): add Connect Apps settings page for compat sign-in Jellyfin-protocol clients offer one username box and one password box and never prompt for a profile, so signing in requires `account#Profile` and `password#PIN`. Nothing in the product taught that syntax, and every way of getting it wrong surfaces as the same "invalid username or password", so it became a recurring support burden. Add Settings -> Connect Apps, which states the credentials for the signed-in account rather than describing them in the abstract. The page is segmented by app type: the two formats are never shown at once, each side names the apps it covers, and the compat side is visually distinct so it cannot be mistaken for the normal Silo login. The compat listener's separate address is shown too, since pointing a client at the Silo address fails identically to a bad password. Backend adds GET /api/v1/compat/connect-info, an account-scoped read of the compat listener's enabled flag, public URL, and server name. The admin status endpoint already covers this ground for operators, but it also reports install paths and version provenance, so it stays admin-only; this returns only what a client learns by connecting anyway. It is auth-only and deliberately not profile-scoped, since it describes how to sign in. ConnectInfoForConfig shares the enabled-flag precedence with WebComponentStatusForConfig via compatEnabled, so the two endpoints cannot disagree about whether compat is on. The page declines to display a username it knows cannot work: profile names permit `#` but the resolver splits at the last one, so `alice#Movie #2` parses as account "alice#Movie". Such profiles get an explanation and no copy button instead of a string that fails to authenticate. Part of #432 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(web): harden Connect Apps against misleading sign-in states Review of the first pass found six ways the page could state something untrue. Each one matters more than usual here, because the page exists specifically to stop people guessing at credentials. Report the running listener, not the stored setting. jellyfin_compat.enabled is restart-required and cmd/silo builds the compat server from the boot config alone, so the stored value describes intent. Reporting it promised credentials for a listener that does not exist yet, or claimed the API was off while the running one kept serving. ConnectInfo now returns the boot-time state plus a pending_restart flag, and the page distinguishes "not running yet" from "turned off". The admin status endpoint keeps reporting configured intent, so the two intentionally diverge until a restart. Stop presenting fetch failures as a disabled compat API. React Query clears isLoading on error, so a failed request rendered "the compatibility API is turned off" and sent users to an admin about a setting that was fine. A failed profile list was worse: it fell through to an empty list and offered the bare account name, which silently drops the profile suffix. Both now withhold credentials and say the load failed. Detect accounts that cannot use password login. Compat login is hardwired to the local provider, which rejects accounts with local_password_login_enabled false before checking any password, so SSO and plugin-provisioned accounts can never authenticate. The page told them to type a password anyway; it now says the compat API cannot accept the account. Flag loopback compat addresses. jellyfin_compat.public_url defaults to http://127.0.0.1:8096, which resolves to the client device on the phones and TVs this page names. An untouched default was offered as the exact address to copy; it is now explained instead of presented as usable. Apply the #-in-profile-name guard to the summary list too. The selected-profile field withheld an unusable username while "Every profile at a glance" reintroduced it two sections below. Read only the three settings this endpoint consumes. GetAll on the encrypted repository decrypted every stored secret on each authenticated page view, and an unrelated decryption failure would have silently dropped valid compat overrides. Invalidate the connect-info cache when an admin saves jellyfin_compat.* settings. The address applies without a restart, but the page cached it for five minutes and kept offering the old value for copying. Part of #432 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
99d205676f |
fix(metadata): prevent stale cross-provider IDs (#480)
* fix(metadata): prevent stale cross-provider IDs * fix(metadata): address stale ID review findings * fix(migrations): build the stale-ID primary key concurrently ALTER TABLE ... ADD PRIMARY KEY builds the index under ACCESS EXCLUSIVE, blocking reads and writes on stale_media_ids for the whole build. Create the wider unique index with CREATE UNIQUE INDEX CONCURRENTLY and attach it with ADD CONSTRAINT ... PRIMARY KEY USING INDEX instead; all three key columns are already NOT NULL, so the attach is metadata-only. Same treatment on the rollback path, plus the repo's INVALID-remnant cleanup so a failed concurrent build is not silently accepted by IF NOT EXISTS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
02203d9e40 |
fix(playback): trust the server's media runtime end to end (#482)
* fix(scanner): reject durations that imply an impossible bitrate The duration-plausibility rule only rejected videos of 10 seconds or less, so a feature film that probed as 61 seconds passed untouched and persisted. Clients then had nothing trustworthy to anchor on: Android's grow-only duration ratchet has no floor to hold when the catalog value is wrong, so the playback engine's growing-HLS-window duration won and a 90-minute movie displayed as ~1 minute. Size and duration together pin an implied bitrate, which separates the two cases the absolute floor conflates. A genuine short clip has an ordinary bitrate; a 100 GB file claiming 61 seconds implies ~13 Gbps. The ceiling sits far above any real medium, so legitimate content cannot trip it — and unlike the absolute floor, it does not false-positive on a genuine high-bitrate short. Also bump the repair-rule revision marker so rows judged by the previous, weaker rule are re-checked once under this one. Without that bump an improved rule never reaches the rows it was written for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(playback): publish source runtime in v3 plans and stop faking the copy seek window Two defects with one root: a v3 plan described where playback sits without ever stating how long the media is. Add source.duration_seconds. It is the file's full runtime, never `total - source_start` and never adjusted by timeline_offset_seconds, and it is omitted rather than null when unknown — clients that coerce null to a numeric default would read it as zero, the exact value this field exists to stop them inventing. It is set in SourceDescriptorFromFileV3, the single place every delivery already flows through, so direct play, progressive remux, HLS remux and HLS transcode all carry it. Until now the v3 plan omitted duration entirely, so clients fell back to the playback engine. On an HLS copy remux the server intentionally serves FFmpeg's still-growing playlist, so the engine reports the length produced so far. With no server-supplied runtime to anchor on, a feature film played back as a couple of minutes. The legacy protocol already answered this correctly via fileDurationSeconds; this restores parity. Separately, the copy branch published seek_window_end_seconds as the media runtime. That made the window look *complete*, which clients read as proof that any target inside it is locally seekable, so they native-seek past the produced head of a growing playlist instead of asking for a reanchor. Leave the end open: an incomplete window plus can_seek_anywhere=false routes every seek through the server, which is what legacy did before v3 added the bound. Advertise plan_source_duration_v1 so a client can distinguish "this server does not populate the field" from "this server knows the runtime is genuinely unknown" — without it, both look like an absent field and a client cannot tell whether its own catalog fallback is still required. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(web): pair the exit position with the media runtime, not the element duration The player's exit state converts its position to media time but took the duration from the video element, which is player-local. On a remux or transcode stream the element only covers the window produced so far, so the two values live in different coordinate systems. Resuming a movie 50 minutes in makes that concrete: the exit position is ~3060s of media time while the element reports ~120s. The progress cache then evaluates `position >= duration`, marks the item completed, latches the watched badge, and — because completion clears the resume point — resets position to 0. Exiting a resumed movie destroyed the resume point and claimed it had been watched. The server's runtime is authoritative and already expressed in media time, so prefer it and fall back to the element only when no server value exists. The rule moves into mediaTimeline.ts next to the coordinate conversions it depends on, which is also what makes it testable — VideoPlayer itself has no test harness. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
10394b0a05 |
fix(scanner): stop hiding media when a library root is offline (#472)
* fix(scanner): stop hiding media when a library root is offline
A scan that cannot read a library root found no files there, so every
cataloged file under it was marked missing. Catalog reads all filter on
missing_since IS NULL, so marking is equivalent to deletion from a user's
point of view: the title leaves browse, search and next-up, and playback
answers "Source media file is missing" for media that is intact on disk.
The dead-root protection already existed but only guarded the destructive
operations. protectedConfiguredRoots was computed *after* the marking loop
in scanPaths, and applyScopedScan received the protected set but applied it
only to its force-delete branch. So an unreachable root could not lose its
rows, but could still have its entire catalog hidden until the next
successful scan.
On a CephFS deployment whose per-library subvolume mounts flap, this marked
190 present files missing in a single day — 15% of all missing-flagged rows
were files sitting untouched on disk, some flagged more than 20 hours after
their last write.
Hoist the probe above the marking loop and skip files under an unreachable
or suspect-empty root in both the folder and scoped paths. Pass the
unreachable set to the walked-scope call too: a nested child mount can die
under a healthy parent, and its rows are inside the parent's scope.
An offline root tells us nothing about whether its files exist. The only
safe reading is to leave them alone and let the next good scan decide.
Genuine deletions under a reachable root are unaffected and still marked
and swept on the same schedule as before.
Report the count as ScanResult.MissingSkippedProtected and log it, so an
operator can tell "my library shrank" from "my mount dropped".
Known gap: a suspect-empty *nested child* root under a healthy parent is
still marked missing, because the suspect set is not resolved until after
the walk loop. Unreachable roots — the case observed in production — are
covered.
* fix(scanner): protect suspect-empty and partially-walked roots too
Addresses review findings on #472. The original change guarded missing-marking
against probe-unreachable roots, but left three ways for a storage fault to
still hide a healthy library.
Suspect-empty detection was reactive. suspectEmptyRoots asked
ListRootsWithOnlyMissingFiles, which returns a root only once it has NO live
rows left. On the first scan after a mount drops — the moment that matters —
the rows are still live, so the root was not classified suspect and the scan
marked everything missing. The protection then engaged on the next scan, in
time to protect the wreckage. Ask ListRootsWithCatalogedFiles instead: any
cataloged row under an empty-but-reachable root is the lost-mount signature.
Intentional emptying is still reachable through the operator's one-time
cleanup allowance, which is the deliberate path for it.
Nested suspect-empty children were unprotected. Root compaction sends only the
populated parent through the walked-scope branch, which received only
unreachableRoots, so an empty child mountpoint had its rows marked missing on
its parent scanning cleanly. Pass the suspect set as well.
Partial walks were treated as authoritative. walkLogicalTree deliberately
swallows per-entry Lstat/ReadDir failures so one bad file cannot abort a scan
of a million, and collectLogicalFilePaths passed nil for the failure counter —
so the video path had no signal at all. A mount dying partway through
traversal produced a short file list indistinguishable from a large deletion.
Thread the counter through, and exclude a scope whose walk came back
incomplete from missing reconciliation, mirroring what the ebook scanner
already does via ebookRootScan.failed.
Also extract the duplicated mark-missing loop into markMissingExcludingProtected
so the folder and scoped paths cannot drift, and correct two comments that
still described the pre-fix "files are marked missing" behaviour — the exact
text a future reader would have trusted when reintroducing this bug.
TestScanFolderNestedSuspectEmptyChildRootProtection asserted the old
behaviour and is updated accordingly.
* fix(scanner): scope walk-failure protection and stop pruning on partial walks
Addresses the second Codex review round on #472. The previous commit's
incomplete-walk protection was too blunt in one direction and applied too late
in another.
Walk failures were counted, not located, and any non-zero count protected the
whole library root. A dangling symlink is both common and permanent, so that
would have suppressed missing-file reconciliation for its entire root on every
future scan — genuinely deleted titles would stay live indefinitely. That is
the same class of bug as the one this PR fixes, pointing the other way.
recordWalkFailure now records the logical path of each unreadable entry, and
only those paths are protected. Per-entry failures record the child path, so a
dangling symlink protects itself and nothing else, while a directory that
cannot be read protects its subtree.
Snapshot and group pruning ran before the protection. reconcileScannedRoots
and reconcileScannedGroups delete whatever the walk did not see, and both run
ahead of the missing-file guard, so a partial walk still dropped root
snapshots, observed locations and group locations for the unread portion —
corrupting later metadata matching even though the media_files rows survived.
Upserting what was seen is always safe; pruning now waits for a scan that read
the whole tree.
The confirmed-cleanup allowance was consumed to no effect for nested suspect
children. The walked-parent branch protected them unconditionally and runs
before the allowance is consumed, and an already-reconciled scope cannot be
revisited — so arming the allowance burned the confirmation while the child's
rows stayed live forever. Read the allowance without consuming it before the
walk loop, and honour it there. Unreachable roots stay protected either way:
an outage is never a confirmation to erase a catalog.
Two new regression tests, plus signature updates in the ebook pipeline, which
already tracked walk failures and now shares the path-based representation.
* fix(scanner): re-probe nested roots and gate group pruning on walk completeness
Third Codex review round on #472; both findings confirmed.
Group pruning ignored walk completeness in the subtree path. scanPaths passed
the completeness decision to reconcileScannedRoots but left
reconcileScannedGroups on !allowEmptyRootGuard, which is always true for
ScanSubtree — so a subtree scan that hit an unreadable directory still replaced
group snapshots and locations from a partial inventory. Same rule now applies
to both.
Nested roots were not re-probed before their parent was reconciled. Root
compaction folds a child mount into its parent for traversal, so a child that
is healthy at the initial probe but drops before the parent is walked leaves no
scope of its own, and the post-walk re-probe only revisits scopes that walked
empty. The parent walks files, looks healthy, and the child's rows are marked
missing on its success. reprobeNestedRoots re-checks this root's configured
children immediately before reconciling, protecting any that have since become
unreachable — or suspect-empty, unless the operator has confirmed cleanup.
Also guard suspectEmptyRoots against a nil file repository, matching
emptyCleanupArmed: without a catalog there is nothing to protect.
* fix(scanner): keep re-probed outages protected through folder-wide cleanup
Fourth Codex review round on #472; both findings confirmed. The first could
destroy data.
reprobeNestedRoots protected a root it found offline only for the scope being
reconciled, then discarded the result. The folder-wide membership reconcile and
the trash sweep afterwards rebuilt their protected set from the initial probe
alone, so rows under a child that dropped mid-scan — already marked missing and
past the removal grace — were hard-deleted by the very scan that noticed the
outage. Accumulate those roots in reprobedRoots, fold them into
protectedScanRoots, and reuse that set for the membership reconcile and sweep
instead of rebuilding. They now also land in ScanResult.UnreachableRoots so the
folder warning reflects the outage rather than presenting a partial scan as
clean.
Snapshot and group pruning was enabled for scopes that were never walked. The
gate was len(walkFailures) == 0, but an unreachable root gets nil walkRoots, so
it has no walk and therefore no failures — and pruning then deleted its
snapshots, observed locations and group locations even though its media rows
were protected. The same held for a suspect-empty child compacted into a
populated parent. Pruning now additionally requires that the scope was actually
walked and contains no protected path.
The new test pins that the sweep honours the protected set it is given. It does
not reproduce the mid-scan race itself: staging that needs the drop to land
between the probe and the walk, which a test cannot reach without hooks. That
path is covered by inspection, and the test comment says so rather than
implying coverage it does not have.
* fix(scanner): route every protection source through one folder-wide set
Fifth Codex review round on #472. Two P1s, one of them the second data-loss
path in this area — and the direct sibling of the one fixed in
|
||
|
|
ee31a1f0e2 |
feat(playback): formalize resumable direct streams and stall observability (#464)
* feat(playback): formalize resumable direct streams and stall observability Implements #443: strong stat-based ETag + If-Range on original-file direct play (via http.ServeContent), stream-end outcome classification in RollingDeadlineWriter (stalled_reap vs client_gone vs completed) with a structured log event and Prometheus counters, the direct_stream_resume_v1 protocol-v3 capability, and a contract doc. Progressive remux is explicitly excluded from the resume contract. Code written by OpenAI Codex CLI (gpt-5.6-sol) from a Claude-authored spec; reviewed and verified by Claude. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(playback): harden direct stream resume contract * test(playback): cover resume platform contracts --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
22fec4ed2d |
feat(metadata): add resilient match queue diagnostics (#463)
* feat(metadata): add resilient match queue diagnostics * fix(metadata): harden match queue lifecycle --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> |
||
|
|
2bc6ebb013 |
Merge pull request #460 from Silo-Server/agent/repair-anchored-identity-links
fix(scanner): repair historical provider-anchored merges |
||
|
|
5e2dd91d6b |
Merge pull request #456 from Silo-Server/feat/admin-settings-contract
fix(admin): enforce settings contracts end to end |
||
|
|
920629e0ad | fix(admin): address settings contract review findings | ||
|
|
e625d574e3 | fix(scanner): repair historical provider-anchored merges | ||
|
|
3c56ed606c | fix(admin): enforce settings contracts end to end | ||
|
|
2e45a9c015 |
Merge remote-tracking branch 'origin/main' into pr-397-devmerge
# Conflicts: # internal/catalog/item_repo.go |
||
|
|
166c5ef32f |
Add reliable Jellycompat watch scrobbling
- Forward start, pause, resume, and stop events with stable media identities - Persist and retry terminal scrobbles across teardown and restart paths - Reject ambiguous playback-report route matches |
||
|
|
9fad08a6fa |
fix(diagnostics): address round-6 review findings on PR #445
- AdminDiagnostics list: fix regression where rows dereferenced the now-omitted manifest for app_build. Project app_build server-side out of manifest JSONB into both list and detail responses (cheap COALESCE(manifest->'report'->>'app_build','')), split the TS type into DiagnosticReportSummary (list, no manifest) and DiagnosticReport (detail, with manifest), and read report.app_build in the row/detail. - embeddedManifestMatches: decode with json.Decoder + UseNumber so large integers above 2^53 (e.g. log_summary.lines) can't collapse to the same float and falsely match; re-assert no-trailing-data strictness. - Quota reservation (SKIP): reserving the client-claimed archive.bytes is sound because archiveMatches requires claimed==actual before MarkReady, so no stored report exceeds its reservation; documented in a code comment. - Multipart parts: reject a wrong-name/wrong-content-type part without calling part.Close(), which would drain up to the bundle limit while holding the in-flight slot; abandon it so malformed uploads fail promptly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh |
||
|
|
dee46f9398 |
fix(diagnostics): address round-5 review findings on PR #445
- service: reject supplied child-profile attribution with a distinct ErrChildProfileForbidden (403 child_profile_forbidden) instead of silently dropping it as if the profile were not found; a profile that is simply not the user's still drops attribution unchanged - repo: add a manifest-free list projection (reportListSelectSQL / scanReportSummary) for admin list and retention/stale cleanup queries so they no longer drag the full manifest JSONB per row; keep the full projection for GetByID/DeleteByID and mark Manifest omitempty - cleanup: delete/mark the DB row before the blob in retention and stale loops so a mid-run DB failure can't leave a ready report pointing at a missing bundle; blob-delete failures are logged with bucket/keys for orphan cleanup to reap rather than aborting the run (shared helper with the admin DeleteReport path) - admin: reject diagnostics settings where max_bytes_per_user would fall below max_bundle_bytes (and the reciprocal), which would make every max-size upload fail quota - router/demo: route POST /diagnostics/reports through DemoGuard and block the reports prefix in demo mode while keeping GET /diagnostics/status available Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh |
||
|
|
1f9bd99990 |
fix(diagnostics): address round-4 review findings on PR #445
- schema: add crash/report.type conditionals (allOf if/then) so a crash/anr/native_crash/hang/abnormal_exit manifest requires `crash` and a `manual` manifest forbids it, matching ValidateManifest. - service: reject uploads where X-Profile-Id and manifest.report.profile_id are both present but differ (new ErrProfileMismatch, mapped to 400 profile_mismatch) instead of silently preferring the header; single-source and matching cases unchanged. Adds service tests for mismatch, match, and header-only attribution. - schema: require manifest.json as the first archive.entries element via prefixItems (contains retained for validators without prefixItems support). - schema: document that maxLength is a character-count bound while the server enforces UTF-8 byte length, via a top-level note and per-field notes on the free-text device_summary and crash fields. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh |
||
|
|
93b851fe81 |
fix(diagnostics): address round-3 review findings on PR #445
- Extend the upload write deadline alongside the read deadline so a slow upload finishing after the integrated server's 120s WriteTimeout can still return its success response instead of timing out a report that succeeded. - Reject child-profile attribution for diagnostics: wire the attribution validator through a shared profile lookup that reports IsChild and drop attribution for child profiles, which must not perform diagnostics actions. - Assert the download test captures the clicked anchor and checks its blob: href and silo-diagnostics-<short_id>.tar.gz filename, not just cleanup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh |
||
|
|
2d5d4980de |
fix(diagnostics): address round-2 review findings on PR #445
- settings.go: cap the parsed cleanup interval at 7 days before converting to time.Duration so a huge configured value can't overflow int64 nanoseconds and wrap into a tiny/negative interval; add boundary tests. - settings.go: propagate genuine settings read failures from LoadSettings (missing/empty -> default, error -> fail) so a transient DB error surfaces retryably instead of silently reporting uploads disabled or wrong quotas. - bundle.go: validate non-manifest bundle entries while streaming with bounded memory -- device.json and crash/*.json must be a single JSON object, logs.jsonl/breadcrumbs.jsonl must be newline-delimited JSON objects with a per-line byte cap (new contract.MaxLogLineBytes); binary members stay opaque. - diagnostics upload handler: extend the read deadline per-route via http.ResponseController.SetReadDeadline (10m) so slow mobile uploads of large bundles aren't cut off by the shared 30s server ReadTimeout. - web admin download: request the ?proxy=1 streaming path directly so downloads work when S3Private is only server-reachable and errors can surface in-page. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012e3QjbPo96ed9Mn2qRiUkh |
||
|
|
75c2260897 | Merge remote-tracking branch 'origin/main' into feat/client-diagnostics-server | ||
|
|
a0851edef0 | fix(watchsync): align MDBList API contracts | ||
|
|
4f249fda8f | fix(watchsync): repair MDBList scrobble lifecycle | ||
|
|
845b96e703 |
fix(playback): preserve remux copy on seek (#422)
* fix(playback): preserve remux copy on seek * fix(playback): harden remux replacement transactions --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> |
||
|
|
e56a1b3e03 |
fix(api): throttle api_keys last_used_at writes in auth middleware (#381)
* fix(api): throttle api_keys last_used_at writes in auth middleware Every API key request spawned a goroutine that ran an UPDATE on api_keys, so a key driving HLS segments or a polling integration hit the table with one write per request, and a stalled database could pile those goroutines up without bound. The jellycompat authenticator already guards this same write with a once-per-minute throttle per key; the main middleware was missing it. Bring the two in line. Track the last write per key ID and only launch the update once a minute has passed, with a timeout on the background write. The map is keyed by key ID so it stays bounded. * fix(auth): bound API key last-used throttling --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> |
||
|
|
4fa84a661a |
feat(diagnostics): client diagnostics server foundation
Implements slice 1 of docs/design/2026-07-19-client-diagnostics.md: the versioned contract (schemas, fixtures, Go validator), storage-validated diagnostics.uploads_enabled gate, account-scoped status endpoint, hardened streaming multipart ingest with quota reservation and a receiving/ready/ failed report state machine, S3 streaming puts, acting-admin report API (list/detail/download/delete with audit events), and the retention + orphan-reconciliation cleanup task. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XppCCycoaskCsW7ja1fZct |
||
|
|
8044eb84dd |
feat(activity): refine play-method tags and add a Jellyfin-client pill (#387)
* feat(activity): refine play-method tags and add a Jellyfin-client pill
Two related tagging improvements to the admin activity views, squashed:
Split audio transcodes into their own tag. The Play Method summary and
Server Activity popover bucketed every session by its raw play_method,
lumping real video transcodes together with video-copy HLS repackages and
having no separate tag for audio-only transcodes. Classify each session by
the per-stream decisions the backend already reports:
- video re-encoded -> "transcode" (yellow)
- only audio re-encoded -> "audio" (red)
- streams only repackaged -> "remux" (blue, incl. video-copy HLS)
- nothing touched -> "direct" (green)
ordered direct -> remux -> transcode -> audio across the distribution bar,
legend, method filter/sort, the per-row badge, and the Server Activity
stream counts.
Add a Jellyfin-client "JF" pill. Sessions from a Jellyfin-ecosystem client
(Jellyfin Web, Findroid, Swiftfin, Infuse, etc.) get a purple "JF" pill
next to the play-method tag. Detection is UI-only: isJellyfinSession()
positively matches client_name (set from the Jellyfin MediaBrowser auth
header) and then the raw user agent against the known Jellyfin client
tokens, mirroring the server's client-labeling list. The pill is orthogonal
to the method classification — a session can be both "transcode" and JF.
Pure UI/presentation change; no backend behavior changes.
AI-use disclosure: implemented with AI assistance (Claude Code).
* fix(web): cache-control on SPA shell so deploys bust stale UI
The frontend handler served index.html with no cache directives, leaving
freshness to browser/CDN heuristics. A stale index.html at a CDN edge kept
serving old content-hashed bundles, so a client-side hard refresh couldn't
recover — one browser would show the new UI while another showed the old.
Apply the standard SPA cache policy:
- index.html (and SPA-route fallbacks): no-cache + a truncated-SHA-256
ETag, so the shell is cached but revalidated on every load and answers
an unchanged request with a cheap 304.
- /assets/* (Vite content-hashed bundles): public, max-age=31536000,
immutable — cached indefinitely; a new build changes the filename hash,
which busts them automatically.
- other stable-named bundled files (sw.js, icons, fonts): no-cache, so a
changed service worker or icon can't stay stuck in a cache.
Caching is preserved (no no-store anywhere); only the tiny HTML shell is
revalidated, which is what busts a stale UI on deploy.
* fix(activity): compute the method bucket server-side and unify every session surface
Review follow-ups for the play-method tags (PR #387):
- The server now emits effective_play_method (additive field) from the same
per-stream decisions that drive the badges, so all consumers — web, realtime
popover, and the Android/Apple admin views later — agree on the bucket
instead of each client re-reducing raw play_method. Rows with an unknown
play_method (stale rows from older nodes) stay unbucketed rather than being
misreported as audio transcodes off the bare transcode_audio flag; the web
fallback classifier mirrors that and reports "unknown".
- Jellyfin-ecosystem detection moved server-side as is_jellyfin_client, owned
next to the client-labeling rules so the two lists cannot drift; the web
token list is gone. Adds kodi/mpv/delfin/finamp, which reach Silo only
through the Jellyfin compat surface.
- The dashboard stream cards, stats session table, and household streams panel
now use the same classification as the activity page and popover — they
previously showed contradictory tags for the same live session.
- One shared method->label/color table in adminActivityPresentation.ts
replaces the four independent copies (METHOD_META + three switches); the
method column sort now uses the shared cost-order comparator instead of
alphabetical; dead "copy"/"hls" order entries removed and the reachable
"unknown" bucket is styled.
* fix(server): make SPA revalidation RFC-compliant and stop rebuilding the shell per request
Review follow-ups for the SPA cache policy (PR #387):
- Stable-URL bundled files (sw.js, icons, vendor bundles) now carry a content
ETag. The embedded FS has no modtimes, so http.FileServer emits no validator
of its own — no-cache alone forced a full re-download of multi-megabyte
vendor trees on every use because there was nothing to revalidate against.
- Shell and favicon conditional requests go through http.ServeContent, which
implements RFC 9110 If-None-Match semantics (weak comparison, ETag lists).
The previous exact string compare never matched once a fronting proxy
compressed the response and weakened the ETag to W/"...", silently killing
the 304 path in the most common deployment topology.
- The rendered shell (index read + branding render + SHA-256) is cached per
branding snapshot via the new Snapshot.RenderKey instead of being rebuilt on
every request — the 304 revalidation that no-cache makes the common case now
costs two header writes. The misnamed weakContentETag (it emits a strong
validator) is renamed contentETag.
* fix(activity): show the JF pill on every session surface, not just the mobile row
Review comments on PR #387: the JF pill only rendered inside Admin
Activity's sm:hidden mobile row, so the desktop table — and the other
session surfaces that now share the method classification — never
identified Jellyfin-compat sessions.
Extract the pill into a shared JellyfinSessionPill component (renders
nothing for native sessions) and drop it into the Admin Activity desktop
client line, the dashboard stream cards, the household streams panel,
and the stats active-session table.
* fix(playback): sync real encode decisions and client identity for compat transcodes
Review comments on PR #387:
- Jellyfin HLS sessions that copy video and re-encode only audio synced as
full video transcodes: ensureUpstreamPlayback resets transcodeAudio for the
transcode transport method, and the TargetCodecVideo "copy" decision lived
only in TranscodeOpts. A new SessionManager.SetTranscodeStreamDetails
mirrors the actual decisions onto the upstream session when the transcode
starts (local and remote-node paths, via an optional interface so test
fakes are unaffected), so these sessions now bucket as "audio"/"remux".
- Transcode recipe cards now record TranscodeAudio derived from the opts
(only an explicit "copy" leaves audio untouched — empty runs ffmpeg's aac
default), so a session rebuilt after a restart keeps the same bucket.
- Recipe cards carry client name/version/user-agent, and reconstruction
restores them, so the admin client label and the JF pill survive server
restarts; the compat fallback card populates them from the live
MediaBrowser request. Deliberately not projected into stream-token claims,
where a user agent would bloat every stream URL.
* feat(api): capability endpoint for the live-session activity fields
Review comment on PR #387: effective_play_method and is_jellyfin_client are
omitempty, so an independently deployed client cannot distinguish an older
server from a supported one reporting an unknown method or a non-Jellyfin
session. GET /admin/sessions/capabilities advertises both fields plus the
closed bucket vocabulary, following the additive capability-endpoint rule
(same pattern as /collections/capabilities).
* fix(playback): treat empty target audio codec as an AAC re-encode in live state
ffmpeg defaults an empty target audio codec to AAC (appendAudioArgs), and the
new recipe logic already records that as an audio transcode — but the live
native path computed transcodeAudio=false for an empty codec, so the running
stream reported remux until a restart flipped it to audio. Extract the
predicate into playback.TranscodesAudio, share it across the live path, the
recipe card, and the compat mirror, and make appendAudioArgs case-insensitive
so the ffmpeg switch agrees with the predicate for any spelling.
Part of #387 review follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(jellycompat): re-sync sessions after recording compat encode decisions
ensureUpstreamPlayback flushes the session (compat_start) before
ensureTranscodeSession / startRemoteTranscode record the actual codec
decisions, and that later mutation triggered no sync — so the admin view
showed a video-copy stream as a full video transcode until the periodic
reconciler ran. Trigger syncSessionsNow after the details are recorded
successfully; the helper is shared, so both the local and remote-node
paths are covered.
Part of #387 review follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
91e1164090 |
feat(metadata): local NFO metadata and sidecar artwork (builtin chain provider) (#390)
* feat(metadata): register builtin NFO provider and broaden parsing Phases A and B of the #216 local-NFO work, implemented test-first. Registration & hint-first identity (Phase A): - Migration seeds a reserved kind='builtin' silo.builtin installation and an 'nfo' metadata capability (default_enabled=false, priority 1 for movie/series) with a partial unique index and documented Down. - In-process builtin provider registry (internal/metadata/builtin.go); buildProviders returns the registered provider for builtin rows. - Guard rails keep the reserved row out of every plugin surface (user plugin-settings, installations list, image resolvers, preload, auto-update, store Delete, mutation handlers -> 409); silo.builtin is a reserved manifest id. - Startup sync materializes legacy content_level='' chains per level, then appends builtin capabilities disabled via AppendProviderToAllChains (idempotent); resolveEnabledProvidersBy priority now respects default_enabled=false. - NFO uniqueids seed the trusted-hint machinery via IdentityHintProvider with per-mode conflict policy (stored IDs win on scheduled refresh, NFO wins on manual refresh, Identify skips NFO); ID-less candidates are excluded from provider-priority tie-breaks and nfo never counts as corroboration. - Web chain-editor empty-state gate is now server-derived so builtin providers are reachable on plugin-less servers. Parser breadth & sidecar hardening (Phase B): - Parser covers the practical Kodi/Jellyfin field set for <movie> and <tvshow>: original title, tagline, runtime, dates, content rating, genres/studios/countries/tags, multi-source ratings with scale normalization, cast with roles/order, director/credits. Empty collections stay nil so merge early-returns apply. - findNFO parses candidates and falls through on read/parse failure or root-type mismatch, so a stray movie.nfo cannot shadow tvshow.nfo; GetMetadata gains the same ContentType guard Search has. - New FieldReleaseDates lock gates Year/ReleaseDate/First+LastAirDate in merge (Go) and the edit-metadata dialog (web), closing the gap where a manual refresh re-applied NFO dates over admin corrections. - Merge-contract tests pin NFO fill semantics, genres whole-list first-provider-wins, and NFO edits propagating on manual refresh only. - Docs: new admin wiki page (supported fields, merge semantics, naming-supplies-structure contract), index bullet, sidecar wording revision, v1-scope feature-detection note. Zero behavior change while the provider is disabled (default); pinned by CI-mode and DB-gated test suites. Part of #216 AI-use disclosure: implemented with Claude Code (Fable 5) via spec-driven TDD and agent-assisted implementation. * feat(metadata): ingest local sidecar artwork and read series-depth NFO Phases C and D of the #216 local-NFO work, implemented test-first, plus the mixed-library use-case pins. Together these deliver the headline case: a series absent from every remote database (e.g. a fitness library) scans into a fully presented show -> named seasons -> titled episodes tree from NFO files and sidecar art alone. Local sidecar artwork through the S3 image cache (Phase C): - The NFO provider implements ImageProvider: poster/backdrop/logo sidecar discovery with a fixed precedence map, symlink/non-regular rejection, an 8 MiB cap, and file:// source URLs at rating 0. Generic filenames apply only via the sidecar search paths, so a shared folder.jpg in a flat multi-movie directory applies to none. - file:// becomes a live local source scheme: routed into *_source_path (never *_path), accepted by every image enqueue gate, attributed as provider "local", excluded from cached-path detection. - The image-cache processor caches local files with lexical-on-logical confinement to the library roots, open-handle reads with re-checks, the same variant widths as remote art, and stable (7-day) failure classification. Keys land under local/{contentType}/{contentID}/{hash8}/{imageType}; superseded prefixes are cleaned on re-cache and item deletion. - applyIfBetter gains a local exemption so rating-0 local art can fill matched items without being stickily displaced; ImageRequest carries additive sidecar path context. Series depth (Phase D): - SeasonsRequest/EpisodesRequest carry additive local path context (series roots, per-season directories, per-episode file paths), derived from naming at match time and reconstructed on refresh. - season.nfo supplies season name/plot; NFO season numbers are advisory (directory-derived number wins with a Warn - naming owns structure). <episodedetails> gains aired/runtime/ratings; <basename>.nfo titles episodes and <basename>-thumb.ext supplies thumbs; filename SxxEyy wins over NFO numbers. - Episode NFOs work without a season.nfo (provider seasons unioned with on-disk seasons); SynthesizeFallbackEpisodes always runs after persist so NFO-less episodes keep synthesized rows. Season/episode file:// art rides the Phase C pipeline unchanged. - Migration adds season:1/episode:1 to the builtin NFO capability's default_priority (still default_enabled=false). Mixed sports-library use case (tests only, no product change): - Pins the classification contract for one library holding movie-shaped and show-shaped content (WWE PPV events as movies next to a "WWE SmackDown" show, NASCAR/F1/FIFA with partial TVDB/TMDB data): naming decides movie-vs-series per file before any provider runs; the NFO supplies metadata/identity but never flips type (ContentType guard); the per-root Type override is the correction path. - NFO-driven type classification at scan time is recorded as an explicit deferred open question. Part of #216 AI-use disclosure: implemented with Claude Code (Fable 5) via spec-driven TDD and agent-assisted implementation. * docs(metadata): document local NFO metadata architecture Add a single as-built architecture page (docs/architecture/local-nfo-metadata.md) for the #216 local-NFO feature: the builtin registration model, hint-first identity semantics, the file:// -> S3 artwork pipeline and its deployment constraint, series depth, the mixed-library classification contract, and known limitations. This replaces the working implementation plan, the per-phase specs, and the narrow sidecar-artwork note, which were planning drafts and are left untracked; admin-facing behavior remains in the wiki. Part of #216 AI-use disclosure: planned, drafted, and consolidated with Claude Code (Fable 5) using multi-agent exploration and adversarial review. * fix(metadata): address PR review findings on NFO builtin provider Fold in the valid, low-risk fixes surfaced by automated review on #390: - imagecache: extract validateCacheRequest so CacheBytes (the local sidecar season/episode path) enforces the same episode-requires-season guard as Cache, preventing distinct episodes' art from colliding under one S3 key. - image_cache_processor: close the sidecar symlink-swap window by rejecting the opened handle unless os.SameFile matches the Lstat'd file, so a leaf swapped to a symlink can't pull an out-of-root target into the public cache. - plugins: guard the reserved builtin installation row in the store's Update, matching Delete, so its version/enabled/capabilities can never be rewritten even if a mutation slips past the HTTP layer. - cmd/silo: bound SyncBuiltinProviderChains with a 30s timeout so a stuck DB round-trip fails fast at startup instead of hanging. - metadata: panic instead of silently no-op'ing on an invalid RegisterBuiltinProvider call (init-time programmer error). - docs: correct the media-folder-and-naming NFO paragraph to state season/episode NFOs and sidecar artwork are actively read. --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> |
||
|
|
1664c60425 |
fix(metadata): publish artwork revisions atomically (#399)
* fix(metadata): publish artwork revisions atomically * fix(metadata): harden artwork revision cleanup * fix(metadata): address artwork revision review findings - restore image applies for all media_items types and reject unsupported target/image combinations with 400 before uploading; episodes coerce to stills and the web dialog no longer offers image tabs episodes can't use - add WHEN clauses to displacement triggers and hoist to_jsonb so bulk catalog upserts that assign unchanged artwork columns skip the trigger - make artworkkey the single variant-ladder owner: imagecache derives its widths from it and triggers store image_type instead of hardcoded variant arrays, expanded by the collector at deletion time - sweep dormant registry rows periodically so references lost through untriggered surfaces degrade to slow cleanup instead of leaking - park just-published revisions dormant, keep dormant rows dormant on re-cache, and batch the GC reference pre-check per run - heal rows re-referencing a just-deleted revision via reconciler-style resets after the deletion commits - share a per-URL image-loaded hook across DetailHero, ItemCard, SectionItemCard, GlobalSearch, and CollectionPosterCard - deduplicate Cache/CacheBytes finalization and drop unused VariantPaths plumbing Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(catalog): cast reused timestamp parameter in revision upsert Postgres cannot deduce one type for $3 used both as a plain value and inside a CASE arm; the dev deploy surfaced it as SQLSTATE 42P08 on every publication. Cast both uses and cover the arm/park/track upserts with database-backed tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(metadata): address artwork revision review comments - keep a durable heal path: deletion marks deleted_at instead of removing the registry row, so a failed post-delete heal retries with backoff and broken references never park; trackers clear the marker on re-upload - never treat bare existence as an immutable-content match; backends without content verification rewrite the object - exercise revisioned cover keys in scanner/enrichment fakes, compare the tracked manifest exactly, and honor cancellation in the blocking test deleter Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b3722dac58 |
fix(playback): open the listener before sweeping stale transcode dirs (#413)
* fix(playback): run orphaned-transcode cleanup in the background at startup The native and Jellyfin-compat routers swept stale per-session transcode dirs synchronously during NewRouter, before the listener bound. On a slow network filesystem this blocked startup for 80+s (64 leftover dirs on the last deploy), so restart-reconnect clients were turned away and the health check reported the server unhealthy the whole time. Move both sweeps into a background goroutine (StartBackgroundOrphanCleanup) so the listener comes up immediately and the cleanup runs concurrently. The delete logic is unchanged: same active-session snapshot and MaxTokenTTL age-sparing, only later. A package-level mutex serializes concurrent sweeps of the shared transcode root so the two background sweeps can't race on os.RemoveAll. Part of #412 * fix(transcode): background the node boot-time transcode-dir sweep A dedicated transcode node swept leftover transcode dirs synchronously in NewServer, before startStandaloneServer bound its listener. On a slow network filesystem that delete blocked the node from coming online at boot, the same startup-stall class as the main server. Move the sweep into the shared StartBackgroundOrphanCleanup goroutine so the node's listener binds immediately. Backgrounding required an age guard: the sweep previously ran as a full wipe (minAge=0) with an empty active-set, which was only safe because it completed before any request could arrive. Run concurrently that would race a token-carried reconstruct writing into TranscodeDir/<sessionID>, deleting segments a fresh ffmpeg is producing. Passing MaxTokenTTL spares any dir younger than the max token lifetime — exactly the ones a still-valid reconnect could reconstruct — while dirs older than any surviving token (never reconstructable) are still reclaimed. Part of #412 * feat(playback): reclaim orphaned transcode dirs periodically, not just at boot The orphaned-transcode sweep only ran at startup on both the central server and transcode nodes, so it only ever reclaimed dirs left by an ungraceful prior shutdown. During a long uptime the in-memory session reapers delete the dirs of sessions they still track, but a dir whose owning session was dropped without its RemoveAll succeeding becomes an "untracked orphan" with no runtime GC — on a box that runs for weeks these accumulate until the next restart. Add StartPeriodicOrphanCleanup: an immediate background sweep followed by an hourly re-run bound to a lifecycle context. Wire it on all three surfaces — native API and Jellyfin-compat (via deps.AppContext) and the transcode node (via a new Server.StartOrphanSweeper(appCtx), replacing its boot-only sweep). When no context is supplied (tests) it degrades to a single boot-time sweep so no ticker goroutine outlives the caller. The sweep stays age-guarded at MaxTokenTTL, so nothing reconstructable is ever reaped. Because the node sweep now runs during live traffic, it snapshots the live job set (Server.activeSessionIDs) and spares those dirs by id rather than by age alone — a long-lived session that only re-serves already-written segments stops advancing its dir mtime, which age could otherwise misclassify. In integrated mode the native and compat sweeps share one TranscodeDir but each snapshots only its own manager's live set; the resulting cross-manager reap of a >24h idle dir is bounded (rebuilds from token/recipe) and documented at both call sites. Part of #412 |
||
|
|
8fc054c15d |
fix(scanner): never purge files under unreachable library roots (#372)
* fix(scanner): never purge files under unreachable library roots An unreachable root is not a removed root. When one root of a multi-root library dies (unmounted share, dead drive) while another root still has files, the whole-library empty-root guard does not fire — the surviving root produced files — so the scan marks everything under the dead root missing_since (desired: hides it from browse/playback) and then, with the default scanner.empty_trash_after_scan=true + 24h file_removal_grace, the next scan after the grace hard-deletes every row under the dead root. A week-long drive outage silently destroys the root's entire catalog state: probe data, intro/credits markers, file hashes. Worse, membership reconciliation immediately purges media_items whose only files lived on the dead root, cascading user collections (library_collection_items has ON DELETE CASCADE) and deleting cached artwork. This change makes "temporarily offline" survivable: - Probe each configured root at scan start (os.Stat + IsDir + ReadDir, factored into the new internal/rootcheck package and shared with the admin mount-check endpoint). Unreachable roots are skipped by the walk but their scopes still reconcile, so files are still marked missing. - The trash sweep (DeleteMissingByFolder) now excludes rows whose path sits under an unreachable root, using the same exact-path + escaped prefix-LIKE matching as ListIDsOutsideRoots (a sibling root that merely shares a string prefix is never protected). With all roots reachable the emitted SQL is unchanged. - Membership removal still happens — browse/home hide items via media_item_libraries, so removal is what keeps a dead-root-only title out of the catalog — but the orphan media_items purge exempts items whose files sit under an unreachable root. Their metadata, artwork, and collection links survive; when the root returns, the upsert clears missing_since and syncPresentLibraryState re-inserts the membership, restoring the item with zero re-probing or re-matching. - The folder surfaces scan_warning_code='dead_root' with a message naming the unreachable roots; a fully healthy scan or a successful mount check clears it, mirroring empty_root. The admin UI shows a badge and banner. - Deliberate deletion is untouched: removing a path from the library config still purges via ListIDsOutsideRoots, files under reachable roots keep the exact 24h-grace purge, the empty-root guard and the autoscan dead-mount guard are unchanged. The audiobook/podcast/ebook reconcile paths share the same folder-wide sweep and orphan purge, so they get the same guard. Covered by tests: an end-to-end two-root scan (root dies -> rows survive a zero-grace sweep and warning is set; root returns -> rows resurrect with their original ids and the warning clears; deleting a file under a reachable root still purges), repo-level sweep-protection and sibling-prefix tests, orphan-purge exemption, and rootcheck unit tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(scanner): probe uncompacted roots and take dead-root path on full outage Review follow-ups: (1) probe every configured path instead of the compacted traversal roots, so a nested child mount that dies under a reachable parent is still protected from the sweep; (2) when every configured root is unreachable, bypass the empty-root confirm flow (without consuming the one-time cleanup allowance), mark files missing, and raise dead_root instead of empty_root; (3) dead_root warning banner no longer shows empty-root confirm-deletion guidance as its fallback hint. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(scanner): simplify dead-root protection plumbing - extract pathscope.CoverageClauses as the single builder for the exact-path + escaped prefix-LIKE root predicate; scanner's rootCoverageClauses delegates to it and catalog's excludeOrphansUnderProtectedPrefixes reuses it instead of hand-rolling the same clause loop - extract Scanner.sweepMissingAndReconcile to replace the identical trash-sweep + membership-reconcile + S3-image-cleanup block that was triplicated across the audiobook, ebook, and podcast scans (callers keep their flavor-specific log lines so messages stay constant) - add unreachableConfiguredRoots helper for the repeated probeUnreachableRoots(ctx, folder.ID, cleanScanRoots(folder.Paths)) expression in scanPaths and ScanFile - drop the unread Path field from rootcheck.Result - move the dead/empty-root warning text constants in AdminLibraries.tsx out of the middle of the import block Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scanner): close dead-root protection gaps found in review Remediates the confirmed findings from the deep review of this PR: - Scoped audiobook scans (autoscan file events, subtree scans) ran the folder-wide sweep while probing only the scoped clone's Paths, so a healthy-subtree event could hard-delete a dead sibling root's rows. sweepMissingAndReconcile now reloads the folder's configured roots from the DB and probes them uncompacted, which also protects nested child mounts in the audiobook/ebook/podcast reconcilers. - A lost mount that leaves an empty, stat-able mountpoint probed as reachable and kept the historical purge timeline. A reachable root that is a literally empty directory while cataloged rows remain under it is now treated as suspect: rows are only marked missing, the sweep and orphan purge exempt it, dead_root is raised, and the mount-check endpoint reports it (additive suspect_empty field) instead of clearing the warning. Arming the one-time empty-cleanup allowance completes the deletion, including in the mixed case where other roots are healthy. Roots that still have directory entries keep the historical grace-then-purge path. - Confirmed empty cleanup (allow_empty_cleanup_once) no longer force-deletes rows under probe-dead roots: an outage is not a confirmation, so a dead sibling root's catalog survives a confirmed cleanout of a reachable empty root. - Root probes are now bounded (rootcheck.ProbeWithTimeout, 5s): a hung network mount degrades into the protected unreachable path with a probe_timeout error code instead of stalling every scan of the folder indefinitely. - Documented the cross-library limitation of the orphan-purge exemption next to the query it applies to. All behavior is pinned by new DB-backed tests (suspect-empty protection + confirmed completion, confirmed-cleanup dead-root survival, scoped/nested-root sweep protection, suspect-root query, probe timeout). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scanner): address dead-root review findings --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> |
||
|
|
1f2125b920 |
feat(plugins): group Apps sidebar by plugin manifest category (#366)
* feat(plugins): group Apps sidebar by plugin manifest category Implements the plugin SDK's documented PluginManifest.category semantics (silo-plugin-sdk proto/silo/plugin/v1/common.proto): a slash-delimited path that groups plugins in the user-facing Apps section, e.g. "Books/Audiobooks" lands under Apps -> Books. The field existed in the manifest proto but silo-server never surfaced it. Server: the user plugin-settings list/detail responses now include an additive-only `category,omitempty` string sourced from the already-loaded manifest via GetCategory(); no new parsing paths. Web: PluginSettingsSummary gains `category?: string`, and AppSidebar groups Apps entries by the FIRST segment of the category path (one level of grouping for now; deeper segments intentionally ignored, documented against the SDK contract). When fewer than 2 distinct categories exist among the visible app links, today's flat list under the single "Apps" header is kept; with 2+ categories, per-category sub-headers render via the existing SidebarSectionHeader (labels hide in the collapsed sidebar the same way other section headers do). Uncategorized plugins fall under "Other", which always sorts last. Tests: Go unit tests for the summary converter (category passthrough and JSON omission when empty) and vitest coverage for the pure groupAppNavLinks helper plus grouped/flat/collapsed sidebar rendering. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(plugins): use generic category examples in comments and tests Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(web): simplify Apps sidebar link list rendering - fold the duplicated <ul> list markup in the grouped and flat Apps branches into a single renderAppNavList helper so the list styling cannot drift between the two render paths Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com> |