Files
silo-server/cmd
CoffeeKnyteandGitHub b3722dac58 fix(playback): open the listener before sweeping stale transcode dirs (#413)
* fix(playback): run orphaned-transcode cleanup in the background at startup

The native and Jellyfin-compat routers swept stale per-session transcode
dirs synchronously during NewRouter, before the listener bound. On a slow
network filesystem this blocked startup for 80+s (64 leftover dirs on the
last deploy), so restart-reconnect clients were turned away and the health
check reported the server unhealthy the whole time.

Move both sweeps into a background goroutine (StartBackgroundOrphanCleanup)
so the listener comes up immediately and the cleanup runs concurrently. The
delete logic is unchanged: same active-session snapshot and MaxTokenTTL
age-sparing, only later. A package-level mutex serializes concurrent sweeps
of the shared transcode root so the two background sweeps can't race on
os.RemoveAll.

Part of #412

* fix(transcode): background the node boot-time transcode-dir sweep

A dedicated transcode node swept leftover transcode dirs synchronously in
NewServer, before startStandaloneServer bound its listener. On a slow
network filesystem that delete blocked the node from coming online at boot,
the same startup-stall class as the main server.

Move the sweep into the shared StartBackgroundOrphanCleanup goroutine so the
node's listener binds immediately. Backgrounding required an age guard: the
sweep previously ran as a full wipe (minAge=0) with an empty active-set,
which was only safe because it completed before any request could arrive.
Run concurrently that would race a token-carried reconstruct writing into
TranscodeDir/<sessionID>, deleting segments a fresh ffmpeg is producing.
Passing MaxTokenTTL spares any dir younger than the max token lifetime —
exactly the ones a still-valid reconnect could reconstruct — while dirs
older than any surviving token (never reconstructable) are still reclaimed.

Part of #412

* feat(playback): reclaim orphaned transcode dirs periodically, not just at boot

The orphaned-transcode sweep only ran at startup on both the central server
and transcode nodes, so it only ever reclaimed dirs left by an ungraceful
prior shutdown. During a long uptime the in-memory session reapers delete the
dirs of sessions they still track, but a dir whose owning session was dropped
without its RemoveAll succeeding becomes an "untracked orphan" with no runtime
GC — on a box that runs for weeks these accumulate until the next restart.

Add StartPeriodicOrphanCleanup: an immediate background sweep followed by an
hourly re-run bound to a lifecycle context. Wire it on all three surfaces —
native API and Jellyfin-compat (via deps.AppContext) and the transcode node
(via a new Server.StartOrphanSweeper(appCtx), replacing its boot-only sweep).
When no context is supplied (tests) it degrades to a single boot-time sweep so
no ticker goroutine outlives the caller. The sweep stays age-guarded at
MaxTokenTTL, so nothing reconstructable is ever reaped.

Because the node sweep now runs during live traffic, it snapshots the live
job set (Server.activeSessionIDs) and spares those dirs by id rather than by
age alone — a long-lived session that only re-serves already-written segments
stops advancing its dir mtime, which age could otherwise misclassify. In
integrated mode the native and compat sweeps share one TranscodeDir but each
snapshots only its own manager's live set; the resulting cross-manager reap of
a >24h idle dir is bounded (rebuilds from token/recipe) and documented at both
call sites.

Part of #412
2026-07-16 14:48:00 -04:00
..