Files
silo-server/internal/worker/cleanup.go
T
9b111649f3 perf+fix(audiobooks): detail page, browse, sessions & scanner (#169)
* perf(catalog): fix audiobook detail N+1 + slow people facets

Audiobook detail pages were slow in proportion to track count (up to 433
files/book). Root causes, found by EXPLAIN ANALYZE on the live DB:

- effectiveAudioSelection ran 3-4 user-store queries (profile, audio pref,
  library pref) per file inside buildPlaybackInfo's loop, though the results
  are invariant across a request. Introduce a request-scoped audioPrefResolver
  that memoizes the store lookups (library prefs keyed by folder); a 400-file
  audiobook now issues each query once instead of per file. Selection logic is
  unchanged (audioPreference returns a copy so the original-language sentinel
  is still resolved per file).
- buildAudiobookExtension ran its four independent related-content queries
  serially; run them concurrently so latency is the slowest, not the sum.
- author/narrator browse facets did a full people-table scan; add a
  (kind, content_id, person_id) index so the facet resolves from an index-only
  scan of just that kind's credits (~112ms -> ~49ms on the live library).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(catalog): cache audiobook author/narrator group browse

The Authors/Narrators audiobook pages were slow on cold load and slow again
after a hard refresh (fast only while the React Query client cache was warm).

Root cause (EXPLAIN ANALYZE on live, 31K-audiobook library): the grouped
browse query is ~234ms/page, there are ~13K distinct authors, and the client
pages through the entire list on every load (sequential 500-row requests). With
no server-side cache each of the ~20 pages re-ran the full aggregation
(COUNT(*) OVER() forces it), so a cold load was ~20x234ms. The client's 60s
staleTime was the only thing making a warm revisit fast; a refresh wiped it.

Fix: AudiobookGroupsCache caches the full sorted group list per (library,
group_by, sort, viewer) for 60s (matching the client staleTime, so no extra
staleness) and serves every page as an in-memory slice — one aggregation per
window instead of one per page, and a refresh is a cache hit. Also raise the
client page size 500->2000 so fewer sequential round-trips are needed now that
a larger page is a cheap slice.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(settings): throttle per-request device last_seen upserts

Device-setting reads (HandleGetDeviceSetting, HandleGetEffectiveSettings,
HandleGetEffectiveSubtitleAppearance) each registered the request's device — an
INSERT ... ON CONFLICT upsert of last_seen_at on a single per-device row. A page
that fetches many settings fired hundreds of these concurrently; they serialized
on that row's lock (observed 100-237ms each, ~250 per page load in the slow
query log), taxing every settings fetch.

Throttle device registration to one upsert per (profile, device) per 5 minutes
via an in-process TTL cache, marking the device seen before the upsert so a
concurrent burst collapses to a single write. last_seen_at stays fresh to within
the window. Reads no longer issue a contended write on the hot path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf+fix(audiobooks): probe-repair, resume position, cache storm, groups reveal, hot-row + stats index

From the full audiobook code review (EXPLAIN + slow-query trace on live):

- #1 (P0, detail-page killer): NeedsCriticalProbeRepair required video codec/
  resolution/tracks, which audio-only files never have, so PlaybackProbeEnsurer
  re-ran ffprobe per file on every detail/watch load (up to N serial spawns for
  an N-track book) and never converged. Gate video-field checks on the file
  actually having a video stream. TDD.
- #3 (P0): abs session-sync rewound the resume cursor — UpdateProgressPosition
  did an unconditional SET with no monotonic guard, ignored its error, and
  no-op'd when no row existed (first-listen resume lost). Now a finish-preserving
  GREATEST upsert; caller logs failures.
- #4 (P0 perf): progress reports fired every ~10s invalidated all of
  catalogKeys.all → refetched every active browse/detail query incl the 13k
  audiobook group lists. Scope invalidation to the reported item's detail.
- #6 (P1 perf): Authors/Narrators page rendered all ~13k groups + cover images
  at once (main-thread freeze). Incremental reveal: render a capped window, grow
  on scroll via IntersectionObserver.
- #10: throttle abs TouchToken last_seen upsert (one per token per 5min) — same
  hot-row contention class as the device fix.
- #8: index abs_playback_sessions (user_id, profile_id, started_at) for the
  listening-stats aggregations.

Deferred (need contract/validation): listening-time idempotency (client delta-vs-
cumulative), scanner deleted-file reconcile, abs session retention job, abs list-
handler batch fetch, scanner-output P2s (need re-backfill).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(audiobooks): batch-fetch abs list/shelf handlers (kill N+1)

handleSimilarItems, handleItemsInProgress, and handleGetMyProgress called
MediaStore.GetAudiobookByID once per row — up to ~500 single fetches (each a
few queries) on app open. Add GetAudiobooksByIDs (one access-scoped fetch +
people/series hydrated once for the whole set) and look results up from the
returned map, preserving order. Underlying primitives (GetByIDsWithAccess,
hydratePeople/Series) were already batch-capable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(audiobooks): reconcile deleted files on scan + prune session history

#5: ScanAudiobookFolder only ever upserted — deleted/renamed books leaked
media_items/media_files/memberships forever. Mirror the ebook reconcile: collect
seenPaths during the walk, MarkMissing files no longer on disk, then
reconcileLibraryMemberships. Safety mirrors ebooks/video: an inaccessible root
(unmounted source) is skipped entirely, and a walk that saw zero files while the
DB has rows only reconciles after operator cleanup confirmation
(ebookEmptyCleanupAllowed) — so a flapping mount can't wipe the catalog. Soft
mark only; the existing grace-period purge hard-deletes later. Reconcile runs
only on a fully-completed (non-cancelled) scan. (#9 coarse case already handled:
audiobookFolderShouldSkip skips unchanged folders; per-file reuse deferred.)

#8-retention: abs_playback_sessions grew unbounded (one row per play-start, never
deleted) and fed every listening-stats scan. Add an hourly sweep in
SessionCleaner: close abandoned open sessions (no /close, stopped syncing >24h)
and delete closed sessions older than 90 days. Mirrors the recommendation_cache /
missing-files prune pattern.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(audiobooks): address max-effort code-review findings

From /code-review max on the pre-PR diff:
- DATA RACE (P0): SessionCleaner.lastABSSessionPrune is read+written by both the
  15s ticker goroutine and the shutdown-path CleanStale call (main.go defers
  Stop() to after that call). Guard the prune-due gate with a mutex. (CleanStale
  was stateless before this branch, so concurrent calls were previously safe.)
- ScanAudiobookFolder hardcoded fullScan=true into the empty-walk cleanup guard,
  but it's also called from ScanSubtree (incremental scans). An empty subtree
  scan would wrongly consume the operator's one-shot empty-cleanup allowance and
  warn. Thread a real fullScan flag (true from ScanFolder, false from the two
  subtree call sites), mirroring the ebook path.
- Revert UpdateProgressPosition to UPDATE-only (drop the INSERT-on-missing):
  keep the monotonic GREATEST + finish guard that fixes the resume rewind, but
  restore the no-op-on-missing contract so a stray sync tick can't resurrect
  just-cleared progress or create a zero-duration continue-listening row.
- Clamp the audiobook-groups handler limit (paging moved into the cache, leaving
  the old 500/page bound stranded); also gofmt the Scanner struct.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(audiobooks): address review feedback for scanner and stats

* fix(audiobooks): address review feedback

* fix(audiobooks): retry failed session prune

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-06-17 09:51:43 -04:00

209 lines
6.9 KiB
Go

package worker
import (
"context"
"fmt"
"log/slog"
"sync"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/Silo-Server/silo-server/internal/cache"
evt "github.com/Silo-Server/silo-server/internal/events"
)
const (
// nodeDeadTimeout is how long a node can go without a heartbeat before
// its sessions are purged.
nodeDeadTimeout = 45 * time.Second
// nodeHeartbeatCleanup is how long before stale heartbeat rows
// themselves are deleted (longer than nodeDeadTimeout to avoid flapping).
nodeHeartbeatCleanup = 5 * time.Minute
// activeSessionGrace is the staleness threshold for active (not paused)
// sessions based on last_sync_at.
activeSessionGrace = 45 * time.Second
// pausedSessionGrace is the staleness threshold for paused sessions.
pausedSessionGrace = 2 * time.Minute
// cleanupInterval is how often the cleanup ticker fires.
cleanupInterval = 15 * time.Second
// absStaleOpenSessionGrace closes audiobook playback sessions that stopped
// syncing without an explicit /close (abandoned playback) so they don't
// linger as "open" forever and inflate listening-stats aggregation.
absStaleOpenSessionGrace = 24 * time.Hour
// absSessionPruneInterval throttles the abandoned-session sweep: it's a slow-moving
// concern, so it runs hourly rather than on every 15s cleanup tick.
absSessionPruneInterval = time.Hour
)
// SessionCleaner removes stale playback sessions and dead node records.
type SessionCleaner struct {
pool *pgxpool.Pool
EventBus cache.EventBus
EventsHub *evt.Hub
stop chan struct{}
// lastABSSessionPrune gates the hourly abs_playback_sessions retention
// sweep. Guarded by absPruneMu because CleanStale is also invoked from the
// shutdown path while the ticker goroutine is still running.
absPruneMu sync.Mutex
lastABSSessionPrune time.Time
}
// NewSessionCleaner creates a SessionCleaner. The graceSeconds parameter is
// accepted for backwards compatibility but ignored — grace periods are now
// fixed at 45s (active) and 2m (paused).
func NewSessionCleaner(pool *pgxpool.Pool, graceSeconds int) *SessionCleaner {
return &SessionCleaner{
pool: pool,
stop: make(chan struct{}),
}
}
// Start begins the background cleanup loop, firing every 15 seconds.
func (c *SessionCleaner) Start() {
go func() {
ticker := time.NewTicker(cleanupInterval)
defer ticker.Stop()
for {
select {
case <-c.stop:
return
case <-ticker.C:
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
if deleted, err := c.CleanStale(ctx); err != nil {
slog.Error("session cleanup error", "error", err)
} else if deleted > 0 {
slog.Debug("cleaned stale sessions", "count", deleted)
}
cancel()
}
}
}()
}
// Stop signals the cleanup loop to stop.
func (c *SessionCleaner) Stop() {
close(c.stop)
}
// CleanStale performs a full cleanup pass:
// 1. Purge sessions from dead nodes (heartbeat stale > 45s)
// 2. Remove stale heartbeat rows (> 5 minutes)
// 3. Remove stale active sessions (last_sync_at > 45s)
// 4. Remove stale paused sessions (last_sync_at > 2 minutes)
func (c *SessionCleaner) CleanStale(ctx context.Context) (int, error) {
var totalDeleted int64
// 1. Purge sessions belonging to dead nodes.
tag, err := c.pool.Exec(ctx, `
DELETE FROM playback_sessions_sync
WHERE reporting_node IN (
SELECT node_id FROM node_heartbeats
WHERE updated_at < NOW() - make_interval(secs => $1::double precision)
)
`, nodeDeadTimeout.Seconds())
if err != nil {
return 0, fmt.Errorf("purging dead node sessions: %w", err)
}
totalDeleted += tag.RowsAffected()
// 2. Clean up stale heartbeat rows.
if _, err := c.pool.Exec(ctx, `
DELETE FROM node_heartbeats
WHERE updated_at < NOW() - make_interval(secs => $1::double precision)
`, nodeHeartbeatCleanup.Seconds()); err != nil {
return int(totalDeleted), fmt.Errorf("cleaning stale heartbeats: %w", err)
}
// 3. Active sessions: 45s grace on last_sync_at.
tag, err = c.pool.Exec(ctx, `
DELETE FROM playback_sessions_sync
WHERE is_paused = FALSE
AND last_sync_at < NOW() - make_interval(secs => $1::double precision)
`, activeSessionGrace.Seconds())
if err != nil {
return int(totalDeleted), fmt.Errorf("cleaning stale active sessions: %w", err)
}
totalDeleted += tag.RowsAffected()
// 4. Paused sessions: 2 minute grace on last_sync_at.
tag, err = c.pool.Exec(ctx, `
DELETE FROM playback_sessions_sync
WHERE is_paused = TRUE
AND last_sync_at < NOW() - make_interval(secs => $1::double precision)
`, pausedSessionGrace.Seconds())
if err != nil {
return int(totalDeleted), fmt.Errorf("cleaning stale paused sessions: %w", err)
}
totalDeleted += tag.RowsAffected()
// 5. Audiobook session cleanup (hourly): close abandoned open sessions.
// Closed rows are retained because the ABS stats endpoint currently has
// all-time semantics and aggregates directly from abs_playback_sessions.
// Kept off totalDeleted so it doesn't trigger the live-session
// invalidation event. The due-check is mutex-guarded so the shutdown-path
// CleanStale and the ticker can't race or double-run it.
c.absPruneMu.Lock()
pruneStartedAt := time.Now()
previousABSSessionPrune := c.lastABSSessionPrune
abndPruneDue := pruneStartedAt.Sub(c.lastABSSessionPrune) >= absSessionPruneInterval
if abndPruneDue {
c.lastABSSessionPrune = pruneStartedAt
}
c.absPruneMu.Unlock()
if abndPruneDue {
if err := c.closeAbandonedABSSessions(ctx); err != nil {
slog.Warn("abs session cleanup failed", "error", err)
c.absPruneMu.Lock()
if c.lastABSSessionPrune.Equal(pruneStartedAt) {
c.lastABSSessionPrune = previousABSSessionPrune
}
c.absPruneMu.Unlock()
}
}
if totalDeleted > 0 && c.EventsHub != nil {
if err := c.EventsHub.PublishJSON(
ctx,
evt.ChannelSessions,
"sessions.replaced",
nil,
evt.PublishOptions{AdminOnly: true},
); err != nil {
return int(totalDeleted), fmt.Errorf("publishing playback cleanup invalidation: %w", err)
}
} else if c.EventBus != nil && totalDeleted > 0 {
if err := c.EventBus.Publish(ctx, cache.ChannelPlayback, cache.Event{
Type: cache.EventPlaybackSessionsChanged,
Payload: "cleanup",
}); err != nil {
return int(totalDeleted), fmt.Errorf("publishing playback cleanup invalidation: %w", err)
}
}
return int(totalDeleted), nil
}
// closeAbandonedABSSessions closes abandoned audiobook playback sessions (no
// explicit /close, stopped syncing). It intentionally does not delete closed
// sessions: AggregateStats currently uses this table for all-time totals.
func (c *SessionCleaner) closeAbandonedABSSessions(ctx context.Context) error {
if _, err := c.pool.Exec(ctx, `
UPDATE abs_playback_sessions
SET closed_at = now()
WHERE closed_at IS NULL
AND COALESCE(last_sync_at, started_at) < NOW() - make_interval(secs => $1::double precision)
`, absStaleOpenSessionGrace.Seconds()); err != nil {
return fmt.Errorf("closing abandoned abs sessions: %w", err)
}
return nil
}