CoffeeKnyteandGitHub a2ef26bece perf: root-cause fixes for endpoints still slow after #292 (NextUp, series badges, resume tail, subtitle fonts) (#350)
* docs(plans): root-cause analysis for endpoints still slow after PR #292

Five endpoint groups stayed slow after the home/Continue Watching/Latest
latency work shipped: Resume (110s p95), NextUp (17s p95), Latest (17s),
/Items, and the home sections routes. The caps and caches from PR #292 are
live in the deployed binary; they bounded how many rows the loops touch but
not what each underlying query costs. Documents the four confirmed root
causes (4.3M stale completed-with-position progress rows + missing resume
index, unbounded next-up anchor scan, per-episode series rollup fanout, two
index-starved history/scanner paths) with live EXPLAIN ANALYZE measurements
and the fix plan implemented by the follow-up commits.

AI-use disclosure: analysis and doc produced with AI (Claude) assistance.

* perf(catalog): bound the global next-up anchor scan to recent completions

The completed_episodes CTE in buildListNextUpQuery derived per-series
anchors from the profile's ENTIRE completed history — DISTINCT ON over 233k
rows joined to episodes for the worst bulk-import profile, then a per-series
LATERAL that scans every episode of a fully-watched series before yielding
nothing. 648 slow executions in a 19h window, 44.7s worst; this drove
/Shows/NextUp (17.1s p95) and the next-up injection on the native home
sections aggregate.

Global queries now derive anchors from the profile's nextUpAnchorMaxRows
(500) most recent completed rows — an ordered index walk on
idx_uwp_profile_completed, with the hidden-items exclusion and date cutoff
applied inside the bounded scan so hidden/old rows never consume the anchor
budget. A next-up rail surfaces ~24 series; the 500 most recent completions
cover every series that can realistically rank on it. Series-scoped calls
(the show-detail tile) keep the unbounded shape: they must anchor on the
series' last completed episode no matter how long ago it was watched, and
are naturally bounded by one series.

Measured on the live worst-case profile with the exact generated SQL:
44.7s worst / ~2.6s avg before; 10ms after (together with the one-time
stale-resume-point data repair applied directly to the deployment DB — see
docs/superpowers/plans/2026-07-06-slow-endpoint-root-causes.md).

AI-use disclosure: implemented with AI (Claude) assistance.

* perf(jellycompat,userstore): aggregate series watch-state rollup in SQL

The series Played/UnplayedItemCount badge on list rails (per-library Latest,
library browse, search results) and series detail pages was computed by
materializing EVERY episode of every series on the page
(episodeRepo.ListBySeriesIDs) and then batching per-episode progress+history
lookups in 500-id chunks. A 50-series page of an episode-heavy library
(Sports) expanded to 32,467 episode rows and ~65 sequential queries —
measured 17-18s per /Items/Latest request, and PR #292's cached Latest fast
path pays it on every response for series libraries. The same fanout made
/Items?searchTerm=... slow whenever the result set was mostly series
(Meilisearch itself answers in milliseconds).

New optional store capability userstore.SeriesEpisodeRollupStore, implemented
by PostgresUserStore as one GROUP BY e.series_id aggregate with semantics
identical to the chunked path (episode availability via episode_libraries,
hidden-items visibility on progress rows, completed-history fold, in-progress
= not watched with position > 0 — verified value-for-value against the old
semantics on a real 1,586-episode series). enrichSeriesListUserData and
enrichDetailUserData use it when present; SQLite-backed stores and rollup
query failures keep the existing chunked path as fallback.
catalog.SeasonUserDataFromCounts pins the counts-to-DTO mapping to
EpisodeRollupUserData.

Measured on the live worst-case profile against the real 50-series Sports
Latest page: ~17s of chunked round-trips before, 119ms in one query after.

Part of docs/superpowers/plans/2026-07-06-slow-endpoint-root-causes.md.

AI-use disclosure: implemented with AI (Claude) assistance.

* perf(catalog): bound superseded-episode completed walk to recent history

The Resume / Continue Watching superseded-episode filter loaded a
profile's *entire* completed history into memory on every request that
contained an in-progress episode: CompletedProgressSnapshots paged
user_watch_progress WHERE completed=TRUE with no upper bound. The
2026-07-06 slow-query comparison showed this surviving as a 60-116s
Resume tail even after the in-progress index landed live, because the
4.3M zeroed Plex-import rows are still completed=TRUE and were re-walked
every load.

A completed episode can only supersede an in-progress one it was
finished more recently than (the query gates on
done_progress.updated_at > ip_progress.updated_at), so only completed
rows newer than the oldest in-progress entry can matter. Compute that
cutoff in SupersededEpisodeProgressIDs and pass it to
CompletedProgressSnapshots, which — since the completed listing is
ordered updated_at DESC — stops paging as soon as it crosses the cutoff.
Import-heavy profiles whose back-catalogue predates their current
in-progress items now stop on the first page instead of paging hundreds
of thousands of irrelevant rows. Correctness is unchanged: no relevant
superseding row is excluded.

* perf(catalog): hard-cap superseded-episode completed walk at 5 pages

The updated_at cutoff added in the previous commit bounds the completed
walk on the relevance axis, but a very old in-progress entry sitting
behind a large volume of newer completions could still page deep. Add a
5-page (2,500-row) hard backstop on top of the cutoff: normal profiles
still stop on page one via the cutoff, and only the adversarial tail hits
the cap. When it engages the tail of the completed set goes unscanned, so
a superseded episode could momentarily survive on Continue Watching — we
log a warning when that happens (with profile_id + rows scanned) rather
than mis-filter silently, and it self-corrects once the stale in-progress
entry ages out of the scanned window.

* perf(playback): extract subtitle fonts in a single ffmpeg pass

Embedded ASS/SSA font extraction spawned one ffmpeg process per font
attachment, each re-opening the (usually CephFS-backed) media file. Anime
releases carry 15-47 fonts, so the per-spawn file-open cost dominated and
pushed GET /api/v1/stream/{sid}/subtitles/{track}/fonts to a 17-60 s plateau
(p95 ~33 s in the live logs).

Collapse the N spawns into one ffmpeg invocation that dumps every attachment
to a temp dir (-dump_attachment:idx path ... -i file -map 0:t? -c copy), then
read the files back. The file is opened once instead of N times, taking p95
from ~30 s to ~1-2 s with no change to output.

Safety is preserved. The 32-attachment / 32 MiB caps still apply: attachment
size is stat'd before read so an over-limit font never enters memory, and a
watchdog polls the dump dir and kills ffmpeg if its on-disk output crosses the
cap -- restoring the hard bound the old pipe-per-attachment reader enforced by
killing at maxBytes+1, so a container with oversized "font" attachments can't
fill the disk.

Part of the slow-endpoint follow-up; see
slow-query-analysis/subtitle-fonts-extraction-findings.md.

* fix(review): report enforced font-byte cap; correct doc subtitle scope

Address PR #350 review:
- dumpFontAttachments reported the maxSubtitleFontBytes package constant in
  both over-limit errors instead of the maxBytes argument the caller passed,
  so the message misstated the enforced bound whenever a different cap was in
  effect (as the tests use). Interpolate maxBytes in both messages.
- The root-cause plan claimed subtitle extraction was 'out of scope' while the
  branch actually optimizes /subtitles/{track}/fonts. Scope the out-of-scope
  note to subtitle *track* conversion and record the fonts single-pass work as
  deliverable 5.
2026-07-09 09:02:45 -04:00
2026-05-22 23:26:56 -04:00
2026-05-22 23:26:56 -04:00
2026-05-22 23:26:56 -04:00
2026-05-22 23:26:56 -04:00
2026-05-22 23:26:56 -04:00

Silo

Silo is a self-hosted media streaming server for your movies, shows, music, and books. Point it at your media folders and stream to your devices — at home or away — with direct play, remuxing, and hardware-accelerated transcoding handled automatically.

Join the community on Discord. If Silo is useful to you, consider sponsoring the project — see Supporting Silo.

Highlights

  • Plays your media, your way — direct play when the device supports it, remux or hardware-accelerated transcode (including NVENC) when it doesn't.
  • Web app included — a full-featured web client and admin interface ship with the server.
  • Works with apps you already use — optional Jellyfin/Emby-compatible API supports clients such as VidHub, Findroid, and Infuse.
  • Household profiles — multiple profiles per account, with per-profile watch state and parental controls.
  • Plugin-driven metadata — match and enrich your libraries with providers like TMDB and TVDB, installed as plugins.
  • Fast setup — one docker compose up -d brings up the whole stack; everything else is configured in the admin UI.

The easiest way to run Silo is with Docker Compose. The default stack assumes you do not already have PostgreSQL and Redis available, so it bundles PostgreSQL, Redis, FFmpeg, and the application for a one-command start.

  1. Create a .env file

    cp .env.example .env
    
  2. Set your media path

    Edit .env and set:

    MEDIA_ROOT=/path/to/your/media
    

    MEDIA_ROOT is the one value most users need to change. You can also override SILO_DATA_ROOT if you do not want bind mounts under /opt/silo, and change ports if the defaults conflict with something else on the host.

  3. Start the default integrated stack

    docker compose up -d
    

    This starts PostgreSQL, Redis, and the integrated Silo server. The app is available at http://localhost:8090. Jellyfin-compatible app support is disabled until an administrator enables it in onboarding or admin settings.

    If you already have PostgreSQL and Redis available, omit those bundled service examples from compose and point Silo at your existing DATABASE_URL and REDIS_URL instead.

    Optional NVIDIA/NVENC

    GPU support is kept out of the default compose file so hosts without NVIDIA drivers work unchanged.

    Install the NVIDIA Container Toolkit and use a Docker Compose version with GPU reservation support before enabling this override.

    Use the optional override file when you want NVENC:

    docker compose -f docker-compose.yml -f docker-compose.nvidia.yml up -d
    

    If you want this controlled from .env, set COMPOSE_FILE:

    COMPOSE_FILE=docker-compose.yml:docker-compose.nvidia.yml
    NVIDIA_GPU_COUNT=1
    

    Windows uses ; instead of : between compose files.

    Then docker compose up -d will include the NVIDIA override automatically.

  4. Configure through the admin UI

    Add libraries, users, metadata providers, and playback settings from the web interface.

Bind Mount Layout

The deploy-oriented compose files use host folder mappings rather than Docker-managed volumes.

By default, data is stored under /opt/silo:

  • /opt/silo/postgres
  • /opt/silo/redis
  • /opt/silo/transcode
  • /opt/silo/catalog-seeds

Media is mounted into the container at /mnt/media from the host path you set in MEDIA_ROOT.

Optional Profiles

The main compose file is integrated-first. These profiles exist for operators testing distributed mode or mirroring a split deployment shape. Most single-host installs should stay on the default integrated service, because it already includes proxying and transcoding.

Profile Command Description
default docker compose up -d Integrated server plus bundled PostgreSQL and Redis
proxy docker compose --profile proxy up -d Start a standalone proxy service for distributed-mode testing
transcode docker compose --profile transcode up -d Start a standalone transcode service for distributed-mode testing

You can enable both optional examples together:

docker compose --profile proxy --profile transcode up -d

If you are splitting workers across multiple hosts, use the separate remote worker example instead of trying to stretch the main compose file across machines.

Advanced Remote Node Example

For a dedicated remote transcode worker, use docker-compose.remote-transcode.yml. That file is intended for a separate worker host that connects back to an existing Silo deployment using shared PostgreSQL and Redis.

Deployment Notes

The default compose stack intentionally bundles PostgreSQL and Redis for ease of setup and assumes a fresh install without those services already available. If you already operate PostgreSQL and Redis, omit those examples from compose and point Silo at your existing infrastructure instead. For serious installs, PostgreSQL is better on a separate VM or a managed service so upgrades, tuning, and backups are isolated from the app host. Redis can stay local for many installs, but externalizing it is also reasonable if you already operate shared infrastructure.

Silo is externally stateful by default rather than fully stateless. Durable application state lives in PostgreSQL. Redis only stores coordination and cache-style data. Silo still writes transient transcode output locally under /tmp/silo-transcode. If you switch userdb.backend=sqlite, Silo also becomes locally stateful at /var/lib/silo/userdb.

Migrating an existing Continuum Docker install should be done with the preflight helper and cutover guide in docs/continuum-to-silo-docker-migration.md.

Configuration

Silo requires only a DATABASE_URL when running from source or against external infrastructure. In the default Docker Compose path, the stack wires the database and Redis URLs for you. All other settings — libraries, metadata providers, transcoding, users — are managed through the admin UI after first launch.

Server Modes

Mode Description
integrated Full server: API + frontend + scanner + transcode (default)
api API server only, no local transcoding
proxy Stream proxy node that connects to the shared deployment database and Redis
transcode HLS transcode worker node that connects to the shared deployment database and Redis

PostgreSQL Auto-Tuning

The default Docker Compose stack does not require a checked-in postgresql.conf. It enables Silo's pgtune-style OLTP tuning by default:

POSTGRES_TUNE: auto

When enabled, Silo connects with DATABASE_URL and applies recommendations with ALTER SYSTEM, which writes to PostgreSQL's postgresql.auto.conf inside the database data directory. Reloadable settings are applied immediately with pg_reload_conf(). Settings that PostgreSQL marks as restart-only are written too, and Silo logs the setting names so you can restart PostgreSQL once:

docker compose restart postgres

The default Compose database user has the required PostgreSQL permissions. If you use an external PostgreSQL server, make sure the configured DATABASE_URL user can run ALTER SYSTEM, or set POSTGRES_TUNE=off and manage PostgreSQL yourself.

For POSTGRES_TUNE_MEMORY=auto, Silo uses the first trustworthy memory source: a finite Docker cgroup limit, the read-only /host/proc/meminfo mount supplied by the bundled Compose file, then /proc/meminfo with container safety guards. Auto-detected memory is treated as a PostgreSQL budget, defaulting to 75% of detected RAM so Silo, Redis, plugins, transcodes, and the OS retain headroom. POSTGRES_TUNE_DB_SIZE=auto queries pg_database_size(current_database()) and classifies the workload by comparing the database size to that memory budget.

Optional tuning overrides:

Variable Default Description
POSTGRES_TUNE_PROFILE oltp Tuning profile. Only oltp is currently supported.
POSTGRES_TUNE_MEMORY auto Server/container RAM, such as 8GB or 32GB; explicit values are used as-is.
POSTGRES_TUNE_MEMORY_BUDGET_PERCENT 75 Percent of auto-detected RAM used for PostgreSQL recommendations.
POSTGRES_TUNE_CPUS auto CPU count used for worker recommendations.
POSTGRES_TUNE_STORAGE ssd One of hdd, ssd, san, or nvme.
POSTGRES_TUNE_DB_SIZE auto Use less_ram when the database comfortably fits in RAM, mid_ram, or greater_ram for very large databases.
POSTGRES_TUNE_CONNECTIONS 100 PostgreSQL max_connections; automatically raised if Silo's app pool is configured higher.
POSTGRES_SHM_SIZE 8gb Docker /dev/shm size for the bundled PostgreSQL container.

Advanced operators can still supply their own PostgreSQL configuration or override these env vars. Set POSTGRES_TUNE=off when you do not want Silo to change PostgreSQL server settings. Settings already written with ALTER SYSTEM remain in postgresql.auto.conf; reset those PostgreSQL parameters if you later move fully to a custom postgresql.conf.

Build from Source

If you prefer running Silo without Docker:

  1. Install prerequisites: Go 1.24+, Bun 1.0+, PostgreSQL 18+, and FFmpeg.

  2. Start PostgreSQL and Redis (skip if you already have them running)

    docker compose up -d postgres redis
    

    The main compose file still expects MEDIA_ROOT to be set even if you only want the bundled PostgreSQL and Redis services, so set that in .env first.

  3. Configure the database connection

    cp .env.example .env
    

    Edit .env and set DATABASE_URL to point to your PostgreSQL instance.

  4. Build and run

    make build
    ./silo
    

    The server starts at http://localhost:8080 by default. All other settings are configured through the admin UI.

Reporting Issues

If you are reporting a bug, install problem, or performance issue, start with the admin workflow and reproduction steps, not Claude/Codex analysis.

Please include:

  • What you were trying to do
  • Exact steps you took
  • What you expected to happen
  • What actually happened
  • What exact action is slow or broken (save, scan, browse, import, playback, etc.)
  • Whether it happens every time or only sometimes
  • The library, media type, filter, setting, or value involved
  • Version, branch, commit, and deployment details if you know them
  • Screenshots, recordings, or log snippets if relevant

If you used Claude/Codex for debugging, put that under Technical notes at the end. Suspected files, SQL output, stack traces, and root-cause theories can be helpful, but only after the workflow and repro steps are clear.

Use this template:

Goal:
Steps:
Expected:
Actual:
What is slow/broken:
Scope:
Version/branch:
Deployment:
Technical notes:

Contributing & Development

Silo is open source and contributions are welcome. See DEVELOPMENT.md for building from source in a dev workflow, running tests, database migrations, and project layout, and CONTRIBUTING.md for contribution expectations, merge request guidance, and the policy for AI-assisted submissions.

Supporting Silo

Silo is an open-source hobby project, developed in spare time and funded out of pocket. If you'd like to support development, you can sponsor via GitHub Sponsors.

Donations go directly toward the costs of building and running the project:

  • AI development tooling subscriptions (Claude, Codex) used to build and maintain Silo
  • Push notification relay infrastructure
  • Future development costs

Sponsoring is entirely optional — Silo is and will remain free and open source. Bug reports, contributions, and feedback are just as valuable.

License & Trademarks

Silo's source code is licensed under the GNU Affero General Public License v3.0 or later (AGPL-3.0-or-later) — see LICENSE.

The Silo name, logo, and wordmark are trademarks of Silo Media L.L.C. and are not covered by the AGPL. You're free to fork and redistribute the code, but forks and redistributions must not use the Silo brand as their identity and must remove or replace the brand assets. Publishing a Silo-branded app to an app store requires written permission. See TRADEMARK.md for what's permitted — including referential use like "compatible with Silo."

S
Description
Self-hosted media streaming server with a Go backend, React web UI, Docker deployment, transcoding, and Jellyfin-compatible APIs.
Readme
340 MiB
Languages
Go 72.6%
TypeScript 26.3%
PLpgSQL 0.4%
CSS 0.3%
Python 0.2%
Other 0.1%