Files
silo-server/docs/superpowers
CoffeeKnyte 2312c0f4be perf(playback): cache embedded text-subtitle extracts
Extracting an embedded subtitle track walks the interleaved container:
subtitle packets sit between video and audio across clusters, so
harvesting a few KB of text means demuxing that stretch of a multi-GB
file off CephFS. Over 41.6h the web player saw p95 120,021ms / max
121,470ms on /api/v1/stream/{session_id}/subtitles/{track}, with 45 of
220 fetches over 10s. The 120s ceiling is the server's absolute
WriteTimeout (cmd/silo/main.go:2389) cutting the body mid-flight -- and
because WriteHeader(200) already ran, those truncations are logged
status=200 and are invisible in error metrics.

The server believes a 600s window bounds this. It does not: -t is passed
as an input option and ffmpeg silently ignores it for these extracts, so
every request runs from the seek point to EOF. Measured against the
production binary, `-ss 4000 -t 30` and `-ss 4000 -t 300` are
byte-identical to passing no -t at all (last cue 02:10:04, end of film).
Only -ss works, so cost tracks (duration - seek) x bitrate.

The window cannot simply be turned on. silo-apple and silo-android both
fetch a track once and depend on receiving the whole thing, so bounding
the output would silently kill subtitles ~10min into every film on both
platforms. The accidental whole-track behaviour is the de-facto contract.

So: keep whole-track delivery, make it cheap. SubtitleCache already had
the right shape for PGS; text was excluded only by the assumption that
"VTT is already windowed and fast", which the inert -t makes false.
Windowing a cached 83KB VTT costs 52ms versus 16s against the original
27GB remux.

Routing on AllowWindow would have poisoned the cache: it is set only in
the PGS branch, so it is always false for text, while streamExtractArgs
applies -ss to any non-ASS/non-PGS source regardless. A seeked subrip
request would take the full-track path, emit seek->EOF, exit cleanly and
publish that partial as canonical -- and every later viewer from 0 would
lose all cues before it (118 of 143 production requests carry a non-zero
seek). Canonicality is now derived from the effective argv instead:
streamExtractPlanFor is the single source of truth for both the argv and
the partial() predicate, so only a seek=0/duration=0 extract can fill.

- key: adds a schema version and the resolved output profile (an ASS
  source is reachable as both .ass and .vtt, so format must be keyed)
- removeStaleSiblings now groups by profile, so committing .vtt no longer
  deletes a valid .ass sibling; cleanup and eviction no longer hardcode
  .sup, so text entries are reclaimed and counted
- InputIsExtractedTrack carries the cached input's format
- text hits use a plain copy with no-store, matching cold-path HTTP
  semantics; ServeContent stays on the PGS path only, since Media3 uses
  range-capable data sources and responses must not vary with cache warmth
- renames SUP-specific identifiers now that the cache carries text

The inert -t is deliberately retained and its comment corrected in place:
removing it or moving it after -i would bound the output and break the
native clients.

Latency-only: for every (source codec, requested format, seek) the bytes
a client receives are unchanged.

Rationale, measurements, four superseded revisions and the dead ends are
recorded in docs/superpowers/plans/2026-07-17-subtitle-extract-cache.md.

AI-use disclosure: investigated and implemented with AI assistance
(Claude Code + Codex gpt-5.6-sol); Codex's review caught the cache-
poisoning bug above.
2026-08-05 04:55:54 +00:00
..