* docs(playback): add v3 neutral-contract finalization plan Supersedes the wire-contract sections of the 2026-07-12 v3 plan: server-owned attempt keys, delivery-keyed negotiation without Media3 engine names, tiered capability evidence, neutral device/output context, track/quality replan operations, audio-only planning, and coordinated no-back-compat rollout across server, Android, Apple, and web. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(playback): make v3 attempt keys server-owned and replace engines with deliveries Contract core of the platform-neutral v3 finalization (plan sections 3.1 and 3.2), breaking on purpose — v3 is dark and all clients move together: - Every PlanV3 now carries plan_attempt_key, an opaque server-computed token clients store and echo in attempted_plan_keys; ReplanRequestV3 gains bounded local_mutations that the replan handler folds into the failed plan's key. Clients never hash anything. - KotlinName() is deleted from DeliveryV3, StreamProtocolV3 and SubtitleModeV3; the attempt-key canonical string now uses lowercase wire tokens, and PlanRecipeVersionV3 bumps to v3.3 so no key or plan ID computed under the old canonicalization can collide. - EngineV3 leaves the wire: ClientPlaybackContextV3.Engines (media3_*) becomes Deliveries keyed original_http|progressive|hls, with EngineCapabilityV3 renamed DeliveryCapabilityV3. PlanV3.Engine is removed; the planner, subtitle policy and quirk registry re-key on delivery class, and the media3_only feature token is deleted. - Validated-claim strings drop the prefix: media3_h264_decode -> h264_decode, media3_audio_decode -> audio_decode. - Golden fixtures in testdata/protocol_v3 are regenerated by Go and are now the cross-repo source of truth. Part of the playback protocol v3 neutral-contract train (steps 2-3 of docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(playback): add v3 evidence tiers and neutral device/output context Implement plan sections 3.3 and 3.4 of the v3 neutral-contract pass: - ClientCodecCapabilitiesV3 gains required video_evidence and audio_evidence closed enums (exact | platform_attested | declared). Planner strictness follows the tier: exact keeps the strict decode-entry validation, platform_attested validates codec/resolution/bit-depth/ frame-rate but skips profile/level matching, declared grants copy routes from the flat codec lists. Only exact audio evidence earns passthrough claims. The detailed_decode_capabilities feature token is deleted (subsumed by video_evidence=exact), and evidence-blocked direct routes carry the new evidence_insufficient_for_direct reason/warning. - DeviceContextV3 is now platform/os_version/manufacturer/model plus a bounded platform_details map (<=16 entries, <=128 chars); the Android Build dump fields are gone. Fire TV quirks keep matching on manufacturer/model (brand fallback removed with the field). - output_route_generation (int64, dual-location) becomes an optional opaque output_context_id string on the output context; the dual-location consistency validation is deleted. Attempt keys, plan invalidation, route events, and the planstore column follow (new Goose migration). - Feature advertisement collapses to the top-level client_features list only; ClientPlaybackContextV3.Features is deleted and ReplanRequestV3 gains an optional client_features refresh. - PlanRecipeVersionV3 bumped v3.3 -> v3.4; fixtures re-keyed. Part of the playback protocol v3 neutral-contract finalization plan (docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(playback): add v3 intent replans, quality menu, and audio-only routes Protocol v3 could only replan after a failure, so changing the audio track or the quality still required the legacy audio PATCH and the client-recipe transcode start — the two endpoints v3 is meant to replace. Clients also had to own a resolution ladder to render a quality menu, and a source with no video track was terminaled by the video/HDR gates, keeping audiobooks on the legacy path. Add track_change and quality_change replan operations. They carry no failure classification and route through the existing replan transaction, so they inherit its idempotency, capacity reservation, and staged-successor commit for free. Because nothing failed, the previous route stays eligible: neither the attempted-key history nor the failed-plan exclusion applies to them. Publish the server ladder on the plan as available_qualities so the quality menu is server-owned; the rungs come from the same resolutionLabelV3 and ladderBitrateKbpsV3 helpers the planner itself uses, not a parallel table. Plan audio-only sources through their own reduced route family: original_http when the client decodes the codec, otherwise a progressive AAC conversion. The plan advertises audio/mp4 for that remux and the transport now serves the same value, because a declared-tier client probes the advertised MIME with isTypeSupported before attaching a source buffer, and "video/mp4" on a stream with no video track is exactly the mismatch that makes the probe lie. Name the protocol's string vocabulary (dynamic ranges, transformations, executors, validated claims, terminal reasons) as constants while touching these lines, so the wire values have one definition. Part of #135 * docs(playback): publish the v3 protocol contract and fix subtitle ordinals Protocol v3 exists only as Go code today, so the Android and Apple ports have no authority to implement against other than reading this repository. Publish the contract as a normative document, machine-checkable schemas, and generated golden fixtures, and fix the one place where the server's own wire output disagreed with the ordinal space it publishes. - docs/architecture/playback-protocol-v3.md is self-contained enough for a third-party client: endpoints and status codes, evidence tiers and their bound-matching rules, delivery classes, the timeline model, replan semantics, registries, track identity, plan identity, quality, and transformations. - docs/design/schemas/playback-v3/ carries JSON Schemas for the five wire shapes plus valid and invalid fixtures, following the client-diagnostics layout. internal/playback/contract validates every fixture against its schema, so a schema that drifts from the Go types fails the Go suite. - cmd/playbackfixtures generates internal/playback/testdata/protocol_v3 from the production planner. `make playback-fixtures` writes them and `make verify-playback-fixtures` (wired into CI) fails when they are stale. These files are what the client ports consume, so drift would otherwise surface as a playback bug on three platforms at once. The subtitle fix: combined ordinals are one dense space over externals, then embedded tracks, then downloaded ones, but the legacy URL builder skipped burn-in-only tracks while assigning indices, so every track after a DVD/DVB track was numbered one too low and resolved to its neighbour. Ordinal assignment now lives in playback.BuildSubtitleInventoryV3 and both the plan inventory and the legacy `subtitle_urls` shape project from it; the legacy shape still filters burn-in-only entries but keeps each track's real index. Part of #135 * feat(web): migrate the players to the neutral playback v3 contract The web player was the last client still speaking the legacy start protocol: it picked its own file version from a codec probe, posted an ffmpeg recipe to start a transcode, PATCHed an endpoint to change audio tracks, and derived its own quality ladder. None of that survives a server-owned plan, and none of it produced telemetry the apps could be compared against. Video player: starts with a v3 request that advertises `declared` evidence from `isTypeSupported` probes and the three delivery classes, then consumes the returned plan for its URL, timeline, tracks and warnings. Quality and track changes become replans (`quality_change`, `track_change`), the quality menu renders `available_qualities` instead of computing rungs, and playback failures emit `route-events` so web failures land in the same diagnostics as Android and Apple. The duration comes from `source.duration_seconds` rather than the playback engine, and the "how was this delivered" overlay reads the plan's delivery and server transformations instead of comparing codec strings. Audiobook player: starts against the audio-only planner path with a single `original` rung, and takes its seek anchor from `timeline.player_start_seconds` so the progressive-remux route (which anchors the stream and restarts the player clock at zero) does not seek twice. Server side, `disable_progress_persistence` left the wire, so the rule it encoded is now derived. Resume state is keyed on the item, but every part of a multipart presentation shares that key while carrying its own file-local clock — persisting part 4's position would store "12 minutes in" as the book's resume point. `PresentationPartTotal > 1` expresses that directly and generalizes to multipart movies and split episodes, and a client can no longer forget to ask or lie about it. `useTranscodeQuality` and the legacy response types are deleted, and `WEBTEST_KNOWN_FAILURES` loses the audiobook entry along with its fix. Part of #135 * feat(playback)!: make v3 the only playback protocol Protocol v3 shipped behind a flag, alongside the legacy start path it was designed to replace. Running both meant every planner change had to be made twice, in two shapes that disagree about who decides the route: the legacy body carried a decision the client had already made, while v3 asks the server to make it. This deletes the legacy half. Removed: - `handleStartPlaybackLegacy` and its request/response bodies. The `POST /playback/start` route stays, but the protocol-version dispatch envelope is now a strict v3 decode — a body that does not declare `protocol_version: 3` gets `426 client_upgrade_required` so an outdated app can render a clear "update required" state instead of misreading a plan. Deliberately not a `400`: the request may be well-formed for the protocol it was written against. - `POST /playback/transcode/start`, superseded by the `quality_change` replan operation, and `PATCH /playback/{session_id}/audio`, superseded by `track_change`. Both mutated a session without re-planning. - The shadow planner and both rollout settings rows. With v3 the only protocol, `playback.protocol_v3_enabled` would mean "no playback at all"; `playback.protocol_v3_shadow_enabled` gated a comparison against a path that no longer exists. `409 protocol_disabled` on route-events goes with them, and capability `enabled` is now constant `true` (the field stays — clients feature-detect against it). - Version-selection helpers in `internal/playback/resolver.go` that only legacy start reached. `Resolve`/`ClientCapabilities`/`PlayDecision` stay: downloads consumes them. `internal/jellycompat` has its own resolution surface and is untouched. Behaviour the legacy handlers owned and v3 now owns explicitly: series version and audio-track preferences are persisted on start and on a `track_change` replan (not on failure recovery, whose forced route is not a user choice); an omitted `start_position` resolves to the profile's saved resume point; and an omitted audio track resolves through the series preference, the profile audio language, then the library override. Both are settled before planning, because the plan's timeline is cut at the start position. Spec §2.2 documents this as "omission is a request, not a default". The encode-target clamp that lived in the deleted transcode handler is already enforced in the planner, twice — `availableQualitiesV3` omits rungs at or above the source height, and the encode path clamps `targetHeight` to it. Unchanged: progress, stop, HLS manifest and segment delivery, the realtime control socket, stream tokens and restart reconstruction, watch together, downloads, jellycompat. Every removal is recorded in the pre-lock removals table in docs/architecture/v1-scope.md. Part of #135 * fix(scanner): stop recording embedded cover art as a video track ffprobe reports embedded cover art as a video stream carrying disposition.attached_pic. convertProbeData appended every "video" stream to VideoTracks without consulting isMainVideoStream, the predicate that already existed for duration decisions, so the picture was persisted as a playable track. That misreports the file twice: - An audio file with a cover picks up a video track, so it no longer satisfies MediaFile.IsAudioOnly and the v3 planner routes an audiobook through the video path instead of planAudioOnlyV3. - When the picture is ordered ahead of the real stream, the flat codec_video/resolution/hdr columns describe the poster: a 954x720 h264 episode was stored as mjpeg 480x480. Filter attached_pic streams out of the track loop. The guard is the disposition flag, not the codec name, so a genuine MJPEG video is still probed as video — the library has one. Already-probed rows self-heal on the next playback: NeedsCriticalProbeRepair already reprobes tracks missing color_range, which covers 21 of the 23 affected rows, and applyProbeData overwrites VideoTracks wholesale. The remaining two need a rescan; nothing persisted records attached_pic, and keying repair off still-image codec names would reprobe the genuine MJPEG file on every playback forever. Part of the playback v3 neutral-contract work: it is what lets Android drop AUDIOBOOK_COVER_ART_CODECS, which fabricated decode support the client cannot honestly claim under video_evidence: "exact". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(playback): publish subtitle URLs even when playback starts with subtitles off The v3 plan's subtitle inventory is the authoritative track list a client builds its subtitle menu from, but the handler only rewrote it with session-scoped URLs when a track was actually selected. A start or replan that resolved to `subtitle.mode: "off"` therefore returned the planner's URL-less inventory, so a client whose picker reads the inventory had a menu it could not fetch anything from. The Cast path hits this every time: it starts with subtitles off and needs the receiver's text tracks up front. attachSubtitleArtifactV3 now scopes and publishes the inventory unconditionally and gates only the artifact stamping on the selection. Spec §8 records that the `url` on a sidecar entry does not depend on the current selection. Part of the v3 neutral-contract finalization. * chore(playback): reconcile neutral v3 with main * fix(playback): preserve subtitle intent across replans * fix(playback): retain subtitle inventory on adapted routes * fix(playback): software-decode High10 AVC for QSV * fix(playback): scale High10 frames before QSV upload * fix(playback): preserve empty subtitle inventories * fix(playback): freeze terminal attempt contract * chore(playback): name fixture contract tokens * fix(playback): close v3 conformance review gaps * chore(playback): name conformance category * fix(playback): complete v3 conformance contract * fix(playback): keep schema fixtures generated * fix(playback): emit schema-valid conformance arrays * fix(playback): omit empty replan failures * fix(web): omit empty replan failures * fix(playback): close neutral v3 contract gaps * fix(playback): harden v3 replan, transcode, and quality-ladder edge cases Review remediation for the neutral v3 cutover, server side: - A failed replan no longer overwrites the durable StartResponse with a terminal or advances the replan request ID; an idempotent start replay of a still-healthy session returns the original plan. - SoftwareVideoDecode is now derived inside the transcode layer from source facts (codec/profile/bit depth) carried on TranscodeOpts, so jellycompat, downloads, recipe-card reconstruction, and transcode nodes get the High10 software-decode fix, not just the v3 handler. video_to_h264 recipe version bumps to 2 so mixed-version node pools that would silently drop the flag fail validation instead. - Local transport startup shares the 30s ManifestStartupTimeout; a timeout with the process still running stays retryable and is no longer persisted as a durable terminal against the attempt. - Sparse replan bodies (failure_recovery et al) no longer reset a user-selected quality preference to auto; the empty-value guard now covers every operation. - availableQualitiesV3 publishes no fixed rungs when the source height is unknown, keeping the no-upscaling ladder contract. - The proxy remux path serves audio-only fMP4 as audio/mp4 via a new additive AudioOnly token claim, matching the integrated path. - Plain text subtitle sidecars accept any requested extension again (served as VTT), restoring the permissive v1 behavior; ASS and bitmap handling is unchanged. - The 4K-disallowed terminal message discloses when a lower-resolution alternate exists but was pinned away by quality "original". Part of #135. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(web): keep playback alive through failed replans and honest audio claims Review remediation for the neutral v3 cutover, web player: - A failed or refused replan no longer unmounts the player: the fatal error screen is reserved for loads with no adopted plan, and replan failures surface through the existing non-fatal replanError path. - changeQuality rolls its optimistic preference back when the replan is refused or errors, so a failed switch is not silently applied by the next unrelated replan and the menu shows the real active rung. - The capability probe now tests mp3/vorbis codecs and mp3/flac/ogg containers (MediaSource with a canPlayType fallback), restoring direct play for mp3 audiobooks instead of per-part AAC re-encodes. - Reanchor seeks issued while a replan is in flight coalesce and run when it settles instead of being silently dropped with the scrubber pinned to a phantom position. - Subtitle refresh/translation replans use the resume anchor while the media element has no metadata, so a subtitle_ready broadcast during startup no longer restarts a resumed stream at 0:00. - An exhausted failure-recovery chain sets a visible error instead of returning silently. Part of #135. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(playback): accept video-only and VP9 probe metadata Treat audio and video probe completeness independently so legitimate video-only assets converge without repeated ffprobe repair. Allow unknown codec profile/level metadata to fall through to server adaptation while preserving exact direct-decode constraints. Fixes #574 * fix(playback): address protocol v3 review findings * fix(playback): harden lease and probe repair decisions * fix(playback): close remaining v3 review gaps * fix(playback): recover failed transcode starts * fix(playback): address remaining review-bot findings on v3 replan and audio planning Server: - The deferred replan lease release is bounded by a 3s timeout so a saturated pool or DB outage cannot wedge a handler goroutine that holds the per-session store lock on an uncancellable context. - planAudioOnlyV3 honors the request bandwidth cap: an over-cap source skips the original_http direct route and converts to AAC with the same bandwidth_cap_applied warning and decision reason the video ladder uses. Unknown source bitrate never triggers the cap. - A copy-audio progressive plan rejected only by a per-delivery audio_decode_codecs subset retries as an AAC conversion instead of returning adaptation_unavailable, and the AAC recipe respects the delivery's max_channels. Web: - failure_recovery replans issued while another replan is in flight queue (superseding a pending seek reanchor) instead of being silently dropped with the fatal overlay already suppressed. - A terminal response to a fresh non-preserving start clears the previous plan and stops its session, so episode navigation cannot keep rendering the prior item under the new title. - A refused recovery replan for a transport-dead plan surfaces the error and re-arms the plan failure key, so transient recovery failures no longer strand an endless spinner; the audiobook player gets the same guard reset. - A track-less subtitle_translation_completed hands off to the refreshed persisted track once the inventory settles, clearing the live overlay, instead of pinning the synthetic live track forever. Part of #135. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(playback): reuse HLS transport for sidecar replans * fix(playback): stabilize copy HLS remount timeline * fix(playback): address v3 review findings * fix(playback): satisfy player contract types --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
586 lines
20 KiB
Go
586 lines
20 KiB
Go
package downloads
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"fmt"
|
|
"log/slog"
|
|
"os"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/Silo-Server/silo-server/internal/config"
|
|
"github.com/Silo-Server/silo-server/internal/idgen"
|
|
"github.com/Silo-Server/silo-server/internal/models"
|
|
"github.com/Silo-Server/silo-server/internal/playback"
|
|
)
|
|
|
|
const (
|
|
artifactLease = 2 * time.Minute
|
|
artifactHeartbeat = 40 * time.Second
|
|
artifactMaxAttempts = 3
|
|
)
|
|
|
|
// EncodePreparer produces a single finalized file for an artifact. The default
|
|
// implementation calls playback.PrepareFile; tests substitute a fake.
|
|
type EncodePreparer interface {
|
|
PrepareFile(ctx context.Context, opts playback.TranscodeOpts, outputPath string) error
|
|
}
|
|
|
|
type playbackPreparer struct{}
|
|
|
|
func (playbackPreparer) PrepareFile(ctx context.Context, opts playback.TranscodeOpts, outputPath string) error {
|
|
return playback.PrepareFile(ctx, opts, outputPath)
|
|
}
|
|
|
|
// NewPlaybackPreparer returns the production EncodePreparer (ffmpeg-backed).
|
|
func NewPlaybackPreparer() EncodePreparer { return playbackPreparer{} }
|
|
|
|
// ArtifactNotifier publishes an event when a linked download changes state.
|
|
type ArtifactNotifier func(ctx context.Context, d *Download)
|
|
|
|
// ArtifactManager owns the durable encode queue: it ensures/deduplicates encode
|
|
// jobs, drains them through a bounded worker pool with leased heartbeats, and
|
|
// recovers stranded jobs on startup.
|
|
type ArtifactManager struct {
|
|
repo *ArtifactRepository
|
|
downloads *Repository
|
|
fileRepo FileResolver
|
|
preparer EncodePreparer
|
|
owner string
|
|
liveCfg func() *config.Config
|
|
notify ArtifactNotifier
|
|
|
|
mu sync.Mutex
|
|
kick func()
|
|
lastDiskSweep time.Time
|
|
lastStaleSweep time.Time
|
|
}
|
|
|
|
// maintenanceInterval spaces the disk-presence and stale-row sweeps: both are
|
|
// O(cache size) (stats / extra queries) and their failure modes self-heal, so
|
|
// running them on every 30s task tick is steady-state waste that grows with
|
|
// the artifact cache. The first run after startup always executes.
|
|
const maintenanceInterval = time.Hour
|
|
|
|
// maintenanceDue reports whether the sweep guarded by last is due, advancing
|
|
// the stamp when it is.
|
|
func (m *ArtifactManager) maintenanceDue(last *time.Time) bool {
|
|
m.mu.Lock()
|
|
defer m.mu.Unlock()
|
|
if !last.IsZero() && time.Since(*last) < maintenanceInterval {
|
|
return false
|
|
}
|
|
*last = time.Now()
|
|
return true
|
|
}
|
|
|
|
// NewArtifactManager constructs an ArtifactManager. liveCfg reads the current
|
|
// config (artifact dir, worker-pool size, byte budget, ffmpeg/hwaccel); owner is
|
|
// this node's id for lease ownership; notify (optional) publishes ready/failed.
|
|
func NewArtifactManager(
|
|
repo *ArtifactRepository,
|
|
downloadRepo *Repository,
|
|
fileRepo FileResolver,
|
|
preparer EncodePreparer,
|
|
owner string,
|
|
liveCfg func() *config.Config,
|
|
notify ArtifactNotifier,
|
|
) *ArtifactManager {
|
|
if preparer == nil {
|
|
preparer = playbackPreparer{}
|
|
}
|
|
if owner == "" {
|
|
owner = "node"
|
|
}
|
|
return &ArtifactManager{
|
|
repo: repo, downloads: downloadRepo, fileRepo: fileRepo, preparer: preparer,
|
|
owner: owner, liveCfg: liveCfg, notify: notify,
|
|
}
|
|
}
|
|
|
|
// SetKick wires a low-latency drain trigger (e.g. taskmanager RunTask) invoked
|
|
// when a new job is enqueued.
|
|
func (m *ArtifactManager) SetKick(kick func()) {
|
|
m.mu.Lock()
|
|
m.kick = kick
|
|
m.mu.Unlock()
|
|
}
|
|
|
|
// Ready returns a ready artifact for serving and bumps its LRU timestamp.
|
|
// Returns ErrDownloadNotActive when the artifact is not yet ready.
|
|
func (m *ArtifactManager) Ready(ctx context.Context, id string) (*Artifact, error) {
|
|
a, err := m.repo.GetByID(ctx, id)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
if a.Status != ArtifactReady {
|
|
return nil, fmt.Errorf("artifact is %s: %w", a.Status, ErrDownloadNotActive)
|
|
}
|
|
_ = m.repo.TouchLastUsed(ctx, id)
|
|
return a, nil
|
|
}
|
|
|
|
func (m *ArtifactManager) downloadConfig() config.DownloadConfig {
|
|
if m.liveCfg != nil {
|
|
if c := m.liveCfg(); c != nil {
|
|
return c.Download
|
|
}
|
|
}
|
|
return config.DownloadConfig{}
|
|
}
|
|
|
|
// artifactDir resolves the effective output directory for prepared artifacts,
|
|
// defaulting under the transcode dir when download.artifact_dir is unset so
|
|
// encodes never write relative to the process working directory.
|
|
func (m *ArtifactManager) artifactDir() string {
|
|
var artifactDir, transcodeDir string
|
|
if m.liveCfg != nil {
|
|
if c := m.liveCfg(); c != nil {
|
|
artifactDir = c.Download.ArtifactDir
|
|
transcodeDir = c.Playback.TranscodeDir
|
|
}
|
|
}
|
|
return effectiveArtifactDir(artifactDir, transcodeDir)
|
|
}
|
|
|
|
// Ensure deduplicates and (when new) enqueues an encode job for file in the
|
|
// given format, returning the current artifact row. The deterministic
|
|
// output_path keeps a reclaimed job idempotent.
|
|
func (m *ArtifactManager) Ensure(ctx context.Context, file *models.MediaFile, format string, target playback.PrepareTarget) (*Artifact, error) {
|
|
hash := paramsHash(format, target.Container, target.CodecVideo, target.CodecAudio, target.Resolution, target.AudioTrackIndex, target.TargetBitrateKbps, false)
|
|
id, err := idgen.NextID()
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
a := &Artifact{
|
|
ID: id,
|
|
MediaFileID: file.ID,
|
|
Format: format,
|
|
ParamsHash: hash,
|
|
Container: target.Container,
|
|
CodecVideo: target.CodecVideo,
|
|
CodecAudio: target.CodecAudio,
|
|
Resolution: target.Resolution,
|
|
AudioTrackIndex: target.AudioTrackIndex,
|
|
TargetBitrateKbps: target.TargetBitrateKbps,
|
|
OutputPath: artifactOutputPath(m.artifactDir(), file.ID, format, hash),
|
|
MaxAttempts: artifactMaxAttempts,
|
|
}
|
|
row, created, err := m.repo.EnsureQueued(ctx, a)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
if row.Status == ArtifactReady {
|
|
_ = m.repo.TouchLastUsed(ctx, row.ID)
|
|
return row, nil
|
|
}
|
|
// A terminally-failed dedup row would otherwise strand every new download
|
|
// linked to it in 'preparing' forever (no drain is triggered for an existing
|
|
// row). Requeue it for a fresh attempt so the new download can resolve — or
|
|
// fail cleanly via reconciliation once the encode is exhausted again.
|
|
if row.Status == ArtifactFailed {
|
|
switch err := m.repo.Requeue(ctx, row.ID); {
|
|
case errors.Is(err, ErrNotFound):
|
|
// The failed row was swept between EnsureQueued and Requeue:
|
|
// create a fresh job instead of linking to a dead artifact id.
|
|
if row, _, err = m.repo.EnsureQueued(ctx, a); err != nil {
|
|
return nil, err
|
|
}
|
|
case err != nil:
|
|
return nil, err
|
|
default:
|
|
row.Status = ArtifactQueued
|
|
}
|
|
m.triggerDrain()
|
|
return row, nil
|
|
}
|
|
if created {
|
|
m.triggerDrain()
|
|
}
|
|
return row, nil
|
|
}
|
|
|
|
func (m *ArtifactManager) triggerDrain() {
|
|
m.mu.Lock()
|
|
kick := m.kick
|
|
m.mu.Unlock()
|
|
if kick != nil {
|
|
// Ensure is called on request goroutines and the kick runs the encode
|
|
// task to completion (the task manager serializes concurrent runs), so
|
|
// it must never execute inline: a POST /downloads would otherwise block
|
|
// on the entire queue drain, ffmpeg encodes included.
|
|
go kick()
|
|
}
|
|
}
|
|
|
|
// RunOnce performs a startup-safe recovery sweep and then drains the queue
|
|
// until empty. It is safe to call concurrently across nodes; FOR UPDATE SKIP
|
|
// LOCKED prevents double-encoding.
|
|
func (m *ArtifactManager) RunOnce(ctx context.Context) error {
|
|
m.recover(ctx)
|
|
return m.drain(ctx)
|
|
}
|
|
|
|
// recover repairs state a crash or lost lease can leave behind: it reclaims
|
|
// expired-lease running rows (back to queued, or failed when attempts are
|
|
// exhausted), reconciles linked downloads against terminal artifact states, and
|
|
// re-queues ready artifacts whose output file is missing on disk. Safe to run
|
|
// repeatedly (each step is idempotent).
|
|
func (m *ArtifactManager) recover(ctx context.Context) {
|
|
if _, err := m.repo.ReclaimExpiredLeases(ctx); err != nil {
|
|
slog.WarnContext(ctx, "download artifact lease reclaim failed", "component", "downloads", "error", err)
|
|
}
|
|
|
|
// Reconcile downloads stranded in 'preparing' against their artifact's
|
|
// terminal state: this closes the non-transactional window between an
|
|
// artifact's MarkReady and its MarkLinkedDownloadsReady, and fails the links
|
|
// of any artifact that reached 'failed' (including the rows just reclaimed to
|
|
// failed above) so a download can never sit 'preparing' forever.
|
|
readyFlipped, failedFlipped, err := m.downloads.ReconcileLinkedDownloads(ctx)
|
|
if err != nil {
|
|
slog.WarnContext(ctx, "reconciling linked downloads failed", "component", "downloads", "error", err)
|
|
} else {
|
|
for _, d := range readyFlipped {
|
|
m.publish(ctx, d)
|
|
}
|
|
for _, d := range failedFlipped {
|
|
m.publish(ctx, d)
|
|
}
|
|
}
|
|
|
|
// Disk-presence sweep: stats every ready file, so it runs on the startup
|
|
// pass and then hourly rather than on every tick.
|
|
if !m.maintenanceDue(&m.lastDiskSweep) {
|
|
return
|
|
}
|
|
ready, err := m.repo.ListReady(ctx)
|
|
if err != nil {
|
|
slog.WarnContext(ctx, "download artifact ready scan failed", "component", "downloads", "error", err)
|
|
return
|
|
}
|
|
for _, a := range ready {
|
|
if a.OutputPath == "" {
|
|
continue
|
|
}
|
|
if _, statErr := os.Stat(a.OutputPath); statErr != nil {
|
|
slog.WarnContext(ctx, "download artifact output missing, re-queuing", "component", "downloads", "artifact_id", a.ID, "path", a.OutputPath)
|
|
if err := m.repo.Requeue(ctx, a.ID); err != nil {
|
|
slog.WarnContext(ctx, "re-queue artifact failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
continue
|
|
}
|
|
m.triggerDrain()
|
|
}
|
|
}
|
|
}
|
|
|
|
// drain claims and encodes jobs through a bounded worker pool until the queue is
|
|
// empty or the context is canceled.
|
|
func (m *ArtifactManager) drain(ctx context.Context) error {
|
|
maxConcurrent := m.downloadConfig().MaxConcurrentPrepares
|
|
if maxConcurrent <= 0 {
|
|
maxConcurrent = 2
|
|
}
|
|
sem := make(chan struct{}, maxConcurrent)
|
|
var wg sync.WaitGroup
|
|
for {
|
|
// Acquire a worker slot BEFORE claiming. A claimed job is leased but only
|
|
// heartbeated once encodeOne runs; claiming first and then blocking for a
|
|
// slot would leave the job leased-but-unattended, so its lease could lapse
|
|
// while it waits — letting another node steal it and encode the same
|
|
// output path concurrently. Reserving the slot first closes that window.
|
|
select {
|
|
case sem <- struct{}{}:
|
|
case <-ctx.Done():
|
|
wg.Wait()
|
|
return ctx.Err()
|
|
}
|
|
job, err := m.repo.ClaimNext(ctx, m.owner, artifactLease)
|
|
if err != nil {
|
|
<-sem // release the slot we reserved but won't use
|
|
if errors.Is(err, ErrNoArtifactJob) {
|
|
break
|
|
}
|
|
wg.Wait()
|
|
return err // includes context cancellation (pgx honors ctx)
|
|
}
|
|
wg.Add(1)
|
|
go func(a *Artifact) {
|
|
defer wg.Done()
|
|
defer func() { <-sem }()
|
|
m.encodeOne(ctx, a)
|
|
}(job)
|
|
}
|
|
wg.Wait()
|
|
return nil
|
|
}
|
|
|
|
// encodeOne runs one claimed job to completion, extending its lease via a
|
|
// heartbeat, and links/notifies the dependent download rows on the outcome.
|
|
func (m *ArtifactManager) encodeOne(ctx context.Context, a *Artifact) {
|
|
hbCtx, cancelHB := context.WithCancel(ctx)
|
|
defer cancelHB()
|
|
// heartbeatLoop cancels hbCtx if the lease is lost; PrepareFile runs on hbCtx
|
|
// so that cancellation aborts ffmpeg, ensuring we never keep writing the
|
|
// output path after another worker has taken the job.
|
|
go m.heartbeatLoop(hbCtx, cancelHB, a.ID)
|
|
|
|
file, err := m.fileRepo.GetByID(ctx, a.MediaFileID)
|
|
if err != nil || file == nil {
|
|
m.failJob(ctx, a, "source media file unavailable")
|
|
return
|
|
}
|
|
|
|
opts := m.buildOpts(file, a)
|
|
if err := m.preparer.PrepareFile(hbCtx, opts, a.OutputPath); err != nil {
|
|
switch {
|
|
case ctx.Err() != nil:
|
|
// Parent shutting down: leave the job 'running'; its lease expires and
|
|
// recovery (here or on another node) reclaims it.
|
|
return
|
|
case hbCtx.Err() != nil:
|
|
// We lost the lease mid-encode; another worker now owns the job.
|
|
slog.WarnContext(ctx, "download artifact encode aborted; lease lost", "component", "downloads", "artifact_id", a.ID)
|
|
return
|
|
default:
|
|
slog.WarnContext(ctx, "download artifact encode failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
m.failJob(ctx, a, err.Error())
|
|
return
|
|
}
|
|
}
|
|
|
|
var size int64
|
|
if fi, statErr := os.Stat(a.OutputPath); statErr == nil {
|
|
size = fi.Size()
|
|
}
|
|
// Fenced on lease ownership: if we lost the lease between encode and commit,
|
|
// applied is false and the current owner is responsible for flipping links —
|
|
// do not flip them here or we would race/duplicate that owner's work.
|
|
applied, err := m.repo.MarkReady(ctx, a.ID, m.owner, a.OutputPath, size)
|
|
if err != nil {
|
|
slog.ErrorContext(ctx, "marking artifact ready failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
return
|
|
}
|
|
if !applied {
|
|
slog.WarnContext(ctx, "download artifact ready skipped; lease lost", "component", "downloads", "artifact_id", a.ID)
|
|
return
|
|
}
|
|
flipped, err := m.downloads.MarkLinkedDownloadsReady(ctx, a.ID, size)
|
|
if err != nil {
|
|
slog.ErrorContext(ctx, "flipping linked downloads ready failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
return
|
|
}
|
|
for _, d := range flipped {
|
|
m.publish(ctx, d)
|
|
}
|
|
}
|
|
|
|
func (m *ArtifactManager) failJob(ctx context.Context, a *Artifact, msg string) {
|
|
terminal, applied, err := m.repo.MarkFailedOrRetry(ctx, a.ID, m.owner, msg, backoffFor(a.Attempts))
|
|
if err != nil {
|
|
slog.ErrorContext(ctx, "marking artifact failed/retry errored", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
return
|
|
}
|
|
if !applied {
|
|
// Lease lost; the current owner is responsible for the job's outcome.
|
|
return
|
|
}
|
|
if terminal {
|
|
m.failLinkedDownloads(ctx, a.ID, msg)
|
|
} else {
|
|
m.triggerDrain()
|
|
}
|
|
}
|
|
|
|
func (m *ArtifactManager) failLinkedDownloads(ctx context.Context, artifactID, msg string) {
|
|
flipped, err := m.downloads.MarkLinkedDownloadsFailed(ctx, artifactID, msg)
|
|
if err != nil {
|
|
slog.ErrorContext(ctx, "flipping linked downloads failed errored", "component", "downloads", "artifact_id", artifactID, "error", err)
|
|
return
|
|
}
|
|
for _, d := range flipped {
|
|
m.publish(ctx, d)
|
|
}
|
|
}
|
|
|
|
func (m *ArtifactManager) publish(ctx context.Context, d *Download) {
|
|
if m.notify != nil {
|
|
m.notify(ctx, d)
|
|
}
|
|
}
|
|
|
|
// heartbeatLoop extends the job's lease until ctx is done. If the lease is lost
|
|
// (another worker stole it, or the row is gone) it calls cancel to abort the
|
|
// encode so two workers never write the same output path. A transient DB error
|
|
// is retried on the next tick rather than aborting a healthy encode.
|
|
func (m *ArtifactManager) heartbeatLoop(ctx context.Context, cancel context.CancelFunc, id string) {
|
|
ticker := time.NewTicker(artifactHeartbeat)
|
|
defer ticker.Stop()
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
return
|
|
case <-ticker.C:
|
|
ok, err := m.repo.Heartbeat(ctx, id, m.owner, artifactLease)
|
|
switch {
|
|
case err != nil && ctx.Err() != nil:
|
|
return // encode finished or shutting down
|
|
case err != nil:
|
|
slog.WarnContext(ctx, "download artifact heartbeat errored", "component", "downloads", "artifact_id", id, "error", err)
|
|
case !ok:
|
|
slog.WarnContext(ctx, "download artifact lease lost; aborting encode", "component", "downloads", "artifact_id", id)
|
|
cancel()
|
|
return
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
func (m *ArtifactManager) buildOpts(file *models.MediaFile, a *Artifact) playback.TranscodeOpts {
|
|
cfg := config.Config{}
|
|
if m.liveCfg != nil {
|
|
if c := m.liveCfg(); c != nil {
|
|
cfg = *c
|
|
}
|
|
}
|
|
sourceVideoCodec, sourceVideoProfile, sourceVideoBitDepth := playback.SourceVideoTranscodeFacts(file)
|
|
return playback.TranscodeOpts{
|
|
InputPath: file.FilePath,
|
|
SourceVideoCodec: sourceVideoCodec,
|
|
SourceVideoProfile: sourceVideoProfile,
|
|
SourceVideoBitDepth: sourceVideoBitDepth,
|
|
TargetCodecVideo: a.CodecVideo,
|
|
TargetCodecAudio: a.CodecAudio,
|
|
TargetResolution: a.Resolution,
|
|
TargetBitrateKbps: a.TargetBitrateKbps,
|
|
AudioTrackIndex: a.AudioTrackIndex,
|
|
SubtitleTrackIndex: -1,
|
|
FFmpegPath: cfg.Playback.FFmpegPath,
|
|
HWAccel: cfg.Playback.HWAccel,
|
|
HWDevice: cfg.Playback.HWDevice,
|
|
TotalDuration: float64(file.Duration),
|
|
}
|
|
}
|
|
|
|
// Hygiene retention windows. These remove only rows nothing can serve again —
|
|
// terminally-failed jobs (linked downloads already flipped to failed by
|
|
// reconciliation) and ready artifacts whose every referencing download row was
|
|
// deleted — plus ephemeral web rows past their convenience-record lifetime.
|
|
// The server-disk *quota* is download.artifact_max_bytes (see the download
|
|
// limits & restrictions design); this sweep is not a quota.
|
|
const (
|
|
failedArtifactRetention = 24 * time.Hour
|
|
unlinkedArtifactRetention = 30 * 24 * time.Hour
|
|
ephemeralDownloadRetention = 7 * 24 * time.Hour
|
|
)
|
|
|
|
// Cleanup runs the hygiene sweep, then evicts ready artifacts (LRU first) once
|
|
// the total exceeds the byte budget, never removing one still linked by any
|
|
// active download row (managed or ephemeral) — only artifacts whose links are
|
|
// all terminal are evictable.
|
|
func (m *ArtifactManager) Cleanup(ctx context.Context) error {
|
|
m.sweepStale(ctx)
|
|
budget := m.downloadConfig().ArtifactMaxBytes
|
|
if budget <= 0 {
|
|
return nil // unlimited
|
|
}
|
|
total, err := m.repo.TotalReadyBytes(ctx)
|
|
if err != nil {
|
|
return err
|
|
}
|
|
if total <= budget {
|
|
return nil
|
|
}
|
|
candidates, err := m.repo.ListReady(ctx) // least-recently-used first
|
|
if err != nil {
|
|
return err
|
|
}
|
|
for _, a := range candidates {
|
|
if total <= budget {
|
|
break
|
|
}
|
|
active, err := m.repo.HasActiveLink(ctx, a.ID)
|
|
if err != nil {
|
|
slog.WarnContext(ctx, "artifact link check failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
continue
|
|
}
|
|
if active {
|
|
continue
|
|
}
|
|
if a.OutputPath != "" {
|
|
if err := os.Remove(a.OutputPath); err != nil && !os.IsNotExist(err) {
|
|
slog.WarnContext(ctx, "removing evicted artifact file failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
}
|
|
}
|
|
if err := m.repo.DeleteArtifact(ctx, a.ID); err != nil {
|
|
slog.WarnContext(ctx, "deleting evicted artifact row failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
continue
|
|
}
|
|
slog.InfoContext(ctx, "evicted download artifact (LRU)", "component", "downloads", "artifact_id", a.ID, "bytes", a.FileSize)
|
|
total -= a.FileSize
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// sweepStale is the age-based hygiene pass: cold terminally-failed artifacts
|
|
// (with their leftover .part files), orphaned ready artifacts no download row
|
|
// references, and expired ephemeral download rows. Best-effort; every step
|
|
// logs and continues.
|
|
func (m *ArtifactManager) sweepStale(ctx context.Context) {
|
|
if !m.maintenanceDue(&m.lastStaleSweep) {
|
|
return
|
|
}
|
|
now := time.Now()
|
|
if failed, err := m.repo.ListFailedBefore(ctx, now.Add(-failedArtifactRetention)); err != nil {
|
|
slog.WarnContext(ctx, "failed-artifact sweep list failed", "component", "downloads", "error", err)
|
|
} else {
|
|
for _, a := range failed {
|
|
m.removeArtifact(ctx, a, "failed")
|
|
}
|
|
}
|
|
if orphans, err := m.repo.ListUnlinkedReadyBefore(ctx, now.Add(-unlinkedArtifactRetention)); err != nil {
|
|
slog.WarnContext(ctx, "unlinked-artifact sweep list failed", "component", "downloads", "error", err)
|
|
} else {
|
|
for _, a := range orphans {
|
|
m.removeArtifact(ctx, a, "unlinked")
|
|
}
|
|
}
|
|
if m.downloads != nil {
|
|
if n, err := m.downloads.PruneEphemeralOlderThan(ctx, now.Add(-ephemeralDownloadRetention)); err != nil {
|
|
slog.WarnContext(ctx, "ephemeral download prune failed", "component", "downloads", "error", err)
|
|
} else if n > 0 {
|
|
slog.InfoContext(ctx, "pruned expired ephemeral downloads", "component", "downloads", "rows", n)
|
|
}
|
|
}
|
|
}
|
|
|
|
// removeArtifact deletes an artifact's output file, its .part leftover, and
|
|
// its row. Used by the hygiene sweep for rows nothing can serve again.
|
|
func (m *ArtifactManager) removeArtifact(ctx context.Context, a *Artifact, reason string) {
|
|
if a.OutputPath != "" {
|
|
if err := os.Remove(a.OutputPath); err != nil && !os.IsNotExist(err) {
|
|
slog.WarnContext(ctx, "removing swept artifact file failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
}
|
|
if err := os.Remove(a.OutputPath + ".part"); err != nil && !os.IsNotExist(err) {
|
|
slog.WarnContext(ctx, "removing swept artifact partial failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
}
|
|
}
|
|
if err := m.repo.DeleteArtifact(ctx, a.ID); err != nil {
|
|
slog.WarnContext(ctx, "deleting swept artifact row failed", "component", "downloads", "artifact_id", a.ID, "error", err)
|
|
return
|
|
}
|
|
slog.InfoContext(ctx, "swept stale download artifact", "component", "downloads", "artifact_id", a.ID, "reason", reason, "bytes", a.FileSize)
|
|
}
|
|
|
|
// backoffFor returns the retry delay for the next attempt after a failure.
|
|
func backoffFor(attempts int) time.Duration {
|
|
if attempts < 1 {
|
|
attempts = 1
|
|
}
|
|
d := time.Duration(attempts) * 30 * time.Second
|
|
if d > 5*time.Minute {
|
|
d = 5 * time.Minute
|
|
}
|
|
return d
|
|
}
|