Commit Graph
5 Commits
Author SHA1 Message Date
881c96864b feat(playback): finalize platform-neutral protocol v3 (#567)
* docs(playback): add v3 neutral-contract finalization plan

Supersedes the wire-contract sections of the 2026-07-12 v3 plan: server-owned
attempt keys, delivery-keyed negotiation without Media3 engine names, tiered
capability evidence, neutral device/output context, track/quality replan
operations, audio-only planning, and coordinated no-back-compat rollout
across server, Android, Apple, and web.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(playback): make v3 attempt keys server-owned and replace engines with deliveries

Contract core of the platform-neutral v3 finalization (plan sections 3.1
and 3.2), breaking on purpose — v3 is dark and all clients move together:

- Every PlanV3 now carries plan_attempt_key, an opaque server-computed
  token clients store and echo in attempted_plan_keys; ReplanRequestV3
  gains bounded local_mutations that the replan handler folds into the
  failed plan's key. Clients never hash anything.
- KotlinName() is deleted from DeliveryV3, StreamProtocolV3 and
  SubtitleModeV3; the attempt-key canonical string now uses lowercase
  wire tokens, and PlanRecipeVersionV3 bumps to v3.3 so no key or plan
  ID computed under the old canonicalization can collide.
- EngineV3 leaves the wire: ClientPlaybackContextV3.Engines (media3_*)
  becomes Deliveries keyed original_http|progressive|hls, with
  EngineCapabilityV3 renamed DeliveryCapabilityV3. PlanV3.Engine is
  removed; the planner, subtitle policy and quirk registry re-key on
  delivery class, and the media3_only feature token is deleted.
- Validated-claim strings drop the prefix: media3_h264_decode ->
  h264_decode, media3_audio_decode -> audio_decode.
- Golden fixtures in testdata/protocol_v3 are regenerated by Go and are
  now the cross-repo source of truth.

Part of the playback protocol v3 neutral-contract train (steps 2-3 of
docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(playback): add v3 evidence tiers and neutral device/output context

Implement plan sections 3.3 and 3.4 of the v3 neutral-contract pass:

- ClientCodecCapabilitiesV3 gains required video_evidence and
  audio_evidence closed enums (exact | platform_attested | declared).
  Planner strictness follows the tier: exact keeps the strict decode-entry
  validation, platform_attested validates codec/resolution/bit-depth/
  frame-rate but skips profile/level matching, declared grants copy routes
  from the flat codec lists. Only exact audio evidence earns passthrough
  claims. The detailed_decode_capabilities feature token is deleted
  (subsumed by video_evidence=exact), and evidence-blocked direct routes
  carry the new evidence_insufficient_for_direct reason/warning.

- DeviceContextV3 is now platform/os_version/manufacturer/model plus a
  bounded platform_details map (<=16 entries, <=128 chars); the Android
  Build dump fields are gone. Fire TV quirks keep matching on
  manufacturer/model (brand fallback removed with the field).

- output_route_generation (int64, dual-location) becomes an optional
  opaque output_context_id string on the output context; the dual-location
  consistency validation is deleted. Attempt keys, plan invalidation,
  route events, and the planstore column follow (new Goose migration).

- Feature advertisement collapses to the top-level client_features list
  only; ClientPlaybackContextV3.Features is deleted and ReplanRequestV3
  gains an optional client_features refresh.

- PlanRecipeVersionV3 bumped v3.3 -> v3.4; fixtures re-keyed.

Part of the playback protocol v3 neutral-contract finalization plan
(docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(playback): add v3 intent replans, quality menu, and audio-only routes

Protocol v3 could only replan after a failure, so changing the audio track
or the quality still required the legacy audio PATCH and the client-recipe
transcode start — the two endpoints v3 is meant to replace. Clients also had
to own a resolution ladder to render a quality menu, and a source with no
video track was terminaled by the video/HDR gates, keeping audiobooks on the
legacy path.

Add track_change and quality_change replan operations. They carry no failure
classification and route through the existing replan transaction, so they
inherit its idempotency, capacity reservation, and staged-successor commit
for free. Because nothing failed, the previous route stays eligible: neither
the attempted-key history nor the failed-plan exclusion applies to them.

Publish the server ladder on the plan as available_qualities so the quality
menu is server-owned; the rungs come from the same resolutionLabelV3 and
ladderBitrateKbpsV3 helpers the planner itself uses, not a parallel table.

Plan audio-only sources through their own reduced route family: original_http
when the client decodes the codec, otherwise a progressive AAC conversion.
The plan advertises audio/mp4 for that remux and the transport now serves the
same value, because a declared-tier client probes the advertised MIME with
isTypeSupported before attaching a source buffer, and "video/mp4" on a stream
with no video track is exactly the mismatch that makes the probe lie.

Name the protocol's string vocabulary (dynamic ranges, transformations,
executors, validated claims, terminal reasons) as constants while touching
these lines, so the wire values have one definition.

Part of #135

* docs(playback): publish the v3 protocol contract and fix subtitle ordinals

Protocol v3 exists only as Go code today, so the Android and Apple ports have
no authority to implement against other than reading this repository. Publish
the contract as a normative document, machine-checkable schemas, and generated
golden fixtures, and fix the one place where the server's own wire output
disagreed with the ordinal space it publishes.

- docs/architecture/playback-protocol-v3.md is self-contained enough for a
  third-party client: endpoints and status codes, evidence tiers and their
  bound-matching rules, delivery classes, the timeline model, replan
  semantics, registries, track identity, plan identity, quality, and
  transformations.
- docs/design/schemas/playback-v3/ carries JSON Schemas for the five wire
  shapes plus valid and invalid fixtures, following the client-diagnostics
  layout. internal/playback/contract validates every fixture against its
  schema, so a schema that drifts from the Go types fails the Go suite.
- cmd/playbackfixtures generates internal/playback/testdata/protocol_v3 from
  the production planner. `make playback-fixtures` writes them and
  `make verify-playback-fixtures` (wired into CI) fails when they are stale.
  These files are what the client ports consume, so drift would otherwise
  surface as a playback bug on three platforms at once.

The subtitle fix: combined ordinals are one dense space over externals, then
embedded tracks, then downloaded ones, but the legacy URL builder skipped
burn-in-only tracks while assigning indices, so every track after a DVD/DVB
track was numbered one too low and resolved to its neighbour. Ordinal
assignment now lives in playback.BuildSubtitleInventoryV3 and both the plan
inventory and the legacy `subtitle_urls` shape project from it; the legacy
shape still filters burn-in-only entries but keeps each track's real index.

Part of #135

* feat(web): migrate the players to the neutral playback v3 contract

The web player was the last client still speaking the legacy start
protocol: it picked its own file version from a codec probe, posted an
ffmpeg recipe to start a transcode, PATCHed an endpoint to change audio
tracks, and derived its own quality ladder. None of that survives a
server-owned plan, and none of it produced telemetry the apps could be
compared against.

Video player: starts with a v3 request that advertises `declared`
evidence from `isTypeSupported` probes and the three delivery classes,
then consumes the returned plan for its URL, timeline, tracks and
warnings. Quality and track changes become replans (`quality_change`,
`track_change`), the quality menu renders `available_qualities` instead
of computing rungs, and playback failures emit `route-events` so web
failures land in the same diagnostics as Android and Apple. The
duration comes from `source.duration_seconds` rather than the playback
engine, and the "how was this delivered" overlay reads the plan's
delivery and server transformations instead of comparing codec strings.

Audiobook player: starts against the audio-only planner path with a
single `original` rung, and takes its seek anchor from
`timeline.player_start_seconds` so the progressive-remux route (which
anchors the stream and restarts the player clock at zero) does not seek
twice.

Server side, `disable_progress_persistence` left the wire, so the rule
it encoded is now derived. Resume state is keyed on the item, but every
part of a multipart presentation shares that key while carrying its own
file-local clock — persisting part 4's position would store "12 minutes
in" as the book's resume point. `PresentationPartTotal > 1` expresses
that directly and generalizes to multipart movies and split episodes,
and a client can no longer forget to ask or lie about it.

`useTranscodeQuality` and the legacy response types are deleted, and
`WEBTEST_KNOWN_FAILURES` loses the audiobook entry along with its fix.

Part of #135

* feat(playback)!: make v3 the only playback protocol

Protocol v3 shipped behind a flag, alongside the legacy start path it was
designed to replace. Running both meant every planner change had to be made
twice, in two shapes that disagree about who decides the route: the legacy
body carried a decision the client had already made, while v3 asks the server
to make it. This deletes the legacy half.

Removed:

- `handleStartPlaybackLegacy` and its request/response bodies. The
  `POST /playback/start` route stays, but the protocol-version dispatch
  envelope is now a strict v3 decode — a body that does not declare
  `protocol_version: 3` gets `426 client_upgrade_required` so an outdated app
  can render a clear "update required" state instead of misreading a plan.
  Deliberately not a `400`: the request may be well-formed for the protocol it
  was written against.
- `POST /playback/transcode/start`, superseded by the `quality_change` replan
  operation, and `PATCH /playback/{session_id}/audio`, superseded by
  `track_change`. Both mutated a session without re-planning.
- The shadow planner and both rollout settings rows. With v3 the only
  protocol, `playback.protocol_v3_enabled` would mean "no playback at all";
  `playback.protocol_v3_shadow_enabled` gated a comparison against a path that
  no longer exists. `409 protocol_disabled` on route-events goes with them, and
  capability `enabled` is now constant `true` (the field stays — clients
  feature-detect against it).
- Version-selection helpers in `internal/playback/resolver.go` that only legacy
  start reached. `Resolve`/`ClientCapabilities`/`PlayDecision` stay: downloads
  consumes them. `internal/jellycompat` has its own resolution surface and is
  untouched.

Behaviour the legacy handlers owned and v3 now owns explicitly: series version
and audio-track preferences are persisted on start and on a `track_change`
replan (not on failure recovery, whose forced route is not a user choice); an
omitted `start_position` resolves to the profile's saved resume point; and an
omitted audio track resolves through the series preference, the profile audio
language, then the library override. Both are settled before planning, because
the plan's timeline is cut at the start position. Spec §2.2 documents this as
"omission is a request, not a default".

The encode-target clamp that lived in the deleted transcode handler is already
enforced in the planner, twice — `availableQualitiesV3` omits rungs at or above
the source height, and the encode path clamps `targetHeight` to it.

Unchanged: progress, stop, HLS manifest and segment delivery, the realtime
control socket, stream tokens and restart reconstruction, watch together,
downloads, jellycompat.

Every removal is recorded in the pre-lock removals table in
docs/architecture/v1-scope.md.

Part of #135

* fix(scanner): stop recording embedded cover art as a video track

ffprobe reports embedded cover art as a video stream carrying
disposition.attached_pic. convertProbeData appended every "video" stream
to VideoTracks without consulting isMainVideoStream, the predicate that
already existed for duration decisions, so the picture was persisted as a
playable track. That misreports the file twice:

  - An audio file with a cover picks up a video track, so it no longer
    satisfies MediaFile.IsAudioOnly and the v3 planner routes an
    audiobook through the video path instead of planAudioOnlyV3.
  - When the picture is ordered ahead of the real stream, the flat
    codec_video/resolution/hdr columns describe the poster: a 954x720
    h264 episode was stored as mjpeg 480x480.

Filter attached_pic streams out of the track loop. The guard is the
disposition flag, not the codec name, so a genuine MJPEG video is still
probed as video — the library has one.

Already-probed rows self-heal on the next playback: NeedsCriticalProbeRepair
already reprobes tracks missing color_range, which covers 21 of the 23
affected rows, and applyProbeData overwrites VideoTracks wholesale. The
remaining two need a rescan; nothing persisted records attached_pic, and
keying repair off still-image codec names would reprobe the genuine MJPEG
file on every playback forever.

Part of the playback v3 neutral-contract work: it is what lets Android
drop AUDIOBOOK_COVER_ART_CODECS, which fabricated decode support the
client cannot honestly claim under video_evidence: "exact".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(playback): publish subtitle URLs even when playback starts with subtitles off

The v3 plan's subtitle inventory is the authoritative track list a client builds
its subtitle menu from, but the handler only rewrote it with session-scoped URLs
when a track was actually selected. A start or replan that resolved to
`subtitle.mode: "off"` therefore returned the planner's URL-less inventory, so a
client whose picker reads the inventory had a menu it could not fetch anything
from. The Cast path hits this every time: it starts with subtitles off and needs
the receiver's text tracks up front.

attachSubtitleArtifactV3 now scopes and publishes the inventory unconditionally
and gates only the artifact stamping on the selection. Spec §8 records that the
`url` on a sidecar entry does not depend on the current selection.

Part of the v3 neutral-contract finalization.

* chore(playback): reconcile neutral v3 with main

* fix(playback): preserve subtitle intent across replans

* fix(playback): retain subtitle inventory on adapted routes

* fix(playback): software-decode High10 AVC for QSV

* fix(playback): scale High10 frames before QSV upload

* fix(playback): preserve empty subtitle inventories

* fix(playback): freeze terminal attempt contract

* chore(playback): name fixture contract tokens

* fix(playback): close v3 conformance review gaps

* chore(playback): name conformance category

* fix(playback): complete v3 conformance contract

* fix(playback): keep schema fixtures generated

* fix(playback): emit schema-valid conformance arrays

* fix(playback): omit empty replan failures

* fix(web): omit empty replan failures

* fix(playback): close neutral v3 contract gaps

* fix(playback): harden v3 replan, transcode, and quality-ladder edge cases

Review remediation for the neutral v3 cutover, server side:

- A failed replan no longer overwrites the durable StartResponse with a
  terminal or advances the replan request ID; an idempotent start replay
  of a still-healthy session returns the original plan.
- SoftwareVideoDecode is now derived inside the transcode layer from
  source facts (codec/profile/bit depth) carried on TranscodeOpts, so
  jellycompat, downloads, recipe-card reconstruction, and transcode
  nodes get the High10 software-decode fix, not just the v3 handler.
  video_to_h264 recipe version bumps to 2 so mixed-version node pools
  that would silently drop the flag fail validation instead.
- Local transport startup shares the 30s ManifestStartupTimeout; a
  timeout with the process still running stays retryable and is no
  longer persisted as a durable terminal against the attempt.
- Sparse replan bodies (failure_recovery et al) no longer reset a
  user-selected quality preference to auto; the empty-value guard now
  covers every operation.
- availableQualitiesV3 publishes no fixed rungs when the source height
  is unknown, keeping the no-upscaling ladder contract.
- The proxy remux path serves audio-only fMP4 as audio/mp4 via a new
  additive AudioOnly token claim, matching the integrated path.
- Plain text subtitle sidecars accept any requested extension again
  (served as VTT), restoring the permissive v1 behavior; ASS and bitmap
  handling is unchanged.
- The 4K-disallowed terminal message discloses when a lower-resolution
  alternate exists but was pinned away by quality "original".

Part of #135.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): keep playback alive through failed replans and honest audio claims

Review remediation for the neutral v3 cutover, web player:

- A failed or refused replan no longer unmounts the player: the fatal
  error screen is reserved for loads with no adopted plan, and replan
  failures surface through the existing non-fatal replanError path.
- changeQuality rolls its optimistic preference back when the replan is
  refused or errors, so a failed switch is not silently applied by the
  next unrelated replan and the menu shows the real active rung.
- The capability probe now tests mp3/vorbis codecs and mp3/flac/ogg
  containers (MediaSource with a canPlayType fallback), restoring
  direct play for mp3 audiobooks instead of per-part AAC re-encodes.
- Reanchor seeks issued while a replan is in flight coalesce and run
  when it settles instead of being silently dropped with the scrubber
  pinned to a phantom position.
- Subtitle refresh/translation replans use the resume anchor while the
  media element has no metadata, so a subtitle_ready broadcast during
  startup no longer restarts a resumed stream at 0:00.
- An exhausted failure-recovery chain sets a visible error instead of
  returning silently.

Part of #135.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): accept video-only and VP9 probe metadata

Treat audio and video probe completeness independently so legitimate video-only assets converge without repeated ffprobe repair. Allow unknown codec profile/level metadata to fall through to server adaptation while preserving exact direct-decode constraints.

Fixes #574

* fix(playback): address protocol v3 review findings

* fix(playback): harden lease and probe repair decisions

* fix(playback): close remaining v3 review gaps

* fix(playback): recover failed transcode starts

* fix(playback): address remaining review-bot findings on v3 replan and audio planning

Server:
- The deferred replan lease release is bounded by a 3s timeout so a
  saturated pool or DB outage cannot wedge a handler goroutine that
  holds the per-session store lock on an uncancellable context.
- planAudioOnlyV3 honors the request bandwidth cap: an over-cap source
  skips the original_http direct route and converts to AAC with the
  same bandwidth_cap_applied warning and decision reason the video
  ladder uses. Unknown source bitrate never triggers the cap.
- A copy-audio progressive plan rejected only by a per-delivery
  audio_decode_codecs subset retries as an AAC conversion instead of
  returning adaptation_unavailable, and the AAC recipe respects the
  delivery's max_channels.

Web:
- failure_recovery replans issued while another replan is in flight
  queue (superseding a pending seek reanchor) instead of being
  silently dropped with the fatal overlay already suppressed.
- A terminal response to a fresh non-preserving start clears the
  previous plan and stops its session, so episode navigation cannot
  keep rendering the prior item under the new title.
- A refused recovery replan for a transport-dead plan surfaces the
  error and re-arms the plan failure key, so transient recovery
  failures no longer strand an endless spinner; the audiobook player
  gets the same guard reset.
- A track-less subtitle_translation_completed hands off to the
  refreshed persisted track once the inventory settles, clearing the
  live overlay, instead of pinning the synthetic live track forever.

Part of #135.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): reuse HLS transport for sidecar replans

* fix(playback): stabilize copy HLS remount timeline

* fix(playback): address v3 review findings

* fix(playback): satisfy player contract types

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 18:14:49 -04:00
28c6ddc237 feat(playback): add per-user transcoding controls (#375)
* feat(playback): add per-user transcoding controls

* fix(playback): enforce forced video transcode permission

* chore: address transcode control review feedback

* fix(playback): recheck transcode permission on audio switch

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-10 22:30:04 -04:00
42602b7896 feat(policy): access groups + embedded OPA policy engine with decision audit log (#282)
* docs(policy): add OPA policy engine design spec and implementation plan

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* build(deps): add OPA v1.18.2 SDK for the policy engine

Pulls github.com/open-policy-agent/opa v1.18.2 (policy engine core for
the upcoming internal/policy subsystem) and the transitive upgrades go
mod tidy applied (otel 1.44, grpc 1.81.1, prometheus/common 0.67.5).
Full build verified.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add OPA engine core, vendor scope policy, and parity suite

New internal/policy package (dead code — nothing wires into request paths
yet): prepared-query Engine with 25ms eval timeout and fail-closed decode,
typed PDP.ResolveViewerScope, go:embed vendor bundle, capabilities lockdown
for future admin-authored Rego, and vendor scope.rego reproducing
access.Resolver.Resolve (library intersection, disabled-library handling,
quality/rating ceilings) with a narrowing-only silo_custom.scope.override
extension hook.

Parity proven by 1368 dual-execution subtests against the real
access.Resolver, including the nil-vs-empty AllowedLibraryIDs battery and
quality/rating variation; rank tables are test-pinned to internal/access.
Rego unit tests run via opa/v1/tester inside go test. Bench:
~106µs/op per scope decision incl. input marshaling.

Also restores the OPA requirement to go.mod (the earlier deps commit ran
go mod tidy before any import existed, so tidy dropped it).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed,
corrected (quality.allowed raw-file-rank divergence), and verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add policy document store, foundation schema, and compile-check

policy_foundation migration: policy_documents (one enabled doc per domain
via partial unique index — two enabled docs would define override twice
and conflict at eval), immutable policy_document_versions, single-row
policy_generation counter, and the partitioned policy_decisions log table
(daily range partitions, no FK, denial partial index).

PolicyStore: transactional version numbering (FOR UPDATE), activation
that verifies compiled_ok and bumps the generation in the same tx,
enable/disable with typed ErrDomainAlreadyEnabled, and a delete guard for
documents with an active version. CompileCheck sandboxes admin Rego:
locked capabilities (no http.send/net.*/opa.runtime), enforced
silo_custom.<domain> package path, vendor+stub layering, 2s budget,
structured row/col errors. Engine gains NewEngineWithCustom /
NewEngineFromStore with WARN-and-skip for invalid custom rows.

DB-backed tests verified against a migrated Postgres (concurrent version
numbering, atomic generation bumps, activation guards).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here (domain constants extracted).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add policy System lifecycle with hot reload and cross-node invalidation

policy.System owns one long-lived Engine and reloads it in place when
policy documents change: EventPolicyChanged on the existing ChannelAdmin
bus (new cache event constant) plus a 60s generation-poll fallback for
Redis-less deployments, with a generation-consistent snapshot read.
Vendor compile failure is startup-fatal; store/custom failures degrade
to vendor-only and the poll loop heals them; runtime reload failures
keep the last known-good engine. NotifyChanged gives the future admin
handlers synchronous local reload + cross-node publish.

Wiring: constructed in integrated/api modes only, PolicySystem field on
api.Dependencies (unused by routes yet), policy.eval_timeout_ms setting
(hot-reloaded via configWatcher.OnChange; default 25ms). Verified by a
full server boot smoke and DB-backed convergence tests (event + poll
paths, degraded boot, last-known-good).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): add async decision logging with sampling, retention, and query repo

DecisionLogger batch-inserts each node's policy decisions straight to
the partitioned policy_decisions table via a non-blocking buffered
channel (drop-and-count on overflow — logging never adds latency to or
fails a decision). Scope decisions sample 1-in-N (default 50, setting
policy.decision_log_scope_sample_rate); denials and eval errors always
log; input/result JSON samples only at policy.decision_log_verbosity=
verbose. Cursor-paginated DecisionRepository backs the upcoming admin
log viewer. Retention via partman (daily partitions) and a
PolicyDecisionLogCleanupTask honoring policy.decision_log_retention_days
(default 14). PDP emits entries per evaluation; the System owns the
logger lifecycle and settings hot-reload.

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here (removed an unused, unsynchronized PDP setter).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(api): add admin policy management API and capability endpoint

/api/v1/policy/capability (authenticated feature detection) plus the
acting-admin /api/v1/admin/policy surface: vendor Rego viewer, document
CRUD with the one-enabled-per-domain conflict mapped to 409, immutable
version creation (compile-checked; failed versions persist as audit
history with structured row/col errors and can never activate),
activate/rollback with synchronous reload + cross-node invalidation via
System.NotifyChanged, stateless validate, throwaway-bundle simulate
(never touches the live engine, never logs decisions), and
cursor-paginated decision-log queries. Routes mount only when the
policy system is wired, keeping proxy/transcode modes untouched.

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here (seeded the FK'd test user; replaced an unchecked
fmt.Sscanf with strconv).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add /admin/policy workspace with Rego editor, simulate, and decision log

New Policy admin page (System nav group): documents list with
one-enabled-per-domain conflict handling, CodeMirror 6 Rego editor
(hand-rolled StreamLanguage mode) with server compile issues rendered as
inline lint diagnostics, explicit Save-version vs Activate flow with
confirm, read-only vendor module viewer, simulate panel with seeded
example inputs, version history with rollback, and a cursor-paginated
decision-log browser. Capability-gated via /policy/capability. Adds the
three decision-log settings to Log Retention. First code-editor
dependency in web/ (@uiw/react-codemirror + @codemirror/*), decided in
the design spec.

Implementation drafted by Codex (GPT-5.5) via codex exec; verified here
(lint, format:check, tsc --noEmit, vitest policy suites).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): make OPA authoritative for viewer scope resolution

policy.ViewerResolver implements the ViewerResolver interface backed by
PDP.ResolveViewerScope and replaces access.Resolver at all five
construction sites: router viewer middleware, notifications scopes, the
reconciler, jellycompat's scope filter, and the ABS resolver (which now
accepts a pre-built resolver, preserving its PIN-at-login semantics).
PIN/profile-token verification and disabled-library loading are
extracted into shared exported helpers used by both implementations, so
the legacy resolver stays compiled as the parity reference with
identical behavior. The adapter lives in internal/policy (which already
depends on internal/access transitively) — direct typed PDP calls, no
new import cycle. Sites without a policy system (proxy modes, bare test
routers) keep the legacy resolver until the cleanup phase.

Verified: full test suite green (jellycompat TestBeginWebOperation* and
one playback GPU test are pre-existing failures, confirmed identical on
main), 1368-case parity suite, dedicated ViewerResolver parity/PIN/
nil-vs-empty/fail-closed tests, and a full server boot smoke.

Implementation drafted by Codex (GPT-5.5) via codex exec; a first-pass
reflection-based adapter was rejected and reworked into the typed
in-policy adapter; reviewed line-by-line and verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): make OPA authoritative for acting-admin and permission gates

vendor/permission.rego reproduces the acting-admin rule (admin role +
primary-profile-or-none), HasEffectivePermission semantics for
marker_edit, and the metadata-curation rule including the subtle
admin-past-refused-bypass case that requires the explicitly ASSIGNED
permission. Policy-backed middleware in policy_gates.go keeps all Go-side
lookups (declared-profile primary check, item->library resolution, the
404-on-unknown-item path) and preserves the legacy status/body taxonomy
exactly — proven by dual-execution middleware tests that run every
scenario through both implementations and assert byte-equal responses.
Permission decisions always log (allowed flag populated); simulate and
the capability endpoint gain the permission domain automatically via the
domain registry. Router swaps behind single constructor choice points
with the legacy gates retained for policy-less wiring.

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed and
verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(policy): make OPA authoritative for download and playback admission decisions

vendor/action.rego decides download eligibility (downloads enabled +
user allowed), download-transcode eligibility (transcode enabled + user
allowed + artifacts available), and playback admission (stream/transcode
counts vs limits, zero = unlimited), with a tightening-only
silo_custom.action override that can also clamp a quality ceiling (never
widen — merged via quality.min). Go keeps everything stateful: config
loading, preset-ladder enumeration, and live session counting.

Downloads consult an optional ActionDecider (nil = legacy logic) mapped
back to the existing sentinel errors and capability response. Playback
gains a minimal AdmissionDecider hook at the exact point of the legacy
limit comparison: counts snapshot under the session mutex, PDP evaluated
OUTSIDE the lock, then revalidated under lock before insert (retry on
count drift) — no admission ever decided on stale counts and no eval
under the mutex. Deny reasons map to the legacy ErrTooManyStreams /
ErrTooManyTranscodes sentinels, pinned by tests.

Parity: combination tables driven against the real PresetsFor /
ensureTranscodeAllowed / SessionLimits math; full suite green (known
pre-existing jellycompat flakes only).

Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed
(locking design verified line-by-line) and verified here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): satisfy tsc -b strict return typing in the Rego stream tokenizer

The production build (tsc -b) rejects assigning CodeMirror's
string | void next() result to string | undefined; tsc --noEmit did not
catch it. Restructured the string-literal loop.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): clearer error when a decision is undefined for partial input

Vendor policies index required input fields directly, so a hand-written
simulate payload missing fields yields an undefined decision. Surface
that as 'decision X is undefined for this input (missing required input
fields?)' instead of 'empty result' — found while exercising the
simulate API against a live server.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(web): set changeOrigin automatically when the API proxy target is remote

Remote dev backends sit behind vhost-routing proxies that reject a
localhost Host header; local targets keep the existing pass-through
behavior. Enables pointing the Vite dev server at a hosted backend via
VITE_API_PROXY_TARGET in web/.env.local.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): redesign the policy workspace around the decision pipeline

The first-pass UI was structurally generic: a five-column document table
squeezed beside the editor, three equal-weight action buttons with
hidden preconditions, raw version IDs, and jargon copy — nothing taught
the model. The page now teaches it:

- A pipeline strip states the mental model up front: Silo decides the
  baseline -> your overrides narrow it -> every decision is logged. Tabs
  renamed to Overrides / Baseline / Decision Log (ids stay stable for
  bookmarked URLs).
- The document table becomes one card per domain (Library visibility /
  Admin & permissions / Downloads & playback) with plain-language
  descriptions, example rules, status pills (Live vN / Draft / Disabled),
  inline creation, and the enable kill-switch in place.
- Selecting an override drills into a full-width editor with a visible
  lifecycle rail (Draft -> Validated -> Saved -> Live) and one contextual
  primary action per step; the unedited live source shows no actions
  until edited. Version comments appear only at the save step.
- Simulate is reframed as 'Test before going live' with a human verdict
  chip (Allowed / Denied — reason / ceiling summary) above the raw JSON;
  internal generation counters no longer surface.
- History uses 'Make live' with plain go-live copy; authors read
  'User N'; the baseline tab explains that upgrades never touch
  overrides.

Hand-written redesign (no Codex); verified via vitest, tsc, eslint,
prettier, and a production build.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(web): present the policy baseline as readable rules, not raw Rego

The Baseline tab dumped five Rego modules into read-only editors. It now
leads with what the rules actually do: one card per domain with
plain-language statements of the shipped behavior and a note on what an
override may change, plus content-rating and playback-quality tier
ladders parsed live from the lib module sources (so the tiers shown are
the ones the server enforces, not a hardcoded copy). The Rego source
stays one click away behind a per-module accordion and remains the
stated source of truth; unrecognized modules fall back to source-only.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(policy): add access-groups design addendum

Groups with permission toggles become the everyday admin surface; the
Rego editor is demoted behind policy.editor_enabled (default off).
Restriction-only composition: group grants are an upper bound, per-user
settings tighten further — same rule as the existing account/profile
merge, one layer up.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): add access groups — group defaults with restriction-only composition

New access_groups table + users.access_group_id (one group per user, NULL
= today's behavior). Group grants are an upper bound composed with the
user's own settings by strictest-wins rules — library intersection,
MinQuality, AND'd booleans, strictest positive stream/transcode limits,
permission-mask intersection, and a requests toggle gating CreateRequest.
The merge happens in Go (access.ApplyGroupPolicy /
EffectivePolicyForUser) before policy inputs are built, so vendor Rego,
the parity suites, and the decision log are untouched; every enforcement
surface (viewer scope in both resolvers, permission gates, downloads,
playback admission, requests) consumes the effective policy and fails
closed on provider errors. Changing a group's quality ceiling bumps its
members' access_policy_revision, mirroring the per-user rule.

Additive admin API: /admin/access-groups CRUD with member counts;
PUT /admin/users/{id} + user DTOs gain access_group_id.

Also demotes the Rego editor: policy.editor_enabled (default off,
hot-reloaded) drives the capability endpoint's editor_available and
403-gates editor endpoints while the engine and decision logging keep
running.

Design: docs/superpowers/specs/2026-07-02-access-groups-design.md.
Implementation drafted by Codex (GPT-5.5) via codex exec; reviewed
(composition core + fail-closed call-site audit) and verified here.
DB-backed group-store tests pending local Postgres recovery.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): add Access Groups admin page and gate the policy editor

New /admin/access-groups: a card grid summarizing each group (member
count + key restrictions), drilling into an editor that reuses the same
LibraryAccessSelector and quality presets as the user editor, with
toggles for downloads/transcoded-downloads/requests, concurrent-stream
and transcode limits, and a permissions mask (all-assignable by default,
narrowable to specific permissions). Delete warns how many members fall
back to the built-in defaults. Copy states the composition rule up front:
a group grants the most a member can do; their own restrictions still
apply on top.

The user editor gains a Group picker and read-only row; the Policy nav
entry is now hidden unless the capability reports the editor enabled.
Plumbing (types, hooks, user-editor picker, nav gating) drafted by Codex
(GPT-5.5); the Groups page hand-built. Verified: 25 tests across the
touched suites, tsc, eslint, prettier, and a production build.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): seed a Default Group and auto-assign newly created users

Adds access_groups.is_default with a partial unique index (one default
at most — the profiles is_primary pattern) and seeds a permissive
'Default Group' whose ceiling is a no-op, so assignment never changes
anyone's effective access until an admin edits it. The seed is guarded
against pre-existing defaults and name collisions; the Down migration
only removes the row if it is still untouched.

Assignment happens at the single INSERT INTO users choke point
(UserRepository.Create): when no explicit group is given, access_group_id
is filled by a scalar subquery on the default flag — NULL when no default
exists. Every creation path (setup, signup, invites, OAuth, admin create)
is covered by construction. Setting a new default via the API atomically
clears the previous one in the same transaction.

Deleting or unsetting the default is legal: new users then start with no
group, which is pre-feature behavior.

Implementation drafted by Codex (GPT-5.5); migration guards and the
choke-point subquery reviewed line-by-line here.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): surface the default access group

Cards show a Default badge; the group editor gains a 'Default for new
users' toggle (with copy noting existing users are never moved); the
delete dialog warns when removing the default that new accounts will
start with no group.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): ship the Default Group with house-rule ceilings

Seed values per product decision: 5 concurrent streams, 5 transcodes,
transcoded downloads off, and a permission mask of marker_edit only
(metadata curation excluded). Plain downloads and requests stay on. The
Down guard matches the new values so it still only removes an untouched
seed row. Only newly created users are affected; existing users are
never assigned.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): retire per-user defaults — the Default Group is the sole default policy

Removes both legacy 'user defaults' mechanisms now that the seeded
Default Group owns new-user policy:

- users.max_streams / max_transcodes column defaults drop from 6/2 to 0
  (= unrestricted at the user layer), so group ceilings apply to new
  signups/invites/OAuth users instead of fighting stale per-user
  numbers. Existing rows keep their stored values — nobody is silently
  uncapped on upgrade.
- The dead defaults.max_playback_quality / defaults.max_profiles
  settings validation goes away with its only writer (the User Defaults
  dialog, removed on the web side).

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): replace the User Defaults dialog with group-governed creation

The Users page's 'User Defaults' dialog (defaults.* server settings)
duplicated what access groups now do properly, and its values were only
ever form prefill — no backend path applied them. The button now links
to Access Groups, and the create-user form seeds unrestricted user-layer
values (0 streams/transcodes, any quality, downloads allowed) so the
member's group governs; per-user fields remain for tightening individual
users.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(access): migrate existing non-admin users into the Default Group

Existing users join the seeded Default Group on upgrade so one policy
source governs the whole instance. Their per-user limits still holding
the retired 6/2 column defaults are normalized to 0 in the same
statement so the group's ceilings actually apply; deliberately
customized values are preserved. Admin accounts stay ungrouped —
scope/action decisions are role-blind, so grouping an admin would cap
the server owner on upgrade.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(access): keep admins out of the Default Group and treat group moves as policy changes

New-user creation now mirrors the migration's admin exclusion: the
default access group is only auto-assigned to non-admin roles, so a
fresh server owner no longer inherits the starter group's transcode
denial and stream caps.

Changing a user's access group now bumps access_policy_revision (the
group carries permissions, quality, and limits, exactly like the
per-user fields that already bump it) and triggers admin session
revocation when the group actually changes.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): enforce marker_edit through the PDP on marker write routes

The Rego permission policy owned marker_edit but no Go caller ever
consulted it: PUT/DELETE /markers went through a handler-local check
that short-circuited admins and read only the user's own permissions,
so group permission masks and custom policy overrides were ignored.

Marker writes are now gated by router middleware like the other
permission surfaces: a PDP-backed RequireMarkerEdit that evaluates the
group-merged effective permissions (plus the legacy variant for
proxy/test wiring without a policy system). The handler-local check and
its user loader are gone.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): assert device/quality policy facts and honor the quality ceiling

The download_transcode action check hard-coded an empty device ID and
never asserted the requested quality, and no caller consumed
ActionDecision.QualityCeiling — custom download policies keyed on those
inputs were silently ineffective.

Resolve now threads the request's device ID and requested quality into
the action input, and a returned quality ceiling downscales the
prepared transcode target (the ceiling applies to what is served,
matching the serve-time rule in serveDownloadBytes). FileQuality and
the content-rating pair stay intentionally empty for downloads —
documented on downloadActionInput: those ceilings are enforced against
the served artifact by the scope-derived access filter, and asserting
the source's quality would wrongly deny capped transcodes.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(access): align the default-group seed assertions with the migration

The DB test still asserted the earlier no-op seed (transcode allowed,
unlimited streams/transcodes, null permissions); the shipped migration
seeds transcode denied, 5/5 limits, and marker_edit-only permissions,
so the test failed on any database with the migration applied.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): lock the Rego sandbox by builtin purity and bound compile work

Exclude every nondeterministic builtin from the admin sandbox instead of
denylisting names, so OPA upgrades cannot silently expose impure builtins
while pure helpers like net.cidr_contains stay usable. Apply the same
capabilities to the runtime engine, cap concurrent compile checks, and
reject oversized sources before they reach the uncancelable compiler.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): require literal booleans in vendor override and input checks

Bare object.get truthiness treated any non-false value as satisfied, so a
malformed override 'allowed' value could fail to tighten a base grant and
hand-crafted simulate input could flip flag predicates. Compare against
literal true so anything else denies.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): surface decision log cleanup failures to the task manager

CleanupDecisionLogsOnce now returns the first error alongside the deleted
count so a broken partition manager or DB outage marks the scheduled task
failed instead of reporting 100% success while policy_decisions grows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(playback): log admission decider errors before failing closed

A policy-evaluation failure was silently mapped to the too-many-streams
denial, making an engine outage indistinguishable from a real limit hit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(access): nil-guard the downloads user and restore the ABS legacy resolver

effectiveDownloadUser dereferenced policy state before its nil-user check,
and the ABS handler lost viewer-scoped filtering entirely when the policy
system was unavailable because no legacy access.NewResolver fallback was
wired like the other resolver paths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): address admin policy review feedback

- invalidate the version query by version_number, the key usePolicyVersion
  actually caches under
- keep the goPrevious cursor-stack updater pure (Strict Mode double-invoke)
- make version history rows keyboard-selectable like the document list
- clamp download_transcode_allowed when downloads are disabled so groups
  cannot save a contradictory record

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(api): cap policy endpoint request bodies at 1 MiB

The policy write endpoints (create document/version, set enabled,
validate, simulate) decoded JSON bodies without a size limit, so an
oversized payload buffered fully in memory before CompileCheck's
256 KiB source cap could reject it. Route all five through a shared
decodePolicyRequest helper that wraps the body in http.MaxBytesReader
and returns 413 with the repo's standard too_large error shape.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UtnZ2Uewzo959hpneLrtRN

* fix(access): forbid deleting or demoting the default access group

Deleting the default group (or unsetting its is_default flag) left the
server with no default: new non-admin users were then created ungrouped
with max_streams/max_transcodes of 0 — unlimited — because the legacy
per-user column defaults were retired in favor of the group's ceilings.

The store now rejects both operations with ErrDefaultGroupRequired
(mapped to 409); promoting another group remains the supported way to
move the default, and atomically clears the previous one. The admin UI
disables the delete button and the default toggle on the default group
and explains the promote-another-group flow.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UtnZ2Uewzo959hpneLrtRN

* fix(web): keep unsaved policy drafts when a newer version activates elsewhere

The editor state was keyed on the active version's id/sha, so a
background refetch after another admin (or another tab) activated a
version remounted the editor and silently discarded the dirty draft.

PolicyEditorPanel now pins the seed it is editing against and only
adopts an incoming seed when nothing can be lost: the editor is clean,
the draft already equals the incoming source (the same-admin activate
flow), or the selection moved to a different document. Otherwise the
pinned editor stays mounted and an inline notice offers an explicit
"Load live version" action.

Part of #272

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UtnZ2Uewzo959hpneLrtRN

* fix(policy): fail reloads on invalid custom sources and surface degraded/apply state

A stored custom source that stops compiling used to be silently skipped on
reload: the bundle widened to vendor-only for that domain while the generation
reported fully applied. Reload is now strict — a bad enabled source fails the
reload and the last known-good engine keeps serving. Boot keeps its vendor
fallback for availability, but skips are recorded on the engine and exposed
(with store-outage reasons) through System.DegradedState and additive
degraded fields on GET /policy/capability. Activate/SetEnabled re-run
CompileCheck instead of trusting the stored compiled_ok flag.

Mutation endpoints also no longer conflate persistence with live apply:
activation/enable responses carry additive applied/failed_step/
loaded_generation fields and return 202 when the store change persisted but
the local reload failed.

Addresses review findings C1, C2, and the degraded-signal gap (6.1).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): type deny reasons across the contract and enforce profile_verified

Deny handling used to branch on exact free-text reason strings in three Go
consumers, and playback reported ANY unrecognized reason — including custom
override free text and engine failures — as a stream-limit error. Decisions
now carry a stable reason_code (custom overrides always get custom_denial);
downloads, the metadata-curation gate, and playback admission switch on codes,
with a new ErrPlaybackNotAllowed -> 403 playback_not_allowed mapping for
non-limit denials. Rego tests pin every vendor code.

The scope contract's tighten-only profile_verified output was also emitted but
never consumed; a policy revocation now surfaces as ErrProfileUnverified (403
profile_unverified) instead of silently proceeding.

Addresses review findings 6.2 and C4.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(catalog): close the dual-library disabled-scope bypass in direct item authorization

EnsureAccessible, EnsureAccessibleIDs, and FilterAccessibleContentIDs gated
library access with allow/deny predicates over a single joined
media_item_libraries row, so an item linked to BOTH a passing library and a
disabled one satisfied the disabled check via the passing row — a direct-ID
bypass of disabled-library scope on the detail, media-file, playback, and
download paths. All library access predicates now share one helper
(libraryAccessConditions) emitting independent EXISTS / NOT EXISTS subqueries,
the semantics GetByIDsWithAccess already used, including the orphan-item
membership guard for disabled-only scopes. SQL-shape tests pin every builder
and a DB-gated regression test covers the dual-library item end to end.

Addresses review finding C3 (plus the same shape in
buildFilterAccessibleContentIDsSQL, which the review did not flag).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): serialize quota check and row creation under a per-user advisory lock

The concurrent-download quota was check-then-insert with nothing serializing
the pair: parallel creates could all observe free quota before any row
existed, bypassing the cap and stacking artifact encode jobs. All four
check->insert spans (ephemeral original, artifact-backed, series batch,
managed batch) now run inside Repository.WithUserQuotaLock — a
pg_advisory_xact_lock keyed by user, so the serialization holds across nodes.
The artifact path keeps the limiter-before-Ensure ordering (a rejected request
must not leave an encode job behind) by holding the lock across Ensure.
Managed-entry replacement stays quota-exempt and lock-free. A DB-gated
barrier test races 8 creates against a cap of 1.

Addresses review finding C5.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(downloads): assert served quality at create time for original and remux downloads

Direct-original and remux downloads serve the source resolution unchanged, but
create-time policy checks left file_quality empty — an over-ceiling source
registered a row serveDownloadBytes could never satisfy. Resolve now runs a
final download action check with FileQuality populated on those two paths
(capped transcodes keep the ceiling-on-artifact behavior), a custom override
ceiling below the served resolution denies, and quality_ceiling_exceeded maps
to ErrQualityUnavailable. The ActionInput contract now documents exactly when
file_quality and the rating facts are supplied so custom policy authors are
not misled.

Addresses review finding C6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(policy): guard activation against slow overrides and make eval timeouts observable

A custom scope override that exceeds the 25ms eval budget compiled fine,
activated fine, and then converted to 500s on every authenticated request —
server-wide lockout authored in the admin editor. Activation and enable now
run GuardEvalCost: the candidate source is evaluated on a throwaway engine
against a canned representative input under the live budget, and a source
that cannot complete is rejected 422 with ErrPolicySlowEval before it goes
live. Runtime timeouts keep failing closed but now carry a distinct
ErrPolicyEvalTimeout sentinel, an Error log, and a per-engine counter exposed
as eval_timeouts on GET /policy/capability so intermittent near-budget
policies are attributable.

Addresses review finding C7.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: gofmt remediation files

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 17:38:19 -04:00
604bbf1a0f feat(playback): unified restart-resilient playback (native + jellycompat) (#174)
* feat(playback): unified restart-resilient playback via shared TranscodeManager

Make direct, remux, and native HLS transcode sessions survive a server
restart through one shared flow instead of per-method paths. A missing
in-memory session becomes a reconstruct trigger, not a 404: the server
rebuilds the session from a tiny durable recipe card plus the position the
client re-supplies on its next request.

- internal/playback/transcode_manager.go: shared TranscodeManager owning the
  transcodes map, recipe-card lifecycle, reconstruct single-flight +
  concurrency cap, LoadOrReconstructSession front door, ReconstructSession /
  ReconstructTranscode, and orphan cleanup. ~90% is logic moved out of the
  native handler (no behavior change), not new surface.
- internal/playback/recipecard.go + recipecard_postgres.go: RecipeCard with a
  PlayMethod discriminator (direct/remux/transcode; empty decodes as transcode
  for back-compat) behind a swappable, nil-safe RecipeStore interface backed by
  transcode_recipes.
- internal/playback/session.go: RegisterReconstructed inserts a rebuilt Session
  under its existing id (no UUID mint, no limit double-count, race-yielding).
- internal/playback/transcode.go: CloseProcess keeps the output dir so a
  reconstruct winner keeps serving; Close removes it.
- internal/api/handlers: drain the transcode lifecycle into the manager; wire
  reconstruct into the stream/segment serve paths; re-bind ownership to the live
  caller (refuse userID==0/mismatch); card-aware orphan cleanup.
- migrations: add transcode_recipes (expires_at TTL, filter-on-read, indexed).

Ownership stays two-factor: an authenticated caller AND a session.UserID that
matches; the card stores no secrets and identity is re-resolved per request.

Tests: recipe-card round-trip/legacy-decode/disabled-noop, RegisterReconstructed
insert/race/concurrency, close-vs-close-process dir semantics, the
LoadOrReconstructSession status matrix, and the reconstruct concurrency cap.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* feat(jellycompat): reconstruct transcodes across restart via shared manager

Bring Jellyfin (jellycompat) HLS playback onto the same restart-resilient flow
as the native path. Previously jellycompat owned a separate PlaybackHandler with
a private transcodes map and a duplicated transcode lifecycle that never grew the
reconstruct half, so an in-flight Jellyfin transcode died on restart and the next
segment request 404'd.

- Embed the shared playback.TranscodeManager and delete the duplicate lifecycle,
  so jellycompat gets reconstruct, the concurrency cap, the node-affinity rule,
  and the card lifecycle for free.
- internal/jellycompat/playback_sessions_postgres.go: DurableCompatPlaybackStore,
  a write-through cache over jellycompat_playback_sessions behind the new
  CompatPlaybackStore interface (nil pool degrades to cache-only). This persists
  the load-bearing PlaySessionId -> UpstreamSessionID mapping (plus media sources,
  route item id, seek) so it survives a restart instead of vanishing with the map.
- Write a recipe card on compat transcode start keyed by the upstream session id,
  using the native StreamAppUserID so the ownership re-bind matches; reconstruct
  the upstream session and the transcode seeked to the requested seg_NNNNN.
- migrations: add jellycompat_playback_sessions (expires_at TTL + compat_token
  index, full PlaybackSession in data JSONB).

Auth is mapped to the native user id before reconstruct so the same two-factor
ownership check and userID==0/mismatch refusal apply unchanged.

Tests: DB-gated (SILO_TEST_DATABASE_URL) durable-store round-trip proving a
session written by one instance reloads in a fresh one (the restart case), plus
a nil-pool cache-only path; existing handler tests updated to the manager.

AI-use: implemented with AI assistance (design, implementation, adversarial review).

* docs(playback): consolidate unified playback reconstruction design

Replace the three overlapping playback docs (the native Postgres
restart-resilience spec, the jellycompat plan, and the unification spec) with a
single self-contained design at
docs/superpowers/specs/unified-playback-reconstruct.md.

The doc leads with the unified design — the one-idea reconstruct model, a strong
visual flow of a restart mid-playback, the shared TranscodeManager + recipe card,
the two swappable durable stores, security, the concurrency cap and node-affinity
constraint, preconditions, and verification. The design history and rationale
(reconstruct-not-rehydrate, phased delivery, Redis-vs-Postgres, token-as-
descriptor, failure analysis) move to an appendix. It references no other md file.

AI-use: written with AI assistance.

* fix(playback): address review on restart-resilient playback

Four fixes from PR review of the unified reconstruction work:

- Rewrite the recipe card on audio-track change. HandleChangeAudioTrack only
  updated the in-memory session/transcode, so after a restart reconstruct
  resumed with the stale AudioTrackIndex/TranscodeAudio (and stale play method)
  from the start-time card. Re-save the card (direct/remux/transcode) with the
  switched state, mirroring the start-card pattern.
- Guard nil TranscodeManager in LoadOrReconstructSession and ReconstructSession.
  StreamHandler.TM is documented optional (tests/minimal setups); a missing
  session previously panicked in recipeEnabled instead of returning
  SessionMissing. ReconstructTranscode already guarded nil; make the two
  siblings consistent.
- Reject direct/remux cards in doReconstructTranscode before spawning ffmpeg, so
  a non-transcode card id can never enter the HLS reconstruction path.
- Log a non-success status from the remote transcode-node DELETE in
  CloseTranscodeSession; a 401/404/500 was previously silent.

AI-use: implemented with AI assistance.

* fix(playback): harden restart-resilient compat sessions

* feat(playback): token-carried reconstruction across restarts

Build on the shared TranscodeManager (introduced earlier in this branch) so a
playback session survives an API-server or transcode-node restart without the
client re-negotiating, and retire the Postgres transcode_recipes store in favor
of a recipe carried inside the signed stream token.

- RecipeCard encodes the byte-affecting encode parameters and rides inside the
  stream token; LoadOrReconstructSession rebuilds the in-memory Session (and,
  for integrated transcodes, the ffmpeg process) on a cold miss, single-flighted
  per session and paced by a spawn semaphore. Removes recipecard_postgres.go and
  the 20260617233705_add_transcode_recipes migration.
- transcodenode reconstructs a lost ffmpeg node-side from the forwarded token.
- TR-lease: proxy/streamauth enforce a revocation deny-marker on every served
  segment, with a 500ms Redis timeout, a bounded per-session "allowed" cache
  (3s TTL, expiry-first graceful eviction), and a degraded-fail-open counter.

Review hardening folded in:
- Manifest/segment handlers do the in-memory session lookup first and only
  verify the stream token on a reconstruct miss (token HMAC was per-segment).
- Copy-mode reconstruct never applies the encoded-only seg*dur seek, at spawn
  time or via the recovery path: RestartSeekTarget reports "unresolved" for a
  copy session whose manifest cannot yet map the segment, so the client retries
  instead of seeking to a fabricated source time.
- Crash teardown is a compare-and-delete (CloseTranscodeSessionIf returns
  whether it matched); the crash closure tears down the playback session only
  when it matched, so a session reconstructed under the same id is not killed.
- Reconstruct enforces the same per-user stream/transcode caps as a fresh start
  (RegisterReconstructedWithLimits), closing a token-replay slot bypass.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* feat(jellycompat): node-side transcode reconstruct via shared recipe store

Make Jellyfin-compat playback sessions survive a server or transcode-node
restart by reusing the shared TranscodeManager reconstruct path and a durable
recipe store, on top of the durable compat session store added earlier in this
branch.

- Node-side transcode reconstruct goes through the shared recipe store; the
  recipe is persisted to the control-plane store (Redis) when a dedicated
  transcode node is used so the node can rebuild ffmpeg after its own restart.
- Adopt the shared manager's API (3-arg OnFFmpegCrash carrying the dead session,
  guarded CloseTranscodeSessionIf, RegisterReconstructedWithLimits).

Review hardening folded in:
- Recipe lifecycle: noderecipe.Store gains Delete, called on deliberate
  teardown (stop, method-switch discard, node stop/force-reload) so a stopped
  session cannot be resurrected by a buffered request after a node restart;
  crash paths intentionally keep the recipe so a resume can reconstruct.
- Crash closure tears down the upstream session only when the guarded transcode
  close matched, so a reconstructed successor is never left orphaned.
- Copy-mode segment recovery surfaces a retryable not-found instead of a
  wrong-position restart, matching the native and node paths.
- Durable Update is now a SELECT ... FOR UPDATE transaction, removing the
  lost-update clobber that could silently drop a transcode recipe.
- Empty-token route resolution no longer falls back to an unbounded full-table
  scan; DB expiry filters bind the injected clock; the redundant re-Get is gone.

AI-use disclosure: implemented with AI assistance (Claude Code), including a
two-round multi-agent adversarial review whose findings drove the hardening.

* docs(playback): consolidate restart-resilient playback design

Replace the superpowers spec with a single architecture record describing the
token-carried recipe card, the shared TranscodeManager reconstruct path for
direct/remux/transcode, the jellycompat durable session + node recipe store, and
the revocation-lease model with its fail-open tradeoff.

AI-use disclosure: written with AI assistance (Claude Code).

* docs(playback): correct jellycompat node-recipe rationale in comments

The noderecipe / transcode-node / jellycompat comments justified the Redis
recipe store with "a Jellyfin client cannot round-trip a token". The real
reason: the node-hop token is server-minted and could carry the recipe, but the
recipe is mutated in place under a stable session id (a /Sessions/Playing/Progress
audio switch restarts ffmpeg without re-minting the client's token) and a
third-party Jellyfin client cannot be driven to refresh a stale token, so the
node must reconstruct from a server-authoritative, node-reachable store.

Aligns the comments with docs/architecture/restart-resilient-playback.md §10.
Comment-only; no behavior change.

* refactor(playback): remove deny-lease revocation, defer to future PR

The deny-lease stream-revocation mechanism (the internal/streamauth
package, its silo:streamauth:<sid> Redis markers, the proxy Allowed()
enforcement, and the admin Stop/Terminate deny write) only ever enforced
on the offload-proxy topology and was a silent no-op on the integrated
single box and the dedicated transcode node. Rather than ship a partial
revocation feature that looks complete but isn't, remove it wholesale and
defer a uniform cross-topology revocation design to a dedicated follow-up.

Removed: internal/streamauth (package + tests); the LeaseDenier field,
StreamLeaseDenier interface, and denyStreamLease helper in playback.go;
the admin deny write; the router/main wiring; and the proxy verifyToken
Allowed() gate. The unified-reconstruct core (recipe-token,
LoadOrReconstructSession) is orthogonal and untouched.

Known limitation (now on every topology): admin Terminate and user Stop
tear down the live in-memory session and ffmpeg producer, but a still-valid
stream token can reconstruct the session until its 24h TTL expires. No
node-side byte-withholding ships in this PR.

docs/architecture/restart-resilient-playback.md is updated to mark the
revocation/deny-lease sections as deferred and to drop the overstated
"instant revocation on admin kill" claim.

* fix(playback): allow zero-caller bearer on transcode reconstruct

The authless HLS transcode delivery routes (master.m3u8 / segment) treat
the session UUID as the bearer credential, so a real request carries
requestUserID == 0. The live serve path already allows this, but
ReconstructSession hard-rejected a zero caller, so a request that worked
before a restart became SessionMissing -> 404 after the in-memory session
was gone, breaking the restart resilience these routes advertise.

Match the live-path contract in LoadOrReconstructSession: allow a zero
caller (UUID-as-bearer) and refuse only a non-zero caller that mismatches
the card owner. The reconstructed session is bound to card.UserID either
way. Adds TestReconstructSession_Ownership covering both cases.

* fix(jellycompat): re-persist recipe on local audio switch

A Jellyfin client switching audio on an integrated/local compat transcode
restarted live ffmpeg with the new track but did not re-persist
PlaybackSession.Recipe. The remote branch already re-persists via
startRemoteTranscode -> persistTranscodeRecipe. After a central restart,
reconstruct rebuilt ffmpeg from the stale Recipe.AudioTrackIndex, so the
integrated session resumed on the original audio track.

Persist the updated recipe (best-effort) after a successful Restart in the
local branch, mirroring the remote branch, so the durable
Recipe.AudioTrackIndex tracks live ffmpeg. Adds a regression test.

* fix(playback): strip stream token from proxied transcode-node URL

proxyToTranscodeNode appended the client's raw query string to the internal
transcode-node URL and logged that URL on transport failure. When a remote
transcode runs without a separate proxy node, that query carries
?st=<signed JWT> — a 24h bearer reconstruction descriptor exposing the
media path and recipe claims — placing the token into internal requests and
error logs.

Strip the "st" param before building targetURL, preserving any other query
params. The token is neither forwarded to the node nor present in the
logged URL. Header-forwarding of the token (so the node can reconstruct) is
a separate follow-up (#6).

* fix(playback): fail open on transient limit-provider error in reconstruct

During the reconstruct wave right after a restart (Postgres under peak
load), a transient limit-provider DB error was collapsed into a hard 404,
permanently stopping playback for a user within their limits. limitsForUser
wrapped any provider error, RegisterReconstructedWithLimits propagated it,
and ReconstructSession mapped every error to SessionMissing -> 404 -
indistinguishable from a genuine over-cap rejection.

Distinguish the two: tag provider errors with a new ErrLimitProviderUnavailable
sentinel and, during reconstruct, fail OPEN on a provider error (admit via
RegisterReconstructed + log a degraded warning) rather than refuse - mirroring
the reliability-first fail-open-on-dependency-error philosophy. A genuine
ErrTooManyStreams / ErrTooManyTranscodes over-cap still refuses. Adds tests
for both the fail-open and still-refused paths.

* fix(playback): forward stream token to transcode node as header

The dedicated transcode node's reconstruct path reads the stream token only
from the X-Silo-Stream-Token header, but proxyToTranscodeNode forwarded only
the node-API bearer token (and #5 now strips st from the URL). So when the
central API proxied to the node and the node self-restarted, it could not
reconstruct from the recipe-complete native token -> 404.

Capture st before stripping it from the URL, verify it at the API boundary
(streamtoken.Verify + SessionID match, mirroring the node's own check), and
forward it as X-Silo-Stream-Token. Best-effort: a missing/invalid token never
blocks the live proxy, and the token is still kept out of the forwarded URL
and logs.

* fix(playback): restart node ffmpeg on native remote audio switch

A native audio-track switch on an offloaded/remote transcode was a no-op at
the node yet returned 200 with a fresh URL: HandleChangeAudioTrack restarted
ffmpeg only when the API owned a LOCAL TranscodeSession, so for an offloaded
transcode the node kept serving the OLD audio (the node consults the token
only on a session miss). The replacement URL was also minted from identity-
only claims, so a later node restart 404'd.

For the offloaded transcode case (detected via session.TranscodeNodeURL),
POST a fresh /transcode/start to the node with the new AudioTrackIndex
(handleStart tears down and restarts ffmpeg) and mint the replacement proxy
URL from a full RecipeCard so reconstruct survives a node restart. The encode
recipe is derived from the durable session target fields plus the file,
mirroring HandleStartTranscode. A concrete SegmentDuration
(playback.DefaultSegmentDuration) is embedded rather than 0: the node's token
completeness gate treats SegmentDuration<=0 as incomplete and falls back to a
recipe store the native path never populates, which would 404 on a node
restart - the exact resilience this path provides. A failed node POST now
surfaces 502 rather than a false 200. Remux and non-offloaded (local)
transcode paths keep their prior identity-claim URLs unchanged.

Known limitation: Session does not persist the original SegmentDuration or
SubtitleTrackIndex/SubtitleBurnIn, so a remote audio switch resets subtitle
selection to none and assumes the default segment length; a client that
started with a non-default segment length will resegment on switch. Making
that state durable on the session is a follow-up.

* docs(playback): scrub stale deny-lease/revalidator comments

The deny-lease revocation mechanism and its "central revalidator" were removed
earlier in this branch, but four comments still described them as live
(transcode_manager.go, noderecipe/store.go, streamtoken/token.go,
proxy/server.go). Reword them to match the shipped behavior: ownership claims
are re-resolved at reconstruct, the noderecipe store shares Redis only with the
node-session tracker, and a sub-TTL hard cut depends on a node-side revocation
mechanism that is deferred to a future PR.

* fix(jellycompat): surface durable playback-session write failures

DurableCompatPlaybackStore.Update applied the in-memory mutation and then
swallowed every Postgres commit-failure path, returning nil. Callers that
promise restart resilience (persistTranscodeRecipe's recipe write, the
upstream-session binds in streams.go) were told the session was durably
persisted when only the cache held it, so a transient DB hiccup could leave
the next restart reloading a stale row (wrong audio track) or 404ing.

updateDB now returns the genuine DB round-trip error (begin/query/unmarshal/
marshal/exec/commit); Update propagates it while still applying the in-memory
mutation so live state stays correct. A nil pool and a genuinely absent/expired
row remain best-effort (return nil) — only real infrastructure failures
propagate, so existing rollback paths fire exactly when durability is lost.

Part of #174

* fix(playback): re-inject stream token into proxied transcode manifests

API-proxied remote transcode manifests dropped the reconstruct token from
their segment URLs, so playback died after a node or API restart. When a
remote transcode has no separate proxy node, the client loads its manifest via
the API-local path; proxyToTranscodeNode strips the signed token ("st") from
the forwarded URL (keeping it off node URLs and logs, forwarded only as the
X-Silo-Stream-Token header), and the node builds relative segment URIs from
that token-less query. The segment URLs the client received carried no token,
and the proxy only re-attached the header when an incoming segment request
already had "st" — which it never did — so a restart made those segments
non-reconstructable and they 404'd.

proxyToTranscodeNode now rewrites the manifest body at the boundary: every
segment and #EXT-X-MAP init URI gets the client-facing, API-verified token
re-appended (new playback.AppendManifestQueryParam helper), so the client's
later segment fetches carry "st" again and reconstruct after a restart. The
token still never reaches the node URL or its logs. Only 200 .m3u8 responses
are rewritten (Content-Length corrected); segments stream through untouched.

Part of #174

* fix(playback): preserve subtitle/cadence recipe across offloaded audio switch

Switching audio on a remote (offloaded) transcode with burned-in subtitles
silently dropped them, and reset a non-default segment cadence. The offloaded
audio-switch restart rebuilt the node start request from Session state, but
Session/SessionStreamState retained no subtitle or segment-duration state
(only the live local ts.Opts() and the RecipeCard did), so the branch
hard-coded SubtitleTrackIndex:-1, SubtitleBurnIn:false and
SegmentDuration:Default — signing that altered recipe into the replacement
stream token. An audio switch then changed bytes beyond audio selection, and
any later reconstruct kept the wrong no-subtitle/wrong-cadence recipe.

Persist the byte-affecting recipe on the session: SubtitleTrackIndex,
SubtitleBurnIn and SegmentDuration are added to Session/SessionStreamState,
populated at start (finalizeTranscodeStart) and on post-restart reconstruct
(ReconstructSession from the card), carried forward on every audio-switch
state update, and read back when rebuilding the offloaded node request and its
recipe card. The restart now reproduces the exact live stream. Also resolves
the M-4b non-default segment_duration reset.

Part of #174

* fix(playback): serialize transcode spawn paths with a per-session lock

Reconstruct was single-flighted only against other reconstructs, so a
restart-driven segment reconstruct racing a quality/seek/audio fresh start
could spawn two ffmpeg processes writing the same output directory at once —
segment corruption, partial-write closes, orphaned processes, and skewed
active-job accounting. The atomic register-after-spawn (GetOrRegister / the
reconstruct compare-on-register) prevented a map leak but not the concurrent
disk writers, because the losing path had already spawned. The dedicated
transcode node had the same split between handleStart and spawnReconstruct.

Add a refcounted per-session lifecycle lock to both TranscodeManager and the
node Server, held across "check existing -> spawn -> register":
- reconstruct (doReconstructTranscode / spawnReconstruct) re-checks under the
  lock and yields to any live session instead of spawning a duplicate;
- the native and jellycompat fresh-start paths take the lock around their
  spawn+register (the native path also closes any session a reconstruct rebuilt
  in the meantime so its fresh ffmpeg is the sole writer);
- the node handleStart holds it across teardown+spawn+register.
The refcount drops the map entry once no path holds/waits, keeping it bounded.
GetOrRegisterTranscodeSession is removed — the lock supersedes it and keeping a
register-after-spawn primitive would invite reintroducing the race.

Part of #174

* fix(playback): serialize restart re-spawn under the session lifecycle lock

TranscodeSession.Restart() releases s.mu across cancel -> wait-for-done ->
re-exec and spawns ffmpeg into opts.OutputDir without holding the per-session
lifecycle lock. LockSessionLifecycle's contract (fresh start, restart,
reconstruct) requires restart to hold it too, but all five callers invoked
Restart unlocked: native audio-switch and segment-recovery, compat
audio-switch and segment-recovery, and the transcode-node segment-recovery.

A restart racing another restart (audio-switch vs segment-recovery) or a
fresh-start/reconstruct could land two ffmpeg processes writing the same
segment directory -- mixed timelines, init.mp4/segment mismatch, and an
orphaned-but-still-writing ffmpeg -- the exact concurrent-writer corruption
the lifecycle lock exists to prevent.

Add RestartSessionLocked (TranscodeManager) and restartSessionLocked (node
Server) that hold LockSessionLifecycle only across the cancel->respawn
transition, re-check that the handle is still the live mapped session under
the lock, and return ErrSessionSuperseded rather than re-spawning a stale
handle. Route all five call sites through them. The lock is released before
callers wait on segments so recovery latency is unchanged.

Tests: gating (restart blocks until the lifecycle lock frees, then spawns),
concurrent-restart serialization, and superseded re-check on both the manager
(covers native + compat) and node lock owners.

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-02 14:23:14 -04:00
Silo Server Migration c085b12fd1 Initial Silo migration 2026-05-22 23:26:56 -04:00