Files
tuliprox/docs/src/operator/dvr.md
T
DarkBreakpoint 67d75d7ab4 docs(dvr): correct the operator reference and repoint the doctor
The documentation described a system that no longer exists, and in two
places described the opposite of what the code now does.

- The layout section documented `<recording-root>/users/<owner-id>/<rel>`
  for private recordings and `shared/<rel>` for shared ones. That resolver
  was deleted: recordings are stored owner-independently at
  `<recording-root>/<rel>`, because one physical file is shared by every
  user who asked for it. The organised layouts and the component
  sanitisation rules are now documented as they are implemented.

- The config reference claimed persisted queue recovery "is tolerant of
  corruption" and "starts with an empty transfer queue instead of aborting
  server boot". The queue now fails closed: a damaged database is rebuilt
  from the recovery history, and a database ahead of every surviving
  history refuses to start. An operator following the old text would have
  expected silent recovery from a condition that is deliberately fatal.

- `download.read` / `download.write` were still listed as grantable
  permissions after their removal, and `recording.write` after its split.

bin/dvr_doctor.sh looked for `downloads_state.json` and summarised it with
jq. That file never existed under this name, and the queue it stood for is
now a B+Tree, so the section printed "(absent)" and skipped its summary
exactly when an operator needed it. It now reports the repository and its
recovery generations: the CURRENT pointer, the retained generation pair,
journal sizes, the fail-closed case where a database has no history, and a
warning when the recovery directory shares a filesystem with the database
— which survives a corrupt file but not the loss of the volume it exists
to protect against.

CHANGELOG records the three breaking changes: the non-migrating queue, the
permission split, and the moved recording files.
2026-09-01 09:23:52 -05:00

719 lines
44 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DVR Operator Reference
This guide covers everything an operator needs to know to deploy, configure, and manage the
DVR. It is the single source of truth for the documentation that
[`config.md`](../configuration/config.md) and [`rest-api-cookbook.md`](../rest-api-cookbook.md)
consume.
## 1. Configuration reference
All fields live under `video.recording` in the config file. Every field is optional;
defaults match the recommended values.
```yaml
video:
recording:
enabled: true # default: true; false stops every DVR supervisor
container_format: mpegts # mpegts (default) | matroska | mp4
directory: recordings/ # default: recordings
timezone: Europe/Berlin # default: UTC (IANA required)
filename_template: "{channel}_{program_title}_{start_time}"
default_pre_roll_secs: 0 # 0..=max_pre_roll_secs
max_pre_roll_secs: 900 # ≤ 900 (15 min)
default_post_roll_secs: 0 # 0..=max_post_roll_secs
max_post_roll_secs: 1800 # ≤ 1800 (30 min)
retention:
keep_last_per_channel: 10 # > 0 when set
delete_after_days: 30 # > 0 when set
sweep_interval_secs: 3600 # default 3600; age/count sweep cadence
disk:
high_water_percent: 85 # 0..=100
low_water_percent: 70 # 0..=100 and < high_water_percent
cleanup_interval_secs: 3600 # > 0; watermark-check cadence
safety_bytes: 1073741824 # > 0 (1 GiB)
quota:
default_private_bytes: 53687091200 # 50 GiB
per_user_bytes:
"web:user-uuid-1": 107374182400 # 100 GiB
shared_bytes: 536870912000 # 500 GiB
notifications:
outbox_buffer: 1024 # default 1024; in-memory queue depth
max_attempts: 6 # default 6; then dead-lettered
backoff_initial_secs: 5 # default 5
backoff_max_secs: 900 # default 900 (15 min)
fallback_bytes_per_minute: 8388608 # 8 MiB/min, > 0
```
| Field | Default | Range | Restart required | Effect |
|--------------------------------------|----------|-----------------------------|------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| `enabled` | `true` | bool | no | `false` skips reconciliation, retention, and the notification outbox at startup, makes running supervisors idle on their next tick, answers every `/api/v1/recording/**` route with `501 recording_disabled`, and hides the DVR entries from the web UI navigation. |
| `container_format` | `mpegts` | `mpegts`, `matroska`, `mp4` | no | The `-f` muxer ffmpeg writes. `mpegts` survives truncation, so a recording killed mid-stream still plays — prefer it unless the source codecs need another container. Applies to recordings that start after the change. |
| `retention.sweep_interval_secs` | `3600` | > 0 | no | Cadence of the age/count sweep. Independent of `disk.cleanup_interval_secs`. |
| `disk.cleanup_interval_secs` | `3600` | > 0 | no | Cadence of the supervisor's tick, and therefore of the watermark check. Measurements are floored at one per 30 s. |
| `notifications.outbox_buffer` | `1024` | ≥ 1 | **yes** | Channel capacity between the recorder and the outbox worker. Fixed when the worker starts. |
| `notifications.max_attempts` | `6` | ≥ 1 | no | Delivery attempts per notification before it is dead-lettered. |
| `notifications.backoff_initial_secs` | `5` | ≥ 1 | no | First retry delay; doubles per attempt. |
| `notifications.backoff_max_secs` | `900` | ≥ `backoff_initial_secs` | no | Ceiling for the doubling. |
Choosing values:
- **Private-only home use** — leave `quota` unset, set `retention.keep_last_per_channel` to taste,
and leave `disk` unset if the recording filesystem is dedicated.
- **Shared family install** — set `quota.shared_bytes` and `disk.{high,low}_water_percent` so a
full disk degrades into retention rather than failed recordings.
- **Multi-tenant** — set `quota.default_private_bytes` plus per-user overrides, and keep
`notifications.max_attempts` low so a broken webhook does not accumulate a backlog.
### 1.0.1 ⚠️ Retention policy warning
**Retention deletes recordings.** If you are upgrading an existing install or reconfiguring
retention, whatever `retention` and `disk` values are in your config take effect on the next
supervisor sweep. If they were set optimistically or copied from an example, the first sweep may
delete recordings you expected to keep.
Before restarting after a configuration change:
1. **Read back your effective policy.** Check `video.recording.retention` and
`video.recording.disk` in `config.yml`.
2. **Work out what would be deleted.** `keep_last_per_channel: N` keeps the *N most recent*
recordings per (owner, channel) and deletes the rest. `delete_after_days: N` deletes anything
whose `completed_at` is more than N days old. The two are a **union**, not an intersection —
a recording matching either policy is deleted.
3. **If you are unsure, start with retention off.** Remove the `retention` block (or the
individual keys) and set `enabled: true` with no policy. Nothing is deleted, and you can
enable a policy deliberately once you have looked at the library.
4. **Back up** `recordings.db`, the recovery generations under `backup_dir`, and the
recording directory.
There is no dry-run mode. Deletions are logged under the `recording::audit` target with a
`recording_retention_delete` line and a reason (`Age`, `Count`, or `watermark`), so
`grep recording_retention_delete` after the first sweep tells you exactly what went.
### 1.1 Validation rules
The validation list, applied at `VideoConfigDto::prepare`:
- `keep_last_per_channel > 0` when set.
- `delete_after_days > 0` when set.
- `cleanup_interval_secs > 0`.
- `low_water_percent < high_water_percent` and both in `0..=100`.
- `safety_bytes > 0`.
- `default_private_bytes`, `per_user_bytes[*]`, `shared_bytes` all `> 0` when set.
- `fallback_bytes_per_minute > 0`.
- `timezone` is a valid IANA zone (`chrono_tz::Tz::parse`).
- `filename_template` contains at least one known placeholder and is no longer than 240 UTF-8 bytes.
Absent quotas and absent retention values disable the corresponding policy (no limit).
### 1.2 Disk watermark semantics
`high_water_percent` and `low_water_percent` are **used-space percentages** on the filesystem
containing the canonical recording root. The retention worker uses them with hysteresis: when
used-space ≥ `high_water_percent`, the worker deletes oldest completed recordings until
used-space ≤ `low_water_percent` (or the eligible list is exhausted). `cleanup_interval_secs` is
the wall-clock interval between passes; the worker uses a cancellation-aware Tokio interval and
will not overlap.
Free space is measured **once per pass**, not once per deletion. The stop condition folds in the
bytes the pass has already reclaimed, so a pass deletes just enough recordings to reach the low
watermark and then stops. Only the recording root's own filesystem is measured — never
`storage_dir` or the generic download directory, which may sit on a different mount.
### 1.3 What changed: supervisors now run
Three background supervisors now actually execute their work. They existed as decision layers
before but were never started, so the DVR worked on the happy path and silently skipped
everything else. The consequences of switching them on are all things this guide describes, but
they happen now where they did not before:
| Supervisor | What now happens |
|--------------------------|-------------------------------------------------------------------------------------------------------------------------------|
| Startup reconciliation | Recordings stuck in `Deleting` from an earlier crash are finished or restored on boot. Orphaned rule tombstones are repaired. |
| Retention | Age, count, and disk-watermark deletion begin. See the retention warning above. |
| Notification outbox | Lifecycle notifications are retried and persisted instead of being dropped on transient error. |
### 1.4 Supervisors
Three background supervisors implement the behaviour described in the rest of this document. All
three are started once the HTTP listener is bound, and all three honour the `downloads`
cancellation token, so a config reload stops and restarts them cleanly.
| Supervisor | Cadence | Responsibility |
|--------------------------|-----------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------|
| Startup reconciliation | once at boot, before the rule scheduler | Finishes or undoes deletions interrupted by a crash; repairs queue/rule-store drift. |
| Retention | `disk.cleanup_interval_secs` tick, policy sweep every `retention.sweep_interval_secs` | Age, count, and watermark deletion — all through the single `system_retention_delete` path. |
| Notification outbox | event-driven, with per-entry retry timers | Durable lifecycle-notification delivery with per-channel retry and dead-lettering. |
Passes never overlap: a tick that arrives while the previous pass is still deleting is skipped.
`GET /api/v1/recording/health` (administrator only) reports each supervisor's last-tick timestamp,
the outbox depth, and the dead-letter count, so liveness can be checked without reading the log:
```json
{
"enabled": true,
"server_time": 1700003600,
"reconciliation_last_run": 1700000000,
"retention_last_tick": 1700003400,
"retention_sweep_interval_secs": 3600,
"notification_last_drain": 1700003100,
"notification_outbox_depth": 0,
"notification_dead_lettered": 0,
"queue_revision": 412
}
```
A `null` timestamp means that supervisor has never completed a pass. Compare `retention_last_tick`
against `server_time` and `disk.cleanup_interval_secs` to detect a stalled sweep.
### 1.5 Support diagnostics
`bin/dvr_doctor.sh` collects everything above plus the on-disk state into one dump suitable for a
support ticket. It is read-only and never touches config, the queue, or a recording.
```bash
bin/dvr_doctor.sh --token "$ADMIN_TOKEN"
bin/dvr_doctor.sh --url https://tuliprox.example --token "$ADMIN_TOKEN" --storage-dir /opt/tuliprox/data
```
It reports supervisor health, the effective `recording` config block, quota, the recording
repository and its recovery generations, and aggregate summaries of `recording_rules.json` and
the notification outbox. The summaries are deliberately aggregate: no titles, filenames, or
owner ids are printed, because a diagnostics dump gets pasted into tickets. The health and
config sections need an administrator token; the on-disk sections work without one, and take
`--backup-dir` because the recovery generations live under `backup_dir`, not `storage_dir`.
## 2. Recording directory layout and immutable IDs
The recording root is the path configured in `recording.directory` (default
`<download-dir>/recordings`). Every recording is stored at `<recording-root>/<rel>`, where
`<rel>` is the collision-safe relative path reserved at the queue-mutation boundary.
The layout carries **no owner or visibility component**. One physical file is shared by every
user who requested it, so keying its directory on an owner would be wrong the moment a second
user attaches, and would force the file to move on disk when the first detaches.
With `organize_into_directories` enabled the layout groups by what the server resolved:
```text
<recording-root>/
<channel>/<file> # live
<title>/<file> # vod
<series>/Season NN/<file> # series episode
<file> # unorganized, or no grouping resolved
```
Path components are sanitised to a single filesystem-safe component: separators, control
characters and the Windows-forbidden set collapse to `_`, leading and trailing dots are
stripped, Windows device names fall back to a placeholder, and each component is capped at 255
bytes on a character boundary. The path is validated against the recording root at open time;
any path that escapes the root (via `..`, symlinks, or absolute paths) is rejected with
`recording_unsafe_path`.
## 3. Filename placeholders
The supported placeholders:
- `{channel}` — channel display name; falls back to the stable channel id when the display name
is missing.
- `{program_title}` — programme title; sanitized via the existing filename sanitizer.
- `{start_time}` — programme start in the configured timezone, rendered as `YYYY-MM-DD_HH-mm`.
- `{end_time}` — same, programme end.
- `{episode}` — renders `SxxExx` when both season and episode numbers exist; otherwise empty.
- `{owner}` — sanitized owner display name. **Filename only; never a directory component.** The
`{owner}` placeholder is the only place the username-derived content may appear in the on-disk
path.
The final stem is capped at 240 UTF-8 bytes without splitting a code point. When the sanitized
stem is empty, the runtime falls back to the recording task id.
## 4. Lifecycle and restart behavior
Every recording goes through three persistence states:
1. **Partial path** — `partial_relative_path` is set; the `O_CREAT | O_EXCL | O_NOFOLLOW` open
guarantees an attacker-prepared symlink at the partial path is rejected.
2. **Finalize** — atomic rename from partial to final; the runtime refuses to clobber any
pre-existing final file (including symlinks).
3. **Complete** — `mutate` sets `Completed`, stamps `measured_bytes`, `completed_at`, and clears
`reserved_bytes`.
A crash at any point is recoverable:
- Active + valid final file → normalize to `Completed`.
- Active + valid partial file → normalize to terminal `Failed`, retain partial.
- Active + no owned file → normalize to terminal `Failed`.
- Unsafe path / type → log a security-category error, normalize to previous terminal state, do
not open.
## 5. Quota charge-by-state
The quota ledger charges the bytes below per `DownloadState`:
| State | Charged bytes |
|---------------------------------------------------------------------------|---------------------------------------------------------:|
| `Scheduled` / `Queued` / `WaitingForCapacity` / `RetryWaiting` / `Paused` | `reserved_bytes` |
| `Downloading` | `max(reserved_bytes, measured_bytes)` |
| `Completed` | final `measured_bytes` |
| `Failed` / `Cancelled` with partial file | partial `measured_bytes` |
A task that is mid-deletion carries `deleting_previous_state = Some(prior)` and is charged the
same bytes as the prior terminal state (`reserved_bytes` or `measured_bytes` depending on what
`prior` was); the field replaces the historical `DownloadState::Deleting` variant, which the
runtime no longer carries. The charge drops to zero only when `finalize_deletion` removes the
task from the queue.
A task is counted exactly once. Private pools key on `RecordingOwner::User(uid)`; shared pools key
on `RecordingVisibility::Shared`; `LegacyAdmin` recordings count toward the shared pool. Per-user
overrides beat the configured default; an absent limit is unlimited.
### 5.1 Active overrun policy
Version one does **not** terminate an active recording because it grew beyond quota. `would_exceed`
is admission-only. When the measured partial size exceeds the reservation, the charge is
`max(reserved, measured)`; the next `would_exceed` call rejects new admissions until the recording
finishes or is deleted. The user-visible DTO surfaces an `Overrun` warning so operators can grant
more quota or delete the recording.
### 5.2 Unknown bitrate
When the bitrate is unknown at create time, the reservation is
`duration_minutes × fallback_bytes_per_minute` (default 8 MiB). The DTO surfaces an
`UnknownBitrate` warning. The runtime re-reserves with the measured rate as soon as the worker
starts.
## 6. Disk admission
The disk admission path:
```text
headroom = free_bytes - safety_bytes - active_disk_reservations
admit = charge <= headroom
```
`free_bytes_for(path)` is `statvfs` (Unix) or `GetDiskFreeSpaceExW` (Windows) keyed on the supplied
path's mount. The pre-start flow always passes the canonical recording root so the measurement is
on the same filesystem the file will live on. Two starts cannot consume the same headroom — the
active reservation is serialized through the queue-mutation boundary.
## 7. Safe deletion guarantees
Deletions use a persisted two-phase operation. The runtime carries the deletion intent in
`recording.deleting_previous_state: Option<DeletingPreviousState>` rather than as a
`DownloadState` variant — every terminal task that is mid-deletion stays in its prior state
(`Completed` / `Failed` / `Cancelled`) but carries the marker, which is what the rest of this
section means by "the task is in the deleting phase".
1. **`begin_deletion`** runs inside the queue-mutation boundary. It stamps
`recording.deleting_previous_state = Some(prior)` (the prior terminal state) and zeros the
byte counts.
2. **`execute_deletion`** runs **after** the boundary. It inspects the path with
`symlink_metadata` (never `metadata`), so a symlink is seen as a symlink and refused rather
than dereferenced, and removes the file. Missing files are idempotent success.
3. **`finalize_deletion`** runs inside a fresh boundary. It removes the task from the queue and
clears `deleting_previous_state`.
Startup recovery (any task whose `deleting_previous_state` is `Some(_)`):
- `deleting_previous_state = Some(_)` + missing file → finish task removal.
- `deleting_previous_state = Some(_)` + existing valid regular file inside the recording root
→ restore the prior terminal state, clear the marker.
- `deleting_previous_state = Some(_)` + unsafe path or non-regular file → restore the prior
state, log `recording_reconciliation_unsafe_path`, leave the file alone.
### 7.1 Portability of the path guarantees
The four guarantees — no symlink is followed, no existing file is clobbered, the publish is
atomic, nothing escapes the recording root — are built from portable primitives and hold
identically on every supported target:
| Guarantee | Primitive | Portable? |
|-------------------------------------|----------------------------------------------------------------------------------------------------------------|-----------|
| No symlink followed on inspection | `symlink_metadata` (never `metadata`) | yes |
| No existing file clobbered | `create_new` → `O_CREAT\|O_EXCL` / `CREATE_NEW`; both fail on an existing entry *including a dangling symlink* | yes |
| Atomic publish | `rename`, after a no-follow existence check on the destination | yes |
| Contained in the recording root | component validation plus an owner-id component check, before any syscall | yes |
Only one call has a platform-specific branch: `open_partial_no_clobber` additionally passes
`O_NOFOLLOW` on Unix. That is **defense in depth, not the mechanism** — the no-clobber property
already comes from `create_new`. `openat2` with `RESOLVE_BENEATH` / `RESOLVE_NO_SYMLINKS` would be
Linux-only and is deliberately not used.
Earlier revisions carried a blanket `#![cfg(unix)]` on the path helper, which removed the module
wholesale on Windows and left every caller with unresolved imports — the DVR did not build on
Windows at all. The gate is now scoped to the single `O_NOFOLLOW` line. Tests that need to
*create* a symlink stay Unix-only (Windows requires developer mode or elevation for that); the
behaviour they cover is asserted portably by the no-clobber tests.
## 8. Authorization matrix
| Operation | Private recording | Shared recording | `LegacyAdmin` | Orphan |
|------------------------------|-----------------------------------|----------------------------------|---------------------------------|--------------|
| Read / Playback / Download | owner with `recording.read` | anyone with `recording.read` | admin only | admin only |
| Create private | user with `recording.create` | n/a | admin only | n/a |
| Create shared | rejected (admin only) | admin + `recording.create` | admin only | n/a |
| Edit / Cancel | owner + `recording.manage` | admin + `recording.manage` | admin only | n/a |
| Delete | owner + `recording.delete` | admin + `recording.delete` | admin only | n/a |
| Manage recurring rule | owner + `recording.manage` | admin + `recording.manage` | admin only | n/a |
| `SystemRetentionDelete` | ownership bypassed; state-gated | ownership bypassed; state-gated | ownership bypassed; state-gated | n/a |
| Orphan catalog | n/a | n/a | n/a | admin only |
Administrators **do not** implicitly receive another regular user's private recording content. The
private owner is the only non-administrator allowed to read it. Administrative access is read-only
for diagnosis; mutations require either the `SystemRetentionDelete` action (which the retention
worker is the only legitimate caller of) or the appropriate recording permission + ownership
combination.
Orphan catalog entries (recordings whose target/input no longer matches a configured source) are
visible only to administrators with `recording.read`. The path is never exposed; an opaque orphan
id is generated per discovery.
## 9. Identity-registry bootstrap
The identity registry is `web_user_ids.json` in the storage directory. The startup sequence is:
1. Pre-scan the recording repository for `RecordingOwner::User(_)` entries (without the
registry loaded).
2. Load the existing registry (if any).
3. Initialize the registry **only** when no persisted real owner exists. New `UserId`s are
generated for any username that lacks one.
4. Fail closed on missing / corrupt registry when real owners exist. The server does not generate
replacement IDs in this case; the operator must restore the registry or run an explicit rename
migration.
5. Sync current principals (insert a new `UserId` for any username that lacks one).
6. Run the full queue load + normalization.
The built-in administrator is the reserved subject id `builtin:admin` (constant). Operators do not
create an entry for it.
## 10. Token refresh on permission schema bump
`Claims` carries `subject_id: Option<UserId>` and `permission_schema_version: u16`. The constant
`CURRENT_PERMISSION_SCHEMA_VERSION` is the source of truth. When the schema changes, bump the
constant; pre-bump tokens become stale:
- `authenticator::validate_token_version` returns `AuthError::StaleSchema` for older versions.
- The HTTP layer emits a 401 with header `X-Token-Refresh: required`.
- The frontend's `RecordingError::TokenRefreshRequired` and the generic auth refresh handler
redirect the user back to the sign-in flow.
Operators do **not** need to manually invalidate tokens on a schema bump. Existing user records in
`web_user_ids.json` are preserved; only the `subject_id` mapping for current usernames is
recomputed if missing.
## 11. Deprecated `/file/record` behavior
The legacy `POST /file/record` route is **deprecated** and delegates to
`RecordingService::create_recording` for administrators only. Non-administrators receive a 403 —
the deprecated route does not bypass the new policy.
The migration:
- **Frontend code**: switch from `downloads_service::queue_recording` to
`recording_service::RecordingService::create_task`. The new client submits
`RecordingSourceInput` (target_id + virtual_id + input_name) and `CreateRecordingTaskRequest`,
never a free-form URL.
- **Operator code**: the legacy route is documented as deprecated and will be removed in the
next major release. New automations should use `/api/v1/recording/tasks` (and
`/api/v1/recording/rules` for recurring rules).
## 12. Scoped REST and WebSocket APIs
The recording surface is exposed under `/api/v1/recording`:
```text
GET /api/v1/recording/tasks
POST /api/v1/recording/tasks
PATCH /api/v1/recording/tasks/{id}
POST /api/v1/recording/tasks/{id}/cancel
DELETE /api/v1/recording/tasks/{id}
POST /api/v1/recording/conflicts/preview
GET /api/v1/recording/quota
GET /api/v1/recording/rules
POST /api/v1/recording/rules
PATCH /api/v1/recording/rules/{id}
DELETE /api/v1/recording/rules/{id}?future=retain|cancel
```
The `tasks` payload is a per-session filtered snapshot. The WebSocket protocol carries
`RecordingSnapshotRequest` and `RecordingSnapshotResponse { revision, tasks }`; there is no
recording delta message — every recording change goes out as a `RecordingChanged` event, and the
client re-requests a filtered snapshot in response. The `revision` field is the monotonic
`QueueRevision`; clients that detect a revision gap must request a fresh filtered snapshot.
Two notifications exist for the recording subsystem:
- `RecordingChanged` (no payload) — broadcast whenever a task mutates the queue (create, edit,
cancel, delete, finalize, retry). Triggers a filtered snapshot refresh on every subscribed
client that holds `recording.read`.
- `RecordingRulesChanged` (no payload) — broadcast on every rule mutation (create, edit,
delete, retain/cancel). Used by the rules view to refresh without polling.
The cancel-recording-task endpoint emits **both** events because cancelling future rule
recordings mutates the queue as well as the rule store.
Filtering is server-side: private events go only to the owner session, shared events go to anyone
with `recording.read`, `LegacyAdmin` events go only to administrator sessions. Generic download
events (`DownloadsResponse`, `DownloadsDeltaResponse`) contain no recording tasks.
## 13. Conflict-preview advisory semantics
The conflict analyzer is **advisory** — runtime capacity is authoritative. The three-bucket
classification:
- `NoKnownConflict` — every segment of the candidate's padded interval is under capacity.
- `PossibleCapacityWait` — some segments are over.
- `LikelyMissedWindow` — every segment is over.
Create / edit operations return the preview's severity as a warning; the request still succeeds
when the hard checks (authorization, source, interval, padding, quota, path reservation) pass.
The preview endpoint accepts the same `CreateRecordingTaskRequest` and returns the preview.
Privacy: the preview never returns another task's id, title, channel, filename, or rule data.
Logs and the response only carry the provider scope, anonymized interval, and severity.
## 14. Recurring-rule matching, DST, and reconciliation
### 14.1 NewEpisode matching
The matching order is stable series id first, normalized title as a fallback. Explicit `Repeat`
airing is excluded when `exclude_repeat = true` (the default). `Unknown` airing is treated as
new. The UI surfaces the title-fallback limitation when the EPG does not publish a stable
series id.
### 14.2 WeeklyTimeslot matching
`weekday` is `1..=7` (Monday = 1, Sunday = 7). `local_start_time` is `HH:MM`. `timezone` is an
IANA zone. The scheduler handles DST:
- Ambiguous local time (fall-back) → the earlier instant.
- Nonexistent local time (spring-forward) → advance 1 hour at a time up to 4 hours.
The UI surfaces the DST + IANA behavior next to the timezone input.
### 14.3 Cross-store reconciliation
Two stores can drift when one commit succeeds and the other fails. The reconciliation pass
produces a list of `ReconcileAction`s the caller applies under the queue-mutation boundary. The
truth table:
| Situation | Action |
|-----------------------------------------------|----------------------------------------------|
| Materialized task, no `Scheduled` tombstone | `AddScheduledTombstone` |
| `Scheduled` tombstone, no task, rule enabled | `Materialize` |
| `Cancelled` tombstone, eligible inactive task | `Finalize` |
| `Cancelled` tombstone, **active** task | `ConflictingIntent` (log, manual resolution) |
| `Completed` tombstone, any task | suppress (no rematerialize) |
| Task terminal, tombstone `Scheduled` | `UpdateTombstone { kind: Completed }` |
| Tombstone past `expires_at` | `PruneTombstone` |
| Disabled rule | skip (no rematerialize) |
| Deleted rule | leave tombstone until expiry |
Tombstones are retained for the longer of the EPG horizon and the 14-day minimum horizon
(`MIN_TOMBSTONE_HORIZON_SECS`). The fixed cross-store lock order is
`queue mutation boundary → rule repository mutation`.
The reconciliation pass runs at startup, before the rule scheduler's first tick, so the scheduler
never plans against half-repaired state. Two notes on how the actions are applied:
- `Materialize` cannot be executed literally — an `occurrence_key` cannot be turned back into a
programme window. The orphan `Scheduled` tombstone is dropped instead, which lets the scheduler
re-plan that occurrence from the rule and the EPG on its next tick.
- `ConflictingIntent` is only logged (`recording::audit`, `warn`). An active recording is never
cancelled because of a stale intent.
Deletions interrupted by a crash are repaired in the same pass. For each task still carrying
`deleting_previous_state`, the physical file decides:
| File state | Action |
|-------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| Gone | Finish the deletion — remove the task from the queue. |
| Present, inside the recording root | Restore the prior terminal state and clear the marker. |
| Present, outside the recording root | Restore the prior state, leave the file alone, and log `recording_reconciliation_unsafe_path` (`recording::audit`, `warn`). Dropping the task instead would orphan a file nothing tracks. |
### 14.4 Delete with retain / cancel
`DELETE /api/v1/recording/rules/{id}?future=retain|cancel`:
- `retain` — set the rule's `enabled = false`. Existing tasks keep their `rule_id` and
`occurrence_key` for historical provenance. The scheduler stops materializing future
occurrences.
- `cancel` — same as `retain` plus cancel only future inactive occurrences. Active recordings
are **never** auto-cancelled by rule deletion; the operator must resolve manually.
`cancel` touches two stores — the queue and the rule repository — and they cannot commit together.
The queue side runs first and hands back a snapshot of every occurrence it cancelled; if the rule
delete then fails, those occurrences are **restored** from that snapshot (original state and
`reserved_bytes` included) before the error is returned. So a failed `future=cancel` leaves the
rule in place *and* its upcoming recordings intact. Only if the restore itself fails does the API
report `PartialOperation { primary: "future_cancelled", secondary: "rule_delete_failed" }`, which
tells the operator exactly which side won.
## 15. At-most-once notification delivery
The notification adapter follows the at-most-once protocol:
1. The queue-mutation boundary persists a `NotificationMarker` for the lifecycle event
(`Started` / `Completed` / `Failed`) in the same transaction as the state transition.
2. After the transaction commits, the adapter hands the notification to the **notification
outbox** instead of delivering it inline. The recorder never blocks on, or waits for, a
messaging provider.
3. The outbox worker persists the entry to `storage_dir/recording_notification_outbox.json`,
then attempts delivery **per channel**.
4. A channel that fails is retried with capped exponential backoff
(`notifications.backoff_initial_secs`, doubling, clamped to `backoff_max_secs`). A channel
that succeeded is removed from the entry, so a retry can never deliver a duplicate to a
channel that already got the message — this is what keeps retries compatible with
at-most-once.
5. After `notifications.max_attempts` the entry is dead-lettered: it is dropped from the outbox
and logged at `error` level under the `recording::audit` target with the event kind, the
attempt count, the channels that never accepted it, and the original enqueue time.
6. On restart the outbox file is reloaded and the backlog resumes from where it stopped. A
corrupt outbox file is logged and skipped rather than blocking startup.
7. A crash between the marker commit and the outbox write can still lose a notification. The
window is one file write wide, and the marker prevents a duplicate on the next boot.
If the in-memory channel to the worker is full (`notifications.outbox_buffer` entries pending),
the notification falls back to a single best-effort direct send. A recording is never delayed
because a messaging provider is down.
The routing decision:
- `Shared` → deliver to global channels.
- `Private` + `LegacyAdmin` owner → deliver.
- `Private` + administrator owner → deliver.
- `Private` + regular user owner → suppress.
Missing messaging configuration is a no-op; the adapter logs the dispatch decision and returns.
## 16. Migration checklist
1. **Stop or quiesce** recording activity. Cancel active recordings and let the queue drain.
2. **Back up** the existing config, `recordings.db` with its recovery generations, user /
auth config, and messaging config.
3. **Deploy** the version with additive normalization. No config changes are required for the
existing flows to keep working.
4. **Back up `web_user_ids.json`** after the first successful start. The bootstrap writes the
file automatically; the operator should preserve it across restarts.
5. **Grant `recording.read`, `recording.create`, `recording.manage` and `recording.delete`** explicitly to the user groups that need them.
The `permissions: 65535` legacy config does **not** implicitly grant the new bits.
6. **Refresh old tokens**. The schema bump forces a token refresh; pre-bump tokens get an
`X-Token-Refresh: required` 401. Users sign in again to receive the current claims.
7. **Verify the recording root and free space** with `statvfs` (Linux) /
`GetDiskFreeSpaceExW` (Windows). Confirm the `safety_bytes` is at least 1 GiB.
8. **Verify legacy recordings and paths**. The pre-Phase-1 `file_dir` / `file_path` fields
normalize to private `LegacyAdmin` recordings. Confirm the existing media files are within
the configured recording root or the legacy download root before enabling retention.
9. **Test one private and one shared recording** end-to-end before enabling the retention
worker in production.
10. **Enable retention / quotas gradually**. Start with `delete_after_days` only; add
`keep_last_per_channel` once the channel count is stable; add disk watermarks once the
free-space baseline is known.
11. **Wire the `/api/v1/recording` routes** in the frontends that need them. The new form
component (`recording_form`) is the single source of truth for both Playlist Explorer and
EPG.
## 17. Acceptance scenarios (full sweep)
The acceptance scenarios the operator should verify before declaring the migration done:
1. Old configuration loads without data loss. The recording queue does **not** migrate: a
pre-existing `recordings_state.json` is ignored and the queue starts empty.
2. Invalid recording kind / metadata combinations fail with `recording_invalid_state` or
`recording_invalid_source`.
3. Queue persistence failure leaves the state and revision unchanged and emits no delta.
4. WebSocket revision gaps trigger a filtered resnapshot.
5. A private recording is invisible to a second user in tasks, deltas, catalog, playback,
conflicts, quota, and logs.
6. Administrators can create shared recordings but do not see another user's private recordings
through ordinary endpoints.
7. New APIs cannot submit raw URLs, owner IDs, absolute paths, or filenames — the wire shape
is server-owned identifiers only.
8. Worker source re-resolution rejects stale / tampered source ownership
(`recording_invalid_source`).
9. Partial files are not clobbered; finalization never overwrites an external file.
10. Crash recovery after each worker lifecycle point is deterministic and does not replay a
missed live window.
11. Completed / failed / cancelled deletion removes only the task-owned final / partial file
safely.
12. A deletion failure preserves the task and file.
13. Successful file removal plus final persistence failure remains recoverable as `Deleting`.
14. Symlink, non-regular, and containment attacks are rejected with `recording_unsafe_path`.
15. Filename templates sanitize, truncate, and reserve collisions safely, including the
`{owner}` filename-only exception.
16. Private and shared quota pools are independent under concurrent create and start operations.
17. Active recordings remain conservatively charged; measured growth cannot create false free
quota.
18. Disk safety and retention use the recording-root filesystem.
19. Retention never deletes generic downloads, active recordings, partials, unsafe legacy files,
or orphan files.
20. DVR entries never enter Movies / Series or global user-independent caches.
21. Every playback / range / download open is authorized again.
22. EPG and Playlist Explorer use one form / one API.
23. Currently-airing scheduling reserves only the remaining duration and rejects elapsed windows.
24. Padding changes execution, quota, and conflicts consistently.
25. Editing succeeds only in allowed upcoming states and rolls back completely on persistence /
quota / path failure.
26. Conflict warnings follow the deterministic classification and redact private metadata.
27. New-episode / weekly rules materialize only in horizon and survive DST / restart.
28. Task / tombstone reconciliation is idempotent after every injected cross-store failure.
29. Deleted, cancelled, and completed occurrences are not recreated within the tombstone horizon.
30. Notifications make no more than one external attempt per committed marker.
31. Missing / disabled messaging never fails a recording.
32. Production Rust remains free of `unwrap`, `expect`, and `panic` additions.
If a command or scenario fails, the operator records the exact command and output, fixes the
smallest root cause, reruns the focused test, and then the full relevant phase gate. Unexplained
known failures do not count as completion.
## 18. Rollback
The DVR is a feature flag, so a rollback does not require a binary downgrade:
```yaml
video:
recording:
enabled: false
```
That stops the supervisors and the rule scheduler, answers `501 recording_disabled` on the
recording routes, serves no recording data over the WebSocket, and hides the sidebar entries.
Existing recordings and the queue are left untouched, so re-enabling resumes where you left off.
Keep the previous binary available as well: the notification outbox writes
`storage_dir/recording_notification_outbox.json`, which an older binary does not know about. It is
ignored rather than misread — an unknown file in `storage_dir` is harmless — but the queued
notifications in it are not delivered until the newer binary runs again.
## 19. Verifying the installation
After deployment or configuration changes, verify supervisor health:
```bash
# Supervisor liveness. Administrator token required.
curl -s -H "Authorization: Bearer $TOKEN" \
http://localhost:8901/api/v1/recording/health | jq
```
A healthy install shows a non-null `reconciliation_last_run` (stamped once at boot) and a
`retention_last_tick` no older than `disk.cleanup_interval_secs`. A `null` value means that
supervisor has never completed a pass.
`bin/dvr_doctor.sh --token "$ADMIN_TOKEN"` wraps this up with the on-disk state — in particular a
`stuck_deleting` count, which should be `0` after a clean boot.
Then check the log for the two lines worth reacting to:
- `recording is enabled with no retention, no disk watermarks, and no quota` — nothing bounds
recording disk usage. Intentional on a dedicated filesystem; a mistake otherwise.
- `enabled NewEpisode recording rule(s) cannot match` — see the limitation below.
### 19.1 Known limitation: `NewEpisode` rules
`NewEpisode` rules do not currently match anything. The scheduler matches them by walking EPG
programmes, and no EPG horizon is supplied to it yet, so only `WeeklyTimeslot` rules materialize.
The condition is logged once per process rather than failing quietly.
Until this is wired, record a recurring programme with a `WeeklyTimeslot` rule, or record
individual programmes from the EPG view.