The documentation described a system that no longer exists, and in two places described the opposite of what the code now does. - The layout section documented `<recording-root>/users/<owner-id>/<rel>` for private recordings and `shared/<rel>` for shared ones. That resolver was deleted: recordings are stored owner-independently at `<recording-root>/<rel>`, because one physical file is shared by every user who asked for it. The organised layouts and the component sanitisation rules are now documented as they are implemented. - The config reference claimed persisted queue recovery "is tolerant of corruption" and "starts with an empty transfer queue instead of aborting server boot". The queue now fails closed: a damaged database is rebuilt from the recovery history, and a database ahead of every surviving history refuses to start. An operator following the old text would have expected silent recovery from a condition that is deliberately fatal. - `download.read` / `download.write` were still listed as grantable permissions after their removal, and `recording.write` after its split. bin/dvr_doctor.sh looked for `downloads_state.json` and summarised it with jq. That file never existed under this name, and the queue it stood for is now a B+Tree, so the section printed "(absent)" and skipped its summary exactly when an operator needed it. It now reports the repository and its recovery generations: the CURRENT pointer, the retained generation pair, journal sizes, the fail-closed case where a database has no history, and a warning when the recovery directory shares a filesystem with the database — which survives a corrupt file but not the loss of the volume it exists to protect against. CHANGELOG records the three breaking changes: the non-migrating queue, the permission split, and the moved recording files.
44 KiB
DVR Operator Reference
This guide covers everything an operator needs to know to deploy, configure, and manage the
DVR. It is the single source of truth for the documentation that
config.md and rest-api-cookbook.md
consume.
1. Configuration reference
All fields live under video.recording in the config file. Every field is optional;
defaults match the recommended values.
video:
recording:
enabled: true # default: true; false stops every DVR supervisor
container_format: mpegts # mpegts (default) | matroska | mp4
directory: recordings/ # default: recordings
timezone: Europe/Berlin # default: UTC (IANA required)
filename_template: "{channel}_{program_title}_{start_time}"
default_pre_roll_secs: 0 # 0..=max_pre_roll_secs
max_pre_roll_secs: 900 # ≤ 900 (15 min)
default_post_roll_secs: 0 # 0..=max_post_roll_secs
max_post_roll_secs: 1800 # ≤ 1800 (30 min)
retention:
keep_last_per_channel: 10 # > 0 when set
delete_after_days: 30 # > 0 when set
sweep_interval_secs: 3600 # default 3600; age/count sweep cadence
disk:
high_water_percent: 85 # 0..=100
low_water_percent: 70 # 0..=100 and < high_water_percent
cleanup_interval_secs: 3600 # > 0; watermark-check cadence
safety_bytes: 1073741824 # > 0 (1 GiB)
quota:
default_private_bytes: 53687091200 # 50 GiB
per_user_bytes:
"web:user-uuid-1": 107374182400 # 100 GiB
shared_bytes: 536870912000 # 500 GiB
notifications:
outbox_buffer: 1024 # default 1024; in-memory queue depth
max_attempts: 6 # default 6; then dead-lettered
backoff_initial_secs: 5 # default 5
backoff_max_secs: 900 # default 900 (15 min)
fallback_bytes_per_minute: 8388608 # 8 MiB/min, > 0
| Field | Default | Range | Restart required | Effect |
|---|---|---|---|---|
enabled |
true |
bool | no | false skips reconciliation, retention, and the notification outbox at startup, makes running supervisors idle on their next tick, answers every /api/v1/recording/** route with 501 recording_disabled, and hides the DVR entries from the web UI navigation. |
container_format |
mpegts |
mpegts, matroska, mp4 |
no | The -f muxer ffmpeg writes. mpegts survives truncation, so a recording killed mid-stream still plays — prefer it unless the source codecs need another container. Applies to recordings that start after the change. |
retention.sweep_interval_secs |
3600 |
> 0 | no | Cadence of the age/count sweep. Independent of disk.cleanup_interval_secs. |
disk.cleanup_interval_secs |
3600 |
> 0 | no | Cadence of the supervisor's tick, and therefore of the watermark check. Measurements are floored at one per 30 s. |
notifications.outbox_buffer |
1024 |
≥ 1 | yes | Channel capacity between the recorder and the outbox worker. Fixed when the worker starts. |
notifications.max_attempts |
6 |
≥ 1 | no | Delivery attempts per notification before it is dead-lettered. |
notifications.backoff_initial_secs |
5 |
≥ 1 | no | First retry delay; doubles per attempt. |
notifications.backoff_max_secs |
900 |
≥ backoff_initial_secs |
no | Ceiling for the doubling. |
Choosing values:
- Private-only home use — leave
quotaunset, setretention.keep_last_per_channelto taste, and leavediskunset if the recording filesystem is dedicated. - Shared family install — set
quota.shared_bytesanddisk.{high,low}_water_percentso a full disk degrades into retention rather than failed recordings. - Multi-tenant — set
quota.default_private_bytesplus per-user overrides, and keepnotifications.max_attemptslow so a broken webhook does not accumulate a backlog.
1.0.1 ⚠️ Retention policy warning
Retention deletes recordings. If you are upgrading an existing install or reconfiguring
retention, whatever retention and disk values are in your config take effect on the next
supervisor sweep. If they were set optimistically or copied from an example, the first sweep may
delete recordings you expected to keep.
Before restarting after a configuration change:
- Read back your effective policy. Check
video.recording.retentionandvideo.recording.diskinconfig.yml. - Work out what would be deleted.
keep_last_per_channel: Nkeeps the N most recent recordings per (owner, channel) and deletes the rest.delete_after_days: Ndeletes anything whosecompleted_atis more than N days old. The two are a union, not an intersection — a recording matching either policy is deleted. - If you are unsure, start with retention off. Remove the
retentionblock (or the individual keys) and setenabled: truewith no policy. Nothing is deleted, and you can enable a policy deliberately once you have looked at the library. - Back up
recordings.db, the recovery generations underbackup_dir, and the recording directory.
There is no dry-run mode. Deletions are logged under the recording::audit target with a
recording_retention_delete line and a reason (Age, Count, or watermark), so
grep recording_retention_delete after the first sweep tells you exactly what went.
1.1 Validation rules
The validation list, applied at VideoConfigDto::prepare:
keep_last_per_channel > 0when set.delete_after_days > 0when set.cleanup_interval_secs > 0.low_water_percent < high_water_percentand both in0..=100.safety_bytes > 0.default_private_bytes,per_user_bytes[*],shared_bytesall> 0when set.fallback_bytes_per_minute > 0.timezoneis a valid IANA zone (chrono_tz::Tz::parse).filename_templatecontains at least one known placeholder and is no longer than 240 UTF-8 bytes.
Absent quotas and absent retention values disable the corresponding policy (no limit).
1.2 Disk watermark semantics
high_water_percent and low_water_percent are used-space percentages on the filesystem
containing the canonical recording root. The retention worker uses them with hysteresis: when
used-space ≥ high_water_percent, the worker deletes oldest completed recordings until
used-space ≤ low_water_percent (or the eligible list is exhausted). cleanup_interval_secs is
the wall-clock interval between passes; the worker uses a cancellation-aware Tokio interval and
will not overlap.
Free space is measured once per pass, not once per deletion. The stop condition folds in the
bytes the pass has already reclaimed, so a pass deletes just enough recordings to reach the low
watermark and then stops. Only the recording root's own filesystem is measured — never
storage_dir or the generic download directory, which may sit on a different mount.
1.3 What changed: supervisors now run
Three background supervisors now actually execute their work. They existed as decision layers before but were never started, so the DVR worked on the happy path and silently skipped everything else. The consequences of switching them on are all things this guide describes, but they happen now where they did not before:
| Supervisor | What now happens |
|---|---|
| Startup reconciliation | Recordings stuck in Deleting from an earlier crash are finished or restored on boot. Orphaned rule tombstones are repaired. |
| Retention | Age, count, and disk-watermark deletion begin. See the retention warning above. |
| Notification outbox | Lifecycle notifications are retried and persisted instead of being dropped on transient error. |
1.4 Supervisors
Three background supervisors implement the behaviour described in the rest of this document. All
three are started once the HTTP listener is bound, and all three honour the downloads
cancellation token, so a config reload stops and restarts them cleanly.
| Supervisor | Cadence | Responsibility |
|---|---|---|
| Startup reconciliation | once at boot, before the rule scheduler | Finishes or undoes deletions interrupted by a crash; repairs queue/rule-store drift. |
| Retention | disk.cleanup_interval_secs tick, policy sweep every retention.sweep_interval_secs |
Age, count, and watermark deletion — all through the single system_retention_delete path. |
| Notification outbox | event-driven, with per-entry retry timers | Durable lifecycle-notification delivery with per-channel retry and dead-lettering. |
Passes never overlap: a tick that arrives while the previous pass is still deleting is skipped.
GET /api/v1/recording/health (administrator only) reports each supervisor's last-tick timestamp,
the outbox depth, and the dead-letter count, so liveness can be checked without reading the log:
{
"enabled": true,
"server_time": 1700003600,
"reconciliation_last_run": 1700000000,
"retention_last_tick": 1700003400,
"retention_sweep_interval_secs": 3600,
"notification_last_drain": 1700003100,
"notification_outbox_depth": 0,
"notification_dead_lettered": 0,
"queue_revision": 412
}
A null timestamp means that supervisor has never completed a pass. Compare retention_last_tick
against server_time and disk.cleanup_interval_secs to detect a stalled sweep.
1.5 Support diagnostics
bin/dvr_doctor.sh collects everything above plus the on-disk state into one dump suitable for a
support ticket. It is read-only and never touches config, the queue, or a recording.
bin/dvr_doctor.sh --token "$ADMIN_TOKEN"
bin/dvr_doctor.sh --url https://tuliprox.example --token "$ADMIN_TOKEN" --storage-dir /opt/tuliprox/data
It reports supervisor health, the effective recording config block, quota, the recording
repository and its recovery generations, and aggregate summaries of recording_rules.json and
the notification outbox. The summaries are deliberately aggregate: no titles, filenames, or
owner ids are printed, because a diagnostics dump gets pasted into tickets. The health and
config sections need an administrator token; the on-disk sections work without one, and take
--backup-dir because the recovery generations live under backup_dir, not storage_dir.
2. Recording directory layout and immutable IDs
The recording root is the path configured in recording.directory (default
<download-dir>/recordings). Every recording is stored at <recording-root>/<rel>, where
<rel> is the collision-safe relative path reserved at the queue-mutation boundary.
The layout carries no owner or visibility component. One physical file is shared by every user who requested it, so keying its directory on an owner would be wrong the moment a second user attaches, and would force the file to move on disk when the first detaches.
With organize_into_directories enabled the layout groups by what the server resolved:
<recording-root>/
<channel>/<file> # live
<title>/<file> # vod
<series>/Season NN/<file> # series episode
<file> # unorganized, or no grouping resolved
Path components are sanitised to a single filesystem-safe component: separators, control
characters and the Windows-forbidden set collapse to _, leading and trailing dots are
stripped, Windows device names fall back to a placeholder, and each component is capped at 255
bytes on a character boundary. The path is validated against the recording root at open time;
any path that escapes the root (via .., symlinks, or absolute paths) is rejected with
recording_unsafe_path.
3. Filename placeholders
The supported placeholders:
{channel}— channel display name; falls back to the stable channel id when the display name is missing.{program_title}— programme title; sanitized via the existing filename sanitizer.{start_time}— programme start in the configured timezone, rendered asYYYY-MM-DD_HH-mm.{end_time}— same, programme end.{episode}— rendersSxxExxwhen both season and episode numbers exist; otherwise empty.{owner}— sanitized owner display name. Filename only; never a directory component. The{owner}placeholder is the only place the username-derived content may appear in the on-disk path.
The final stem is capped at 240 UTF-8 bytes without splitting a code point. When the sanitized stem is empty, the runtime falls back to the recording task id.
4. Lifecycle and restart behavior
Every recording goes through three persistence states:
- Partial path —
partial_relative_pathis set; theO_CREAT | O_EXCL | O_NOFOLLOWopen guarantees an attacker-prepared symlink at the partial path is rejected. - Finalize — atomic rename from partial to final; the runtime refuses to clobber any pre-existing final file (including symlinks).
- Complete —
mutatesetsCompleted, stampsmeasured_bytes,completed_at, and clearsreserved_bytes.
A crash at any point is recoverable:
- Active + valid final file → normalize to
Completed. - Active + valid partial file → normalize to terminal
Failed, retain partial. - Active + no owned file → normalize to terminal
Failed. - Unsafe path / type → log a security-category error, normalize to previous terminal state, do not open.
5. Quota charge-by-state
The quota ledger charges the bytes below per DownloadState:
| State | Charged bytes |
|---|---|
Scheduled / Queued / WaitingForCapacity / RetryWaiting / Paused |
reserved_bytes |
Downloading |
max(reserved_bytes, measured_bytes) |
Completed |
final measured_bytes |
Failed / Cancelled with partial file |
partial measured_bytes |
A task that is mid-deletion carries deleting_previous_state = Some(prior) and is charged the
same bytes as the prior terminal state (reserved_bytes or measured_bytes depending on what
prior was); the field replaces the historical DownloadState::Deleting variant, which the
runtime no longer carries. The charge drops to zero only when finalize_deletion removes the
task from the queue.
A task is counted exactly once. Private pools key on RecordingOwner::User(uid); shared pools key
on RecordingVisibility::Shared; LegacyAdmin recordings count toward the shared pool. Per-user
overrides beat the configured default; an absent limit is unlimited.
5.1 Active overrun policy
Version one does not terminate an active recording because it grew beyond quota. would_exceed
is admission-only. When the measured partial size exceeds the reservation, the charge is
max(reserved, measured); the next would_exceed call rejects new admissions until the recording
finishes or is deleted. The user-visible DTO surfaces an Overrun warning so operators can grant
more quota or delete the recording.
5.2 Unknown bitrate
When the bitrate is unknown at create time, the reservation is
duration_minutes × fallback_bytes_per_minute (default 8 MiB). The DTO surfaces an
UnknownBitrate warning. The runtime re-reserves with the measured rate as soon as the worker
starts.
6. Disk admission
The disk admission path:
headroom = free_bytes - safety_bytes - active_disk_reservations
admit = charge <= headroom
free_bytes_for(path) is statvfs (Unix) or GetDiskFreeSpaceExW (Windows) keyed on the supplied
path's mount. The pre-start flow always passes the canonical recording root so the measurement is
on the same filesystem the file will live on. Two starts cannot consume the same headroom — the
active reservation is serialized through the queue-mutation boundary.
7. Safe deletion guarantees
Deletions use a persisted two-phase operation. The runtime carries the deletion intent in
recording.deleting_previous_state: Option<DeletingPreviousState> rather than as a
DownloadState variant — every terminal task that is mid-deletion stays in its prior state
(Completed / Failed / Cancelled) but carries the marker, which is what the rest of this
section means by "the task is in the deleting phase".
begin_deletionruns inside the queue-mutation boundary. It stampsrecording.deleting_previous_state = Some(prior)(the prior terminal state) and zeros the byte counts.execute_deletionruns after the boundary. It inspects the path withsymlink_metadata(nevermetadata), so a symlink is seen as a symlink and refused rather than dereferenced, and removes the file. Missing files are idempotent success.finalize_deletionruns inside a fresh boundary. It removes the task from the queue and clearsdeleting_previous_state.
Startup recovery (any task whose deleting_previous_state is Some(_)):
deleting_previous_state = Some(_)+ missing file → finish task removal.deleting_previous_state = Some(_)+ existing valid regular file inside the recording root → restore the prior terminal state, clear the marker.deleting_previous_state = Some(_)+ unsafe path or non-regular file → restore the prior state, logrecording_reconciliation_unsafe_path, leave the file alone.
7.1 Portability of the path guarantees
The four guarantees — no symlink is followed, no existing file is clobbered, the publish is atomic, nothing escapes the recording root — are built from portable primitives and hold identically on every supported target:
| Guarantee | Primitive | Portable? |
|---|---|---|
| No symlink followed on inspection | symlink_metadata (never metadata) |
yes |
| No existing file clobbered | create_new → O_CREAT|O_EXCL / CREATE_NEW; both fail on an existing entry including a dangling symlink |
yes |
| Atomic publish | rename, after a no-follow existence check on the destination |
yes |
| Contained in the recording root | component validation plus an owner-id component check, before any syscall | yes |
Only one call has a platform-specific branch: open_partial_no_clobber additionally passes
O_NOFOLLOW on Unix. That is defense in depth, not the mechanism — the no-clobber property
already comes from create_new. openat2 with RESOLVE_BENEATH / RESOLVE_NO_SYMLINKS would be
Linux-only and is deliberately not used.
Earlier revisions carried a blanket #![cfg(unix)] on the path helper, which removed the module
wholesale on Windows and left every caller with unresolved imports — the DVR did not build on
Windows at all. The gate is now scoped to the single O_NOFOLLOW line. Tests that need to
create a symlink stay Unix-only (Windows requires developer mode or elevation for that); the
behaviour they cover is asserted portably by the no-clobber tests.
8. Authorization matrix
| Operation | Private recording | Shared recording | LegacyAdmin |
Orphan |
|---|---|---|---|---|
| Read / Playback / Download | owner with recording.read |
anyone with recording.read |
admin only | admin only |
| Create private | user with recording.create |
n/a | admin only | n/a |
| Create shared | rejected (admin only) | admin + recording.create |
admin only | n/a |
| Edit / Cancel | owner + recording.manage |
admin + recording.manage |
admin only | n/a |
| Delete | owner + recording.delete |
admin + recording.delete |
admin only | n/a |
| Manage recurring rule | owner + recording.manage |
admin + recording.manage |
admin only | n/a |
SystemRetentionDelete |
ownership bypassed; state-gated | ownership bypassed; state-gated | ownership bypassed; state-gated | n/a |
| Orphan catalog | n/a | n/a | n/a | admin only |
Administrators do not implicitly receive another regular user's private recording content. The
private owner is the only non-administrator allowed to read it. Administrative access is read-only
for diagnosis; mutations require either the SystemRetentionDelete action (which the retention
worker is the only legitimate caller of) or the appropriate recording permission + ownership
combination.
Orphan catalog entries (recordings whose target/input no longer matches a configured source) are
visible only to administrators with recording.read. The path is never exposed; an opaque orphan
id is generated per discovery.
9. Identity-registry bootstrap
The identity registry is web_user_ids.json in the storage directory. The startup sequence is:
- Pre-scan the recording repository for
RecordingOwner::User(_)entries (without the registry loaded). - Load the existing registry (if any).
- Initialize the registry only when no persisted real owner exists. New
UserIds are generated for any username that lacks one. - Fail closed on missing / corrupt registry when real owners exist. The server does not generate replacement IDs in this case; the operator must restore the registry or run an explicit rename migration.
- Sync current principals (insert a new
UserIdfor any username that lacks one). - Run the full queue load + normalization.
The built-in administrator is the reserved subject id builtin:admin (constant). Operators do not
create an entry for it.
10. Token refresh on permission schema bump
Claims carries subject_id: Option<UserId> and permission_schema_version: u16. The constant
CURRENT_PERMISSION_SCHEMA_VERSION is the source of truth. When the schema changes, bump the
constant; pre-bump tokens become stale:
authenticator::validate_token_versionreturnsAuthError::StaleSchemafor older versions.- The HTTP layer emits a 401 with header
X-Token-Refresh: required. - The frontend's
RecordingError::TokenRefreshRequiredand the generic auth refresh handler redirect the user back to the sign-in flow.
Operators do not need to manually invalidate tokens on a schema bump. Existing user records in
web_user_ids.json are preserved; only the subject_id mapping for current usernames is
recomputed if missing.
11. Deprecated /file/record behavior
The legacy POST /file/record route is deprecated and delegates to
RecordingService::create_recording for administrators only. Non-administrators receive a 403 —
the deprecated route does not bypass the new policy.
The migration:
- Frontend code: switch from
downloads_service::queue_recordingtorecording_service::RecordingService::create_task. The new client submitsRecordingSourceInput(target_id + virtual_id + input_name) andCreateRecordingTaskRequest, never a free-form URL. - Operator code: the legacy route is documented as deprecated and will be removed in the
next major release. New automations should use
/api/v1/recording/tasks(and/api/v1/recording/rulesfor recurring rules).
12. Scoped REST and WebSocket APIs
The recording surface is exposed under /api/v1/recording:
GET /api/v1/recording/tasks
POST /api/v1/recording/tasks
PATCH /api/v1/recording/tasks/{id}
POST /api/v1/recording/tasks/{id}/cancel
DELETE /api/v1/recording/tasks/{id}
POST /api/v1/recording/conflicts/preview
GET /api/v1/recording/quota
GET /api/v1/recording/rules
POST /api/v1/recording/rules
PATCH /api/v1/recording/rules/{id}
DELETE /api/v1/recording/rules/{id}?future=retain|cancel
The tasks payload is a per-session filtered snapshot. The WebSocket protocol carries
RecordingSnapshotRequest and RecordingSnapshotResponse { revision, tasks }; there is no
recording delta message — every recording change goes out as a RecordingChanged event, and the
client re-requests a filtered snapshot in response. The revision field is the monotonic
QueueRevision; clients that detect a revision gap must request a fresh filtered snapshot.
Two notifications exist for the recording subsystem:
RecordingChanged(no payload) — broadcast whenever a task mutates the queue (create, edit, cancel, delete, finalize, retry). Triggers a filtered snapshot refresh on every subscribed client that holdsrecording.read.RecordingRulesChanged(no payload) — broadcast on every rule mutation (create, edit, delete, retain/cancel). Used by the rules view to refresh without polling.
The cancel-recording-task endpoint emits both events because cancelling future rule recordings mutates the queue as well as the rule store.
Filtering is server-side: private events go only to the owner session, shared events go to anyone
with recording.read, LegacyAdmin events go only to administrator sessions. Generic download
events (DownloadsResponse, DownloadsDeltaResponse) contain no recording tasks.
13. Conflict-preview advisory semantics
The conflict analyzer is advisory — runtime capacity is authoritative. The three-bucket classification:
NoKnownConflict— every segment of the candidate's padded interval is under capacity.PossibleCapacityWait— some segments are over.LikelyMissedWindow— every segment is over.
Create / edit operations return the preview's severity as a warning; the request still succeeds
when the hard checks (authorization, source, interval, padding, quota, path reservation) pass.
The preview endpoint accepts the same CreateRecordingTaskRequest and returns the preview.
Privacy: the preview never returns another task's id, title, channel, filename, or rule data. Logs and the response only carry the provider scope, anonymized interval, and severity.
14. Recurring-rule matching, DST, and reconciliation
14.1 NewEpisode matching
The matching order is stable series id first, normalized title as a fallback. Explicit Repeat
airing is excluded when exclude_repeat = true (the default). Unknown airing is treated as
new. The UI surfaces the title-fallback limitation when the EPG does not publish a stable
series id.
14.2 WeeklyTimeslot matching
weekday is 1..=7 (Monday = 1, Sunday = 7). local_start_time is HH:MM. timezone is an
IANA zone. The scheduler handles DST:
- Ambiguous local time (fall-back) → the earlier instant.
- Nonexistent local time (spring-forward) → advance 1 hour at a time up to 4 hours.
The UI surfaces the DST + IANA behavior next to the timezone input.
14.3 Cross-store reconciliation
Two stores can drift when one commit succeeds and the other fails. The reconciliation pass
produces a list of ReconcileActions the caller applies under the queue-mutation boundary. The
truth table:
| Situation | Action |
|---|---|
Materialized task, no Scheduled tombstone |
AddScheduledTombstone |
Scheduled tombstone, no task, rule enabled |
Materialize |
Cancelled tombstone, eligible inactive task |
Finalize |
Cancelled tombstone, active task |
ConflictingIntent (log, manual resolution) |
Completed tombstone, any task |
suppress (no rematerialize) |
Task terminal, tombstone Scheduled |
UpdateTombstone { kind: Completed } |
Tombstone past expires_at |
PruneTombstone |
| Disabled rule | skip (no rematerialize) |
| Deleted rule | leave tombstone until expiry |
Tombstones are retained for the longer of the EPG horizon and the 14-day minimum horizon
(MIN_TOMBSTONE_HORIZON_SECS). The fixed cross-store lock order is
queue mutation boundary → rule repository mutation.
The reconciliation pass runs at startup, before the rule scheduler's first tick, so the scheduler never plans against half-repaired state. Two notes on how the actions are applied:
Materializecannot be executed literally — anoccurrence_keycannot be turned back into a programme window. The orphanScheduledtombstone is dropped instead, which lets the scheduler re-plan that occurrence from the rule and the EPG on its next tick.ConflictingIntentis only logged (recording::audit,warn). An active recording is never cancelled because of a stale intent.
Deletions interrupted by a crash are repaired in the same pass. For each task still carrying
deleting_previous_state, the physical file decides:
| File state | Action |
|---|---|
| Gone | Finish the deletion — remove the task from the queue. |
| Present, inside the recording root | Restore the prior terminal state and clear the marker. |
| Present, outside the recording root | Restore the prior state, leave the file alone, and log recording_reconciliation_unsafe_path (recording::audit, warn). Dropping the task instead would orphan a file nothing tracks. |
14.4 Delete with retain / cancel
DELETE /api/v1/recording/rules/{id}?future=retain|cancel:
retain— set the rule'senabled = false. Existing tasks keep theirrule_idandoccurrence_keyfor historical provenance. The scheduler stops materializing future occurrences.cancel— same asretainplus cancel only future inactive occurrences. Active recordings are never auto-cancelled by rule deletion; the operator must resolve manually.
cancel touches two stores — the queue and the rule repository — and they cannot commit together.
The queue side runs first and hands back a snapshot of every occurrence it cancelled; if the rule
delete then fails, those occurrences are restored from that snapshot (original state and
reserved_bytes included) before the error is returned. So a failed future=cancel leaves the
rule in place and its upcoming recordings intact. Only if the restore itself fails does the API
report PartialOperation { primary: "future_cancelled", secondary: "rule_delete_failed" }, which
tells the operator exactly which side won.
15. At-most-once notification delivery
The notification adapter follows the at-most-once protocol:
- The queue-mutation boundary persists a
NotificationMarkerfor the lifecycle event (Started/Completed/Failed) in the same transaction as the state transition. - After the transaction commits, the adapter hands the notification to the notification outbox instead of delivering it inline. The recorder never blocks on, or waits for, a messaging provider.
- The outbox worker persists the entry to
storage_dir/recording_notification_outbox.json, then attempts delivery per channel. - A channel that fails is retried with capped exponential backoff
(
notifications.backoff_initial_secs, doubling, clamped tobackoff_max_secs). A channel that succeeded is removed from the entry, so a retry can never deliver a duplicate to a channel that already got the message — this is what keeps retries compatible with at-most-once. - After
notifications.max_attemptsthe entry is dead-lettered: it is dropped from the outbox and logged aterrorlevel under therecording::audittarget with the event kind, the attempt count, the channels that never accepted it, and the original enqueue time. - On restart the outbox file is reloaded and the backlog resumes from where it stopped. A corrupt outbox file is logged and skipped rather than blocking startup.
- A crash between the marker commit and the outbox write can still lose a notification. The window is one file write wide, and the marker prevents a duplicate on the next boot.
If the in-memory channel to the worker is full (notifications.outbox_buffer entries pending),
the notification falls back to a single best-effort direct send. A recording is never delayed
because a messaging provider is down.
The routing decision:
Shared→ deliver to global channels.Private+LegacyAdminowner → deliver.Private+ administrator owner → deliver.Private+ regular user owner → suppress.
Missing messaging configuration is a no-op; the adapter logs the dispatch decision and returns.
16. Migration checklist
- Stop or quiesce recording activity. Cancel active recordings and let the queue drain.
- Back up the existing config,
recordings.dbwith its recovery generations, user / auth config, and messaging config. - Deploy the version with additive normalization. No config changes are required for the existing flows to keep working.
- Back up
web_user_ids.jsonafter the first successful start. The bootstrap writes the file automatically; the operator should preserve it across restarts. - Grant
recording.read,recording.create,recording.manageandrecording.deleteexplicitly to the user groups that need them. Thepermissions: 65535legacy config does not implicitly grant the new bits. - Refresh old tokens. The schema bump forces a token refresh; pre-bump tokens get an
X-Token-Refresh: required401. Users sign in again to receive the current claims. - Verify the recording root and free space with
statvfs(Linux) /GetDiskFreeSpaceExW(Windows). Confirm thesafety_bytesis at least 1 GiB. - Verify legacy recordings and paths. The pre-Phase-1
file_dir/file_pathfields normalize to privateLegacyAdminrecordings. Confirm the existing media files are within the configured recording root or the legacy download root before enabling retention. - Test one private and one shared recording end-to-end before enabling the retention worker in production.
- Enable retention / quotas gradually. Start with
delete_after_daysonly; addkeep_last_per_channelonce the channel count is stable; add disk watermarks once the free-space baseline is known. - Wire the
/api/v1/recordingroutes in the frontends that need them. The new form component (recording_form) is the single source of truth for both Playlist Explorer and EPG.
17. Acceptance scenarios (full sweep)
The acceptance scenarios the operator should verify before declaring the migration done:
- Old configuration loads without data loss. The recording queue does not migrate: a
pre-existing
recordings_state.jsonis ignored and the queue starts empty. - Invalid recording kind / metadata combinations fail with
recording_invalid_stateorrecording_invalid_source. - Queue persistence failure leaves the state and revision unchanged and emits no delta.
- WebSocket revision gaps trigger a filtered resnapshot.
- A private recording is invisible to a second user in tasks, deltas, catalog, playback, conflicts, quota, and logs.
- Administrators can create shared recordings but do not see another user's private recordings through ordinary endpoints.
- New APIs cannot submit raw URLs, owner IDs, absolute paths, or filenames — the wire shape is server-owned identifiers only.
- Worker source re-resolution rejects stale / tampered source ownership
(
recording_invalid_source). - Partial files are not clobbered; finalization never overwrites an external file.
- Crash recovery after each worker lifecycle point is deterministic and does not replay a missed live window.
- Completed / failed / cancelled deletion removes only the task-owned final / partial file safely.
- A deletion failure preserves the task and file.
- Successful file removal plus final persistence failure remains recoverable as
Deleting. - Symlink, non-regular, and containment attacks are rejected with
recording_unsafe_path. - Filename templates sanitize, truncate, and reserve collisions safely, including the
{owner}filename-only exception. - Private and shared quota pools are independent under concurrent create and start operations.
- Active recordings remain conservatively charged; measured growth cannot create false free quota.
- Disk safety and retention use the recording-root filesystem.
- Retention never deletes generic downloads, active recordings, partials, unsafe legacy files, or orphan files.
- DVR entries never enter Movies / Series or global user-independent caches.
- Every playback / range / download open is authorized again.
- EPG and Playlist Explorer use one form / one API.
- Currently-airing scheduling reserves only the remaining duration and rejects elapsed windows.
- Padding changes execution, quota, and conflicts consistently.
- Editing succeeds only in allowed upcoming states and rolls back completely on persistence / quota / path failure.
- Conflict warnings follow the deterministic classification and redact private metadata.
- New-episode / weekly rules materialize only in horizon and survive DST / restart.
- Task / tombstone reconciliation is idempotent after every injected cross-store failure.
- Deleted, cancelled, and completed occurrences are not recreated within the tombstone horizon.
- Notifications make no more than one external attempt per committed marker.
- Missing / disabled messaging never fails a recording.
- Production Rust remains free of
unwrap,expect, andpanicadditions.
If a command or scenario fails, the operator records the exact command and output, fixes the smallest root cause, reruns the focused test, and then the full relevant phase gate. Unexplained known failures do not count as completion.
18. Rollback
The DVR is a feature flag, so a rollback does not require a binary downgrade:
video:
recording:
enabled: false
That stops the supervisors and the rule scheduler, answers 501 recording_disabled on the
recording routes, serves no recording data over the WebSocket, and hides the sidebar entries.
Existing recordings and the queue are left untouched, so re-enabling resumes where you left off.
Keep the previous binary available as well: the notification outbox writes
storage_dir/recording_notification_outbox.json, which an older binary does not know about. It is
ignored rather than misread — an unknown file in storage_dir is harmless — but the queued
notifications in it are not delivered until the newer binary runs again.
19. Verifying the installation
After deployment or configuration changes, verify supervisor health:
# Supervisor liveness. Administrator token required.
curl -s -H "Authorization: Bearer $TOKEN" \
http://localhost:8901/api/v1/recording/health | jq
A healthy install shows a non-null reconciliation_last_run (stamped once at boot) and a
retention_last_tick no older than disk.cleanup_interval_secs. A null value means that
supervisor has never completed a pass.
bin/dvr_doctor.sh --token "$ADMIN_TOKEN" wraps this up with the on-disk state — in particular a
stuck_deleting count, which should be 0 after a clean boot.
Then check the log for the two lines worth reacting to:
recording is enabled with no retention, no disk watermarks, and no quota— nothing bounds recording disk usage. Intentional on a dedicated filesystem; a mistake otherwise.enabled NewEpisode recording rule(s) cannot match— see the limitation below.
19.1 Known limitation: NewEpisode rules
NewEpisode rules do not currently match anything. The scheduler matches them by walking EPG
programmes, and no EPG horizon is supplied to it yet, so only WeeklyTimeslot rules materialize.
The condition is logged once per process rather than failing quietly.
Until this is wired, record a recurring programme with a WeeklyTimeslot rule, or record
individual programmes from the EPG view.