* docs(playback): add v3 neutral-contract finalization plan Supersedes the wire-contract sections of the 2026-07-12 v3 plan: server-owned attempt keys, delivery-keyed negotiation without Media3 engine names, tiered capability evidence, neutral device/output context, track/quality replan operations, audio-only planning, and coordinated no-back-compat rollout across server, Android, Apple, and web. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(playback): make v3 attempt keys server-owned and replace engines with deliveries Contract core of the platform-neutral v3 finalization (plan sections 3.1 and 3.2), breaking on purpose — v3 is dark and all clients move together: - Every PlanV3 now carries plan_attempt_key, an opaque server-computed token clients store and echo in attempted_plan_keys; ReplanRequestV3 gains bounded local_mutations that the replan handler folds into the failed plan's key. Clients never hash anything. - KotlinName() is deleted from DeliveryV3, StreamProtocolV3 and SubtitleModeV3; the attempt-key canonical string now uses lowercase wire tokens, and PlanRecipeVersionV3 bumps to v3.3 so no key or plan ID computed under the old canonicalization can collide. - EngineV3 leaves the wire: ClientPlaybackContextV3.Engines (media3_*) becomes Deliveries keyed original_http|progressive|hls, with EngineCapabilityV3 renamed DeliveryCapabilityV3. PlanV3.Engine is removed; the planner, subtitle policy and quirk registry re-key on delivery class, and the media3_only feature token is deleted. - Validated-claim strings drop the prefix: media3_h264_decode -> h264_decode, media3_audio_decode -> audio_decode. - Golden fixtures in testdata/protocol_v3 are regenerated by Go and are now the cross-repo source of truth. Part of the playback protocol v3 neutral-contract train (steps 2-3 of docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(playback): add v3 evidence tiers and neutral device/output context Implement plan sections 3.3 and 3.4 of the v3 neutral-contract pass: - ClientCodecCapabilitiesV3 gains required video_evidence and audio_evidence closed enums (exact | platform_attested | declared). Planner strictness follows the tier: exact keeps the strict decode-entry validation, platform_attested validates codec/resolution/bit-depth/ frame-rate but skips profile/level matching, declared grants copy routes from the flat codec lists. Only exact audio evidence earns passthrough claims. The detailed_decode_capabilities feature token is deleted (subsumed by video_evidence=exact), and evidence-blocked direct routes carry the new evidence_insufficient_for_direct reason/warning. - DeviceContextV3 is now platform/os_version/manufacturer/model plus a bounded platform_details map (<=16 entries, <=128 chars); the Android Build dump fields are gone. Fire TV quirks keep matching on manufacturer/model (brand fallback removed with the field). - output_route_generation (int64, dual-location) becomes an optional opaque output_context_id string on the output context; the dual-location consistency validation is deleted. Attempt keys, plan invalidation, route events, and the planstore column follow (new Goose migration). - Feature advertisement collapses to the top-level client_features list only; ClientPlaybackContextV3.Features is deleted and ReplanRequestV3 gains an optional client_features refresh. - PlanRecipeVersionV3 bumped v3.3 -> v3.4; fixtures re-keyed. Part of the playback protocol v3 neutral-contract finalization plan (docs/superpowers/plans/2026-07-30-playback-protocol-v3-neutral-contract.md). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(playback): add v3 intent replans, quality menu, and audio-only routes Protocol v3 could only replan after a failure, so changing the audio track or the quality still required the legacy audio PATCH and the client-recipe transcode start — the two endpoints v3 is meant to replace. Clients also had to own a resolution ladder to render a quality menu, and a source with no video track was terminaled by the video/HDR gates, keeping audiobooks on the legacy path. Add track_change and quality_change replan operations. They carry no failure classification and route through the existing replan transaction, so they inherit its idempotency, capacity reservation, and staged-successor commit for free. Because nothing failed, the previous route stays eligible: neither the attempted-key history nor the failed-plan exclusion applies to them. Publish the server ladder on the plan as available_qualities so the quality menu is server-owned; the rungs come from the same resolutionLabelV3 and ladderBitrateKbpsV3 helpers the planner itself uses, not a parallel table. Plan audio-only sources through their own reduced route family: original_http when the client decodes the codec, otherwise a progressive AAC conversion. The plan advertises audio/mp4 for that remux and the transport now serves the same value, because a declared-tier client probes the advertised MIME with isTypeSupported before attaching a source buffer, and "video/mp4" on a stream with no video track is exactly the mismatch that makes the probe lie. Name the protocol's string vocabulary (dynamic ranges, transformations, executors, validated claims, terminal reasons) as constants while touching these lines, so the wire values have one definition. Part of #135 * docs(playback): publish the v3 protocol contract and fix subtitle ordinals Protocol v3 exists only as Go code today, so the Android and Apple ports have no authority to implement against other than reading this repository. Publish the contract as a normative document, machine-checkable schemas, and generated golden fixtures, and fix the one place where the server's own wire output disagreed with the ordinal space it publishes. - docs/architecture/playback-protocol-v3.md is self-contained enough for a third-party client: endpoints and status codes, evidence tiers and their bound-matching rules, delivery classes, the timeline model, replan semantics, registries, track identity, plan identity, quality, and transformations. - docs/design/schemas/playback-v3/ carries JSON Schemas for the five wire shapes plus valid and invalid fixtures, following the client-diagnostics layout. internal/playback/contract validates every fixture against its schema, so a schema that drifts from the Go types fails the Go suite. - cmd/playbackfixtures generates internal/playback/testdata/protocol_v3 from the production planner. `make playback-fixtures` writes them and `make verify-playback-fixtures` (wired into CI) fails when they are stale. These files are what the client ports consume, so drift would otherwise surface as a playback bug on three platforms at once. The subtitle fix: combined ordinals are one dense space over externals, then embedded tracks, then downloaded ones, but the legacy URL builder skipped burn-in-only tracks while assigning indices, so every track after a DVD/DVB track was numbered one too low and resolved to its neighbour. Ordinal assignment now lives in playback.BuildSubtitleInventoryV3 and both the plan inventory and the legacy `subtitle_urls` shape project from it; the legacy shape still filters burn-in-only entries but keeps each track's real index. Part of #135 * feat(web): migrate the players to the neutral playback v3 contract The web player was the last client still speaking the legacy start protocol: it picked its own file version from a codec probe, posted an ffmpeg recipe to start a transcode, PATCHed an endpoint to change audio tracks, and derived its own quality ladder. None of that survives a server-owned plan, and none of it produced telemetry the apps could be compared against. Video player: starts with a v3 request that advertises `declared` evidence from `isTypeSupported` probes and the three delivery classes, then consumes the returned plan for its URL, timeline, tracks and warnings. Quality and track changes become replans (`quality_change`, `track_change`), the quality menu renders `available_qualities` instead of computing rungs, and playback failures emit `route-events` so web failures land in the same diagnostics as Android and Apple. The duration comes from `source.duration_seconds` rather than the playback engine, and the "how was this delivered" overlay reads the plan's delivery and server transformations instead of comparing codec strings. Audiobook player: starts against the audio-only planner path with a single `original` rung, and takes its seek anchor from `timeline.player_start_seconds` so the progressive-remux route (which anchors the stream and restarts the player clock at zero) does not seek twice. Server side, `disable_progress_persistence` left the wire, so the rule it encoded is now derived. Resume state is keyed on the item, but every part of a multipart presentation shares that key while carrying its own file-local clock — persisting part 4's position would store "12 minutes in" as the book's resume point. `PresentationPartTotal > 1` expresses that directly and generalizes to multipart movies and split episodes, and a client can no longer forget to ask or lie about it. `useTranscodeQuality` and the legacy response types are deleted, and `WEBTEST_KNOWN_FAILURES` loses the audiobook entry along with its fix. Part of #135 * feat(playback)!: make v3 the only playback protocol Protocol v3 shipped behind a flag, alongside the legacy start path it was designed to replace. Running both meant every planner change had to be made twice, in two shapes that disagree about who decides the route: the legacy body carried a decision the client had already made, while v3 asks the server to make it. This deletes the legacy half. Removed: - `handleStartPlaybackLegacy` and its request/response bodies. The `POST /playback/start` route stays, but the protocol-version dispatch envelope is now a strict v3 decode — a body that does not declare `protocol_version: 3` gets `426 client_upgrade_required` so an outdated app can render a clear "update required" state instead of misreading a plan. Deliberately not a `400`: the request may be well-formed for the protocol it was written against. - `POST /playback/transcode/start`, superseded by the `quality_change` replan operation, and `PATCH /playback/{session_id}/audio`, superseded by `track_change`. Both mutated a session without re-planning. - The shadow planner and both rollout settings rows. With v3 the only protocol, `playback.protocol_v3_enabled` would mean "no playback at all"; `playback.protocol_v3_shadow_enabled` gated a comparison against a path that no longer exists. `409 protocol_disabled` on route-events goes with them, and capability `enabled` is now constant `true` (the field stays — clients feature-detect against it). - Version-selection helpers in `internal/playback/resolver.go` that only legacy start reached. `Resolve`/`ClientCapabilities`/`PlayDecision` stay: downloads consumes them. `internal/jellycompat` has its own resolution surface and is untouched. Behaviour the legacy handlers owned and v3 now owns explicitly: series version and audio-track preferences are persisted on start and on a `track_change` replan (not on failure recovery, whose forced route is not a user choice); an omitted `start_position` resolves to the profile's saved resume point; and an omitted audio track resolves through the series preference, the profile audio language, then the library override. Both are settled before planning, because the plan's timeline is cut at the start position. Spec §2.2 documents this as "omission is a request, not a default". The encode-target clamp that lived in the deleted transcode handler is already enforced in the planner, twice — `availableQualitiesV3` omits rungs at or above the source height, and the encode path clamps `targetHeight` to it. Unchanged: progress, stop, HLS manifest and segment delivery, the realtime control socket, stream tokens and restart reconstruction, watch together, downloads, jellycompat. Every removal is recorded in the pre-lock removals table in docs/architecture/v1-scope.md. Part of #135 * fix(scanner): stop recording embedded cover art as a video track ffprobe reports embedded cover art as a video stream carrying disposition.attached_pic. convertProbeData appended every "video" stream to VideoTracks without consulting isMainVideoStream, the predicate that already existed for duration decisions, so the picture was persisted as a playable track. That misreports the file twice: - An audio file with a cover picks up a video track, so it no longer satisfies MediaFile.IsAudioOnly and the v3 planner routes an audiobook through the video path instead of planAudioOnlyV3. - When the picture is ordered ahead of the real stream, the flat codec_video/resolution/hdr columns describe the poster: a 954x720 h264 episode was stored as mjpeg 480x480. Filter attached_pic streams out of the track loop. The guard is the disposition flag, not the codec name, so a genuine MJPEG video is still probed as video — the library has one. Already-probed rows self-heal on the next playback: NeedsCriticalProbeRepair already reprobes tracks missing color_range, which covers 21 of the 23 affected rows, and applyProbeData overwrites VideoTracks wholesale. The remaining two need a rescan; nothing persisted records attached_pic, and keying repair off still-image codec names would reprobe the genuine MJPEG file on every playback forever. Part of the playback v3 neutral-contract work: it is what lets Android drop AUDIOBOOK_COVER_ART_CODECS, which fabricated decode support the client cannot honestly claim under video_evidence: "exact". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(playback): publish subtitle URLs even when playback starts with subtitles off The v3 plan's subtitle inventory is the authoritative track list a client builds its subtitle menu from, but the handler only rewrote it with session-scoped URLs when a track was actually selected. A start or replan that resolved to `subtitle.mode: "off"` therefore returned the planner's URL-less inventory, so a client whose picker reads the inventory had a menu it could not fetch anything from. The Cast path hits this every time: it starts with subtitles off and needs the receiver's text tracks up front. attachSubtitleArtifactV3 now scopes and publishes the inventory unconditionally and gates only the artifact stamping on the selection. Spec §8 records that the `url` on a sidecar entry does not depend on the current selection. Part of the v3 neutral-contract finalization. * chore(playback): reconcile neutral v3 with main * fix(playback): preserve subtitle intent across replans * fix(playback): retain subtitle inventory on adapted routes * fix(playback): software-decode High10 AVC for QSV * fix(playback): scale High10 frames before QSV upload * fix(playback): preserve empty subtitle inventories * fix(playback): freeze terminal attempt contract * chore(playback): name fixture contract tokens * fix(playback): close v3 conformance review gaps * chore(playback): name conformance category * fix(playback): complete v3 conformance contract * fix(playback): keep schema fixtures generated * fix(playback): emit schema-valid conformance arrays * fix(playback): omit empty replan failures * fix(web): omit empty replan failures * fix(playback): close neutral v3 contract gaps * fix(playback): harden v3 replan, transcode, and quality-ladder edge cases Review remediation for the neutral v3 cutover, server side: - A failed replan no longer overwrites the durable StartResponse with a terminal or advances the replan request ID; an idempotent start replay of a still-healthy session returns the original plan. - SoftwareVideoDecode is now derived inside the transcode layer from source facts (codec/profile/bit depth) carried on TranscodeOpts, so jellycompat, downloads, recipe-card reconstruction, and transcode nodes get the High10 software-decode fix, not just the v3 handler. video_to_h264 recipe version bumps to 2 so mixed-version node pools that would silently drop the flag fail validation instead. - Local transport startup shares the 30s ManifestStartupTimeout; a timeout with the process still running stays retryable and is no longer persisted as a durable terminal against the attempt. - Sparse replan bodies (failure_recovery et al) no longer reset a user-selected quality preference to auto; the empty-value guard now covers every operation. - availableQualitiesV3 publishes no fixed rungs when the source height is unknown, keeping the no-upscaling ladder contract. - The proxy remux path serves audio-only fMP4 as audio/mp4 via a new additive AudioOnly token claim, matching the integrated path. - Plain text subtitle sidecars accept any requested extension again (served as VTT), restoring the permissive v1 behavior; ASS and bitmap handling is unchanged. - The 4K-disallowed terminal message discloses when a lower-resolution alternate exists but was pinned away by quality "original". Part of #135. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(web): keep playback alive through failed replans and honest audio claims Review remediation for the neutral v3 cutover, web player: - A failed or refused replan no longer unmounts the player: the fatal error screen is reserved for loads with no adopted plan, and replan failures surface through the existing non-fatal replanError path. - changeQuality rolls its optimistic preference back when the replan is refused or errors, so a failed switch is not silently applied by the next unrelated replan and the menu shows the real active rung. - The capability probe now tests mp3/vorbis codecs and mp3/flac/ogg containers (MediaSource with a canPlayType fallback), restoring direct play for mp3 audiobooks instead of per-part AAC re-encodes. - Reanchor seeks issued while a replan is in flight coalesce and run when it settles instead of being silently dropped with the scrubber pinned to a phantom position. - Subtitle refresh/translation replans use the resume anchor while the media element has no metadata, so a subtitle_ready broadcast during startup no longer restarts a resumed stream at 0:00. - An exhausted failure-recovery chain sets a visible error instead of returning silently. Part of #135. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(playback): accept video-only and VP9 probe metadata Treat audio and video probe completeness independently so legitimate video-only assets converge without repeated ffprobe repair. Allow unknown codec profile/level metadata to fall through to server adaptation while preserving exact direct-decode constraints. Fixes #574 * fix(playback): address protocol v3 review findings * fix(playback): harden lease and probe repair decisions * fix(playback): close remaining v3 review gaps * fix(playback): recover failed transcode starts * fix(playback): address remaining review-bot findings on v3 replan and audio planning Server: - The deferred replan lease release is bounded by a 3s timeout so a saturated pool or DB outage cannot wedge a handler goroutine that holds the per-session store lock on an uncancellable context. - planAudioOnlyV3 honors the request bandwidth cap: an over-cap source skips the original_http direct route and converts to AAC with the same bandwidth_cap_applied warning and decision reason the video ladder uses. Unknown source bitrate never triggers the cap. - A copy-audio progressive plan rejected only by a per-delivery audio_decode_codecs subset retries as an AAC conversion instead of returning adaptation_unavailable, and the AAC recipe respects the delivery's max_channels. Web: - failure_recovery replans issued while another replan is in flight queue (superseding a pending seek reanchor) instead of being silently dropped with the fatal overlay already suppressed. - A terminal response to a fresh non-preserving start clears the previous plan and stops its session, so episode navigation cannot keep rendering the prior item under the new title. - A refused recovery replan for a transport-dead plan surfaces the error and re-arms the plan failure key, so transient recovery failures no longer strand an endless spinner; the audiobook player gets the same guard reset. - A track-less subtitle_translation_completed hands off to the refreshed persisted track once the inventory settles, clearing the live overlay, instead of pinning the synthetic live track forever. Part of #135. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(playback): reuse HLS transport for sidecar replans * fix(playback): stabilize copy HLS remount timeline * fix(playback): address v3 review findings * fix(playback): satisfy player contract types --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
42 KiB
Client Diagnostics: crash reports and debug log upload
Status: Draft (spec + cross-repo plan)
Date: 2026-07-19
Repos affected: silo-server, silo-android, silo-apple
Summary
Add an opt-in, self-hosted diagnostics pipeline: Silo clients detect crashes and abnormal exits, collect structured debug logs (playback, TV focus, network, lifecycle) plus a hardware snapshot (device model, display/HDMI, audio path, codec capabilities), and — only with explicit user consent — upload a compressed diagnostics bundle to the user's own Silo server. Admins review bundles through new acting-admin API endpoints alongside the existing server logs tooling.
Everything ships disabled by default:
- Server: a
diagnostics.uploads_enabledserver setting gates the whole feature (default off), and it cannot be enabled unless the server's private object storage is configured and validated. - Client: verbose "debug logging" is a per-device setting (default off). Crash reporting defaults to Ask — after an abnormal exit the app shows an incident-specific prompt (e.g. "Silo closed unexpectedly during playback. Send a diagnostic report to your server?") with Send / Don't send / Always send. Nothing uploads without one of: an explicit prompt confirmation, a prior "Always send" election bound to this server and account, or the user explicitly sending a manual capture.
Motivation
Today we are blind to client-side failures:
- No crash reporting anywhere. Neither the Android repo nor the Apple repo installs an uncaught-exception handler, uses MetricKit, or bundles a crash SDK (verified: no handler/SDK references in either repo; the only Firebase dependency is messaging). A user saying "the app crashed" is a dead end.
- Diagnostics don't survive to the report. Android logs go to logcat (raw
android.util.Login ~40 files, no buffer or file sink); the Apple player's rich[CMP]/[CMP-DIAG]diagnostics go throughprint()(PlayerLog.swift/PlayerCore.swift) with no app-owned persistent sink. By the time a user reports a problem, the evidence is gone. - The hardest bugs are device-specific. Shield HLG quirks, Fire TV passthrough suppression, HDMI mode switches, TV focus traps — reproducing them requires knowing the exact device, panel, audio chain, and codec path. Users can't reliably report any of that.
- The server already has half the story.
operational_logsrecordsplayback_session_id(migration028_log_partitioning.sql), the admin logs API filters by it, and the web admin has a logs viewer. What's missing is the client half, and a way to join the two.
Goals
- Capture crashes and abnormal exits (JVM/native/ANR on Android; crash and hang diagnostics on iOS; abnormal-exit detection on tvOS) with enough context to act on: stack or exit trace where the platform provides one, recent structured logs, and a device snapshot.
- Capture deep opt-in debug logs for playback, TV focus behavior, networking, and navigation/lifecycle — everything needed to troubleshoot user issues.
- Always include a hardware snapshot: device (Shield / Fire TV / onn /
phone), OS, app build, display and HDMI mode with HDR capabilities, audio
output path and passthrough capabilities, decoder inventory — with honest
unknown/not_collectedstates where a platform can't provide a field. - One payload contract across Android and Apple; one server endpoint; one admin surface. The contract is a checked-in versioned JSON Schema with golden fixtures consumed by all three repos' CI.
- Explicit, server-bound consent for every byte uploaded; feature fully off by default on both server and client.
- Bounded cost: size caps, per-account quotas, concurrency limits, and server-side retention.
Non-goals (v1)
- Live/continuous log streaming from clients to the server.
- Screenshots or screen recordings in bundles.
- Native/tvOS crash stack capture via an in-process crash reporter (PLCrashReporter/KSCrash-style). tvOS v1 detects abnormal exits and attaches breadcrumbs, not stacks (see Apple design). Adopting an in-process crash reporter is a separate decision with its own signal-safety and review implications.
- Symbolication. Bundles store what the OS gives us; a server-side symbolication job can come later.
- Native-app admin viewer screens (per current product policy, admin surfaces beyond STATS on the clients are a separate product decision; v1 admin consumption is via API + web admin).
- macOS. The Apple repo has a macOS target, but v1 scopes to
android,android-tv,ios,tvos; the platform enum reservesmacoswithout accepting it. - Third-party telemetry, metrics aggregation, or fleet analytics. This is a support/debugging tool, not analytics.
Why this route
Upload to the user's own Silo server, not a third-party crash service. Crashlytics/Sentry SaaS would send user data to a third party, require vendor accounts/keys we'd have to ship, and put the reports in front of us rather than the person who can actually act — the self-hosting admin. Requiring every admin to also run self-hosted Sentry is far too heavy. The Silo server is the natural, privacy-aligned destination, and it already has auth, admin tooling, object storage, and retention patterns to build on.
Bundle-on-event, not streaming. A bundle captured at the moment of failure (crash, or a user-triggered capture) contains exactly the context needed: recent structured logs, the playback stats timeline, and the hardware snapshot. Continuous streaming would cost bandwidth and privacy while mostly capturing noise, and would still miss pre-crash buffers.
Hybrid storage: Postgres row + object-store blob. Report metadata (a
small manifest: device summary, truncated crash summary, type, timestamps)
goes in a queryable table; the payload itself (potentially megabytes of logs
and crash artifacts) goes to the S3-compatible private bucket as a single
archive. Ingesting client log lines into the partitioned operational_logs
table was considered and rejected: bundles are consumed as whole documents by
a human, not queried row-wise; client lines are untrusted input with
different retention needs; and multi-MB uploads would bloat a table tuned for
the server's own high-churn logs. The server does emit normalized
operational events about diagnostics activity (report accepted / rejected /
downloaded / deleted, with report id, account, sizes, result) so ingest and
admin actions are auditable and correlate with existing logs — raw client
lines never enter operational_logs.
A versioned archive, not a bare gzipped log stream. Crash evidence is heterogeneous: JSONL logs, ANR trace text, Android native tombstones (protobuf on API 31+), MetricKit diagnostic JSON. A tar.gz archive with named entries carries each artifact losslessly with a declared media type, keeps large traces out of Postgres, and can evolve by adding entries.
Build small in-house rather than adopt ACRA/other SDKs. The surface we need on Android (one exception handler writing a bounded marker, a ring buffer, one upload call) is a few hundred lines; ACRA would add a dependency and its own report format without covering Apple. On Apple, the OS provides MetricKit natively (iOS; not tvOS — see Apple design).
Architecture overview
Android / Apple client silo-server
┌────────────────────────────────┐ ┌──────────────────────────────────────┐
│ Safe-logging facade → ring │ │ GET /api/v1/diagnostics/status │
│ └ debug mode: rotating │ │ └ available|disabled| │
│ segment files │ │ storage_unavailable + limits │
│ Crash/abnormal-exit capture │ │ POST /api/v1/diagnostics/reports │
│ ├ UEH marker (Android) │──▶│ ├ gate + storage check │
│ ├ ApplicationExitInfo (30+) │ │ ├ streaming, bounded multipart │
│ ├ MetricKit (iOS) │ │ ├ account quotas + concurrency │
│ └ exit sentinel (tvOS) │ │ ├ manifest → client_diagnostic_ │
│ Pending reports (server-bound) │ │ │ reports (state machine) │
│ DeviceSnapshot (probes+new │ │ └ archive → S3 private bucket │
│ accessors) │ │ Admin (acting admin): list/detail/ │
│ Consent UI (ask/always/never, │ │ download/delete + audit events │
│ per server+account) │ │ Retention + orphan reconciler task │
└────────────────────────────────┘ └──────────────────────────────────────┘
Consent model
Consent is scoped, versioned, and destination-bound — not a global boolean. Both clients support multiple servers and accounts, so a device-wide "Always send" would let a bundle full of server A's titles and URLs upload to server B after a switch. Rules:
- Consent (including "Always send") is account-scoped: every consent
record and every pending report is bound to
(server_instance_id, account_user_id, consent_notice_version), whereserver_instance_idis a stable opaque ID returned by the status endpoint. The capturingprofile_idis recorded on reports for attribution, not as a consent scope — quotas and elections follow the account. - A pending report is only ever uploaded to the server+account it was captured under. Reports are never retargeted. Switching server or account neither uploads nor migrates pending reports; signing out of an account or removing a server deletes its pending reports.
- "Always send" is an election for one server+account. Bumping
consent_notice_version(when we materially change what bundles contain) invalidates prior elections and reverts to Ask. - Setting crash reports to Never deletes pending reports for that server+account and stops persistent capture (breadcrumbs/debug files); the in-memory ring keeps running (memory-only, never uploaded without a later explicit action).
- Child/managed profiles: the Diagnostics section is not shown to
profiles with
is_child, and every diagnostics action — changing settings, electing Always, the crash prompt, and manual capture/send — requires a non-child profile. A crash captured while a child profile was active is held pending and prompted for the next time a non-child profile is active. Reports are attributed to account + capturing profile. - The crash prompt and the manual flow both offer View report — a human-readable summary (incident, categories and line counts, device identity, destination server) before anything is sent.
- Uploads carry
consent: {mode, notice_version}in the manifest so the server can reject stale-consent uploads outright.
Bundle format (shared contract)
An upload is multipart/form-data with exactly two parts, in order:
manifest—application/json, ≤ 64 KiB. Queryable summary; stored in Postgres.bundle—application/gzip: a tar.gz archive,manifest.jsonfirst. Entry names come from a fixed allowlist, and each name implies its media type — there is no per-entry type metadata to drift:
manifest.json (manifest minus the `archive` object; see hashing note)
device.json (full device snapshot)
logs.jsonl (structured log lines, newest-last)
crash/summary.json (normalized crash info incl. provenance)
crash/stack.txt (JVM stack / ANR trace, when present)
crash/tombstone.pb (Android native tombstone, opaque bytes, API 31+)
crash/metrickit.json (raw MetricKit diagnostic JSON, iOS)
breadcrumbs.jsonl (persisted pre-crash context journal, when present)
Hashing note: archive.sha256 and the size fields in part 1 describe the
finalized bundle part bytes as transmitted. Since the archive is hashed
after it is built, the embedded manifest.json is the manifest without
the archive object; the server verifies part 1's archive fields against
the received blob. archive.bytes is the transmitted (gzip) size;
archive.uncompressed_bytes is the total decompressed tar stream size,
including tar headers, the end-of-archive marker, and record padding —
i.e. the byte count between the client's tar writer and gzip writer, and
what gzip -l reports. Standard end-of-archive zero padding (up to 64 KiB)
is accepted; any other trailing data is rejected.
The contract lives in this repo as a versioned JSON Schema
(docs/design/schemas/client-diagnostics/v1/) with canonical valid and
invalid fixtures; Go, Kotlin, and Swift CI all validate against the same
fixtures. Servers accept all schema versions they support and ignore unknown
additive fields.
Manifest (part 1)
{
"schema_version": 1,
"report": {
"type": "crash", // crash | anr | native_crash | hang | abnormal_exit | manual
"captured_at": "2026-07-19T18:22:31Z",
"capture_session_id": "run_c7d…", // UUID per app run, also stamped on log lines
"app_version": "1.4.2",
"app_build": "20841",
"platform": "android-tv", // android | android-tv | ios | tvos (macos reserved)
"os_version": "11 (API 30)",
"profile_id": "prof_…" // capturing profile, attribution only
},
"destination": { "server_instance_id": "srv_…" }, // consent binding; server rejects mismatches
"consent": { "mode": "prompt", "notice_version": 1 }, // prompt | always | manual
"crash": { // absent for manual
"summary": "NullPointerException in PlaybackSessionManager.start",
"stack_excerpt": "…truncated ≤ 8 KiB; full artifact in archive…",
"thread": "main",
"foreground": true,
"source": "ueh", // ueh | exit_info | metrickit | exit_sentinel
"provenance": "pre_failure", // pre_failure | post_restart | metric_reporting_period
"occurred_at": "2026-07-19T18:22:31Z" // best known; MetricKit reports carry period bounds
},
"device_summary": { "manufacturer":"NVIDIA", "model":"SHIELD Android TV", "os":"11", "form_factor":"tv" },
"playback_session_ids": ["ps_9f2…"],
"log_summary": { "lines": 3841, "bytes_gz": 412034, "dropped_lines": 12, "categories": ["playback","focus","network"], "debug_logging": true },
"archive": { "entries": ["manifest.json","device.json","logs.jsonl","crash/summary.json","crash/stack.txt"], "bytes": 498112, "uncompressed_bytes": 3110400, "sha256": "…" }
}
Log line schema (logs.jsonl)
{"ts":"2026-07-19T18:22:29.412Z","run":"run_c7d…","lvl":"W","cat":"playback","tag":"AudioCapabilityMgr","msg":"passthrough suppressed: HDMI sink lost TrueHD","attrs":{"sink":"HDMI","fmt":"truehd"}}
lvl:V|D|I|W|E.cat:playback,focus,network,lifecycle,browse,cast,download,crash,other.run(capture_session_id) distinguishes lines from prior app runs when breadcrumbs or debug files span sessions.- Playback additionally emits periodic (~5s) player-stats snapshot lines
(
"cat":"playback","tag":"StatsSnapshot") built from the existing stats models — a timeline of decoder, resolution, HDR mode, bitrate, dropped frames, and audio underruns around the failure.
Safe logging contract (redaction happens at collection time)
Best-effort scrubbing at package time is not enough: today's client logs
already emit full URLs with query strings, token suffixes, and raw exception
strings (e.g. the Apple HTTPClient response logging and player URL prints;
Android exception messages containing backend URLs). Therefore:
- The facade accepts typed attributes, not preformatted strings, for
anything risky: URL values are normalized (userinfo, query, and fragment
always stripped; path preserved), exceptions pass through a sanitizing
formatter, and free-text
msgis authored, not interpolated from network data. - Attribute keys are normatively registered: a per-category key registry (checked in next to the JSON Schema) declares each key and its value type (scalar types plus a few enumerated units). Unregistered keys fail debug builds/CI and are dropped in production builds.
- Existing call sites are migrated curated, not mechanically — each
migrated line is reviewed against the contract. Unmigrated
Log.*/printcall sites simply never enter bundles. - Defense in depth at bundle time: values matching the account's current access/profile tokens are replaced before packaging.
- Golden leak-fixture tests (JWTs, query tokens, signed URLs, cookies, token fragments) run in all three repos' CI against the facade and the bundle builder.
- The network category logs
method, templated path, status, duration; never headers, bodies, query strings, or full URLs.
Device snapshot (device.json)
Fields are tri-state where a platform can't provide them: a value,
"unknown" (probe exists, couldn't determine), or "not_collected"
(no probe on this platform). Every snapshot carries captured_at and
provenance (pre_failure for snapshots persisted with the incident,
post_restart for snapshots taken at bundle-build time).
{
"captured_at": "2026-07-19T18:22:31Z",
"provenance": "pre_failure",
"identity": { "manufacturer":"NVIDIA", "model":"SHIELD Android TV", "device":"mdarcy", "form_factor":"tv", "device_id":"…" },
"display": { "mode":"3840x2160@59.94", "modes_supported":[…], "hdr_types":["HDR10","HLG","DV"], "wide_gamut":true },
"audio": { "outputs":[{"type":"HDMI","encodings":["AC3","EAC3","TRUEHD"],"channels":8}], "passthrough":["AC3","EAC3","TRUEHD"], "suppressions":[…] },
"video_codecs": [ { "mime":"video/hevc", "hw":true, "profiles":["Main10","DV P8"], "max":"4096x2160@60" } ],
"network": { "transport":"ethernet" }
}
Android sources: AndroidDeviceMetadataProvider (id/model/platform/app
version), MediaCodecCapabilitiesProbe (codec + HDR decode caps),
AudioCapabilityManager (aggregate passthrough). New accessors required
(the probes exist but don't currently export everything): current + supported
Display.Mode list (DisplayHdrProbe exposes HDR support only), exact
identity fields (Build.DEVICE, form factor), AudioManager.getDevices()
output enumeration, and a snapshot/export API on
PassthroughSuppressionRegistry (its set is currently private).
Apple sources: AppleDeviceIdentity.current (persisted id, device name,
platform) and ApplePlaybackV3Capabilities.snapshot() (codec/HDR/device
context). Honesty constraints: the current outputSnapshot() helper is
private, reports route UIDs/types rather than full output capabilities, and
its passthrough payload is a placeholder — so Apple v1 reports audio as
route metadata with passthrough: "unknown" until a real probe exists, and
the private helper is refactored into an internal shared accessor so
diagnostics can reuse it without duplicating route logic. Audio route UIDs
are persistent peripheral identifiers and are one-way hashed before
inclusion.
Playback session correlation
Clients include recent server playback session IDs so admins can pivot into
the existing server-side logs (operational_logs.playback_session_id +
the admin logs filter). Both clients currently retain only the active
session ID (Android PlaybackSessionManager, Apple PlaybackSessionBridge),
so each client adds a small bounded, persisted recent-session tracker
(last ~10 session IDs with timestamps) that survives restarts — required for
next-launch crash reports.
Server design (silo-server)
Feature gate
Borrows the behavior of the requests gate (status endpoint + per-call
guard that rereads settings) but stores keys in the general server_settings
store — requests uses its own dedicated table, which diagnostics doesn't
need. Key naming follows existing namespaces (opslog.*,
audiobooks.enabled, playback.*):
diagnostics.uploads_enabled(bool, defaultfalse)diagnostics.max_bundle_bytes(default 10 MiB compressed)diagnostics.max_uncompressed_bytes(default 64 MiB)diagnostics.max_reports_per_user_per_day(default 20)diagnostics.retention_days(default 30)diagnostics.max_bytes_per_user(default 200 MiB)
Enabling diagnostics.uploads_enabled requires a storage validation
(private object store configured and writable); the admin settings API
refuses to enable it otherwise. S3Private is a nullable dependency in the
router and is only populated when configured, so the feature must degrade
predictably, not 500.
Status endpoint
GET /api/v1/diagnostics/status — authenticated (RequireAuth), no
profile required: crashes during login/profile selection must still be
reportable, so capability discovery and upload are account-scoped with
optional profile attribution.
{
"status": "available", // available | disabled | storage_unavailable
"server_instance_id": "srv_…", // stable; binds consent + pending reports
"accepted_schema_versions": [1],
"max_bundle_bytes": 10485760,
"max_manifest_bytes": 65536,
"retention_days": 30,
"consent_notice_version": 1
}
Clients cache this alongside existing server-driven config refresh.
Ingest endpoint
POST /api/v1/diagnostics/reports — RequireAuth; access-token
Authorization header only (no query credentials; sa_ API keys rejected
for diagnostics). Optional validated profile attribution. Hardened parsing —
the avatar handler is the precedent (multipart + PutObject), but its
buffer-everything approach (ParseMultipartForm + io.ReadAll) is not
reused:
- Guard:
diagnostics.uploads_enabled+ storage present, else stable error codes (disabled,storage_unavailable). http.MaxBytesReaderwithmax_bundle_bytes+ manifest overhead before any parsing; request deadline; per-account concurrency cap (1 in-flight) and a server-global concurrency ceiling, on top of the existing rate-limit middleware.- Streaming
multipart.Reader: exactly the two named parts in order; manifest capped at 64 KiB and schema-validated; unknown parts rejected. - Quota reservation: in one transaction holding a per-account advisory
lock (
pg_advisory_xact_lock(user_id)— serializing concurrent uploads from multiple devices on one account), check the per-day count and per-user byte quota and insert the report row in statereceiving(server-generated UUID +short_id). Failures returnquota_exceeded(429, withretry_after) ortoo_large(413) — machine-readable so clients back off and keep the report pending locally. - Stream the
bundlepart directly to its final key (diagnostics/{user_id}/{report_id}.tar.gz) while hashing and counting bytes; enforce the compressed cap during streaming. Validate the gzip/tar structure and per-entry limits (entry count, name allowlist, uncompressed total ≤max_uncompressed_bytes, compression ratio bound) in the same bounded streaming pass. Reject on violation (invalid_bundle,unsupported_schema). There is no staging/rename step — S3-style stores have no cheap rename, and the row'sstateis the source of truth: a blob is not a report until its row isready. - On success, update the row to
readywith sizes + sha256. On any failure: delete the object, mark the rowfailed(compensation); the reconciler (below) cleans up stragglers from process death mid-request. - Respond
201 {"report_id":"…","short_id":"SILO-…"}and emit a normalized operational event (accepted/rejected, account, sizes, result).
short_id is 12 Crockford-Base32 characters, server-generated with
collision retry, case-insensitively unique, shown to the user for support
conversations, filterable in the admin API — and explicitly not an
authorization secret.
New code: internal/diagnostics/ (service + store interface), handler in
internal/api/handlers/diagnostics.go, wired through Dependencies in
internal/api/router.go. The store interface is diagnostics-owned (the
avatar handler's is package-private and presign-only) and covers everything
the design uses: a streaming put (io.Reader — the existing s3client
PutObject takes []byte, so this is a small s3client addition),
GetObject (streamed admin download fallback), DeleteObject,
ListObjects(prefix) (reconciler), and PresignGetURL.
Storage
CREATE TABLE client_diagnostic_reports (
id UUID PRIMARY KEY,
short_id TEXT NOT NULL, -- unique, case-insensitive (expression index)
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
profile_id TEXT, -- capturing profile, nullable
state TEXT NOT NULL, -- receiving | ready | failed
captured_at TIMESTAMPTZ NOT NULL, -- client claim
received_at TIMESTAMPTZ NOT NULL DEFAULT now(), -- server truth; retention keys off this
report_type TEXT NOT NULL, -- crash|anr|native_crash|hang|abnormal_exit|manual
platform TEXT NOT NULL,
app_version TEXT NOT NULL,
crash_summary TEXT, -- truncated ≤ 8 KiB
manifest JSONB NOT NULL, -- part 1 verbatim (≤ 64 KiB by construction)
playback_session_ids TEXT[] NOT NULL DEFAULT '{}',
blob_bucket TEXT,
blob_key TEXT,
blob_bytes BIGINT,
uncompressed_bytes BIGINT,
blob_sha256 TEXT
);
CREATE UNIQUE INDEX ON client_diagnostic_reports (lower(short_id));
CREATE INDEX ON client_diagnostic_reports (user_id, received_at DESC);
CREATE INDEX ON client_diagnostic_reports (received_at);
Plain table, not partitioned — volume is orders of magnitude below
operational_logs; the byte-heavy payload lives in object storage. Full
stacks/traces live only in the archive; the row holds a truncated summary.
Admin API
Under the acting-admin subtree (same placement as the admin_logs.go
routes):
GET /api/v1/admin/diagnostics/reports— list; filters: user, platform, type, date range,short_id(exact, case-insensitive); paginated.GET /api/v1/admin/diagnostics/reports/{id}— row incl. manifest.GET /api/v1/admin/diagnostics/reports/{id}/download— presigned GET URL (primary, matching avatar behavior) with a streamedGetObjectproxy fallback for deployments where presigned URLs aren't reachable — the fallback is new behavior, not something the avatar path provides.DELETE /api/v1/admin/diagnostics/reports/{id}— row + blob.
Admin download and delete emit audit/operational events. Web admin gets a list/detail/download page next to the logs viewer, with a playback-session-id link into the existing filtered logs view. Native apps add no admin screens in v1 (product policy).
Retention + reconciliation
A taskmanager task internal/taskmanager/tasks/cleanup_client_diagnostics.go
mirroring cleanup_operational_log.go, registered in cmd/silo/main.go
with the other cleanup tasks:
- Delete reports older than
diagnostics.retention_days(byreceived_at), then enforcediagnostics.max_bytes_per_useroldest-first. Blob deleted first (tolerating already-missing objects), then the row; deletes are idempotent. - Reconcile: remove
receiving/failedrows older than a grace window, and (viaListObjects) anydiagnostics/*objects whose row is absent,failed, or stale-receiving— so partial failures and process deaths mid-request never leak storage. - Account deletion cascades rows (FK) ; the reconciler then removes orphaned blobs.
Android design (silo-android)
New package
android-shared/src/androidMain/kotlin/org/siloserver/silo/common/diagnostics/.
Logging
SiloLog.kt— safe-logging facade per the contract above:SiloLog.d(cat, tag, msg, attrs)(+i/w/e, sanitizedThrowable). Always forwards toandroid.util.Log(today's behavior unchanged) and offers the entry to the ring buffer.LogRing.kt— always-on in-memory ring: fixed-size array of pre-rendered compact entries, non-blocking writes (atomic index; a torn final entry during crash snapshot is acceptable), ~4000 entries / ~1.5 MB cap, lazy attribute formatting off the hot path, dropped-entry counter. Must add no measurable contention to playback; a playback benchmark gate accompanies the PR that wires player events in.DiagnosticsFileLogger.kt— only while debug logging is on: a single-writer coroutine draining a bounded channel (drop + count under backpressure) into append-only framed segment files (5 × 2 MB, app-privatefiles/diagnostics/logs/,noBackup), compressed only at bundle-build time — a process death never corrupts more than the tail of the active segment (unlike a single long-lived gzip stream).- Playback wiring: each
PlaybackAnalyticsListenerevent plus a ~5s periodicPlayerStatsSnapshotline while debug logging is on (reuses the existing reducer). TV focus logging hooks the TV app's focus observers undercat=focus. Network: a Ktor hook loggingmethod, templated path, status, duration_msonly.
Crash and abnormal-exit capture
CrashCapture.kt— installed as early as possible in bothApplicationclasses; idempotent; captures the previousThread.getDefaultUncaughtExceptionHandleronce. On crash it does the minimum on the dying thread: render the stack (bounded), copy what the non-blocking ring exposes within a hard time budget, write one bounded marker file (≤ ~512 KiB) via temp-write + atomic rename, and invoke the prior handler in afinally. No locks shared with the ring's write path, no coroutines, no allocation-heavy JSON building (preformatted line buffer). Bundle assembly happens on next launch, never in the handler. At most 3 pending reports are kept (oldest dropped) as a crash-loop guard.ExitInfoCollector.kt— API 30+ enhancement (minSdk is 24; many Fire TV devices ship below API 30, where the handler is the only source): readsActivityManager.getHistoricalProcessExitReasons()on launch.REASON_CRASH,REASON_CRASH_NATIVE,REASON_ANRbecome pending reports. Dedupe uses a persisted bounded fingerprint set (process name, pid, timestamp, reason, status, trace hash) — not a raw timestamp watermark, which loses same-stamp events and breaks on clock changes.getTraceInputStream()is treated as optional opaque bytes with a bounded read: ANR text →crash/stack.txt; API 31+ native tombstone protobuf →crash/tombstone.pb(never decoded into the manifest). UEH-captured events are correlated by fingerprint so the same death isn't reported twice.- Next foreground launch, per the consent model: Ask → aggregated, incident-specific prompt (one prompt covering N pending reports; default focus Don't send on TV; "Always send" requires a second confirmation; full-screen review flow on TV rather than a small dialog); Always → silent upload, at most one bundle per crash fingerprint per day; Never → purge. Per-fingerprint prompt suppression with backoff prevents crash-loop prompt fatigue. Pending reports expire after 7 days — visibly (the Diagnostics screen shows pending reports with expiry), not silently.
Device snapshot
DeviceSnapshotCollector.kt assembles device.json from the existing
probes plus the new accessors listed in the contract section (display mode
list, Build.DEVICE/form factor, AudioManager.getDevices() enumeration,
suppression-registry export). A compact snapshot is persisted with each
crash marker (provenance: pre_failure) since post-restart audio/display
state can differ from crash-time state.
Settings + consent UI
DiagnosticsSettingsStore.kt— DataStore. Deliberately not the per-profile pattern ofAndroidPlayerSettingsStore: consent and pending reports are keyed by server+account (per the consent model), and debug logging is per-device. Holds consent records, crash-report mode, debug-logging flag, fingerprint set, session tracker, and status cache.- Phone: "Diagnostics" section in
SettingsScreen.kt; TV: rail category inTvSettingsScreen.kt. For non-child profiles the section is always visible and state-aware rather than hidden when the server gate is off — it distinguishesavailable/disabled by server/storage unavailable/offline, and always shows local pending reports (count, size, expiry) with review/delete controls. Uploads are simply unavailable in non-available states; a report captured while the gate was off is not auto-uploaded later unless consent for that server+account exists. - Manual capture is the primary support flow and is stronger than a bare "send now": Start diagnostic capture (turns on verbose capture, shows an indicator) → user reproduces the issue → Stop & review → summary (categories, time range, sizes, destination) → send. A one-shot "Send diagnostics now" remains for quick cases but warns when only the in-memory ring is available ("this bundle contains only the last few minutes of basic logs").
- After upload, the
short_idis displayed and kept in a small sent history so a user can read it to their admin later.
Upload
shared/commonMain/…/network/api/DiagnosticsApi.kt—getStatus()anduploadReport(manifestJson, bundleBytes)via KtorMultiPartFormDataContent(first multipart use in the client — verified none exists — written as a small reusable helper), using the existingsafeApiCall/ApiResultpattern, wired through Koin like other*Apiclasses.DiagnosticsUploader.kt— builds the tar.gz (marker + ring/segment files overlapping the incident window + breadcrumbs + snapshot), streams it to the endpoint, deletes the pending report only on 201, recordsshort_id. Onquota_exceeded/too_large/disabled/storage_unavailableit keeps the pending report and backs off perretry_after; onunsupported_schemait surfaces "server needs an update" rather than silently retrying until expiry.
Apple design (silo-apple)
Mirrors Android with platform-native mechanisms, and with tvOS scoped honestly — MetricKit diagnostic payloads (crash/hang) are iOS-only; the SDK marks them unavailable on tvOS, and delivery is delayed (typically daily, while the app runs), not "next launch".
- Ring + facade:
DiagLogimplementing the same safe-logging contract. The player's existingcmpLoghelper (PlayerLog.swift) is extended into a tee — it keepsprint()ing (preservingdevicectl --consolecapture on tvOS) and also feeds the ring; remaining directprint()call sites inScreens/Player/migrate to it curated, per the redaction contract. OSLog-based subsystems are harvested at bundle-build time viaOSLogStore(scope: .currentProcessIdentifier)(deployment targets are iOS/tvOS 26, far above the OSLogStore 15+ floor), merged intologs.jsonlwith provenance. - Breadcrumb journal: because MetricKit reports arrive in a later
process, the in-memory ring is gone by delivery time. While crash
reporting is enabled, a small append-only breadcrumb journal (bounded,
rotating, crash-safe appends: lifecycle transitions, playback start/stop,
route changes, screen changes) persists across runs and becomes
breadcrumbs.jsonl, matched to a diagnostic bycapture_session_idand time window. Snapshots taken at bundle-build time are labeledprovenance: post_restart; the design never presents post-restart state as crash-time state. - iOS crash/hang detection: an
MXMetricManagersubscriber. Each receivedMXCrashDiagnostic/MXHangDiagnosticbecomes its own pending report: raw diagnostic JSON preserved ascrash/metrickit.json, deduped by a canonical hash of the diagnostic (not timestamps), carrying its own app version and reporting-period bounds. Prompt copy is type-specific ("crashed" vs "was not responding") and honest about timing ("yesterday during playback"). Delivery is verified on physical devices; parsing is tested via injected fixtures. - tvOS abnormal-exit detection: no crash stacks in v1. A dirty-exit
sentinel (marker written at launch, cleared on clean
background/termination) plus the breadcrumb journal yields
type: abnormal_exitreports — "Silo did not shut down cleanly last time" — with breadcrumbs and device snapshot but no stack. An in-process crash reporter for tvOS is explicitly deferred (see Non-goals). - Device snapshot:
AppleDeviceIdentity.current+ApplePlaybackV3Capabilities.snapshot()for identity/codec/HDR context; audio reported as route metadata with hashed route UIDs andpassthrough: "unknown"until a real output-capability probe exists (the current one is a private placeholder). - Settings/consent: Diagnostics section in
SettingsView(iOS) andTVSettingsView(tvOS) with the same state-aware presentation, manual capture flow, and consent rules. - Upload: a small multipart body writer added to
HTTPClient(verified: JSON-only today, no multipart/uploadTask), same status/error handling as Android. - Housekeeping note: product bundle IDs are already
org.siloserver.silo; only legacy keychain/logging identifiers (com.continuum.*) remain, so nothing here blocks on rebranding.
Privacy summary
- Consent boundaries: no network transmission of diagnostic data without prompt confirmation, a server+account-bound Always election, or an explicit manual send. The always-on ring is memory-only; persistent capture (segments, breadcrumbs, markers) is app-private, size-capped, backup-excluded, and purged by Never/sign-out/server removal.
- Never collected: credentials,
Authorization/profile-token values, request bodies, query strings, full URLs (normalized at collection), unsanitized exception dumps. Bearer-token matching at bundle time as defense in depth. Golden leak fixtures in CI. - Deliberately collected (disclosed in View report): device identity incl. device name, media titles appearing in playback logs, templated server paths. The recipient is the user's own server admin, who already has server-side access to this information; that is the trust model this feature is scoped to — and why third-party destinations are out.
- Persistent identifiers: the existing device IDs (already sent to the server today) plus hashed audio-route UIDs; each is enumerated in the schema docs with its purpose and retention.
- User control: state-aware Diagnostics screen with pending list, review, delete, sent history; admin-side deletion; account deletion cascades reports; admin access is acting-admin-only and audited.
- Bounded retention: defaults 30 days / 200 MiB per user, admin-tunable, enforced server-side.
Rollout plan (vertical slices, each independently shippable)
- Server foundation — settings + storage-validated gate, status endpoint, hardened ingest with state machine, schema + fixtures, admin API, retention/reconciler task, operational audit events. Feature stays off; nothing user-visible.
- Contract fixtures in clients — JSON Schema + golden fixtures wired into Android/Apple CI; safe-logging facades + rings land dark (no UI, no capture persistence, no prompts).
- Android end-to-end slice — consent model + Diagnostics settings UI + crash marker/ExitInfo capture + device snapshot + uploader + manual capture flow, with the initial curated logging migration (player tags + TV focus + network hook + stats snapshots). First release where a user can be asked for, and send, a useful report.
- Android coverage expansion — remaining curated call-site migration, breadcrumbs-on-Android if wanted, prompt-fatigue tuning from the field.
- Apple slice — cmpLog tee + OSLogStore harvest + breadcrumbs + MetricKit (iOS) + exit sentinel (tvOS) + settings/consent + multipart upload.
- Web admin viewer + correlation UX — list/detail/download page, playback-session pivot into the logs viewer.
- Later / separate decisions — symbolication job, tvOS in-process crash reporter, native-app admin viewer, macOS.
Slices 3 and 5 ship complete consent→capture→upload flows; intermediate releases never prompt for reports they can't send.
Open questions
- Should users be able to delete their uploaded reports from the client (self-service deletion), or is pending-report control + admin deletion enough for v1?
- tvOS: is abnormal-exit detection (no stacks) acceptable for v1, or does the Shield/Fire TV debugging need justify evaluating an in-process crash reporter (PLCrashReporter/KSCrash) with its signal-safety and review trade-offs?
- Do we want breadcrumbs on Android in v1 (cheap once the facade exists), or is the UEH marker + ExitInfo enough there?
- Localization of prompt/consent strings at launch vs English-first.
- Exact quota defaults (20/day, 200 MiB/user, 10 MiB/bundle) — tune before enabling anywhere real.
- Whether the status payload should also fold into an existing server-driven config refresh response later (clients poll both today); the dedicated endpoint ships first either way.
Review notes
This spec was adversarially reviewed before the draft PR: two independent review passes (a fact-check of every codebase claim against the three repos, and a design attack across consent, crash-capture correctness, ingest robustness, storage consistency, UX, and phasing). Material corrections adopted include: tvOS MetricKit unavailability and honest MetricKit timing; server+account-bound consent; storage-validated gating (private object store is optional in real deployments); streaming bounded ingest with a report state machine and orphan reconciliation; the archive container replacing a bare log gzip; collection-time typed redaction replacing scrub-at-package; fingerprint-based ExitInfo dedupe; non-blocking ring + bounded crash marker; recent-session trackers (only the active playback session ID exists today); and honest Apple snapshot capabilities (passthrough currently unknown).