Commit Graph
9 Commits
Author SHA1 Message Date
ceb3393992 fix(historyimport): use standard Jellyfin authorization header (#572)
Co-authored-by: OpenAI Codex (GPT-5) <codex@openai.com>
2026-08-11 10:19:56 -04:00
255b1be89c fix(history-import): import Emby favorites (#378)
* fix(history-import): import Emby favorites

* fix(history-import): tolerate Emby favorite errors

* fix(history-import): count atomic favorite inserts

---------

Co-authored-by: Quick104 <31828688+Quick104@users.noreply.github.com>
2026-07-16 14:34:35 -04:00
203a18ae83 feat(observability): OpenTelemetry logs+traces with secret redaction and slog standardization (#290)
* feat(observability): OpenTelemetry logs+traces with secret redaction

Part of #265. Adds opt-in OpenTelemetry (logs + traces) alongside the existing
stderr + opslog pipeline, plus secret redaction on all sinks. Default-off: with
no OTEL_* / SILO_OTEL_ENABLED config, behavior is unchanged.

Bootstrap (internal/telemetry):
- Setup() builds one shared resource, a TracerProvider (parent-based trace-id
  ratio sampler), a LoggerProvider, and the W3C TraceContext+Baggage propagator
  from env. It installs NO MeterProvider — metrics stay on Prometheus, and the
  built-in no-op global MeterProvider keeps the trace instrumentation libs from
  double-emitting. Shutdown is deferred with a flush timeout.
- Logs are bridged via otelslog fan-out (slog.MultiHandler), level-gated by the
  shared LevelVar and best-effort so a failing collector can't break the console
  or DB branches. stderr + opslog stay untouched.

Secret redaction (internal/logredact):
- A slog.Handler masks secret-keyed attributes (password, token, api_key,
  authorization, cookie, ...) — including .With-bound attrs, nested groups,
  secret-keyed group subtrees, and values behind a LogValuer — on the console
  and OTLP sinks, with a no-op fast path when a record has no secret keys.
  opslog.shouldRedact delegates to logredact.SecretKey so all sinks share one
  marker list.

Rotation is infra-managed (no custom file sink): container runtime for stderr,
collector/backend for OTLP, opslog partition-pruning for the DB. Documented in
docs/architecture/observability.md.

Verification: go build ./..., go vet, gofmt -l — clean; go test
./internal/telemetry/ ./internal/logredact/ -race pass.

AI-use disclosure: implemented with AI assistance (Claude Code), including
adversarial reviews that hardened the bootstrap and fixed two redaction leak
paths; reviewed by the author.

* refactor(observability): slog context+component sweep, sloglint gate (phase 3)

Part of #265. Builds on the OTel bootstrap + redaction commit.

Standardizes every log call site onto the context-carrying slog variants so
records correlate with the active OpenTelemetry trace, and locks the standard
in with a machine gate so future code (human- or AI-authored) can't drift back.

- Call-site sweep: converted the remaining slog.<Level>(...) calls to the
  slog.<Level>Context(ctx, ...) form wherever a context.Context is in scope
  (background/init calls with no ctx are left as-is), across 183 files. Applied
  via a type-aware AST codemod. Log levels and message strings are preserved
  verbatim; a component attr (canonical per-package name) is added to direct
  package-level slog calls. Bound-logger calls keep their existing .With
  bindings. The main.go and telemetry package conversions rode with their file
  in the previous commit to keep each file within a single commit.
- Enforcement (.golangci.yml): enable sloglint with context=scope, static-msg,
  key-naming-case=snake, no-mixed-args. After the sweep all four report zero
  violations repo-wide (tests included), so make lint / CI now blocks any
  regression to the non-context form. The gate ships with the sweep because it
  cannot be green until the legacy sites are converted.

Metrics remain on Prometheus; no behavior change to /metrics or Grafana.

Verification: go build ./..., go vet ./..., gofmt -l — clean; sloglint (all 4
rules) 0 violations repo-wide; log levels verified unchanged.

AI-use disclosure: implemented with AI assistance (Claude Code), including the
codemod; reviewed by the author.

* fix(observability): honor per-signal OTLP protocol and secret WithGroup names

Two Codex review findings on PR #290:

- telemetry: OTEL_EXPORTER_OTLP_{TRACES,LOGS}_PROTOCOL now override the
  generic OTEL_EXPORTER_OTLP_PROTOCOL per signal, so mixed collector
  setups (e.g. HTTP logs + gRPC traces) build the right exporter.
- logredact: entering a group whose name is secret-bearing (e.g.
  WithGroup("authorization")) now masks every leaf in that subtree,
  matching how slog.Group("authorization", ...) is masked as a whole.

* fix(observability): address review feedback on telemetry bootstrap

- Telemetry setup failure no longer kills boot: Setup returns usable
  no-op providers alongside the error and main logs and continues with
  telemetry disabled, honoring the best-effort contract.
- Honor OTEL_TRACES_SAMPLER (always_on/off, traceidratio, parentbased_*
  variants); unsupported values fall back to parentbased_traceidratio.
- Attach node identity as semconv service.instance.id instead of the
  non-semconv node.name.
- Rename opslog retention-scope log attrs to target_component/target_level
  so they no longer collide with the canonical component routing key, and
  tag those lines with component=opslog.
- Fix stale levelGated comment casing; use WarnContext in the telemetry
  shutdown defer; document the LogValuer double-resolve on the redaction
  slow path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 08:53:52 -04:00
QuickandClaude Fable 5 8bb88658b1 fix(historyimport): use a discover-safe page size for the Plex watchlist
The discover API rejects the PMS page size (500) with 400 "Invalid value
provided for x-plex-container-size!"; page the watchlist at 100 instead.

Part of #245

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 13:21:43 -04:00
QuickandClaude Fable 5 9ead29bfcc fix(historyimport): make Plex watchlist import actually work
The browser Plex OAuth flow only sent the PMS server access token, which
the discover API rejects (401), so the watchlist step always failed with
a buried warning. The web client now forwards the plex.tv account token
via a new additive plex_account_token field.

The discover watchlist listing also ignores includeGuids, so items
arrived without external ids and could only exact-title/year match.
FetchWatchlist now resolves ids per item from the discover metadata
endpoint, decoding both Metadata- and Video-keyed containers, and
degrades to a title/year fallback warning instead of dropping items.

Part of #245

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 13:11:12 -04:00
a5fb16a5c5 feat(historyimport): import the Plex account watchlist alongside watch history (#285)
* feat(historyimport): import the Plex account watchlist alongside watch history

The Plex import migrated only watch history; the user's saved watchlist
had to be rebuilt by hand (#245).

- PlexClient gains FetchWatchlist: pages the account-level watchlist on
  the Plex discover API (discover.provider.plex.tv). It authenticates
  with the plex.tv ACCOUNT token — the PIN/OAuth session token, which
  resolvePlexAuth now threads through plexAuth.AccountToken (manual-token
  imports pass the user token, which doubles as the account token).
- Watchlist entries become import Records flagged Watchlisted, carrying
  identity only (movie/show → KindMovie/KindSeries, guids parsed) and no
  watch state. They ride the existing matcher (series matching already
  exists), and matched entries are added to the importing profile's
  watchlist via the idempotent AddToWatchlistAt — re-imports do not
  duplicate. A watchlist fetch failure downgrades to a run warning so the
  history import still completes.
- Run summaries gain a watchlist_added counter (new column + repo
  plumbing + client/admin UI cards).

Fixes #245

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(historyimport): count only newly inserted watchlist rows in WatchlistAdded

AddToWatchlistAt now reports whether a row was actually inserted (the
insert is ON CONFLICT DO NOTHING / INSERT OR IGNORE), and the import
summary increments WatchlistAdded only for genuine inserts, so
re-importing the same Plex account no longer inflates the count.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
2026-07-05 00:19:31 -04:00
46540dfec3 fix(progress): track resume points independently of watched state (#117)
* fix(progress): track resume points independently of watched state

Re-watching a finished item never re-entered Continue Watching: completion
latched completed = TRUE one-way, pinned position_seconds to the duration,
and the resume query filtered on completed = FALSE — so a rewatch heartbeat
could never surface the item again (and releasing the latch would have
erased the watched state clients display).

Adopt the Jellyfin invariant instead of guard heuristics:

- Completion resets position_seconds to 0 (UpdateProgress, SetProgress,
  SetProgressAt, SetProgressIfNewer, MarkWatched, MarkProgressBatch), so
  position_seconds > 0 now means "live resume point".
- completed stays a pure one-way watched latch; rewatch heartbeats re-enter
  Continue Watching through plain GREATEST/MAX while the watched flag and
  PlayCount survive (matching Plex and Jellyfin master).
- ListProgress("in_progress") keys on position_seconds > 0 in both stores;
  the SQLite store also gains the min-resume floor the Postgres store had.
- jellycompat reports Played=true with live PositionTicks during a rewatch
  (resumePositionTicks no longer zeroes played items) — the DTO shape real
  Jellyfin emits since jellyfin/jellyfin#15762.
- Web mirrors the latch (playbackProgressCache), resumes rewatches at their
  stored position, and shows progress bars on rewatched episodes.
- ABS audiobook surfaces keep today's behavior: finished books report 100%
  via the completed flag and Continue Listening still excludes them.
- Migrations reset legacy completed rows (position pinned to duration) to
  0: a Goose migration for Postgres and a user_version-gated one-time fix
  for the per-user SQLite DBs.

Replaces the guard-based approach of #109, whose restart detection
(50% fraction + 60s time gap) could never release the latch for immediate
rewatches (blocked heartbeats refreshed updated_at, re-arming the gap) and
un-watched items on position-0 heartbeats.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(progress): address review — migration gate, one-way latch, missed writers/readers

Review fixes for the position-based watch-progress model:

- The per-user SQLite data fix is now migrateToV11 in the existing
  versioned runMigrations chain (schemaVersion 11). The previous
  standalone PRAGMA gate compared against 1, but existing DBs already
  sit at user_version 10, so the reset never ran for them — and the
  gate would have rewound the version. Fresh DBs short-circuit to the
  current version as before.
- `completed` is now one-way across every playback/sync writer:
  SetProgress (the RecordPlaybackStop path — stopping a rewatch below
  the watched threshold no longer clears the watched state),
  SetProgressAt, SetProgressIfNewer (both stores), and the history
  import upsert, which also stops pinning completed imports to
  position = duration. Mark-unwatched still releases the latch via
  ClearProgress/ClearProgressBatch.
- MarkProgressBatch regains its freshness guard: a delayed batch mark
  carrying an old timestamp can no longer zero a newer rewatch resume
  point (the position-reset now rides the original updated_at check).
- Catalog read paths align with the new in-progress definition
  (position_seconds > 0, completed-agnostic): smart-collection
  in_progress filter, progress sort ratio, episode progress CTE, and
  both next-up predicates.
- jellycompat derives PlayedPercentage and PlaybackPositionTicks from
  the same clamped position; a played item at rest reports 100 (as the
  old model did) while a rewatch reports its live fraction.
- ABS audiobook UpsertProgress stores position 0 on finish so finished
  books can't surface as phantom resume entries; re-listens still move
  position forward from 0 with the latch intact.
- The web optimistic cache zeroes the resume point on completion,
  mirroring the server invariant until the refetch lands.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 19:44:40 -04:00
9e29e7b330 feat(security): encrypt server-owned credentials at rest (#45) (#95)
* feat(security): encrypt server-owned credentials at rest

Introduce AES-256-GCM at-rest encryption (HKDF-derived from a required
SECRET_KEY) for server-owned credentials, with row-bound AAD, a versioned
enc:v1: envelope, and an idempotent startup backfill.

- internal/secret: cipher + RowAAD/SettingsAAD + the startup backfill engine.
- SECRET_KEY required at bootstrap; cipher threaded as an explicit dependency.
- server_settings: EncryptedSettingsRepo decorator over the audited
  SensitiveSettingKeys (also drives admin redaction); the config watcher and
  watch-sync settings reads decrypt too.
- Arr keys inline-encrypted; the ambiguous SecretResolver indirection removed
  from requests/autoscan.
- Per-table columns encrypted: subtitles, watch-sync, webhook-sync (not
  webhook_secret), history-import, and the jellycompat session's bridged Silo
  access/refresh tokens.
- Startup backfill (resolve-then-encrypt for arr refs) is best-effort and
  primary-node gated.

Equality-looked-up secrets and plugin_runtime_configs.config_value are out of
scope (need hashing / cross-repo design) — see
docs/architecture/secret-encryption.md.

Refs #45

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(compose): require SECRET_KEY in docker-compose

The server now fatals without SECRET_KEY, so the integrated service (and the
commented distributed proxy/transcode examples) pass it through with a
fail-fast guard matching the existing MEDIA_ROOT pattern. Distributed worker
nodes must use the SAME key as the primary to decrypt shared data.
Generate with: openssl rand -base64 48.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(security): encrypt history import session credentials

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 15:25:48 -04:00
Silo Server Migration c085b12fd1 Initial Silo migration 2026-05-22 23:26:56 -04:00