* feat(observability): OpenTelemetry logs+traces with secret redaction
Part of #265. Adds opt-in OpenTelemetry (logs + traces) alongside the existing
stderr + opslog pipeline, plus secret redaction on all sinks. Default-off: with
no OTEL_* / SILO_OTEL_ENABLED config, behavior is unchanged.
Bootstrap (internal/telemetry):
- Setup() builds one shared resource, a TracerProvider (parent-based trace-id
ratio sampler), a LoggerProvider, and the W3C TraceContext+Baggage propagator
from env. It installs NO MeterProvider — metrics stay on Prometheus, and the
built-in no-op global MeterProvider keeps the trace instrumentation libs from
double-emitting. Shutdown is deferred with a flush timeout.
- Logs are bridged via otelslog fan-out (slog.MultiHandler), level-gated by the
shared LevelVar and best-effort so a failing collector can't break the console
or DB branches. stderr + opslog stay untouched.
Secret redaction (internal/logredact):
- A slog.Handler masks secret-keyed attributes (password, token, api_key,
authorization, cookie, ...) — including .With-bound attrs, nested groups,
secret-keyed group subtrees, and values behind a LogValuer — on the console
and OTLP sinks, with a no-op fast path when a record has no secret keys.
opslog.shouldRedact delegates to logredact.SecretKey so all sinks share one
marker list.
Rotation is infra-managed (no custom file sink): container runtime for stderr,
collector/backend for OTLP, opslog partition-pruning for the DB. Documented in
docs/architecture/observability.md.
Verification: go build ./..., go vet, gofmt -l — clean; go test
./internal/telemetry/ ./internal/logredact/ -race pass.
AI-use disclosure: implemented with AI assistance (Claude Code), including
adversarial reviews that hardened the bootstrap and fixed two redaction leak
paths; reviewed by the author.
* refactor(observability): slog context+component sweep, sloglint gate (phase 3)
Part of #265. Builds on the OTel bootstrap + redaction commit.
Standardizes every log call site onto the context-carrying slog variants so
records correlate with the active OpenTelemetry trace, and locks the standard
in with a machine gate so future code (human- or AI-authored) can't drift back.
- Call-site sweep: converted the remaining slog.<Level>(...) calls to the
slog.<Level>Context(ctx, ...) form wherever a context.Context is in scope
(background/init calls with no ctx are left as-is), across 183 files. Applied
via a type-aware AST codemod. Log levels and message strings are preserved
verbatim; a component attr (canonical per-package name) is added to direct
package-level slog calls. Bound-logger calls keep their existing .With
bindings. The main.go and telemetry package conversions rode with their file
in the previous commit to keep each file within a single commit.
- Enforcement (.golangci.yml): enable sloglint with context=scope, static-msg,
key-naming-case=snake, no-mixed-args. After the sweep all four report zero
violations repo-wide (tests included), so make lint / CI now blocks any
regression to the non-context form. The gate ships with the sweep because it
cannot be green until the legacy sites are converted.
Metrics remain on Prometheus; no behavior change to /metrics or Grafana.
Verification: go build ./..., go vet ./..., gofmt -l — clean; sloglint (all 4
rules) 0 violations repo-wide; log levels verified unchanged.
AI-use disclosure: implemented with AI assistance (Claude Code), including the
codemod; reviewed by the author.
* fix(observability): honor per-signal OTLP protocol and secret WithGroup names
Two Codex review findings on PR #290:
- telemetry: OTEL_EXPORTER_OTLP_{TRACES,LOGS}_PROTOCOL now override the
generic OTEL_EXPORTER_OTLP_PROTOCOL per signal, so mixed collector
setups (e.g. HTTP logs + gRPC traces) build the right exporter.
- logredact: entering a group whose name is secret-bearing (e.g.
WithGroup("authorization")) now masks every leaf in that subtree,
matching how slog.Group("authorization", ...) is masked as a whole.
* fix(observability): address review feedback on telemetry bootstrap
- Telemetry setup failure no longer kills boot: Setup returns usable
no-op providers alongside the error and main logs and continues with
telemetry disabled, honoring the best-effort contract.
- Honor OTEL_TRACES_SAMPLER (always_on/off, traceidratio, parentbased_*
variants); unsupported values fall back to parentbased_traceidratio.
- Attach node identity as semconv service.instance.id instead of the
non-semconv node.name.
- Rename opslog retention-scope log attrs to target_component/target_level
so they no longer collide with the canonical component routing key, and
tag those lines with component=opslog.
- Fix stale levelGated comment casing; use WarnContext in the telemetry
shutdown defer; document the LogValuer double-resolve on the redaction
slow path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(metadata): stop manual rematch from resurrecting recorded stale IDs
The Apply Match flow (ModeIdentify) re-injected durable provider IDs into
the identify request without checking stale_media_ids, so a known-dead
tmdb ID rode along, 404ed again during the Phase-2 fetch, and was
re-recorded with a fresh last_seen_at — the item never left the Stale
External IDs list and jumped back to the top after every rematch.
Filter recorded-stale IDs out of the injected durable set in
prepareProcessRequest. Caller-supplied IDs are untouched, so an admin
deliberately re-selecting a previously-stale ID still retries it (which
is also why the ModeIdentify suppression guard in processInternal stays).
Fixes#268
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(metadata): normalize provider-id keys so stale-ID suppression can't be bypassed by casing
Review on PR #276 flagged that suppressRecordedStaleProviderIDs lowercases
and trims the stored stale row's provider before looking it up in the
incoming map, while the map keys are used verbatim, and that
HandleApplyItemMatch passes req.ProviderIDs from the JSON body straight
into metadata.Process without the normalization the search endpoint
applies. A caller-supplied key like "TMDB" or " tmdb " therefore defeated
the suppression. The same normalization gap was previously flagged on
PR #182.
Fix both layers:
- HandleApplyItemMatch now runs req.ProviderIDs through
normalizeMatchProviderIDs (same semantics as the search endpoint) and
returns 400 when no non-blank entries remain, mirroring the existing
empty-map rejection.
- suppressRecordedStaleProviderIDs now indexes the incoming map by
normalized key and deletes the matching original keys, so suppression
is robust regardless of caller casing or padding.
Adds regression tests at both layers: a metadata-level case where the
durable row arrives as "TMDB " while the stale row records "tmdb", and
handler-level cases asserting apply normalizes keys/values and rejects
all-blank provider-id maps.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com>