* feat(observability): OpenTelemetry logs+traces with secret redaction Part of #265. Adds opt-in OpenTelemetry (logs + traces) alongside the existing stderr + opslog pipeline, plus secret redaction on all sinks. Default-off: with no OTEL_* / SILO_OTEL_ENABLED config, behavior is unchanged. Bootstrap (internal/telemetry): - Setup() builds one shared resource, a TracerProvider (parent-based trace-id ratio sampler), a LoggerProvider, and the W3C TraceContext+Baggage propagator from env. It installs NO MeterProvider — metrics stay on Prometheus, and the built-in no-op global MeterProvider keeps the trace instrumentation libs from double-emitting. Shutdown is deferred with a flush timeout. - Logs are bridged via otelslog fan-out (slog.MultiHandler), level-gated by the shared LevelVar and best-effort so a failing collector can't break the console or DB branches. stderr + opslog stay untouched. Secret redaction (internal/logredact): - A slog.Handler masks secret-keyed attributes (password, token, api_key, authorization, cookie, ...) — including .With-bound attrs, nested groups, secret-keyed group subtrees, and values behind a LogValuer — on the console and OTLP sinks, with a no-op fast path when a record has no secret keys. opslog.shouldRedact delegates to logredact.SecretKey so all sinks share one marker list. Rotation is infra-managed (no custom file sink): container runtime for stderr, collector/backend for OTLP, opslog partition-pruning for the DB. Documented in docs/architecture/observability.md. Verification: go build ./..., go vet, gofmt -l — clean; go test ./internal/telemetry/ ./internal/logredact/ -race pass. AI-use disclosure: implemented with AI assistance (Claude Code), including adversarial reviews that hardened the bootstrap and fixed two redaction leak paths; reviewed by the author. * refactor(observability): slog context+component sweep, sloglint gate (phase 3) Part of #265. Builds on the OTel bootstrap + redaction commit. Standardizes every log call site onto the context-carrying slog variants so records correlate with the active OpenTelemetry trace, and locks the standard in with a machine gate so future code (human- or AI-authored) can't drift back. - Call-site sweep: converted the remaining slog.<Level>(...) calls to the slog.<Level>Context(ctx, ...) form wherever a context.Context is in scope (background/init calls with no ctx are left as-is), across 183 files. Applied via a type-aware AST codemod. Log levels and message strings are preserved verbatim; a component attr (canonical per-package name) is added to direct package-level slog calls. Bound-logger calls keep their existing .With bindings. The main.go and telemetry package conversions rode with their file in the previous commit to keep each file within a single commit. - Enforcement (.golangci.yml): enable sloglint with context=scope, static-msg, key-naming-case=snake, no-mixed-args. After the sweep all four report zero violations repo-wide (tests included), so make lint / CI now blocks any regression to the non-context form. The gate ships with the sweep because it cannot be green until the legacy sites are converted. Metrics remain on Prometheus; no behavior change to /metrics or Grafana. Verification: go build ./..., go vet ./..., gofmt -l — clean; sloglint (all 4 rules) 0 violations repo-wide; log levels verified unchanged. AI-use disclosure: implemented with AI assistance (Claude Code), including the codemod; reviewed by the author. * fix(observability): honor per-signal OTLP protocol and secret WithGroup names Two Codex review findings on PR #290: - telemetry: OTEL_EXPORTER_OTLP_{TRACES,LOGS}_PROTOCOL now override the generic OTEL_EXPORTER_OTLP_PROTOCOL per signal, so mixed collector setups (e.g. HTTP logs + gRPC traces) build the right exporter. - logredact: entering a group whose name is secret-bearing (e.g. WithGroup("authorization")) now masks every leaf in that subtree, matching how slog.Group("authorization", ...) is masked as a whole. * fix(observability): address review feedback on telemetry bootstrap - Telemetry setup failure no longer kills boot: Setup returns usable no-op providers alongside the error and main logs and continues with telemetry disabled, honoring the best-effort contract. - Honor OTEL_TRACES_SAMPLER (always_on/off, traceidratio, parentbased_* variants); unsupported values fall back to parentbased_traceidratio. - Attach node identity as semconv service.instance.id instead of the non-semconv node.name. - Rename opslog retention-scope log attrs to target_component/target_level so they no longer collide with the canonical component routing key, and tag those lines with component=opslog. - Fix stale levelGated comment casing; use WarnContext in the telemetry shutdown defer; document the LogValuer double-resolve on the redaction slow path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
9.5 KiB
Observability (OpenTelemetry logs + traces)
Silo emits structured logs and distributed traces via OpenTelemetry (OTLP),
in addition to the existing stderr and opslog database pipeline. The feature is
opt-in and default-off: with no OTEL_* / SILO_OTEL_ENABLED configuration the
server behaves exactly as before (stderr + opslog only).
Metrics are not part of OpenTelemetry here. They remain on Prometheus
(client_golang, the /metrics endpoint, and the domain metrics in
internal/api/middleware). See Metrics below.
Enabling it
Telemetry turns on when either SILO_OTEL_ENABLED is truthy or
OTEL_EXPORTER_OTLP_ENDPOINT is set.
| Variable | Purpose | Default |
|---|---|---|
SILO_OTEL_ENABLED |
Master gate (1/true/yes/on). |
off |
OTEL_EXPORTER_OTLP_ENDPOINT |
Collector endpoint; also implicitly enables telemetry. | — |
OTEL_EXPORTER_OTLP_PROTOCOL |
grpc (default) or http/protobuf. |
grpc |
OTEL_SERVICE_NAME |
service.name resource attribute. |
silo-server |
OTEL_SERVICE_VERSION |
service.version resource attribute. |
unset |
OTEL_TRACES_SAMPLER |
always_on, always_off, traceidratio, parentbased_always_on, parentbased_always_off, or parentbased_traceidratio. Unsupported values (e.g. jaeger_remote) fall back to the default. |
parentbased_traceidratio |
OTEL_TRACES_SAMPLER_ARG |
Trace-id ratio for the ratio-based samplers (0–1; clamps >1 to 1). | 1.0 |
The node identity is attached as the service.instance.id resource attribute, so
multiple Silo nodes sharing one service.name stay distinguishable in the backend.
All other OTEL_EXPORTER_OTLP_* knobs (headers, TLS, per-signal endpoints) are read
directly from the environment by the OTLP exporters — the environment is the single
source of truth for exporter wiring.
The endpoint/exporter connection is lazy and non-blocking: an unreachable collector
does not delay or crash startup. Setup failure is also non-fatal: if the OTEL_*
environment is malformed (e.g. an unparseable endpoint URL), the server logs an error
and keeps running with telemetry disabled rather than crash-looping.
Architecture
The bootstrap lives in internal/telemetry (Setup in telemetry.go). When enabled
it builds one shared resource.Resource, a TracerProvider (sampler per
OTEL_TRACES_SAMPLER, batched OTLP exporter), a LoggerProvider (batched OTLP exporter), and the W3C
TraceContext + Baggage propagator. Shutdown is deferred in cmd/silo/main.go with a
flush timeout so buffered spans/logs drain on exit.
Log handler chain
Logs are bridged, not rerouted. The OTel otelslog handler is added to the existing
slog handler chain via fan-out, so every record still reaches stderr and the
opslog DB/admin-stream pipeline unchanged:
slog.<Level>Context(ctx, …)
→ opslog.Handler (DB capture + admin stream, when level ≥ capture)
→ logfilter.Handler (drops quieted subsystem prefixes)
→ slog.MultiHandler
├── stderr (json/text)
└── otelslog → OTLP (level-gated + best-effort)
Three properties matter:
- Level-gated.
slog.MultiHandler.EnabledORs its children, so the OTel branch is wrapped in a level gate bound to the sharedLevelVar; console and OTLP share one verbosity knob. Without this, Debug records would be built and exported even atlog_level=info. - Best-effort. The OTel branch never propagates an export error, so a failing collector cannot break the console or DB branches.
- Redacted. The whole fan-out is wrapped in
internal/logredact, so console and OTLP emit secret-masked output. Redaction is key-based (password,secret,token,api_key,authorization,cookie, …); a matching attribute's value becomes[REDACTED], including attrs bound via.With(...)and nested groups. The marker list is shared with the opslog DB path (opslog.shouldRedact→logredact.SecretKey) so all sinks agree. Limitation: values are not scanned, so a secret embedded in a free-text message or under a non-secret key is not caught.
Logging conventions (enforced)
Use the context-carrying slog variants and tag the subsystem:
slog.InfoContext(ctx, "scanner: starting", "component", "scanner", "folder_id", id)
Rules, enforced by sloglint in make lint (.golangci.yml):
slog.<Level>Context(ctx, …)wherever acontext.Contextis in scope (so records carry the activetrace_id/span_id). The plainslog.Info(…)form is only allowed where noctxexists (early boot, top-level goroutines).- Static message — the message is a constant; move dynamic parts to attributes.
- snake_case attribute keys —
component,request_id,trace_id,user_id, … - No mixed args — don't combine key-value pairs and
slog.Attrin one call.
Component classification
Every direct slog.*Context call carries a component attribute. opslog classifies
records by that attribute first, falling back to opslog.InferComponent (the subsystem:
prefix in the message) when no attribute is present — bound loggers that already set
component via .With(...) are left as-is.
Limitation: sloglint enforces the shape (context variant, snake keys, static
message) but cannot enforce that a component attribute is present. That convention
rests on this doc, the canonical list below, and review.
Canonical component values (first path segment under internal/; app for cmd/silo):
activitylog, adminjob, ai, api, app, audiobooks, auth, autoscan,
catalog, chapterthumbs, downloads, ebooks, historyimport, jellycompat,
libraryingest, manga, metadata, nodeconfig, nodepool, noderecipe,
nodesessions, notifications, opslog, playback, plugins, policy, proxy, ratelimit,
recommendations, requests, scanner, scanqueue, sections, taskmanager,
telemetry, transcodenode, watchlist, watchsync, webhooksync, worker.
Metrics stay on Prometheus
OpenTelemetry here installs no MeterProvider. Metrics continue to flow through the
existing Prometheus client_golang instrumentation to /metrics; Grafana scraping is
unaffected.
This is deliberate and guarded: the trace-instrumentation libraries (otelhttp, etc.,
added in later phases) also emit metrics through the global MeterProvider. Because none
is set, that global stays the built-in no-op, so those metric calls are silently
discarded — no double-counting into Prometheus, no second /metrics source. A guard test
asserts otel.GetMeterProvider() remains the no-op after Setup. Migrating metrics to
the OTel metrics SDK is out of scope.
Local collector (example)
# docker-compose.yml (dev)
otel-collector:
image: otel/opentelemetry-collector:latest
command: ["--config=/etc/otelcol/config.yaml"]
volumes:
- ./otelcol.yaml:/etc/otelcol/config.yaml
ports:
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
# otelcol.yaml — minimal debug pipeline
receivers:
otlp:
protocols:
grpc: {}
http: {}
exporters:
debug:
verbosity: detailed
service:
pipelines:
traces: { receivers: [otlp], exporters: [debug] }
logs: { receivers: [otlp], exporters: [debug] }
Run Silo with SILO_OTEL_ENABLED=1 OTEL_EXPORTER_OTLP_ENDPOINT=localhost:4317 and watch
logs + traces arrive at the collector. Swap debug for Loki/Tempo (or a vendor OTLP
endpoint) for real storage; Grafana can then show traces alongside the unchanged
Prometheus metrics.
Log retention & rotation
Silo does not write or rotate its own log files. Rotation/retention is handled by the layer that owns each sink, which is the intended cloud-native split:
-
Console (stderr) is owned by the container runtime. Under Docker, configure the
json-filedriver to cap and rotate on-disk logs, and mount the log directory on a volume so history survives container recreation (the problem that motivated this work):# docker-compose.yml services: silo: logging: driver: json-file options: max-size: "50m" # rotate at 50 MB max-file: "5" # keep 5 rotated filesFor a non-Docker deployment, run under a supervisor/journald or pipe to
logrotate. -
OTLP is the durable, queryable path: the collector/backend (Loki, a vendor, …) owns retention and is the recommended place to keep searchable history. Enable it (see above) and set retention on the backend.
-
opslog (Postgres) self-prunes: daily partitions with per-component retention and size caps (
internal/opslog/cleanup.go). No operator action needed.
An in-process rotating file sink (e.g. lumberjack) was deliberately not added — it would re-introduce a custom sink the OTLP + runtime split already covers.
Known limitations
log_quietalso suppresses OTel logs. The fan-out sits below the quiet filter, so quieting a subsystem empties its OTLP stream too (OTLP mirrors the console stream).- Redaction is key-based, not value-based. Secrets under a recognized key are masked on every sink, but a secret embedded in a message string or under an unrecognized key is not caught (see Redacted).
- Early-boot logs are not exported. Records emitted before the handler is installed
(DB connect, migrations, tuning) reach stderr only, matching existing
opslogbehavior. - Per-subsystem trace propagation into plugins is a follow-up owned by
silo-plugin-sdk; this repo instruments only the host side.