Files
silo-server/internal/api/handlers/nodes.go
T
CoffeeKnyte ecb4555eec fix(playback): make edge tracking lifecycle-safe and session identity canonical
Five defects that all produced a WRONG over-cap count, which is why they land
before the revocation batch: decision A1 raises the over-cap revocation TTL from
5m to ~24h, removing the self-healing that currently limits the damage of a
miscount. A false positive after A1 blocks a legitimate stream for a day, so the
count has to be trustworthy first.

#1 -- overlapping edge requests deleted a live stream. Tracker.sessions was a
set and Remove tore down all state plus the Redis key, while both proxy pour
handlers deferred removal unconditionally. Two overlapping Range GETs on one
session id -- ordinary seek behaviour -- meant the first to finish deleted the
record while the second was still pouring, and later AddBytes calls were then
dropped because AddBytes ignores bytes for a session with no live record. The
stream went invisible to authoritative monitoring while still serving.

Track now returns a Lease that the request-scoped caller releases exactly once;
teardown happens when the last live lease is released. A plain refcount would
have been wrong: Track(A) -> Remove -> Track(B) -> Release(A) decrements B, and
clamping at zero does not help because the count legitimately belongs to B. That
is not hypothetical -- the transcode node deliberately replaces sessions under
the same id so a quality switch does not orphan ffmpeg, and it calls
unconditional Remove from its reaper and stop paths. So each generation carries
an epoch, Remove and Cleanup bump it, and a release from a superseded generation
is a logged no-op. Lease identity is a set rather than a counter, which makes a
duplicate release detectable instead of silently destructive.

The transcode node keeps using Remove: its Track calls are not request-scoped and
are correctly owned by session lifecycle. "Every Track needs a paired Release" is
true only of the request-scoped callers.

#8 -- async transcode tracking could leave a permanent ghost. The tracking write
ran as a bare goroutine with a WithoutCancel context, so if stop won the race the
delayed Track recreated the record after cleanup -- and because it landed in
sessions, Snapshot treated it as live until Remove and it NEVER idle-expired. A
permanent phantom inflating its owner's count, able to trigger false over-cap
kills of that user's real streams. The write now takes the per-session lifecycle
lock that stop and reap already hold, and re-checks session pointer identity
before writing, so a stopped or replaced generation cannot resurrect a record.
Pointer identity rather than id equality is what makes same-id replacement safe.
The write stays off the request path -- the API server and the playback client
are blocked on the 202.

#9 + M3 -- protocol-v3 counted one stream twice. The stream token carries a
transport id distinct from the logical session id, and the node tracked under the
transport id while the API/proxy record used the logical one, so mergeStreams saw
two streams. M3 was the reason this had not yet bitten: the v3 fresh-start caller
sent no owner attribution at all, so the transport record landed under user 0,
which the enforcer skips -- silently exempting the stream from the cap entirely.
Fresh v3 starts now carry the logical session id and full owner attribution
(both were already in scope at the call site), and merging is keyed on logical
identity where present via one shared helper used by both merge functions, which
had already drifted apart once.

The enforcer view resolves SessionID to the logical id so a kill targets the real
session rather than a replaceable transport generation. The raw admin view keeps
the transport id and exposes logical_session_id as an additive omitempty field,
advertised on the node-sessions capability endpoint, so the v1 response shape is
unchanged.

GAP-15 -- edge transcode liveness was request-observed. touchTranscodeSession
fired before proxying, so hammering dead segment URLs advanced LastServedAt with
zero bytes served. Visibility and liveness are now separate operations:
EnsureEphemeral makes a session visible without claiming bytes were served, and
served-byte liveness advances only from a 2xx/206 upstream response. Previously
the proxy metered every upstream body regardless of status, so a node 404's error
body counted as served bytes -- moving the touch later would not have fixed it.

S4 -- LiveLocalSessions moved from the HTTP handlers package to streammonitor,
which owns monitoring. A background enforcer importing api/handlers was
backwards. Pure move; its existing mapping assertions moved with it. The
LastActivityAt fallback inside it is left as-is -- decision A5 removes it in the
liveness batch.

Verified with go test -race across nodesessions, proxy and transcodenode; the
overlap regression test was confirmed to fail under the old unconditional
teardown.

Part of #305.
2026-07-30 09:09:35 +00:00

463 lines
16 KiB
Go

package handlers
import (
"context"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"errors"
"fmt"
"log/slog"
"net/http"
"strconv"
"sync"
"time"
"github.com/Silo-Server/silo-server/internal/cache"
"github.com/Silo-Server/silo-server/internal/nodepool"
"github.com/Silo-Server/silo-server/internal/nodesessions"
"github.com/Silo-Server/silo-server/internal/playback"
"github.com/Silo-Server/silo-server/internal/streammonitor"
"github.com/Silo-Server/silo-server/internal/transfers"
"github.com/go-chi/chi/v5"
"github.com/redis/go-redis/v9"
)
// NodeRepository defines the operations the NodeHandler needs on the node store.
type NodeRepository interface {
List(ctx context.Context) ([]*nodepool.Node, error)
GetByID(ctx context.Context, id int) (*nodepool.Node, error)
Create(ctx context.Context, input nodepool.CreateNodeInput) (*nodepool.Node, error)
Update(ctx context.Context, id int, input nodepool.UpdateNodeInput) (*nodepool.Node, error)
Delete(ctx context.Context, id int) error
UpdateHealth(ctx context.Context, id int, healthy bool, activeJobs, egressKbps int) error
}
// NodeListEnabled queries enabled nodes by type for pool reload.
type NodeListEnabled interface {
ListEnabled(ctx context.Context, nodeType string) ([]*nodepool.Node, error)
}
// TransferSnapshotSource is the narrow process-local monitoring view needed by
// the admin endpoint.
type TransferSnapshotSource interface {
Snapshot() []transfers.Transfer
}
// NodeHandler handles CRUD operations and health checks for stream nodes.
type NodeHandler struct {
repo NodeRepository
proxyPool *nodepool.ProxyPool
transcodePool *nodepool.TranscodePool
lister NodeListEnabled
eventBus cache.EventBus
redisClient *redis.Client // for reading session keys
jwtSecret string // for bearer auth when calling force-reload on nodes
// sessionMgr + localNodeName let the session list include integrated
// single-node streams (which never write Redis). Optional; nil in edge modes.
sessionMgr *playback.SessionManager
localNodeName string
transfers TransferSnapshotSource
}
// SetLocalSessionSource wires the in-process session manager so HandleListSessions
// can union integrated streams with the Redis-backed edge records.
func (h *NodeHandler) SetLocalSessionSource(sm *playback.SessionManager, nodeName string) {
h.sessionMgr = sm
h.localNodeName = nodeName
}
// SetTransferSource wires process-local download monitoring into the unfiltered
// admin response. Transfers do not belong to an edge node or live-stream source.
func (h *NodeHandler) SetTransferSource(source TransferSnapshotSource) {
h.transfers = source
}
// NewNodeHandler creates a new NodeHandler.
func NewNodeHandler(repo NodeRepository, proxyPool *nodepool.ProxyPool, transcodePool *nodepool.TranscodePool, lister NodeListEnabled, eventBus cache.EventBus, redisClient *redis.Client, jwtSecret string) *NodeHandler {
return &NodeHandler{
repo: repo,
proxyPool: proxyPool,
transcodePool: transcodePool,
lister: lister,
eventBus: eventBus,
redisClient: redisClient,
jwtSecret: jwtSecret,
}
}
// ForceReloadResult represents the result of a force-reload on a single node.
type ForceReloadResult struct {
NodeID int `json:"node_id"`
NodeName string `json:"node_name"`
Status string `json:"status"`
Error string `json:"error,omitempty"`
}
// checkNodeResult is the JSON response for a node health check.
type checkNodeResult struct {
Healthy bool `json:"healthy"`
ActiveJobs int `json:"active_jobs"`
EgressKbps int `json:"egress_kbps"`
}
// HandleListNodes handles GET /admin/nodes.
func (h *NodeHandler) HandleListNodes(w http.ResponseWriter, r *http.Request) {
nodes, err := h.repo.List(r.Context())
if err != nil {
slog.ErrorContext(r.Context(), "listing nodes", "component", "api", "error", err)
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to list nodes")
return
}
writeJSON(w, http.StatusOK, nodes)
}
// HandleCreateNode handles POST /admin/nodes.
func (h *NodeHandler) HandleCreateNode(w http.ResponseWriter, r *http.Request) {
var input nodepool.CreateNodeInput
if err := json.NewDecoder(r.Body).Decode(&input); err != nil {
writeError(w, http.StatusBadRequest, "bad_request", "Invalid request body")
return
}
node, err := h.repo.Create(r.Context(), input)
if err != nil {
// Validation errors from CreateNodeInput.Validate() are treated as 400.
if !errors.Is(err, nodepool.ErrNodeNotFound) {
// Check if it's a validation error (non-sentinel, non-wrapped).
// The repository calls input.Validate() which returns plain errors.
writeError(w, http.StatusBadRequest, "bad_request", err.Error())
return
}
slog.ErrorContext(r.Context(), "creating node", "component", "api", "error", err)
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to create node")
return
}
writeJSON(w, http.StatusCreated, node)
h.reloadPools(r.Context())
}
// HandleUpdateNode handles PUT /admin/nodes/{id}.
func (h *NodeHandler) HandleUpdateNode(w http.ResponseWriter, r *http.Request) {
id, err := parseIDParam(r)
if err != nil {
writeError(w, http.StatusBadRequest, "bad_request", "Invalid node ID")
return
}
var input nodepool.UpdateNodeInput
if err := json.NewDecoder(r.Body).Decode(&input); err != nil {
writeError(w, http.StatusBadRequest, "bad_request", "Invalid request body")
return
}
node, err := h.repo.Update(r.Context(), id, input)
if err != nil {
if errors.Is(err, nodepool.ErrNodeNotFound) {
writeError(w, http.StatusNotFound, "not_found", "Node not found")
return
}
slog.ErrorContext(r.Context(), "updating node", "component", "api", "error", err)
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to update node")
return
}
writeJSON(w, http.StatusOK, node)
h.reloadPools(r.Context())
}
// HandleDeleteNode handles DELETE /admin/nodes/{id}.
func (h *NodeHandler) HandleDeleteNode(w http.ResponseWriter, r *http.Request) {
id, err := parseIDParam(r)
if err != nil {
writeError(w, http.StatusBadRequest, "bad_request", "Invalid node ID")
return
}
err = h.repo.Delete(r.Context(), id)
if err != nil {
if errors.Is(err, nodepool.ErrNodeNotFound) {
writeError(w, http.StatusNotFound, "not_found", "Node not found")
return
}
slog.ErrorContext(r.Context(), "deleting node", "component", "api", "error", err)
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to delete node")
return
}
w.WriteHeader(http.StatusNoContent)
h.reloadPools(r.Context())
}
// HandleCheckNode handles POST /admin/nodes/{id}/check.
func (h *NodeHandler) HandleCheckNode(w http.ResponseWriter, r *http.Request) {
id, err := parseIDParam(r)
if err != nil {
writeError(w, http.StatusBadRequest, "bad_request", "Invalid node ID")
return
}
node, err := h.repo.GetByID(r.Context(), id)
if err != nil {
if errors.Is(err, nodepool.ErrNodeNotFound) {
writeError(w, http.StatusNotFound, "not_found", "Node not found")
return
}
slog.ErrorContext(r.Context(), "fetching node for check", "component", "api", "error", err)
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to fetch node")
return
}
healthy, activeJobs, egressKbps := nodepool.CheckNode(r.Context(), node)
if err := h.repo.UpdateHealth(r.Context(), id, healthy, activeJobs, egressKbps); err != nil {
slog.ErrorContext(r.Context(), "persisting health check result", "component", "api", "node_id", id, "error", err)
}
writeJSON(w, http.StatusOK, checkNodeResult{
Healthy: healthy,
ActiveJobs: activeJobs,
EgressKbps: egressKbps,
})
}
// HandleForceReloadNodes handles POST /admin/nodes/force-reload — sends a
// force-reload signal to every enabled node in parallel.
func (h *NodeHandler) HandleForceReloadNodes(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
allNodes, err := h.repo.List(ctx)
if err != nil {
writeError(w, http.StatusInternalServerError, "internal_error", err.Error())
return
}
var nodes []*nodepool.Node
for _, n := range allNodes {
if n.Enabled {
nodes = append(nodes, n)
}
}
results := make([]ForceReloadResult, len(nodes))
var wg sync.WaitGroup
for i, n := range nodes {
wg.Add(1)
go func(idx int, node *nodepool.Node) {
defer wg.Done()
result := ForceReloadResult{NodeID: node.ID, NodeName: node.Name}
client := &http.Client{Timeout: 10 * time.Second}
req, err := http.NewRequestWithContext(ctx, http.MethodPost, node.URL+"/admin/force-reload", nil)
if err != nil {
result.Status = "error"
result.Error = err.Error()
results[idx] = result
return
}
req.Header.Set("Authorization", "Bearer "+h.jwtSecret)
resp, err := client.Do(req)
if err != nil {
result.Status = "error"
result.Error = err.Error()
results[idx] = result
return
}
resp.Body.Close()
if resp.StatusCode == http.StatusNoContent || resp.StatusCode == http.StatusOK {
result.Status = "ok"
} else {
result.Status = "error"
result.Error = fmt.Sprintf("unexpected status %d", resp.StatusCode)
}
results[idx] = result
}(i, n)
}
wg.Wait()
type forceReloadResponse struct {
Results []ForceReloadResult `json:"results"`
}
writeJSON(w, http.StatusOK, forceReloadResponse{Results: results})
}
// HandleForceReloadNode handles POST /admin/nodes/{id}/force-reload — sends a
// force-reload signal to a single node identified by its ID.
func (h *NodeHandler) HandleForceReloadNode(w http.ResponseWriter, r *http.Request) {
id, err := strconv.Atoi(chi.URLParam(r, "id"))
if err != nil {
writeError(w, http.StatusBadRequest, "invalid_id", "node ID must be an integer")
return
}
node, err := h.repo.GetByID(r.Context(), id)
if err != nil {
writeError(w, http.StatusNotFound, "not_found", "node not found")
return
}
client := &http.Client{Timeout: 10 * time.Second}
req, err := http.NewRequestWithContext(r.Context(), http.MethodPost, node.URL+"/admin/force-reload", nil)
if err != nil {
writeError(w, http.StatusInternalServerError, "internal_error", err.Error())
return
}
req.Header.Set("Authorization", "Bearer "+h.jwtSecret)
resp, err := client.Do(req)
if err != nil {
type forceReloadResponse struct {
Results []ForceReloadResult `json:"results"`
}
writeJSON(w, http.StatusOK, forceReloadResponse{Results: []ForceReloadResult{{
NodeID: node.ID, NodeName: node.Name,
Status: "error", Error: err.Error(),
}}})
return
}
resp.Body.Close()
status := "ok"
if resp.StatusCode != http.StatusNoContent && resp.StatusCode != http.StatusOK {
status = "error"
}
type forceReloadResponse struct {
Results []ForceReloadResult `json:"results"`
}
writeJSON(w, http.StatusOK, forceReloadResponse{Results: []ForceReloadResult{{
NodeID: node.ID, NodeName: node.Name, Status: status,
}}})
}
// HandleListSessions handles GET /admin/nodes/sessions — lists active playback
// sessions, unioning the Redis-backed edge records with the in-process
// integrated sessions (which never write Redis) so a single-node deployment is
// not blind. The union is deduped by session id (the same stream is tracked by
// the central manager AND by the edge serving it — one row per stream, like the
// enforcer's monitoring picture, so an operator never sees a single stream
// counted twice against a cap). Optionally filtered by node_id: a filter
// targets a specific edge node, so integrated sessions are only included in the
// unfiltered listing.
func (h *NodeHandler) HandleListSessions(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
nodeFilter := r.URL.Query().Get("node_id")
var infos []nodesessions.SessionInfo
if h.redisClient != nil {
pattern := nodesessions.KeyPrefix + "*"
if nodeFilter != "" {
if nodeID, err := strconv.Atoi(nodeFilter); err == nil {
if node, err := h.repo.GetByID(ctx, nodeID); err == nil {
hashBytes := sha256.Sum256([]byte(node.URL))
nodeHash := hex.EncodeToString(hashBytes[:4])
pattern = nodesessions.KeyPrefix + nodeHash + ":*"
}
}
}
var cursor uint64
for {
keys, next, err := h.redisClient.Scan(ctx, cursor, pattern, 100).Result()
if err != nil {
writeError(w, http.StatusInternalServerError, "redis_error", err.Error())
return
}
for _, key := range keys {
val, err := h.redisClient.Get(ctx, key).Result()
if err != nil {
continue
}
var info nodesessions.SessionInfo
if err := json.Unmarshal([]byte(val), &info); err != nil {
slog.Debug("skip malformed session record", "key", key, "error", err)
continue
}
infos = append(infos, info)
}
cursor = next
if cursor == 0 {
break
}
}
}
// Integrated single-node streams live only in the in-process session manager.
// Include them in the unfiltered listing (a node_id filter targets an edge).
if h.sessionMgr != nil && nodeFilter == "" {
infos = append(infos, streammonitor.LiveLocalSessions(h.sessionMgr, h.localNodeName)...)
}
sessions := []json.RawMessage{}
for _, info := range streammonitor.DedupeSessionInfos(infos) {
if data, err := json.Marshal(info); err == nil {
sessions = append(sessions, json.RawMessage(data))
}
}
type sessionsResponse struct {
Sessions []json.RawMessage `json:"sessions"`
Transfers []transfers.Transfer `json:"transfers"`
}
activeTransfers := []transfers.Transfer{}
if nodeFilter == "" && h.transfers != nil {
activeTransfers = h.transfers.Snapshot()
}
writeJSON(w, http.StatusOK, sessionsResponse{Sessions: sessions, Transfers: activeTransfers})
}
// nodeSessionsCapabilitiesResponse advertises the additive surface of
// GET /admin/node-sessions so independently deployed clients (Android, Apple)
// can feature-detect it instead of sniffing the server version.
//
// Schema support and runtime availability are deliberately separate booleans.
// The transfers key is always present in the response shape once this endpoint
// exists, but the process-local registry behind it is optional wiring — an edge
// deployment can serve the field as an empty list forever. Collapsing the two
// would advertise download monitoring that is not actually running.
type nodeSessionsCapabilitiesResponse struct {
// LogicalSessionID reports that session records may carry the stable logical
// identity associated with a replaceable transport generation.
LogicalSessionID bool `json:"logical_session_id"`
// Transfers reports that the node-sessions payload carries a transfers array.
Transfers bool `json:"transfers"`
// TransfersActive reports that a transfer registry is wired on this server,
// so the array reflects real in-flight download-class pours.
TransfersActive bool `json:"transfers_active"`
}
// HandleGetNodeSessionsCapabilities exposes additive feature support for the
// live node-sessions payload (GET /admin/node-sessions/capabilities). It is
// mounted alongside the endpoint it describes, so its presence tracks that
// endpoint's availability rather than being advertised from an unrelated route.
func (h *NodeHandler) HandleGetNodeSessionsCapabilities(w http.ResponseWriter, _ *http.Request) {
writeJSON(w, http.StatusOK, nodeSessionsCapabilitiesResponse{
LogicalSessionID: true,
Transfers: true,
TransfersActive: h.transfers != nil,
})
}
// reloadPools refreshes the in-memory proxy and transcode pools from the database.
func (h *NodeHandler) reloadPools(ctx context.Context) {
if h.lister == nil {
return
}
proxyNodes, proxyErr := h.lister.ListEnabled(ctx, nodepool.NodeTypeProxy)
transcodeNodes, tcErr := h.lister.ListEnabled(ctx, nodepool.NodeTypeTranscode)
if proxyErr != nil || tcErr != nil {
slog.WarnContext(ctx, "node pool reload failed", "component", "api", "proxy_err", proxyErr, "transcode_err", tcErr)
return
}
if h.proxyPool != nil {
h.proxyPool.SetNodes(proxyNodes)
}
if h.transcodePool != nil {
h.transcodePool.SetNodes(transcodeNodes)
}
if h.eventBus != nil {
_ = h.eventBus.Publish(ctx, cache.ChannelAdmin, cache.Event{Type: cache.EventNodePoolChanged})
}
}