* fix(playback): route v3 direct play and remux through proxy nodes
Protocol v3 consulted the node planner only for the HLS deliveries, so
`original_http` and `server_remux_progressive` sessions returned an
API-local `/stream/{session_id}` URL and the API node served the bytes —
ServeDirectPlay for direct play, a locally spawned ffmpeg for the remux.
An operator running dedicated proxy nodes still saw all of that egress on
the API node.
The capability already existed: the proxy implements /stream/direct and
/stream/remux, and the Jellyfin-compat transport already plans a proxy for
exactly these two methods. Native v3 was the only surface skipping it, so
Jellyfin clients routed correctly on a deployment where Silo's own clients
did not. This wires the same shape into the v3 identity transport rather
than inventing a second selection path.
The proxy serves from the stream token alone, so the token now carries the
media path, the file's Dolby Vision profile (a P7 remux must strip the
dangling RPU) and the audio-only flag (which picks audio/mp4 over
video/mp4, the MIME the plan promised). RecipeCard models none of the
three; a missing claim would not fail loudly, it would serve a subtly
different stream than the plan promised.
Two related fixes:
- Proxy direct play served via http.ServeFile, which sets no strong ETag.
direct_stream_resume_v1 depends on the ETag ServeDirectPlay sets before
ServeContent, so routing direct play to a proxy without this would have
silently broken resumable direct streams: If-Range never validates and a
resumed range restarts at 200. The proxy now uses the same serve path.
- playback.local_transcode_fallback was only checked in the HLS branch, so
a progressive remux that converts audio still spawned ffmpeg locally on
an API-only node with the setting disabled. Identity deliveries now
honor the gate too — direct play still falls back locally, since moving
bytes is not transcode work and single-node deployments must keep
working.
Falling back to the API-local path when no proxy is eligible preserves
single-node behavior, and a planner reservation is released whenever the
session does not actually reach a proxy.
Closes #619
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(playback): validate proxy recipes and keep proxy sessions alive
Addresses three P1 findings on the proxy-transport change.
Proxies do run ffmpeg — /stream/remux converts audio and strips Dolby
Vision RPUs — but they exposed no capability endpoint, so unlike the HLS
offload path nothing checked that the selected proxy could execute the
transformations a plan froze. A pool whose proxies carry a different
ffmpeg build (rolling upgrade, custom image) would fail at stream time: a
missing aac encoder 500s, a missing dovi_rpu filter is refused outright by
the remux itself. Proxies now serve /hw-capabilities in the same shape and
at the same path as a transcode node, and identity planning validates the
frozen recipe against the selected proxy, falling back to a node that can
do the work. A proxy that does not answer is treated as incapable rather
than assumed good: an older proxy predating the endpoint is exactly the
mismatched build the check exists to catch. Direct play copies bytes and
needs no recipe, so it skips the probe entirely.
meteredResponseWriter implemented neither Unwrap nor SetWriteDeadline, so
RollingDeadlineWriter could not install its stall deadline on any proxy
stream. With the standalone proxy running WriteTimeout 0 there was no
server-level guard behind it, so a client that stopped reading without
closing its connection would block a write forever, holding the session,
the file, the goroutine and the connection.
A proxy-served session never produces a transport request on the API node,
so activeTransportCount — what protects a local stream from the idle
reaper — stays zero and a heartbeat gap longer than the active grace would
reap a healthy stream, after which progress, stop and replan all fail with
session-not-found while bytes still flow. Sessions are now marked as
remotely transported, which widens their idle windows rather than granting
immunity: this manager has no absolute session lifetime, so unconditional
immunity would leak a session forever when a client disappears without
stopping. The mark is always set on commit, so a re-plan that moves a
session back onto the API clears a stale one.
Also adopts the exported transformation constants in the tests and covers
the effective-recipe bitrate branch, per review.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(playback): pick capable sibling proxies and refresh transport locality
Narrow proxy selection by capability *before* selection rather than
rejecting a single round-robin pick afterwards. Abandoning the pool on one
mismatch meant a capable proxy with free capacity sat unused while
playback either ran ffmpeg on the API node or, with
playback.local_transcode_fallback disabled, was refused outright — the
exact api/proxy split this branch targets, during exactly the rolling
ffmpeg upgrade the capability check exists for. PlanSessionWith now
applies its eligibility predicate to the proxy on proxy-only plans (the
proxy is the executor there), mirroring how HLS filters transcode nodes,
and the planner grows ProxyNodeURLs to match TranscodeNodeURLs. Direct
play still skips the probe: it copies bytes and needs no recipe.
Every committed route now records transport locality, not just the
identity-proxy one. A session replanned from a proxy onto the integrated
transcoder previously kept a stale remote-transport mark, and the widened
idle grace it grants would hold that session's stream and transcode slots
for five minutes after the local stream disconnected without an explicit
stop. The remote HLS route sets it too — it also hands the client an
absolute proxy URL that never reaches this server.
The proxy's CORS config exposed no response headers, so cross-origin
JavaScript could send the If-Range/Range request headers it already allows
but never read the ETag, Accept-Ranges or Content-Range needed to build
them. direct_stream_resume_v1 silently degraded to a full restart whenever
the proxy was on a different origin than the web app, which is the normal
deployment.
Also regenerates internal/playback/testdata/protocol_v3 and the schema
fixtures, which were stale for output_change_v1 since #613/#617 and failed
CI on every branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(playback): restore the alternate-version fallback for burn-in refusals
#617 renamed the terminal a burn-in-forced adaptation reports: when the
subtitle burn requirement is the sole trigger, an HDR source that cannot
be re-encoded now returns subtitle_conversion_unsupported instead of
hdr_transcode_unsupported, so the refusal names the thing the viewer can
actually act on.
terminalAllowsAlternateFileV3 was not updated to match, and it gates the
alternate-version retry on the old reason strings. That silently retired
the fallback for exactly the case its own comment describes — a bitmap
subtitle needing burn-in that an HDR source cannot support while an SDR
alternate can. Playback was refused outright instead of switching to the
version that can serve it.
Adds the new reason to the gate and covers it directly, so a future
rename of a refusal reason fails on the gate rather than only on the
end-to-end replan test.
Also drops debug instrumentation that was committed by mistake in
TestHandleReplanPlaybackV3BitmapSubtitleFallsBackFromHDRToSDRVersion; the
assertion is back to its original form and now passes on the merits.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
524 lines
17 KiB
Go
524 lines
17 KiB
Go
package nodepool
|
|
|
|
import (
|
|
"context"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/google/uuid"
|
|
)
|
|
|
|
// Plan is the result of a node selection for one playback session.
|
|
// Either field may be nil when no suitable node exists.
|
|
type Plan struct {
|
|
TranscodeNode *Node
|
|
ProxyNode *Node
|
|
}
|
|
|
|
// SessionPlanner selects transcode and proxy nodes for playback sessions.
|
|
// Implemented by *Planner; defined as an interface so handlers can be tested
|
|
// without a real pool.
|
|
type SessionPlanner interface {
|
|
PlanSession(sessionID, currentTranscodeURL string, needsTranscode bool, estBitrateKbps int) Plan
|
|
}
|
|
|
|
// DownloadPlanner selects proxy nodes for unbounded file delivery. Downloads
|
|
// have no predictable bitrate, so implementations must not admit them onto a
|
|
// proxy with a configured bandwidth cap.
|
|
type DownloadPlanner interface {
|
|
PlanDownload(sessionID string, preferredGroup ...string) Plan
|
|
ReleaseSession(sessionID string)
|
|
}
|
|
|
|
// TranscodeWorkPlanner reserves capacity for non-streaming GPU work such as a
|
|
// prepared download. The returned release function must be called after the
|
|
// remote operation ends or falls back locally.
|
|
type TranscodeWorkPlanner interface {
|
|
ReserveTranscodeWork(workID string) (*Node, func())
|
|
TranscodeNode(nodeID int) (*Node, bool)
|
|
}
|
|
|
|
// reservation bridges the gap between assigning a session to a node and the
|
|
// node's health reports reflecting that session.
|
|
//
|
|
// The job count stops counting toward a node's effective load as soon as the
|
|
// node delivers a health report newer than the reservation (the node's own
|
|
// count then includes the session), or after maxReservationAge as a safety
|
|
// net. The bandwidth estimate instead counts for a fixed bandwidthBridgeAge
|
|
// regardless of health freshness: a proxy's measured egress is a rolling
|
|
// average that only converges on the new stream's rate gradually, so an
|
|
// early health report would otherwise drop the estimate before the meter
|
|
// reflects it.
|
|
type reservation struct {
|
|
transcodeURL string
|
|
proxyURL string
|
|
kbps int // estimated stream bitrate, counted against the proxy
|
|
createdAt time.Time
|
|
}
|
|
|
|
const (
|
|
maxReservationAge = 90 * time.Second
|
|
bandwidthBridgeAge = 60 * time.Second // matches the proxy egress meter window
|
|
)
|
|
|
|
// Planner makes group- and capacity-aware node selections on top of the
|
|
// existing pools.
|
|
//
|
|
// Grouping: nodes sharing a group label are co-located (same host/LAN). A
|
|
// group is eligible only while every enabled member is healthy. A transcode
|
|
// node from group G is always paired with a proxy from G so transcoded bytes
|
|
// never cross the LAN twice (round-robin when G has several proxies).
|
|
// Ungrouped nodes keep the historical behavior: least-connections transcode
|
|
// selection and global round-robin proxy selection.
|
|
//
|
|
// Capacity: a node with MaxJobs set is skipped once its effective load
|
|
// (health-reported active jobs plus unexpired reservations) reaches the cap.
|
|
// A proxy with MaxBandwidthKbps set is skipped once its measured egress plus
|
|
// the estimated bitrate of recently admitted streams would exceed the cap.
|
|
type Planner struct {
|
|
proxies *ProxyPool
|
|
transcodes *TranscodePool
|
|
|
|
mu sync.Mutex
|
|
rr map[string]int // per-group round-robin cursor; "" = global
|
|
reserved map[string]*reservation // keyed by playback session ID
|
|
now func() time.Time // overridable for tests
|
|
}
|
|
|
|
// NewPlanner creates a planner over the given pools.
|
|
func NewPlanner(proxies *ProxyPool, transcodes *TranscodePool) *Planner {
|
|
return &Planner{
|
|
proxies: proxies,
|
|
transcodes: transcodes,
|
|
rr: make(map[string]int),
|
|
reserved: make(map[string]*reservation),
|
|
now: time.Now,
|
|
}
|
|
}
|
|
|
|
// PlanSession picks the nodes serving one playback session.
|
|
//
|
|
// When needsTranscode is true it selects a transcode node (soft affinity to
|
|
// currentTranscodeURL, matching the historical quality-switch behavior) and a
|
|
// proxy from the same group. When false (direct play / proxy-side remux) it
|
|
// selects only a proxy. estBitrateKbps is the expected stream bitrate (target
|
|
// bitrate for transcodes, source bitrate otherwise; 0 = unknown), used for
|
|
// bandwidth-cap admission. Re-planning the same session replaces its previous
|
|
// reservation, so quality switches don't double-count.
|
|
func (p *Planner) PlanSession(sessionID, currentTranscodeURL string, needsTranscode bool, estBitrateKbps int) Plan {
|
|
return p.PlanSessionWith(sessionID, currentTranscodeURL, needsTranscode, estBitrateKbps, nil)
|
|
}
|
|
|
|
// TranscodeNode returns the current pool record for a persistent node id. It
|
|
// lets durable artifact locators follow an administrator-edited node URL/group.
|
|
func (p *Planner) TranscodeNode(nodeID int) (*Node, bool) {
|
|
if p == nil || p.transcodes == nil || nodeID == 0 {
|
|
return nil, false
|
|
}
|
|
for _, node := range p.transcodes.Nodes() {
|
|
if node != nil && node.ID == nodeID && node.Enabled {
|
|
return node, true
|
|
}
|
|
}
|
|
return nil, false
|
|
}
|
|
|
|
// PlanDownload picks a healthy proxy for an unbounded file transfer. A
|
|
// configured proxy bandwidth cap cannot be reserved accurately without a known
|
|
// transfer rate, so capped proxies are excluded instead of being oversubscribed
|
|
// during the egress meter's convergence window.
|
|
func (p *Planner) PlanDownload(sessionID string, preferredGroup ...string) Plan {
|
|
if p == nil || p.proxies == nil || sessionID == "" {
|
|
return Plan{}
|
|
}
|
|
p.mu.Lock()
|
|
defer p.mu.Unlock()
|
|
|
|
now := p.now()
|
|
p.pruneReservations(now)
|
|
delete(p.reserved, sessionID)
|
|
|
|
group := ""
|
|
if len(preferredGroup) > 0 {
|
|
group = preferredGroup[0]
|
|
}
|
|
var candidates, fallback []*Node
|
|
for _, node := range p.proxies.Nodes() {
|
|
if node == nil || !node.Enabled || !node.Healthy || !p.underCap(node, now) {
|
|
continue
|
|
}
|
|
if node.MaxBandwidthKbps != nil && *node.MaxBandwidthKbps > 0 {
|
|
continue
|
|
}
|
|
fallback = append(fallback, node)
|
|
if group != "" && node.Group != nil && *node.Group == group {
|
|
candidates = append(candidates, node)
|
|
}
|
|
}
|
|
if group == "" || len(candidates) == 0 {
|
|
candidates = fallback
|
|
}
|
|
if len(candidates) == 0 {
|
|
return Plan{}
|
|
}
|
|
rrKey := "download:" + group
|
|
proxy := candidates[p.rr[rrKey]%len(candidates)]
|
|
p.rr[rrKey]++
|
|
p.reserved[sessionID] = &reservation{proxyURL: proxy.URL, createdAt: now}
|
|
return Plan{ProxyNode: proxy}
|
|
}
|
|
|
|
// PlanSessionWith behaves like PlanSession but restricts transcode-node
|
|
// selection to nodes accepted by eligible (nil accepts every node). Capability
|
|
// -aware playback planning uses it so a recipe that only some pooled nodes
|
|
// can execute is never load-balanced onto a node that cannot. The predicate
|
|
// runs under the planner lock and must be cheap and non-blocking (a set
|
|
// lookup, never a network call). Group health is still computed over the
|
|
// full pool: eligibility narrows selection, not co-location semantics.
|
|
func (p *Planner) PlanSessionWith(sessionID, currentTranscodeURL string, needsTranscode bool, estBitrateKbps int, eligible func(*Node) bool) Plan {
|
|
if p == nil {
|
|
return Plan{}
|
|
}
|
|
p.mu.Lock()
|
|
defer p.mu.Unlock()
|
|
|
|
now := p.now()
|
|
p.pruneReservations(now)
|
|
// Drop this session's own reservation before computing loads so a
|
|
// re-plan doesn't count the session against its current node.
|
|
delete(p.reserved, sessionID)
|
|
|
|
if estBitrateKbps < 0 {
|
|
estBitrateKbps = 0
|
|
}
|
|
proxies := p.proxies.Nodes()
|
|
transcodes := p.transcodes.Nodes()
|
|
// Group health is computed over the full pool before any narrowing:
|
|
// eligibility restricts what may be selected, not co-location semantics.
|
|
groupHealthy := groupHealth(proxies, transcodes)
|
|
|
|
var plan Plan
|
|
if needsTranscode {
|
|
if eligible != nil {
|
|
transcodes = filterNodes(transcodes, eligible)
|
|
}
|
|
plan.TranscodeNode = p.pickTranscode(transcodes, proxies, groupHealthy, currentTranscodeURL, estBitrateKbps, now)
|
|
if plan.TranscodeNode != nil {
|
|
plan.ProxyNode = p.pickProxy(proxies, groupHealthy, plan.TranscodeNode.Group, estBitrateKbps, now)
|
|
}
|
|
} else {
|
|
// A proxy-only plan has no transcode node, so the predicate applies to
|
|
// the proxy: it is the node that will execute the recipe. Filtering
|
|
// before selection means a capability mismatch skips to a capable
|
|
// sibling instead of abandoning the pool after one round-robin pick.
|
|
if eligible != nil {
|
|
proxies = filterNodes(proxies, eligible)
|
|
}
|
|
plan.ProxyNode = p.pickProxy(proxies, groupHealthy, nil, estBitrateKbps, now)
|
|
}
|
|
|
|
if plan.TranscodeNode != nil || plan.ProxyNode != nil {
|
|
res := &reservation{createdAt: now}
|
|
if plan.TranscodeNode != nil {
|
|
res.transcodeURL = plan.TranscodeNode.URL
|
|
}
|
|
if plan.ProxyNode != nil {
|
|
res.proxyURL = plan.ProxyNode.URL
|
|
res.kbps = estBitrateKbps
|
|
}
|
|
p.reserved[sessionID] = res
|
|
}
|
|
return plan
|
|
}
|
|
|
|
// filterNodes returns the nodes accepted by keep, preserving pool order so
|
|
// round-robin cursors stay meaningful across selections.
|
|
func filterNodes(nodes []*Node, keep func(*Node) bool) []*Node {
|
|
filtered := make([]*Node, 0, len(nodes))
|
|
for _, node := range nodes {
|
|
if keep(node) {
|
|
filtered = append(filtered, node)
|
|
}
|
|
}
|
|
return filtered
|
|
}
|
|
|
|
// ProxyNodeURLs lists the URLs of every enabled pooled proxy node, healthy or
|
|
// not, mirroring TranscodeNodeURLs. Capability planning wants the deployment's
|
|
// toolchain; an unreachable node excludes itself when its capability fetch
|
|
// fails.
|
|
func (p *Planner) ProxyNodeURLs() []string {
|
|
if p == nil || p.proxies == nil {
|
|
return nil
|
|
}
|
|
nodes := p.proxies.Nodes()
|
|
urls := make([]string, 0, len(nodes))
|
|
for _, node := range nodes {
|
|
if node != nil && node.URL != "" {
|
|
urls = append(urls, node.URL)
|
|
}
|
|
}
|
|
return urls
|
|
}
|
|
|
|
// TranscodeNodeURLs lists the URLs of every enabled pooled transcode node,
|
|
// healthy or not: capability planning wants the deployment's toolchain, and
|
|
// an unreachable node excludes itself when its capability fetch fails. An
|
|
// empty slice means no nodes are pooled.
|
|
func (p *Planner) TranscodeNodeURLs() []string {
|
|
if p == nil || p.transcodes == nil {
|
|
return nil
|
|
}
|
|
nodes := p.transcodes.Nodes()
|
|
urls := make([]string, 0, len(nodes))
|
|
for _, node := range nodes {
|
|
if node != nil && node.URL != "" {
|
|
urls = append(urls, node.URL)
|
|
}
|
|
}
|
|
return urls
|
|
}
|
|
|
|
// ReleaseSession removes a provisional node reservation when playback setup
|
|
// fails or falls back locally before a node health report can account for it.
|
|
func (p *Planner) ReleaseSession(sessionID string) {
|
|
if p == nil {
|
|
return
|
|
}
|
|
p.mu.Lock()
|
|
delete(p.reserved, sessionID)
|
|
p.mu.Unlock()
|
|
}
|
|
|
|
// ReserveTranscodeWork selects the least-loaded healthy transcode node while
|
|
// sharing the same health-bridging reservation accounting as playback. Unlike
|
|
// a playback session it does not require a proxy partner: the completed file
|
|
// is written to the configured shared artifact store and served later.
|
|
func (p *Planner) ReserveTranscodeWork(workID string) (*Node, func()) {
|
|
if p == nil || p.transcodes == nil || workID == "" {
|
|
return nil, func() {}
|
|
}
|
|
p.mu.Lock()
|
|
now := p.now()
|
|
p.pruneReservations(now)
|
|
reservationID := workID + "-" + uuid.NewString()
|
|
|
|
var best *Node
|
|
for _, node := range p.transcodes.Nodes() {
|
|
if node == nil || !node.Enabled || !node.Healthy || !p.underCap(node, now) {
|
|
continue
|
|
}
|
|
if best == nil || p.effectiveJobs(node, now) < p.effectiveJobs(best, now) {
|
|
best = node
|
|
}
|
|
}
|
|
if best != nil {
|
|
p.reserved[reservationID] = &reservation{transcodeURL: best.URL, createdAt: now}
|
|
}
|
|
p.mu.Unlock()
|
|
if best == nil {
|
|
return nil, func() {}
|
|
}
|
|
|
|
var once sync.Once
|
|
return best, func() {
|
|
once.Do(func() { p.ReleaseSession(reservationID) })
|
|
}
|
|
}
|
|
|
|
// groupHealth reports, for every group label present in either pool, whether
|
|
// all of its enabled members are healthy. Pools only hold enabled nodes, so
|
|
// disabled nodes never count against a group.
|
|
func groupHealth(proxies, transcodes []*Node) map[string]bool {
|
|
health := make(map[string]bool)
|
|
for _, nodes := range [][]*Node{proxies, transcodes} {
|
|
for _, n := range nodes {
|
|
if n.Group == nil {
|
|
continue
|
|
}
|
|
healthy, seen := health[*n.Group]
|
|
if !seen {
|
|
healthy = true
|
|
}
|
|
health[*n.Group] = healthy && n.Healthy
|
|
}
|
|
}
|
|
return health
|
|
}
|
|
|
|
// pickTranscode returns the eligible transcode node with the fewest effective
|
|
// jobs, keeping the session on currentURL unless a candidate has at least two
|
|
// fewer jobs (the historical soft-affinity rule).
|
|
func (p *Planner) pickTranscode(transcodes, proxies []*Node, groupHealthy map[string]bool, currentURL string, estKbps int, now time.Time) *Node {
|
|
var best, current *Node
|
|
for _, n := range transcodes {
|
|
if !p.transcodeEligible(n, proxies, groupHealthy, estKbps, now) {
|
|
continue
|
|
}
|
|
if n.URL == currentURL {
|
|
current = n
|
|
}
|
|
if best == nil || p.effectiveJobs(n, now) < p.effectiveJobs(best, now) {
|
|
best = n
|
|
}
|
|
}
|
|
if current == nil || best == nil || current == best {
|
|
return best
|
|
}
|
|
if p.effectiveJobs(best, now)+2 <= p.effectiveJobs(current, now) {
|
|
return best
|
|
}
|
|
return current
|
|
}
|
|
|
|
// transcodeEligible reports whether a transcode node may take a new session:
|
|
// it must be healthy and under cap, and a grouped node additionally requires
|
|
// its whole group healthy and — when the group contains proxies — at least
|
|
// one of them with job and bandwidth headroom (a group's capacity is bounded
|
|
// by its proxies).
|
|
func (p *Planner) transcodeEligible(n *Node, proxies []*Node, groupHealthy map[string]bool, estKbps int, now time.Time) bool {
|
|
if !n.Healthy || !n.Enabled || !p.underCap(n, now) {
|
|
return false
|
|
}
|
|
if n.Group == nil {
|
|
return true
|
|
}
|
|
if !groupHealthy[*n.Group] {
|
|
return false
|
|
}
|
|
groupHasProxy := false
|
|
for _, proxy := range proxies {
|
|
if proxy.Group == nil || *proxy.Group != *n.Group {
|
|
continue
|
|
}
|
|
groupHasProxy = true
|
|
if proxy.Healthy && proxy.Enabled && p.underCap(proxy, now) && p.underBandwidthCap(proxy, estKbps, now) {
|
|
return true
|
|
}
|
|
}
|
|
// A group without proxies pins nothing; its transcode nodes fall back
|
|
// to global proxy selection.
|
|
return !groupHasProxy
|
|
}
|
|
|
|
// pickProxy selects a proxy round-robin. When group is set and contains
|
|
// proxies, only that group's proxies are considered (keeping transcoded
|
|
// traffic on the group's LAN); otherwise any healthy proxy qualifies.
|
|
func (p *Planner) pickProxy(proxies []*Node, groupHealthy map[string]bool, group *string, estKbps int, now time.Time) *Node {
|
|
var candidates []*Node
|
|
rrKey := ""
|
|
if group != nil {
|
|
for _, n := range proxies {
|
|
if n.Group != nil && *n.Group == *group && n.Healthy && n.Enabled &&
|
|
groupHealthy[*group] && p.underCap(n, now) && p.underBandwidthCap(n, estKbps, now) {
|
|
candidates = append(candidates, n)
|
|
}
|
|
}
|
|
rrKey = *group
|
|
}
|
|
if len(candidates) == 0 {
|
|
if group != nil {
|
|
groupHasProxy := false
|
|
for _, n := range proxies {
|
|
if n.Group != nil && *n.Group == *group {
|
|
groupHasProxy = true
|
|
break
|
|
}
|
|
}
|
|
// Strict pinning: a group that has proxies but none usable
|
|
// never spills onto other LANs. (Unreachable from PlanSession
|
|
// for transcode plans — transcodeEligible already requires a
|
|
// usable group proxy — but enforced here for safety.)
|
|
if groupHasProxy {
|
|
return nil
|
|
}
|
|
}
|
|
rrKey = ""
|
|
for _, n := range proxies {
|
|
if n.Healthy && n.Enabled && p.underCap(n, now) && p.underBandwidthCap(n, estKbps, now) {
|
|
candidates = append(candidates, n)
|
|
}
|
|
}
|
|
}
|
|
if len(candidates) == 0 {
|
|
return nil
|
|
}
|
|
idx := p.rr[rrKey] % len(candidates)
|
|
p.rr[rrKey]++
|
|
return candidates[idx]
|
|
}
|
|
|
|
// underCap reports whether a node can take one more job.
|
|
func (p *Planner) underCap(n *Node, now time.Time) bool {
|
|
return n.MaxJobs == nil || p.effectiveJobs(n, now) < *n.MaxJobs
|
|
}
|
|
|
|
// underBandwidthCap reports whether a proxy has bandwidth headroom for a
|
|
// stream of the given estimated bitrate. With an unknown bitrate (0) the
|
|
// node only needs to be below its cap.
|
|
func (p *Planner) underBandwidthCap(n *Node, estKbps int, now time.Time) bool {
|
|
if n.MaxBandwidthKbps == nil {
|
|
return true
|
|
}
|
|
egress := p.effectiveEgressKbps(n, now)
|
|
if estKbps <= 0 {
|
|
return egress < *n.MaxBandwidthKbps
|
|
}
|
|
return egress+estKbps <= *n.MaxBandwidthKbps
|
|
}
|
|
|
|
// effectiveEgressKbps is the node's health-reported egress plus the estimated
|
|
// bitrate of streams admitted within the bandwidth bridge window, which the
|
|
// rolling egress average doesn't fully reflect yet.
|
|
func (p *Planner) effectiveEgressKbps(n *Node, now time.Time) int {
|
|
egress := n.EgressKbps
|
|
for _, res := range p.reserved {
|
|
if res.proxyURL != n.URL || res.kbps <= 0 {
|
|
continue
|
|
}
|
|
if now.Sub(res.createdAt) >= bandwidthBridgeAge {
|
|
continue
|
|
}
|
|
egress += res.kbps
|
|
}
|
|
return egress
|
|
}
|
|
|
|
// effectiveJobs is the node's health-reported job count plus reservations the
|
|
// health checker hasn't had a chance to observe yet.
|
|
func (p *Planner) effectiveJobs(n *Node, now time.Time) int {
|
|
jobs := n.ActiveJobs
|
|
for _, res := range p.reserved {
|
|
if res.transcodeURL != n.URL && res.proxyURL != n.URL {
|
|
continue
|
|
}
|
|
if n.LastHealthCheck != nil && n.LastHealthCheck.After(res.createdAt) {
|
|
continue // a newer health report already reflects this session
|
|
}
|
|
jobs++
|
|
}
|
|
return jobs
|
|
}
|
|
|
|
func (p *Planner) pruneReservations(now time.Time) {
|
|
for id, res := range p.reserved {
|
|
if now.Sub(res.createdAt) > maxReservationAge {
|
|
delete(p.reserved, id)
|
|
}
|
|
}
|
|
}
|
|
|
|
// LocalTranscodeFallbackAllowed reports whether the API server may transcode
|
|
// locally when no eligible transcode node exists, based on the
|
|
// playback.local_transcode_fallback setting. Defaults to allowed so
|
|
// deployments without the setting keep the historical behavior.
|
|
func LocalTranscodeFallbackAllowed(ctx context.Context, settings interface {
|
|
Get(ctx context.Context, key string) (string, error)
|
|
}) bool {
|
|
if settings == nil {
|
|
return true
|
|
}
|
|
v, _ := settings.Get(ctx, "playback.local_transcode_fallback")
|
|
return v != "false"
|
|
}
|