`transfers.Registry` caps at 10,000 entries. Past that `Begin` returned
`ErrRegistryFull`, and every call site logged at Debug and served the file
anyway with its byte updates discarded. With no connection cap anywhere, one
actor could pin the registry and blind download-class monitoring for everyone —
which is the monitoring the abuse controls depend on.
Per decision A7, saturation now fails closed and is unreachable by a single
actor in the first place:
- A per-user concurrent-transfer cap (`playback.max_user_concurrent_transfers`,
default 24) checked before the global limit, so one actor exhausts its own
budget rather than the shared registry. 24 leaves headroom for
multi-connection downloaders, which open 4-8 sockets per file and take a
transfer id per concurrent Range request. `MaxPerUser` is a *pointer* option
because a plain int cannot distinguish "unset, use the default" from
"explicitly 0, unlimited".
- All five download-class call sites refuse rather than serve unmonitored:
429 `transfer_limit_exceeded` for the per-user cap, 503
`monitoring_unavailable` for a full registry, both with `Retry-After`.
The handoff listed four sites; `internal/api/handlers/ebook_reader.go` was
the missing fifth.
- The ABS file handler now admits the transfer *before* setting
Content-Disposition and the audio Content-Type, so a rejection is not
mislabeled as an audio attachment.
Two correctness fixes fall out of doing this properly:
- `End` now decrements the per-user count only when it actually removed an
entry, and drops the map entry at zero so neither counts nor warning
timestamps grow unbounded.
- The two download handlers registered `defer transfers.End(transfer.ID)`
*before* the service called `Begin`, so a duplicate id would have removed
another request's live record. `Begin` moves into the handler beside its
`End`, and the service keeps its late DownloadID/MediaFileID enrichment
through a new `Registry.Annotate`.
Saturation is logged at Warn inside the registry, rate-limited per user and
carrying the route and user id, rather than once per rejected request at each
call site. A7 asked for Warn at the call sites, but an actor parked at its cap
would then generate unbounded warning traffic — a log-amplification vector of
its own. No information is lost: route and user id are already on the Transfer.
429/503 are new failure statuses on existing endpoints, not repurposed ones.
silo-android and silo-apple need follow-up to retry with backoff on
`Retry-After`; the cap and the fail-closed behavior are advertised on
`GET /downloads/capability` so clients can feature-detect instead of probing.
The jellycompat call site is verified by reading only: that test package does
not compile on origin/main (pre-existing, unrelated).
Part 3 of 3 for the Batch 4 liveness/replica work.
Part of #305