Files
silo-server/internal/imagecache/imagecache.go
T
eb6024573e feat(audiobooks): make audiobook libraries first-class catalog items (#73)
* docs(audiobooks): design spec for plugin absorption

Plan to absorb silo-plugin-audiobooks into silo-server as a first-party
feature. Audiobooks land in silo's existing SPA; ABS clients connect
directly. Hard constraints: reuse existing tables (media_items,
media_files, user_watch_progress, user_playback_sessions, people,
item_people, library_collections); only two new tables (abs_sessions,
podcast_feeds) and at most one column add (media_libraries.kind);
silo's main :8080 listener handles ABS Socket.io natively. Out of
scope: audiobook requests flow, smart collections, share links,
external recommender, custom metadata providers, separate audiobook
SPA.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): implementation plan sub-plan 1 (discovery + schema)

First of six sub-plans for the absorption. Six tasks: a discovery
audit that resolves the spec's Risk questions, four idempotent SQL
migrations (abs_sessions, podcast_feeds, media_libraries.kind,
audiobooks.enabled feature flag), and an empty-but-compiling
internal/audiobooks package scaffolded into cmd/silo. Lands as a
strict no-op for users (feature flag defaults to false).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): discovery findings for absorption sub-plan 1

Locks schema/code decisions for migrations 139-142 and downstream
sub-plans. Resolves open Risk questions from the absorption design spec.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): migration 139 add abs_sessions table

Parallel of jellycompat_sessions for Audiobookshelf-compatible clients.
Lets ABS mobile/desktop apps maintain a device-bound session that
silo's audiobooks/abs handlers will validate.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(audiobooks): match codebase conventions in migration 139

Lowercases type keywords in the abs_sessions CREATE TABLE body to
match neighboring migrations, fixes the client_version column
alignment, and replaces the misleading "parallel to
jellycompat_sessions" header comment with a more accurate
description of the table's role.

Cosmetic only — the running schema is unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): migration 140 add podcast_feeds table

Side table on media_items for RSS-subscribed podcasts. Holds feed URL,
ETag/Last-Modified for conditional fetches, last-refresh timestamp, and
the per-feed refresh interval consumed by the upcoming
podcastfeed.Refresher scheduled task.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(audiobooks): uppercase PRIMARY KEY in migration 140

Aligns with the codebase convention (type keywords lowercase,
constraint keywords uppercase) established in migration 139's
post-style-fix form. Cosmetic only — running schema is unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(audiobooks): migration 141 no-op for media_folders.type

Sub-plan 1 originally reserved migration 141 to add a 'kind' column to
media_libraries discriminating audiobook/podcast libraries. Discovery
audit (sub-plan 1 Task 1) found that the actual table is media_folders
and it already has a type text NOT NULL column with no CHECK constraint
or enum, so 'audiobooks' and 'podcasts' can be added as future values
without DDL.

Landing this migration as a documented no-op preserves the version
numbering audit trail and pins the decision in git history. The
matching down migration is also a no-op.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): migration 142 add audiobooks.enabled flag

Server-settings row that gates the absorbed audiobooks feature.
Defaults to 'false' so sub-plan 1 lands as a strict no-op; subsequent
sub-plans branch on this flag and operators flip it to 'true' at
cutover.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): scaffold internal/audiobooks package

Empty-but-compiling Service that reads the audiobooks.enabled feature
flag from server_settings. Wired into cmd/silo so the package is
referenced from the binary; no routes mounted, no scheduled tasks
registered, no DB writes. Subsequent sub-plans hang scanner branches,
ABS handlers, Socket.io, podcast refresher, and SPA pages off this
Service.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(audiobooks): cosmetic cleanups in scaffolded package

Two pre-emptive cleanups flagged by code review before sub-plan 2
copies the patterns:

  1. Sort the internal/audiobooks import after internal/adminjob in
     cmd/silo/main.go (alphabetical).
  2. Drop the redundant "audiobooks: " prefix from the Enabled() error
     wrap; matches how every other top-level service package
     (watchstate, scanqueue, metadata, etc.) formats errors.

No behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): implementation plan sub-plan 2 (scanner)

Second of six sub-plans. 10 tasks: PersonKind constants for Author and
Narrator, audio-extension recognizer, library-type helpers, a
walkLogicalTree refactor (movieLibrary bool -> typed walkMode), chapter
extraction via ffprobe, single-file and multi-file audiobook parsers,
scanner write path producing media_items.type='audiobook', and a
filesystem podcast parser (RSS deferred to sub-plan 5).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): add Author and Narrator PersonKind constants

Discovery audit confirmed item_people.kind is unconstrained smallint
with values 1-6 in use. Reserve 7 = Author, 8 = Narrator for audiobook
people-links written by the upcoming scanner branches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): add audio-extension recognizer for scanner

Mirrors the existing videoExtensions/SupportsVideoFile pair. Used by
upcoming audiobook and podcast scanner branches to filter directory
walks.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): library-type recognizers for scanner dispatch

isAudiobookLibraryType and isPodcastLibraryType match singular and
plural forms case-insensitively, mirroring isMovieLibraryType. Used by
upcoming scanner walk branches (Task 4) that filter audio files into
audiobook and podcast libraries.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(scanner): replace movieLibrary bool with typed walkMode

Lets walkLogicalTree dispatch on multiple library shapes (video, movie,
audiobook, podcast) without proliferating boolean flags. Behavior for
existing video and movie libraries is unchanged; audiobook and podcast
modes will be consumed by the upcoming audiobook.go and podcast.go
parsers in later tasks of this sub-plan.

walkModeFor() derives the mode from a media_folders.type string;
unknown types default to walkModeVideo to preserve prior behavior for
any caller still passing a raw type.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): expose ffprobe format tags on ProbeData

The audiobook scanner needs format-level tags (title, artist, album,
date) for media_items metadata; ffprobe already parses them in
ffprobeFormat.Tags but ProbeData previously discarded them. Add
FormatTags map[string]string to ProbeData, populate it in
convertProbeData via a new normalizeFormatTags helper that lowercases
keys and trims values.

Adds a fixture audiobook .m4b with embedded chapters (Intro/Outro) and
format tags, and a test that verifies ProbeFile() returns both
correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): parser for single-file audiobook folders

parseAudiobookFolder reads tags + chapters via the existing ProbeFile
(now that Task 5 exposes FormatTags on ProbeData) and produces a
parsedAudiobook struct. Title falls back from "title" tag to "album";
author from "artist" -> "album_artist" -> "composer"; series from
"album" -> "series" -> "mvnm" (Movement Name, used by some MP4 tools).
Year parsed from "date" or "year" tags, tolerating ISO dates and
parenthesized forms.

Single-file case only; multi-file folders (one audio file per chapter)
return a placeholder error and arrive in Task 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): multi-file audiobook folder support

Folders containing N audio files (one per chapter/part) get one
parsedAudiobookFile per file; each file's chapter list is synthesized
as a single chapter with title = filename stem. Title/author/series/
year come from the first file's tags.

Also drops the duplicate pickFirstNonEmpty helper added in Task 6 in
favor of the existing firstNonEmpty already in probe.go.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): scanner write path produces audiobook media_items

ScanAudiobookFolder walks an audiobooks-typed media folder and treats
each immediate subdirectory as one audiobook. For each parsed audiobook
it upserts:
  - one media_items row with type='audiobook'
  - one media_files row per audio file (with chapters JSONB)
  - author/narrator links in item_people (kind=7, kind=8)

Adds itemRepo and personRepo to the Scanner struct, wired from
fileRepo.Pool() in NewScanner — no constructor signature change needed.

ScanFolder dispatches to this path when folder.Type='audiobooks',
bypassing the per-file movie/TV pipeline because audiobooks are
folder-scoped entities.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): filesystem podcast scanner

ScanPodcastFolder walks a podcasts-typed media folder, treating each
subdirectory as a podcast show and each audio file inside as an
episode. Writes media_items.type='podcast' + episodes rows + media_files
rows. RSS-subscribed feeds (podcast_feeds table) arrive in sub-plan 5;
this task covers filesystem-only ingestion.

ScanFolder dispatches to this path when folder.Type='podcasts'.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): implementation plan sub-plan 5 (podcasts)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): expose audiobooks/podcasts library types in admin UI

Adds 'Audiobooks' and 'Podcasts' options to the library-type dropdown
in the admin libraries page so operators can flag a folder as an
audiobook or podcast library. Extends contentLevelsForType() so the
admin UI's downstream filtering treats those types correctly
(audiobook -> ['audiobook'], podcasts -> ['podcast',
'podcast_episode']).

Backend scanner branches for these types were already wired in
sub-plan 2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 139_abs_sessions to 147 for origin/main merge

origin/main adds 139_media_requests at the same number our local
audiobook branch had used for abs_sessions. Renumber ours to 147 to
free up 139 for the upstream migration. The schema_versions row is
updated in lockstep on the running database so the migrator sees the
abs_sessions migration as already applied at its new version.

Migrations 140-146 (podcast feeds, media_folders kind noop, audiobook
feature flag, abs playback sessions, podcast episode guid, audiobook
series, audiobook title cleanup) stay where they are — they don't
collide with anything on origin/main.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 140_podcast_feeds to 157 for origin/main merge

origin/main added 140_user_permissions at the same version this branch
had used for podcast_feeds. Renumber ours to 157 (next free above the
collections-unify migration at 156) so 140 is free for the upstream
migration. schema_versions on the running database is updated in lockstep
so the migrator sees podcast_feeds as already applied at its new version.

Same pattern as d59c1cb (renumber 139_abs_sessions to 147 for the prior
main merge). Pending migrations after this rename: 132 (downloaded
subtitles admin index, main), 140 (user_permissions, main), and 156
(unify_user_collections, this branch).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 141_media_folders_kind_noop to 159 for origin/main merge

Same shape as eb8f67d (the 140→157 renumber from the previous main
merge). origin/main added 141_episode_title_sort_index at the same
version this branch had used for media_folders_kind_noop. Renumber
ours to 159 (next free above the audiobook_series truncate at 158) so
141 is open for the upstream migration. schema_versions on the
running database is updated in lockstep so the migrator sees
media_folders_kind_noop as already applied at its new version.

Pending migrations on silo-prod after this rename: 141
(episode_title_sort_index, main) and any other newer ones from main
that the branch hasn't picked up yet.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 142_audiobooks_feature_flag to 160 for origin/main merge

Companion to 3c6f062's 141 renumber — origin/main also added
142_episode_catalog_entries (alongside 141_episode_title_sort_index)
at a version this branch had used for the audiobooks feature flag.
Renumber ours to 160 so 142 is open for the upstream migration;
schema_versions on silo-prod is updated in lockstep so the migrator
sees audiobooks_feature_flag as already applied at its new version.

This was the only remaining collision (verified by checking for
duplicate version prefixes across migrations/).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(audiobooks): address foundation review comments

* fix(audiobooks): tighten scanner identity handling

* fix(audiobooks): propagate scanner cancellation

* chore(audiobooks): adopt goose migration layout

* docs(audiobooks): implementation plan sub-plan 3 (API + frontend MVP)

Third of six sub-plans. 9 tasks: three REST endpoints (list/detail/
progress), TanStack Query hooks + types, three React pages
(Library/Detail/Player), and navigation integration. Scoped to MVP —
author/series indices, smart collections, share links, and other
nice-to-haves from the spec are deferred. Streaming reuses silo's
existing /api/v1/stream/{session_id}; no new transcode code.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): list endpoint at GET /api/v1/audiobooks

Paginated list of media_items with type='audiobook' scoped to the
caller's accessible libraries via the existing access filter.
Mirrors silo's existing list-style handlers for movies and series.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): detail endpoint at GET /api/v1/audiobooks/{id}

Returns the media_items row, its media_files (with chapters JSONB),
author/narrator extracted from item_people (kinds 7/8), and the
caller's per-profile listening progress from user_watch_progress.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): progress endpoint at POST /api/v1/audiobooks/{id}/progress

UPSERTs user_watch_progress for the caller's (user_id, profile_id,
content_id). Body carries position_seconds; clients are expected to
post every 5-10s during playback plus on pause/seek (matching silo's
existing video progress cadence).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): frontend types and TanStack Query hooks

TypeScript types match the JSON shapes from the new
/api/v1/audiobooks endpoints (list, detail, progress). Three hooks:
useAudiobookLibrary (list), useAudiobook (detail), and
useReportAudiobookProgress (mutation that invalidates the detail
query on success so progress updates reflect immediately).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): library grid page at /audiobooks

Renders a paginated grid of audiobook cards using the
useAudiobookLibrary hook. Each card links to /audiobooks/book/{id}.
Cards show poster, title, and year; falls back to a "No cover"
placeholder when the audiobook has no poster_url. Empty state hints
to operators that they need to set a library's type to 'audiobooks'.

Routes themselves are wired in Task 8 (navigation integration).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): detail page with chapter list

Renders cover, title, author, narrator, year, and overview alongside a
chapter list. Clicking a chapter opens an inline sticky
AudiobookPlayer at that chapter's start. A "Resume" button restarts
playback at the saved progress position if present.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): HTML5 audio player with chapter navigation

Single-file audiobook playback for MVP. Multi-file queuing arrives in
a follow-up. Streams via the existing /api/v1/direct-download GET
endpoint. Position is reported to /api/v1/audiobooks/{id}/progress
every 10s while playing plus on pause/seek/end. Skip-30s, playback
rate select, chapter list panel.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): wire navigation and routes

Adds an Audiobooks entry to the sidebar and registers the two new
routes (/audiobooks for the library grid, /audiobooks/book/:id for
detail). The player renders inline inside the detail page; no
dedicated player route is required for MVP.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(audiobooks): address native API review comments

* feat(audiobooks): add ABS compatibility and polish

* fix(audiobooks): stabilize ABS playback progress reporting

* fix(audiobooks): clean up ABS branch review fixes

* chore(audiobooks): adopt goose layout for ABS migrations

* fix(audiobooks): align player seek bar props

* feat(audiobooks): make libraries first-class catalog items

* feat(admin): add server restart endpoint

* fix(audiobooks): address review comment findings

---------

Co-authored-by: RXWatcher <14085001+RXWatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-07 15:57:05 -04:00

421 lines
13 KiB
Go

// Package imagecache downloads images from URLs, generates sized variants,
// computes thumbhashes, and uploads all variants to S3.
package imagecache
import (
"context"
"fmt"
"io"
"net"
"net/http"
"net/netip"
"net/url"
"strings"
"sync"
"time"
"github.com/Silo-Server/silo-server/internal/imageutil"
"github.com/Silo-Server/silo-server/internal/metadata"
)
const (
maxDownloadBytes = 10 * 1024 * 1024 // 10 MB
downloadTimeout = 30 * time.Second
)
// ObjectPutter is the S3 interface required by Cacher.
type ObjectPutter interface {
PutObject(ctx context.Context, bucket, key string, data []byte) error
Bucket() string
}
// ImageURLResolver resolves plugin:// paths to HTTP URLs.
type ImageURLResolver interface {
ResolveImageURL(ctx context.Context, path string, variant string) string
}
// CacheRequest describes a single image to cache. For season posters and
// episode stills, ContentID is the parent series's provider ID and the
// SeasonNumber / EpisodeNumber fields scope the S3 key so siblings do not
// collide. Both pointers are nil for item-level images.
type CacheRequest struct {
SourceURL string
ProviderID string
ContentType string // "movies" or "series"
ContentID string
ImageType metadata.ImageType
SeasonNumber *int
EpisodeNumber *int
ImageResolver ImageURLResolver // optional; used when SourceURL is a plugin:// path
}
// CacheResult is returned by Cache on success.
type CacheResult struct {
BasePath string // S3 key prefix, e.g. "tmdb/movies/550/poster"
Thumbhash string // base64-encoded
Ext string // file extension including dot (e.g. ".jpg", ".png")
}
// Cacher downloads and stores image variants to S3.
type Cacher struct {
s3 ObjectPutter
httpClient *http.Client
enforcePublicURLs bool
}
// New creates a new Cacher backed by the given ObjectPutter.
func New(s3 ObjectPutter) *Cacher {
return &Cacher{s3: s3, httpClient: newSecureHTTPClient(), enforcePublicURLs: true}
}
func newWithHTTPClient(s3 ObjectPutter, client *http.Client) *Cacher {
if client == nil {
client = http.DefaultClient
}
return &Cacher{s3: s3, httpClient: client}
}
// CacheImage implements metadata.ImageCacher using the internal Cache method.
func (c *Cacher) CacheImage(ctx context.Context, req metadata.CacheImageRequest) (*metadata.CacheImageResult, error) {
result, err := c.Cache(ctx, CacheRequest{
SourceURL: req.SourceURL,
ProviderID: req.ProviderID,
ContentType: req.ContentType,
ContentID: req.ContentID,
ImageType: req.ImageType,
SeasonNumber: req.SeasonNumber,
EpisodeNumber: req.EpisodeNumber,
})
if err != nil {
return nil, err
}
return &metadata.CacheImageResult{
BasePath: result.BasePath,
Thumbhash: result.Thumbhash,
Ext: result.Ext,
}, nil
}
// CacheAudiobookCover is a thin convenience over CacheBytes specifically
// for the audiobook scanner. Avoids exporting the imagecache request
// struct to the scanner package (which would create an import cycle
// scanner -> imagecache -> metadata -> scanner). Stores under
// "local/audiobooks/{contentID}/poster/...".
func (c *Cacher) CacheAudiobookCover(ctx context.Context, data []byte, contentID string) (basePath string, ext string, thumbhash string, err error) {
res, err := c.CacheBytes(ctx, data, CacheRequest{
ProviderID: "local",
ContentType: "audiobooks",
ContentID: contentID,
ImageType: metadata.ImagePoster,
})
if err != nil {
return "", "", "", err
}
return res.BasePath, res.Ext, res.Thumbhash, nil
}
// CacheBytes performs the same variant generation, thumbhash, and S3 upload as
// Cache but starts from raw image bytes already in hand. Used by the
// audiobook scanner to push embedded M4B cover art into S3 without round-
// tripping through HTTP.
func (c *Cacher) CacheBytes(ctx context.Context, data []byte, req CacheRequest) (*CacheResult, error) {
if strings.TrimSpace(req.ProviderID) == "" {
return nil, fmt.Errorf("imagecache: provider ID is required")
}
if strings.TrimSpace(req.ContentType) == "" {
return nil, fmt.Errorf("imagecache: content type is required")
}
if strings.TrimSpace(req.ContentID) == "" {
return nil, fmt.Errorf("imagecache: content ID is required")
}
if len(data) == 0 {
return nil, fmt.Errorf("imagecache: image data is empty")
}
thumbhash, err := imageutil.Thumbhash(data)
if err != nil {
return nil, fmt.Errorf("imagecache: thumbhash: %w", err)
}
widths := variantWidths(req.ImageType)
result, err := imageutil.GenerateVariants(data, widths)
if err != nil {
return nil, fmt.Errorf("imagecache: generate variants: %w", err)
}
basePath := buildBasePath(req.ProviderID, req.ContentType, req.ContentID, req.ImageType, req.SeasonNumber, req.EpisodeNumber)
bucket := c.s3.Bucket()
var wg sync.WaitGroup
uploadErrs := make([]error, len(result.Variants))
for i, v := range result.Variants {
wg.Add(1)
go func(idx int, variant imageutil.Variant) {
defer wg.Done()
key := basePath + "/" + variant.Key + result.Ext
if err := c.s3.PutObject(ctx, bucket, key, variant.Data); err != nil {
uploadErrs[idx] = fmt.Errorf("imagecache: upload %s: %w", key, err)
}
}(i, v)
}
wg.Wait()
for _, err := range uploadErrs {
if err != nil {
return nil, err
}
}
return &CacheResult{BasePath: basePath, Thumbhash: thumbhash, Ext: result.Ext}, nil
}
// Cache downloads the image at req.SourceURL, generates variants, computes a
// thumbhash, uploads all variants to S3, and returns the base path and thumbhash.
func (c *Cacher) Cache(ctx context.Context, req CacheRequest) (*CacheResult, error) {
if strings.TrimSpace(req.ProviderID) == "" {
return nil, fmt.Errorf("imagecache: provider ID is required")
}
if strings.TrimSpace(req.ContentType) == "" {
return nil, fmt.Errorf("imagecache: content type is required")
}
if strings.TrimSpace(req.ContentID) == "" {
return nil, fmt.Errorf("imagecache: content ID is required")
}
if req.EpisodeNumber != nil && req.SeasonNumber == nil {
return nil, fmt.Errorf("imagecache: episode number requires a season number")
}
url := req.SourceURL
// Resolve non-HTTP paths (e.g. plugin_id://path) via the resolver.
if !strings.HasPrefix(url, "http://") && !strings.HasPrefix(url, "https://") {
if req.ImageResolver == nil {
return nil, fmt.Errorf("imagecache: non-HTTP URL %q requires ImageResolver", url)
}
url = req.ImageResolver.ResolveImageURL(ctx, url, "original")
if url == "" {
return nil, fmt.Errorf("imagecache: resolver returned empty URL for %q", req.SourceURL)
}
}
data, err := c.downloadImage(ctx, url)
if err != nil {
return nil, fmt.Errorf("imagecache: download %s: %w", url, err)
}
// Compute thumbhash from the original downloaded data (JPEG/PNG) before
// converting to WebP, since Go's image.Decode doesn't support WebP.
thumbhash, err := imageutil.Thumbhash(data)
if err != nil {
return nil, fmt.Errorf("imagecache: thumbhash: %w", err)
}
widths := variantWidths(req.ImageType)
result, err := imageutil.GenerateVariants(data, widths)
if err != nil {
return nil, fmt.Errorf("imagecache: generate variants: %w", err)
}
basePath := buildBasePath(req.ProviderID, req.ContentType, req.ContentID, req.ImageType, req.SeasonNumber, req.EpisodeNumber)
bucket := c.s3.Bucket()
// Upload all variants concurrently.
var wg sync.WaitGroup
uploadErrs := make([]error, len(result.Variants))
for i, v := range result.Variants {
wg.Add(1)
go func(idx int, variant imageutil.Variant) {
defer wg.Done()
key := basePath + "/" + variant.Key + result.Ext
if err := c.s3.PutObject(ctx, bucket, key, variant.Data); err != nil {
uploadErrs[idx] = fmt.Errorf("imagecache: upload %s: %w", key, err)
}
}(i, v)
}
wg.Wait()
for _, err := range uploadErrs {
if err != nil {
return nil, err
}
}
return &CacheResult{
BasePath: basePath,
Thumbhash: thumbhash,
Ext: result.Ext,
}, nil
}
// variantWidths returns the resize widths for the given image type.
func variantWidths(t metadata.ImageType) []int {
switch t {
case metadata.ImagePoster:
return []int{500, 300}
case metadata.ImageBackdrop:
return []int{1920, 1280, 300}
case metadata.ImageLogo:
return []int{500}
case metadata.ImageStill:
return []int{500, 300}
default:
return []int{500, 300}
}
}
// buildBasePath constructs the S3 key prefix for a given image. Season
// posters and episode stills nest under their parent series so a single
// DeletePrefix on the series prefix cascades to all child images.
//
// item-level: {provider}/{type}/{id}/{imageType}
// season: {provider}/{type}/{id}/seasons/{n}/{imageType}
// episode: {provider}/{type}/{id}/seasons/{n}/episodes/{m}/{imageType}
func buildBasePath(providerID, contentType, contentID string, t metadata.ImageType, seasonNumber, episodeNumber *int) string {
imageTypeName := imageTypeName(t)
base := fmt.Sprintf("%s/%s/%s", providerID, contentType, contentID)
if seasonNumber != nil {
base = fmt.Sprintf("%s/seasons/%d", base, *seasonNumber)
if episodeNumber != nil {
base = fmt.Sprintf("%s/episodes/%d", base, *episodeNumber)
}
}
return base + "/" + imageTypeName
}
// imageTypeName returns the lowercase string name for an ImageType.
func imageTypeName(t metadata.ImageType) string {
switch t {
case metadata.ImagePoster:
return "poster"
case metadata.ImageBackdrop:
return "backdrop"
case metadata.ImageLogo:
return "logo"
case metadata.ImageStill:
return "still"
default:
return "unknown"
}
}
// downloadImage fetches the image at the given URL, enforcing size, timeout,
// and public-network limits.
func (c *Cacher) downloadImage(ctx context.Context, rawURL string) ([]byte, error) {
parsed, err := url.Parse(rawURL)
if err != nil {
return nil, fmt.Errorf("parse URL: %w", err)
}
if c.enforcePublicURLs {
if err := validatePublicImageURL(parsed); err != nil {
return nil, err
}
}
ctx, cancel := context.WithTimeout(ctx, downloadTimeout)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, rawURL, nil)
if err != nil {
return nil, fmt.Errorf("build request: %w", err)
}
client := c.httpClient
if client == nil {
client = newSecureHTTPClient()
}
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("http get: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("unexpected status %d", resp.StatusCode)
}
limited := io.LimitReader(resp.Body, maxDownloadBytes+1)
data, err := io.ReadAll(limited)
if err != nil {
return nil, fmt.Errorf("read body: %w", err)
}
if int64(len(data)) > maxDownloadBytes {
return nil, fmt.Errorf("image exceeds %d byte limit", maxDownloadBytes)
}
return data, nil
}
func newSecureHTTPClient() *http.Client {
transport := &http.Transport{
Proxy: nil,
DialContext: secureImageDialContext,
TLSHandshakeTimeout: 10 * time.Second,
}
return &http.Client{
Transport: transport,
CheckRedirect: func(req *http.Request, via []*http.Request) error {
if len(via) >= 10 {
return http.ErrUseLastResponse
}
return validatePublicImageURL(req.URL)
},
}
}
func validatePublicImageURL(u *url.URL) error {
if u == nil {
return fmt.Errorf("empty URL")
}
if u.Scheme != "http" && u.Scheme != "https" {
return fmt.Errorf("unsupported URL scheme %q", u.Scheme)
}
host := u.Hostname()
if host == "" {
return fmt.Errorf("URL host is required")
}
if addr, err := netip.ParseAddr(host); err == nil && !isPublicAddr(addr) {
return fmt.Errorf("private image host %q is not allowed", host)
}
return nil
}
func secureImageDialContext(ctx context.Context, network, address string) (net.Conn, error) {
host, port, err := net.SplitHostPort(address)
if err != nil {
return nil, err
}
addr, err := resolvePublicAddr(ctx, host)
if err != nil {
return nil, err
}
dialer := &net.Dialer{Timeout: downloadTimeout}
return dialer.DialContext(ctx, network, net.JoinHostPort(addr.String(), port))
}
func resolvePublicAddr(ctx context.Context, host string) (netip.Addr, error) {
if addr, err := netip.ParseAddr(host); err == nil {
if isPublicAddr(addr) {
return addr, nil
}
return netip.Addr{}, fmt.Errorf("private image host %q is not allowed", host)
}
ips, err := net.DefaultResolver.LookupIPAddr(ctx, host)
if err != nil {
return netip.Addr{}, fmt.Errorf("resolve image host %q: %w", host, err)
}
for _, ip := range ips {
addr, ok := netip.AddrFromSlice(ip.IP)
if ok && isPublicAddr(addr) {
return addr, nil
}
}
return netip.Addr{}, fmt.Errorf("image host %q did not resolve to a public address", host)
}
func isPublicAddr(addr netip.Addr) bool {
if addr.Is4In6() {
addr = addr.Unmap()
}
return addr.IsGlobalUnicast() &&
!addr.IsPrivate() &&
!addr.IsLoopback() &&
!addr.IsLinkLocalUnicast() &&
!addr.IsLinkLocalMulticast() &&
!addr.IsMulticast() &&
!addr.IsUnspecified()
}