Files
silo-server/internal/scanner/probe.go
T
eb6024573e feat(audiobooks): make audiobook libraries first-class catalog items (#73)
* docs(audiobooks): design spec for plugin absorption

Plan to absorb silo-plugin-audiobooks into silo-server as a first-party
feature. Audiobooks land in silo's existing SPA; ABS clients connect
directly. Hard constraints: reuse existing tables (media_items,
media_files, user_watch_progress, user_playback_sessions, people,
item_people, library_collections); only two new tables (abs_sessions,
podcast_feeds) and at most one column add (media_libraries.kind);
silo's main :8080 listener handles ABS Socket.io natively. Out of
scope: audiobook requests flow, smart collections, share links,
external recommender, custom metadata providers, separate audiobook
SPA.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): implementation plan sub-plan 1 (discovery + schema)

First of six sub-plans for the absorption. Six tasks: a discovery
audit that resolves the spec's Risk questions, four idempotent SQL
migrations (abs_sessions, podcast_feeds, media_libraries.kind,
audiobooks.enabled feature flag), and an empty-but-compiling
internal/audiobooks package scaffolded into cmd/silo. Lands as a
strict no-op for users (feature flag defaults to false).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): discovery findings for absorption sub-plan 1

Locks schema/code decisions for migrations 139-142 and downstream
sub-plans. Resolves open Risk questions from the absorption design spec.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): migration 139 add abs_sessions table

Parallel of jellycompat_sessions for Audiobookshelf-compatible clients.
Lets ABS mobile/desktop apps maintain a device-bound session that
silo's audiobooks/abs handlers will validate.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(audiobooks): match codebase conventions in migration 139

Lowercases type keywords in the abs_sessions CREATE TABLE body to
match neighboring migrations, fixes the client_version column
alignment, and replaces the misleading "parallel to
jellycompat_sessions" header comment with a more accurate
description of the table's role.

Cosmetic only — the running schema is unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): migration 140 add podcast_feeds table

Side table on media_items for RSS-subscribed podcasts. Holds feed URL,
ETag/Last-Modified for conditional fetches, last-refresh timestamp, and
the per-feed refresh interval consumed by the upcoming
podcastfeed.Refresher scheduled task.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(audiobooks): uppercase PRIMARY KEY in migration 140

Aligns with the codebase convention (type keywords lowercase,
constraint keywords uppercase) established in migration 139's
post-style-fix form. Cosmetic only — running schema is unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(audiobooks): migration 141 no-op for media_folders.type

Sub-plan 1 originally reserved migration 141 to add a 'kind' column to
media_libraries discriminating audiobook/podcast libraries. Discovery
audit (sub-plan 1 Task 1) found that the actual table is media_folders
and it already has a type text NOT NULL column with no CHECK constraint
or enum, so 'audiobooks' and 'podcasts' can be added as future values
without DDL.

Landing this migration as a documented no-op preserves the version
numbering audit trail and pins the decision in git history. The
matching down migration is also a no-op.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): migration 142 add audiobooks.enabled flag

Server-settings row that gates the absorbed audiobooks feature.
Defaults to 'false' so sub-plan 1 lands as a strict no-op; subsequent
sub-plans branch on this flag and operators flip it to 'true' at
cutover.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): scaffold internal/audiobooks package

Empty-but-compiling Service that reads the audiobooks.enabled feature
flag from server_settings. Wired into cmd/silo so the package is
referenced from the binary; no routes mounted, no scheduled tasks
registered, no DB writes. Subsequent sub-plans hang scanner branches,
ABS handlers, Socket.io, podcast refresher, and SPA pages off this
Service.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style(audiobooks): cosmetic cleanups in scaffolded package

Two pre-emptive cleanups flagged by code review before sub-plan 2
copies the patterns:

  1. Sort the internal/audiobooks import after internal/adminjob in
     cmd/silo/main.go (alphabetical).
  2. Drop the redundant "audiobooks: " prefix from the Enabled() error
     wrap; matches how every other top-level service package
     (watchstate, scanqueue, metadata, etc.) formats errors.

No behavior change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): implementation plan sub-plan 2 (scanner)

Second of six sub-plans. 10 tasks: PersonKind constants for Author and
Narrator, audio-extension recognizer, library-type helpers, a
walkLogicalTree refactor (movieLibrary bool -> typed walkMode), chapter
extraction via ffprobe, single-file and multi-file audiobook parsers,
scanner write path producing media_items.type='audiobook', and a
filesystem podcast parser (RSS deferred to sub-plan 5).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): add Author and Narrator PersonKind constants

Discovery audit confirmed item_people.kind is unconstrained smallint
with values 1-6 in use. Reserve 7 = Author, 8 = Narrator for audiobook
people-links written by the upcoming scanner branches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): add audio-extension recognizer for scanner

Mirrors the existing videoExtensions/SupportsVideoFile pair. Used by
upcoming audiobook and podcast scanner branches to filter directory
walks.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): library-type recognizers for scanner dispatch

isAudiobookLibraryType and isPodcastLibraryType match singular and
plural forms case-insensitively, mirroring isMovieLibraryType. Used by
upcoming scanner walk branches (Task 4) that filter audio files into
audiobook and podcast libraries.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(scanner): replace movieLibrary bool with typed walkMode

Lets walkLogicalTree dispatch on multiple library shapes (video, movie,
audiobook, podcast) without proliferating boolean flags. Behavior for
existing video and movie libraries is unchanged; audiobook and podcast
modes will be consumed by the upcoming audiobook.go and podcast.go
parsers in later tasks of this sub-plan.

walkModeFor() derives the mode from a media_folders.type string;
unknown types default to walkModeVideo to preserve prior behavior for
any caller still passing a raw type.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): expose ffprobe format tags on ProbeData

The audiobook scanner needs format-level tags (title, artist, album,
date) for media_items metadata; ffprobe already parses them in
ffprobeFormat.Tags but ProbeData previously discarded them. Add
FormatTags map[string]string to ProbeData, populate it in
convertProbeData via a new normalizeFormatTags helper that lowercases
keys and trims values.

Adds a fixture audiobook .m4b with embedded chapters (Intro/Outro) and
format tags, and a test that verifies ProbeFile() returns both
correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): parser for single-file audiobook folders

parseAudiobookFolder reads tags + chapters via the existing ProbeFile
(now that Task 5 exposes FormatTags on ProbeData) and produces a
parsedAudiobook struct. Title falls back from "title" tag to "album";
author from "artist" -> "album_artist" -> "composer"; series from
"album" -> "series" -> "mvnm" (Movement Name, used by some MP4 tools).
Year parsed from "date" or "year" tags, tolerating ISO dates and
parenthesized forms.

Single-file case only; multi-file folders (one audio file per chapter)
return a placeholder error and arrive in Task 7.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): multi-file audiobook folder support

Folders containing N audio files (one per chapter/part) get one
parsedAudiobookFile per file; each file's chapter list is synthesized
as a single chapter with title = filename stem. Title/author/series/
year come from the first file's tags.

Also drops the duplicate pickFirstNonEmpty helper added in Task 6 in
favor of the existing firstNonEmpty already in probe.go.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): scanner write path produces audiobook media_items

ScanAudiobookFolder walks an audiobooks-typed media folder and treats
each immediate subdirectory as one audiobook. For each parsed audiobook
it upserts:
  - one media_items row with type='audiobook'
  - one media_files row per audio file (with chapters JSONB)
  - author/narrator links in item_people (kind=7, kind=8)

Adds itemRepo and personRepo to the Scanner struct, wired from
fileRepo.Pool() in NewScanner — no constructor signature change needed.

ScanFolder dispatches to this path when folder.Type='audiobooks',
bypassing the per-file movie/TV pipeline because audiobooks are
folder-scoped entities.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): filesystem podcast scanner

ScanPodcastFolder walks a podcasts-typed media folder, treating each
subdirectory as a podcast show and each audio file inside as an
episode. Writes media_items.type='podcast' + episodes rows + media_files
rows. RSS-subscribed feeds (podcast_feeds table) arrive in sub-plan 5;
this task covers filesystem-only ingestion.

ScanFolder dispatches to this path when folder.Type='podcasts'.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(audiobooks): implementation plan sub-plan 5 (podcasts)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): expose audiobooks/podcasts library types in admin UI

Adds 'Audiobooks' and 'Podcasts' options to the library-type dropdown
in the admin libraries page so operators can flag a folder as an
audiobook or podcast library. Extends contentLevelsForType() so the
admin UI's downstream filtering treats those types correctly
(audiobook -> ['audiobook'], podcasts -> ['podcast',
'podcast_episode']).

Backend scanner branches for these types were already wired in
sub-plan 2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 139_abs_sessions to 147 for origin/main merge

origin/main adds 139_media_requests at the same number our local
audiobook branch had used for abs_sessions. Renumber ours to 147 to
free up 139 for the upstream migration. The schema_versions row is
updated in lockstep on the running database so the migrator sees the
abs_sessions migration as already applied at its new version.

Migrations 140-146 (podcast feeds, media_folders kind noop, audiobook
feature flag, abs playback sessions, podcast episode guid, audiobook
series, audiobook title cleanup) stay where they are — they don't
collide with anything on origin/main.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 140_podcast_feeds to 157 for origin/main merge

origin/main added 140_user_permissions at the same version this branch
had used for podcast_feeds. Renumber ours to 157 (next free above the
collections-unify migration at 156) so 140 is free for the upstream
migration. schema_versions on the running database is updated in lockstep
so the migrator sees podcast_feeds as already applied at its new version.

Same pattern as d59c1cb (renumber 139_abs_sessions to 147 for the prior
main merge). Pending migrations after this rename: 132 (downloaded
subtitles admin index, main), 140 (user_permissions, main), and 156
(unify_user_collections, this branch).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 141_media_folders_kind_noop to 159 for origin/main merge

Same shape as eb8f67d (the 140→157 renumber from the previous main
merge). origin/main added 141_episode_title_sort_index at the same
version this branch had used for media_folders_kind_noop. Renumber
ours to 159 (next free above the audiobook_series truncate at 158) so
141 is open for the upstream migration. schema_versions on the
running database is updated in lockstep so the migrator sees
media_folders_kind_noop as already applied at its new version.

Pending migrations on silo-prod after this rename: 141
(episode_title_sort_index, main) and any other newer ones from main
that the branch hasn't picked up yet.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(migrations): renumber 142_audiobooks_feature_flag to 160 for origin/main merge

Companion to 3c6f062's 141 renumber — origin/main also added
142_episode_catalog_entries (alongside 141_episode_title_sort_index)
at a version this branch had used for the audiobooks feature flag.
Renumber ours to 160 so 142 is open for the upstream migration;
schema_versions on silo-prod is updated in lockstep so the migrator
sees audiobooks_feature_flag as already applied at its new version.

This was the only remaining collision (verified by checking for
duplicate version prefixes across migrations/).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(audiobooks): address foundation review comments

* fix(audiobooks): tighten scanner identity handling

* fix(audiobooks): propagate scanner cancellation

* chore(audiobooks): adopt goose migration layout

* docs(audiobooks): implementation plan sub-plan 3 (API + frontend MVP)

Third of six sub-plans. 9 tasks: three REST endpoints (list/detail/
progress), TanStack Query hooks + types, three React pages
(Library/Detail/Player), and navigation integration. Scoped to MVP —
author/series indices, smart collections, share links, and other
nice-to-haves from the spec are deferred. Streaming reuses silo's
existing /api/v1/stream/{session_id}; no new transcode code.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): list endpoint at GET /api/v1/audiobooks

Paginated list of media_items with type='audiobook' scoped to the
caller's accessible libraries via the existing access filter.
Mirrors silo's existing list-style handlers for movies and series.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): detail endpoint at GET /api/v1/audiobooks/{id}

Returns the media_items row, its media_files (with chapters JSONB),
author/narrator extracted from item_people (kinds 7/8), and the
caller's per-profile listening progress from user_watch_progress.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): progress endpoint at POST /api/v1/audiobooks/{id}/progress

UPSERTs user_watch_progress for the caller's (user_id, profile_id,
content_id). Body carries position_seconds; clients are expected to
post every 5-10s during playback plus on pause/seek (matching silo's
existing video progress cadence).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): frontend types and TanStack Query hooks

TypeScript types match the JSON shapes from the new
/api/v1/audiobooks endpoints (list, detail, progress). Three hooks:
useAudiobookLibrary (list), useAudiobook (detail), and
useReportAudiobookProgress (mutation that invalidates the detail
query on success so progress updates reflect immediately).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): library grid page at /audiobooks

Renders a paginated grid of audiobook cards using the
useAudiobookLibrary hook. Each card links to /audiobooks/book/{id}.
Cards show poster, title, and year; falls back to a "No cover"
placeholder when the audiobook has no poster_url. Empty state hints
to operators that they need to set a library's type to 'audiobooks'.

Routes themselves are wired in Task 8 (navigation integration).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): detail page with chapter list

Renders cover, title, author, narrator, year, and overview alongside a
chapter list. Clicking a chapter opens an inline sticky
AudiobookPlayer at that chapter's start. A "Resume" button restarts
playback at the saved progress position if present.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): HTML5 audio player with chapter navigation

Single-file audiobook playback for MVP. Multi-file queuing arrives in
a follow-up. Streams via the existing /api/v1/direct-download GET
endpoint. Position is reported to /api/v1/audiobooks/{id}/progress
every 10s while playing plus on pause/seek/end. Skip-30s, playback
rate select, chapter list panel.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(audiobooks): wire navigation and routes

Adds an Audiobooks entry to the sidebar and registers the two new
routes (/audiobooks for the library grid, /audiobooks/book/:id for
detail). The player renders inline inside the detail page; no
dedicated player route is required for MVP.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(audiobooks): address native API review comments

* feat(audiobooks): add ABS compatibility and polish

* fix(audiobooks): stabilize ABS playback progress reporting

* fix(audiobooks): clean up ABS branch review fixes

* chore(audiobooks): adopt goose layout for ABS migrations

* fix(audiobooks): align player seek bar props

* feat(audiobooks): make libraries first-class catalog items

* feat(admin): add server restart endpoint

* fix(audiobooks): address review comment findings

---------

Co-authored-by: RXWatcher <14085001+RXWatcher@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-07 15:57:05 -04:00

500 lines
14 KiB
Go

package scanner
import (
"context"
"encoding/json"
"fmt"
"os/exec"
"slices"
"strconv"
"strings"
"github.com/Silo-Server/silo-server/internal/lang"
)
// ffprobeOutput represents the top-level JSON output from ffprobe.
type ffprobeOutput struct {
Format ffprobeFormat `json:"format"`
Streams []ffprobeStream `json:"streams"`
Chapters []ffprobeChapter `json:"chapters"`
}
// ffprobeScalarString accepts ffprobe fields that may be emitted as either
// JSON strings or numbers depending on codec/container details.
type ffprobeScalarString string
func (s *ffprobeScalarString) UnmarshalJSON(data []byte) error {
if string(data) == "null" {
*s = ""
return nil
}
var str string
if err := json.Unmarshal(data, &str); err == nil {
*s = ffprobeScalarString(str)
return nil
}
var num json.Number
if err := json.Unmarshal(data, &num); err == nil {
*s = ffprobeScalarString(num.String())
return nil
}
return fmt.Errorf("unsupported ffprobe scalar %s", string(data))
}
// ffprobeFormat represents the "format" section of ffprobe JSON output.
type ffprobeFormat struct {
Filename string `json:"filename"`
FormatName string `json:"format_name"`
FormatLongName string `json:"format_long_name"`
Duration string `json:"duration"`
Size string `json:"size"`
BitRate string `json:"bit_rate"`
Tags map[string]string `json:"tags"`
}
// ffprobeStream represents a single stream entry in ffprobe JSON output.
type ffprobeStream struct {
Index int `json:"index"`
CodecName string `json:"codec_name"`
CodecLongName string `json:"codec_long_name"`
CodecType string `json:"codec_type"`
Profile string `json:"profile"`
Level int `json:"level"`
Width int `json:"width"`
Height int `json:"height"`
DisplayAspectRatio string `json:"display_aspect_ratio"`
FieldOrder string `json:"field_order"`
AvgFrameRate string `json:"avg_frame_rate"`
BitRate string `json:"bit_rate"`
ColorTransfer string `json:"color_transfer"`
ColorPrimaries string `json:"color_primaries"`
ColorSpace string `json:"color_space"`
PixFmt string `json:"pix_fmt"`
Refs int `json:"refs"`
BitsPerRawSample ffprobeScalarString `json:"bits_per_raw_sample"`
BitsPerSample ffprobeScalarString `json:"bits_per_sample"`
Channels int `json:"channels"`
ChannelLayout string `json:"channel_layout"`
SampleRate string `json:"sample_rate"`
Disposition ffprobeDisp `json:"disposition"`
Tags map[string]string `json:"tags"`
SideDataList []ffprobeSideData `json:"side_data_list"`
}
type ffprobeChapter struct {
ID int `json:"id"`
Start ffprobeScalarString `json:"start"`
End ffprobeScalarString `json:"end"`
TimeBase string `json:"time_base"`
StartTime ffprobeScalarString `json:"start_time"`
EndTime ffprobeScalarString `json:"end_time"`
Tags map[string]string `json:"tags"`
}
type ffprobeSideData struct {
SideDataType string `json:"side_data_type"`
DVProfile int `json:"dv_profile"`
DVBlPresent int `json:"dv_bl_present"`
DVElPresent int `json:"dv_el_present"`
}
// ffprobeDisp represents the disposition flags on a stream.
type ffprobeDisp struct {
Default int `json:"default"`
Forced int `json:"forced"`
}
// ProbeFile runs ffprobe on the given file and returns parsed ProbeData.
// ffprobePath is the path to the ffprobe binary. filePath is the media file to probe.
func ProbeFile(ctx context.Context, ffprobePath string, filePath string) (*ProbeData, error) {
cmd := exec.CommandContext(ctx, ffprobePath,
"-v", "quiet",
"-print_format", "json",
"-show_format",
"-show_streams",
"-show_chapters",
filePath,
)
output, err := cmd.Output()
if err != nil {
return nil, fmt.Errorf("ffprobe failed for %s: %w", filePath, err)
}
var raw ffprobeOutput
if err := json.Unmarshal(output, &raw); err != nil {
return nil, fmt.Errorf("ffprobe JSON parse failed for %s: %w", filePath, err)
}
return convertProbeData(&raw), nil
}
// FFprobePathFromFFmpeg derives the sibling ffprobe binary path from a configured ffmpeg path.
func FFprobePathFromFFmpeg(ffmpegPath string) string {
if i := strings.LastIndex(ffmpegPath, "ffmpeg"); i >= 0 {
ffprobePath := ffmpegPath[:i] + "ffprobe" + ffmpegPath[i+len("ffmpeg"):]
if ffprobePath != "" && ffprobePath != ffmpegPath {
return ffprobePath
}
}
return "ffprobe"
}
// convertProbeData transforms raw ffprobe JSON output into ProbeData.
func convertProbeData(raw *ffprobeOutput) *ProbeData {
pd := &ProbeData{
Container: detectContainer(raw.Format.FormatName),
}
// Parse duration from format (ffprobe reports seconds).
// Some containers (notably MKV) may produce duration in microseconds;
// detect and normalise to seconds when the value is unreasonably large.
if raw.Format.Duration != "" {
if dur, err := strconv.ParseFloat(raw.Format.Duration, 64); err == nil {
const maxReasonableSec = 100_000 // ~27.8 hours
if dur > maxReasonableSec {
dur = dur / 1_000_000 // microseconds → seconds
}
pd.Duration = int(dur)
}
}
// Parse bitrate from format (bps to kbps).
if raw.Format.BitRate != "" {
if br, err := strconv.Atoi(raw.Format.BitRate); err == nil {
pd.Bitrate = br / 1000
}
}
for _, s := range raw.Streams {
switch s.CodecType {
case "video":
track := VideoTrackInfo{
Title: firstNonEmpty(s.Tags["title"], s.CodecLongName, strings.ToUpper(s.CodecName)),
Codec: s.CodecName,
DolbyVision: dolbyVisionProfile(s.SideDataList),
Profile: s.Profile,
Level: s.Level,
Width: s.Width,
Height: s.Height,
AspectRatio: s.DisplayAspectRatio,
Interlaced: isInterlaced(s.FieldOrder),
FrameRate: normalizeFrameRate(s.AvgFrameRate),
Bitrate: parseNumeric(s.BitRate) / 1000,
VideoRange: videoRangeLabel(s),
ColorPrimaries: s.ColorPrimaries,
ColorSpace: s.ColorSpace,
ColorTransfer: s.ColorTransfer,
BitDepth: parseBitDepth(s),
PixelFormat: s.PixFmt,
ReferenceFrames: s.Refs,
}
pd.VideoTracks = append(pd.VideoTracks, track)
if pd.CodecVideo == "" {
pd.CodecVideo = s.CodecName
pd.Resolution = mapResolution(s.Width, s.Height)
pd.HDR = isHDR(s.ColorTransfer)
}
case "audio":
track := AudioTrackInfo{
Title: firstNonEmpty(s.Tags["title"], s.CodecLongName, strings.ToUpper(s.CodecName)),
EmbeddedTitle: s.Tags["title"],
Language: lang.Canonical(s.Tags["language"]),
Codec: s.CodecName,
Layout: s.ChannelLayout,
Channels: s.Channels,
Bitrate: parseNumeric(s.BitRate) / 1000,
SampleRate: parseNumeric(s.SampleRate),
BitDepth: parseBitDepth(s),
Default: s.Disposition.Default == 1,
}
pd.AudioTracks = append(pd.AudioTracks, track)
if pd.CodecAudio == "" {
pd.CodecAudio = s.CodecName
pd.AudioChannels = s.Channels
}
case "subtitle":
track := SubtitleTrackInfo{
Index: s.Index,
Codec: s.CodecName,
Language: lang.Canonical(s.Tags["language"]),
Title: firstNonEmpty(s.Tags["title"], strings.ToUpper(s.CodecName)),
EmbeddedTitle: s.Tags["title"],
Resolution: subtitleResolutionLabel(s),
Forced: s.Disposition.Forced == 1,
Default: s.Disposition.Default == 1,
HearingImpaired: dispositionFlag(s.Tags, "hearing_impaired"),
}
pd.SubtitleTracks = append(pd.SubtitleTracks, track)
}
}
pd.Chapters = normalizeChapters(raw.Chapters, pd.Duration)
pd.FormatTags = normalizeFormatTags(raw.Format.Tags)
return pd
}
func parseNumeric(raw string) int {
if raw == "" {
return 0
}
v, err := strconv.Atoi(raw)
if err != nil {
return 0
}
return v
}
func parseFloat(raw string) float64 {
if raw == "" {
return 0
}
value, err := strconv.ParseFloat(raw, 64)
if err != nil {
return 0
}
return value
}
func normalizeChapters(raw []ffprobeChapter, durationSeconds int) []ChapterInfo {
if len(raw) == 0 {
return []ChapterInfo{}
}
limit := float64(durationSeconds)
type chapterRange struct {
title string
start float64
end float64
}
ranges := make([]chapterRange, 0, len(raw))
for _, chapter := range raw {
start := parseFloat(string(chapter.StartTime))
end := parseFloat(string(chapter.EndTime))
if end <= 0 {
end = parseFloat(string(chapter.End))
}
if start <= 0 {
start = parseFloat(string(chapter.Start))
}
if limit > 0 {
if start < 0 {
start = 0
}
if end > limit {
end = limit
}
}
if end <= start {
continue
}
title := strings.TrimSpace(firstNonEmpty(
chapter.Tags["title"],
chapter.Tags["TITLE"],
))
ranges = append(ranges, chapterRange{
title: title,
start: start,
end: end,
})
}
if len(ranges) == 0 {
return []ChapterInfo{}
}
slices.SortStableFunc(ranges, func(a, b chapterRange) int {
switch {
case a.start < b.start:
return -1
case a.start > b.start:
return 1
case a.end < b.end:
return -1
case a.end > b.end:
return 1
default:
return 0
}
})
chapters := make([]ChapterInfo, 0, len(ranges))
for i, chapter := range ranges {
end := chapter.end
if i+1 < len(ranges) && ranges[i+1].start < end {
end = ranges[i+1].start
}
if end <= chapter.start {
continue
}
title := chapter.title
if title == "" {
title = fmt.Sprintf("Chapter %02d", len(chapters)+1)
}
chapters = append(chapters, ChapterInfo{
Index: len(chapters),
Title: title,
StartSeconds: chapter.start,
EndSeconds: end,
Source: "embedded",
})
}
return chapters
}
func parseBitDepth(s ffprobeStream) int {
if v := parseNumeric(string(s.BitsPerRawSample)); v > 0 {
return v
}
return parseNumeric(string(s.BitsPerSample))
}
func normalizeFrameRate(raw string) string {
if raw == "" || raw == "0/0" {
return ""
}
parts := strings.SplitN(raw, "/", 2)
if len(parts) != 2 {
return raw
}
num, err1 := strconv.ParseFloat(parts[0], 64)
den, err2 := strconv.ParseFloat(parts[1], 64)
if err1 != nil || err2 != nil || den == 0 {
return raw
}
return strconv.FormatFloat(num/den, 'f', 3, 64)
}
func isInterlaced(fieldOrder string) bool {
switch strings.ToLower(fieldOrder) {
case "tt", "bb", "tb", "bt":
return true
default:
return false
}
}
func videoRangeLabel(s ffprobeStream) string {
if dv := dolbyVisionProfile(s.SideDataList); dv != "" {
return "DolbyVision"
}
if isHDR(s.ColorTransfer) {
return "HDR"
}
return ""
}
func dolbyVisionProfile(sideData []ffprobeSideData) string {
for _, data := range sideData {
if strings.EqualFold(data.SideDataType, "DOVI configuration record") && data.DVProfile > 0 {
return fmt.Sprintf("Profile %d", data.DVProfile)
}
}
return ""
}
func subtitleResolutionLabel(s ffprobeStream) string {
if s.Width <= 0 || s.Height <= 0 {
return ""
}
return fmt.Sprintf("%dx%d", s.Width, s.Height)
}
func dispositionFlag(tags map[string]string, key string) bool {
if tags == nil {
return false
}
value := strings.TrimSpace(strings.ToLower(tags[key]))
return value == "1" || value == "true" || value == "yes"
}
func firstNonEmpty(values ...string) string {
for _, value := range values {
if strings.TrimSpace(value) != "" {
return value
}
}
return ""
}
// mapResolution converts video dimensions to a standard resolution string.
// Uses upper-bound bucketing (similar to Jellyfin) checking both width and
// height, which correctly handles ultra-wide and non-standard aspect ratios.
func mapResolution(width, height int) string {
switch {
case width <= 0 && height <= 0:
return ""
case width <= 854 && height <= 480:
return "480p"
case width <= 1280 && height <= 962:
return "720p"
case width <= 2560 && height <= 1440:
return "1080p"
case width <= 4096 && height <= 3072:
return "2160p"
case width <= 8192 && height <= 6144:
return "4320p"
default:
return "2160p"
}
}
// isHDR checks whether the color transfer characteristic indicates HDR content.
func isHDR(colorTransfer string) bool {
ct := strings.ToLower(colorTransfer)
return strings.Contains(ct, "smpte2084") || strings.Contains(ct, "arib-std-b67")
}
// normalizeFormatTags lowercases tag keys so callers can look up
// "title", "artist", "album" without worrying about ffprobe's mixed-case
// output. Trims whitespace from values.
func normalizeFormatTags(raw map[string]string) map[string]string {
if len(raw) == 0 {
return nil
}
out := make(map[string]string, len(raw))
for k, v := range raw {
out[strings.ToLower(strings.TrimSpace(k))] = strings.TrimSpace(v)
}
return out
}
// detectContainer maps ffprobe format names to common container names.
func detectContainer(formatName string) string {
// ffprobe format_name can contain multiple names separated by commas
// e.g. "mov,mp4,m4a,3gp,3g2,mj2"
parts := strings.Split(formatName, ",")
for _, p := range parts {
p = strings.TrimSpace(p)
switch p {
case "matroska", "webm":
return "mkv"
case "mov", "mp4", "m4a":
return "mp4"
case "avi":
return "avi"
case "mpegts":
return "ts"
case "flv":
return "flv"
case "ogg":
return "ogg"
case "wmv", "asf":
return "wmv"
}
}
// Fallback: return first part
if len(parts) > 0 && parts[0] != "" {
return strings.TrimSpace(parts[0])
}
return formatName
}