Files
silo-server/internal/access/resolver.go
T
39ba284c9d feat(ai): shared AI core — metadata translation, Whisper ASR, per-profile language, on-view translation (#127)
* docs: design + plan for shared AI core, metadata translation, Whisper ASR

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(ai): shared LLM client, segment translator, and job runner packages

internal/ai/llm: OpenAI-compatible chat client moved out of subtitles/ai,
plus /v1/audio/transcriptions (verbose_json) for the ASR work; one shared
retry/backoff loop for both. internal/ai/translate: the batched indexed-JSON
translation protocol generalized to text segments. internal/ai/jobrunner:
dispatch/heartbeat/reaper/cancel lifecycle extracted behind a minimal store
interface, with a semaphore shareable across job services.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(subtitles): consume shared AI core

LLMTranslator becomes a thin cue<->segment adapter over aitranslate; the
service delegates dispatch/heartbeat/reaper/cancel to jobrunner; the local
OpenAI client is gone in favor of internal/ai/llm. Behavior (prompts, wire
protocol, job rows, recovery semantics) is unchanged. NewService now takes
the dispatch semaphore so all AI job services can share one bound.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(config): shared ai.* settings, metadata translation job table, localization provenance columns

ai.* connection keys (chat + optional separate ASR endpoint) load with a
fallback to the legacy subtitle_ai.* rows — those are never renamed in SQL
because encrypted values are GCM-bound to their setting key. New toggles:
subtitle_ai.transcribe_enabled, metadata_ai.enabled. Migration adds
metadata_translation_jobs, per-field provenance (provider|ai|manual) on the
localization tables, and media_folders.auto_translate_metadata.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(catalog): localization field provenance with provider/ai/manual precedence

Provider upserts keep manual values and never blank a field with an empty
incoming value; new UpsertAITranslation/UpsertAIOverview methods write AI
fields only over empty or ai-sourced values (force adds provider, never
manual) — all enforced in single-statement SQL. Serving now merges only
non-empty localized fields onto the base item, since localization rows are
legitimately partial (AI rows carry no titles/artwork).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(metadata): AI translation service, refresh auto-fallback, and admin API

internal/metadata/translation: job service over the shared AI core that
expands an item to its season/episode overviews, skips already-localized
fields (zero model calls on repeat runs), batches paragraphs through the
generic translator, and persists per batch with provenance-aware upserts.
MetadataService gains an AutoTranslator seam invoked after each refresh for
libraries with auto_translate_metadata. Admin endpoints under the metadata
curation guard: enqueue, list (poll), cancel; plus a status probe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(subtitles): Whisper ASR transcribe and transcribe_translate jobs

New WhisperTranscriber: one ffmpeg pass extracts the audio track to 10-min
16kHz mono WAV chunks (temp dir cleaned on every exit path), each chunk goes
to the OpenAI-compatible /v1/audio/transcriptions endpoint (verbose_json,
per-request timeout sized to 3x chunk duration), segment timestamps are
offset and built into wrapped cues. Chunks process playhead-first and stream
live to the requesting session. The transcript is stored as an ordinary
downloaded subtitle (provider 'transcribed'); transcribe_translate chains
the existing translator and stores the translated track as the job result.
Enqueue accepts an optional kind; status reports transcribe_enabled.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): AI services settings, metadata translate action, library auto-translate, generate-from-audio

New AI Services admin page hosts the shared endpoint config (reads fall back
to legacy subtitle_ai.* values, writes target ai.*) and the three feature
toggles; the AI card moves out of Subtitles settings. The metadata editor
gains a Translate-with-AI panel with job polling and force/re-translate. The
library form gains the auto-translate toggle (threaded through the libraries
API). The player translate modal gains a From-audio mode that lists audio
tracks and submits transcribe / transcribe_translate jobs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* style: gofmt import grouping in router and translation tests

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(catalog): per-profile metadata language and viewer-triggered description translation

user_profiles.preferred_metadata_language threads through the access scope
into catalog serving: presentation language now resolves explicit param ->
profile preference -> library metadata language (native API and jellycompat).
ItemDetail gains pending_translation_language when the viewer's language is
missing a localized overview. New metadata_ai.on_view setting (off|button|
auto) gates POST /items/{id}/translate-description: any profile with item
access may request its language, with in-flight dedup and a 15-minute
failure cooldown so page views never hammer a broken endpoint.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): on-view description translation with per-profile metadata language

Profile playback settings gain a Metadata language picker (library default
inherit). Detail pages: when the server reports pending_translation_language
and metadata_ai.on_view is 'auto', the description translates on view with a
pulse animation until the refetched detail comes back localized (45s
timeout); in 'button' mode a small Translate chip triggers the same flow.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(web): expose metadata_ai.on_view in AI Services settings

The on-view translation mode had no UI control, so it could only ever be
'off' — viewers got neither the auto translation nor the fallback button.
Adds the off/button/auto selector to the Features card, and the config
loader now warns and falls back to 'off' on a bad row instead of refusing
to start.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ai): clear configuration hint when the transcription endpoint is chat-only

A blank Transcription base URL falls back to the chat endpoint; chat-only
gateways reject the multipart upload with an opaque 400 that reads like a
pipeline bug. 400/404/405 transcription failures now carry a hint to set a
Whisper-compatible endpoint in AI Services.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(subtitles): wrap ASR cue text by rune count, not bytes

Arabic/Cyrillic/Greek text is 2+ bytes per character in UTF-8, so byte-based
wrapping broke lines at roughly half the intended visual width.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(web): steer transcription base URL hint away from chat-only gateways

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(ai): block chat-only gateways for transcription, add endpoint presets

llm.IsChatOnlyGateway (OpenRouter et al — no timestamped transcription API)
is enforced in three layers: the settings API rejects ai.asr_base_url values
pointing at one, the router disables ASR with a warning when the blank-URL
fallback would land on one, and llm.Transcribe refuses outright. The AI
Services page gains one-click transcription presets (Groq turbo/accurate,
OpenAI, self-hosted speaches) plus the mirrored client-side check, and the
settings API now also validates metadata_ai.on_view.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(subtitles): tighten ASR subtitle sync

Three systematic timing-error sources addressed: cue offsets now use the
segment muxer's exact per-chunk start times (segment_list CSV) instead of
assuming index*chunk_seconds; the audio stream's start delay relative to the
container timeline (common in TS remuxes) is probed via ffprobe and added to
every cue; and the chunk length is now operator-tunable via
subtitle_ai.asr_chunk_seconds (60-600s, default 600) since shorter chunks
bound Whisper's within-chunk timestamp drift. Playhead-first ordering now
pivots on real chunk starts, and a beyond-end playhead starts at the final
chunk instead of restarting from zero.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(ai): tolerate base URLs that already include the /v1 segment

Providers like DeepInfra expose their OpenAI-compatible API under a base
that contains the version segment (api.deepinfra.com/v1/openai); always
appending /v1/... mangled those. endpointURL now appends bare paths when
the base already carries /v1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(web): prefer self-hosted transcription in presets and hints

Preset order becomes self-hosted (recommended) -> Groq turbo -> Groq
large-v3 -> OpenAI, and the settings hint plus the job-error hint lead with
the self-hosted option. The self-hosted preset now fills the turbo CT2 model
to match the recommended speaches setup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(subtitles): request VAD and word timestamps for ASR cue accuracy

Without vad_filter, faster-whisper servers report wall-to-wall segment
times: cues linger on screen through silence (verified up to 91s) and
paragraph-length segments become single 400+ char cues. Request
vad_filter=true (skipped for hosted providers that reject non-OpenAI
fields and run VAD server-side) plus timestamp_granularities word+segment,
and rebuild cues from word timings: split at speech pauses, sentence ends,
text capacity, and a 7s max duration; cap word-less segments instead of
trusting their reported end; stretch sub-second cues to a readable minimum.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 14:58:54 -04:00

207 lines
5.6 KiB
Go

package access
import (
"context"
"encoding/json"
"fmt"
"sort"
"github.com/Silo-Server/silo-server/internal/models"
"github.com/Silo-Server/silo-server/internal/userstore"
)
// settingKeyDisabledLibraryIDs is the user-settings key that stores a JSON
// array of library IDs the user has chosen to hide.
const settingKeyDisabledLibraryIDs = "disabled_library_ids"
// UserRepository loads account-level access settings.
type UserRepository interface {
GetByID(ctx context.Context, id int) (*models.User, error)
}
// ProfileTokenValidator validates short-lived profile verification tokens.
type ProfileTokenValidator interface {
Validate(tokenStr string) (*ProfileTokenClaims, error)
}
// Resolver resolves a viewer request into an effective access scope.
type Resolver struct {
users UserRepository
storeFactory userstore.UserStoreProvider
tokens ProfileTokenValidator
}
// NewResolver creates a new scope resolver.
func NewResolver(users UserRepository, storeFactory userstore.UserStoreProvider, tokens ProfileTokenValidator) *Resolver {
return &Resolver{
users: users,
storeFactory: storeFactory,
tokens: tokens,
}
}
// Resolve computes the effective viewer scope for the request.
func (r *Resolver) Resolve(ctx context.Context, input ResolveInput) (Scope, error) {
user, err := r.users.GetByID(ctx, input.UserID)
if err != nil {
return Scope{}, fmt.Errorf("loading user %d: %w", input.UserID, err)
}
scope := Scope{
UserID: user.ID,
ProfileID: input.ProfileID,
AllowedLibraryIDs: cloneInts(user.LibraryIDs),
LibrariesRestricted: user.LibraryIDs != nil,
MaxPlaybackQuality: NormalizePlaybackQuality(user.MaxPlaybackQuality),
PolicyRevision: user.AccessPolicyRevision,
ProfileVerified: input.ProfileID == "",
}
store, err := r.storeFactory.ForUser(ctx, input.UserID)
if err != nil {
return Scope{}, fmt.Errorf("opening user store for %d: %w", input.UserID, err)
}
if input.ProfileID != "" {
profile, err := store.GetProfile(ctx, input.ProfileID)
if err != nil {
return Scope{}, fmt.Errorf("loading profile %s: %w", input.ProfileID, err)
}
if profile == nil {
return Scope{}, ErrProfileNotFound
}
scope.MaxContentRating = profile.MaxContentRating
scope.MaxPlaybackQuality = MinQuality(scope.MaxPlaybackQuality, NormalizePlaybackQuality(profile.MaxPlaybackQuality))
scope.PreferredMetadataLanguage = profile.PreferredMetadataLanguage
scope.AllowedLibraryIDs, scope.LibrariesRestricted = effectiveLibraries(user.LibraryIDs, profile)
scope.ProfileVerified = profile.PINHash == "" || input.SkipPINVerification
if profile.PINHash != "" && !input.SkipPINVerification {
if r.tokens == nil {
return Scope{}, ErrProfileUnverified
}
claims, err := r.tokens.Validate(input.ProfileToken)
if err != nil {
return Scope{}, err
}
if claims.UserID != user.ID || claims.SessionID != input.SessionID || claims.ProfileID != profile.ID || claims.PolicyRevision != user.AccessPolicyRevision {
return Scope{}, ErrProfileUnverified
}
scope.ProfileVerified = true
}
}
// Apply user-level disabled library IDs setting.
disabled := loadDisabledLibraryIDs(ctx, store)
if len(disabled) > 0 {
if scope.AllowedLibraryIDs != nil {
// Restricted user: subtract disabled IDs from the allowed set.
scope.AllowedLibraryIDs = subtractInts(scope.AllowedLibraryIDs, disabled)
} else {
// Unrestricted user: pass disabled IDs through so query layer
// can apply a NOT IN filter.
scope.DisabledLibraryIDs = disabled
}
}
return scope, nil
}
// loadDisabledLibraryIDs reads and parses the disabled_library_ids user setting.
func loadDisabledLibraryIDs(ctx context.Context, store userstore.UserStore) []int {
raw, err := store.GetSetting(ctx, settingKeyDisabledLibraryIDs)
if err != nil || raw == "" {
return nil
}
var ids []int
if err := json.Unmarshal([]byte(raw), &ids); err != nil {
return nil
}
// Filter out invalid values.
n := 0
for _, id := range ids {
if id > 0 {
ids[n] = id
n++
}
}
return ids[:n]
}
// subtractInts removes all values in exclude from src, preserving order.
func subtractInts(src, exclude []int) []int {
if len(exclude) == 0 {
return src
}
set := make(map[int]struct{}, len(exclude))
for _, id := range exclude {
set[id] = struct{}{}
}
out := make([]int, 0, len(src))
for _, id := range src {
if _, ok := set[id]; !ok {
out = append(out, id)
}
}
return out
}
func effectiveLibraries(accountLibraryIDs []int, profile *userstore.Profile) ([]int, bool) {
accountRestricted := accountLibraryIDs != nil
profileRestricted := profile.LibraryRestrictionsEnabled
switch {
case accountRestricted && profileRestricted:
return intersectInts(accountLibraryIDs, profile.AllowedLibraryIDs), true
case accountRestricted:
return cloneInts(accountLibraryIDs), true
case profileRestricted:
return sortedUniqueInts(profile.AllowedLibraryIDs), true
default:
return nil, false
}
}
func intersectInts(left, right []int) []int {
if len(left) == 0 || len(right) == 0 {
return []int{}
}
set := make(map[int]struct{}, len(left))
for _, id := range left {
set[id] = struct{}{}
}
var out []int
for _, id := range right {
if _, ok := set[id]; ok {
out = append(out, id)
}
}
return sortedUniqueInts(out)
}
func sortedUniqueInts(values []int) []int {
if len(values) == 0 {
return []int{}
}
out := cloneInts(values)
sort.Ints(out)
n := 1
for i := 1; i < len(out); i++ {
if out[i] != out[i-1] {
out[n] = out[i]
n++
}
}
return out[:n]
}
func cloneInts(values []int) []int {
if values == nil {
return nil
}
out := make([]int, len(values))
copy(out, values)
return out
}