* docs: design + plan for shared AI core, metadata translation, Whisper ASR Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(ai): shared LLM client, segment translator, and job runner packages internal/ai/llm: OpenAI-compatible chat client moved out of subtitles/ai, plus /v1/audio/transcriptions (verbose_json) for the ASR work; one shared retry/backoff loop for both. internal/ai/translate: the batched indexed-JSON translation protocol generalized to text segments. internal/ai/jobrunner: dispatch/heartbeat/reaper/cancel lifecycle extracted behind a minimal store interface, with a semaphore shareable across job services. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(subtitles): consume shared AI core LLMTranslator becomes a thin cue<->segment adapter over aitranslate; the service delegates dispatch/heartbeat/reaper/cancel to jobrunner; the local OpenAI client is gone in favor of internal/ai/llm. Behavior (prompts, wire protocol, job rows, recovery semantics) is unchanged. NewService now takes the dispatch semaphore so all AI job services can share one bound. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(config): shared ai.* settings, metadata translation job table, localization provenance columns ai.* connection keys (chat + optional separate ASR endpoint) load with a fallback to the legacy subtitle_ai.* rows — those are never renamed in SQL because encrypted values are GCM-bound to their setting key. New toggles: subtitle_ai.transcribe_enabled, metadata_ai.enabled. Migration adds metadata_translation_jobs, per-field provenance (provider|ai|manual) on the localization tables, and media_folders.auto_translate_metadata. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(catalog): localization field provenance with provider/ai/manual precedence Provider upserts keep manual values and never blank a field with an empty incoming value; new UpsertAITranslation/UpsertAIOverview methods write AI fields only over empty or ai-sourced values (force adds provider, never manual) — all enforced in single-statement SQL. Serving now merges only non-empty localized fields onto the base item, since localization rows are legitimately partial (AI rows carry no titles/artwork). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(metadata): AI translation service, refresh auto-fallback, and admin API internal/metadata/translation: job service over the shared AI core that expands an item to its season/episode overviews, skips already-localized fields (zero model calls on repeat runs), batches paragraphs through the generic translator, and persists per batch with provenance-aware upserts. MetadataService gains an AutoTranslator seam invoked after each refresh for libraries with auto_translate_metadata. Admin endpoints under the metadata curation guard: enqueue, list (poll), cancel; plus a status probe. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(subtitles): Whisper ASR transcribe and transcribe_translate jobs New WhisperTranscriber: one ffmpeg pass extracts the audio track to 10-min 16kHz mono WAV chunks (temp dir cleaned on every exit path), each chunk goes to the OpenAI-compatible /v1/audio/transcriptions endpoint (verbose_json, per-request timeout sized to 3x chunk duration), segment timestamps are offset and built into wrapped cues. Chunks process playhead-first and stream live to the requesting session. The transcript is stored as an ordinary downloaded subtitle (provider 'transcribed'); transcribe_translate chains the existing translator and stores the translated track as the job result. Enqueue accepts an optional kind; status reports transcribe_enabled. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(web): AI services settings, metadata translate action, library auto-translate, generate-from-audio New AI Services admin page hosts the shared endpoint config (reads fall back to legacy subtitle_ai.* values, writes target ai.*) and the three feature toggles; the AI card moves out of Subtitles settings. The metadata editor gains a Translate-with-AI panel with job polling and force/re-translate. The library form gains the auto-translate toggle (threaded through the libraries API). The player translate modal gains a From-audio mode that lists audio tracks and submits transcribe / transcribe_translate jobs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * style: gofmt import grouping in router and translation tests Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(catalog): per-profile metadata language and viewer-triggered description translation user_profiles.preferred_metadata_language threads through the access scope into catalog serving: presentation language now resolves explicit param -> profile preference -> library metadata language (native API and jellycompat). ItemDetail gains pending_translation_language when the viewer's language is missing a localized overview. New metadata_ai.on_view setting (off|button| auto) gates POST /items/{id}/translate-description: any profile with item access may request its language, with in-flight dedup and a 15-minute failure cooldown so page views never hammer a broken endpoint. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(web): on-view description translation with per-profile metadata language Profile playback settings gain a Metadata language picker (library default inherit). Detail pages: when the server reports pending_translation_language and metadata_ai.on_view is 'auto', the description translates on view with a pulse animation until the refetched detail comes back localized (45s timeout); in 'button' mode a small Translate chip triggers the same flow. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(web): expose metadata_ai.on_view in AI Services settings The on-view translation mode had no UI control, so it could only ever be 'off' — viewers got neither the auto translation nor the fallback button. Adds the off/button/auto selector to the Features card, and the config loader now warns and falls back to 'off' on a bad row instead of refusing to start. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai): clear configuration hint when the transcription endpoint is chat-only A blank Transcription base URL falls back to the chat endpoint; chat-only gateways reject the multipart upload with an opaque 400 that reads like a pipeline bug. 400/404/405 transcription failures now carry a hint to set a Whisper-compatible endpoint in AI Services. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(subtitles): wrap ASR cue text by rune count, not bytes Arabic/Cyrillic/Greek text is 2+ bytes per character in UTF-8, so byte-based wrapping broke lines at roughly half the intended visual width. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(web): steer transcription base URL hint away from chat-only gateways Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ai): block chat-only gateways for transcription, add endpoint presets llm.IsChatOnlyGateway (OpenRouter et al — no timestamped transcription API) is enforced in three layers: the settings API rejects ai.asr_base_url values pointing at one, the router disables ASR with a warning when the blank-URL fallback would land on one, and llm.Transcribe refuses outright. The AI Services page gains one-click transcription presets (Groq turbo/accurate, OpenAI, self-hosted speaches) plus the mirrored client-side check, and the settings API now also validates metadata_ai.on_view. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(subtitles): tighten ASR subtitle sync Three systematic timing-error sources addressed: cue offsets now use the segment muxer's exact per-chunk start times (segment_list CSV) instead of assuming index*chunk_seconds; the audio stream's start delay relative to the container timeline (common in TS remuxes) is probed via ffprobe and added to every cue; and the chunk length is now operator-tunable via subtitle_ai.asr_chunk_seconds (60-600s, default 600) since shorter chunks bound Whisper's within-chunk timestamp drift. Playhead-first ordering now pivots on real chunk starts, and a beyond-end playhead starts at the final chunk instead of restarting from zero. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai): tolerate base URLs that already include the /v1 segment Providers like DeepInfra expose their OpenAI-compatible API under a base that contains the version segment (api.deepinfra.com/v1/openai); always appending /v1/... mangled those. endpointURL now appends bare paths when the base already carries /v1. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(web): prefer self-hosted transcription in presets and hints Preset order becomes self-hosted (recommended) -> Groq turbo -> Groq large-v3 -> OpenAI, and the settings hint plus the job-error hint lead with the self-hosted option. The self-hosted preset now fills the turbo CT2 model to match the recommended speaches setup. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(subtitles): request VAD and word timestamps for ASR cue accuracy Without vad_filter, faster-whisper servers report wall-to-wall segment times: cues linger on screen through silence (verified up to 91s) and paragraph-length segments become single 400+ char cues. Request vad_filter=true (skipped for hosted providers that reject non-OpenAI fields and run VAD server-side) plus timestamp_granularities word+segment, and rebuild cues from word timings: split at speech pauses, sentence ends, text capacity, and a 7s max duration; cap word-less segments instead of trusting their reported end; stretch sub-second cues to a readable minimum. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
213 lines
7.6 KiB
Go
213 lines
7.6 KiB
Go
package handlers
|
|
|
|
import (
|
|
"encoding/json"
|
|
"errors"
|
|
"log/slog"
|
|
"net/http"
|
|
"strconv"
|
|
|
|
"github.com/go-chi/chi/v5"
|
|
|
|
"github.com/Silo-Server/silo-server/internal/access"
|
|
apimw "github.com/Silo-Server/silo-server/internal/api/middleware"
|
|
"github.com/Silo-Server/silo-server/internal/catalog"
|
|
"github.com/Silo-Server/silo-server/internal/metadata/translation"
|
|
)
|
|
|
|
// MetadataAIHandler exposes AI translation of catalog descriptions into the
|
|
// localization tables. The admin routes are mounted under the per-item
|
|
// metadata curation guard; the on-view route is viewer-facing and enforces
|
|
// item access itself.
|
|
type MetadataAIHandler struct {
|
|
service *translation.Service
|
|
// ItemAccess authorizes the viewer-facing on-view route; nil disables it.
|
|
ItemAccess *catalog.ItemRepository
|
|
}
|
|
|
|
// NewMetadataAIHandler creates a handler backed by the given service.
|
|
func NewMetadataAIHandler(service *translation.Service) *MetadataAIHandler {
|
|
return &MetadataAIHandler{service: service}
|
|
}
|
|
|
|
// HandleStatus reports whether metadata AI translation is available and the
|
|
// viewer-facing on-view mode, so the metadata editor and detail pages can
|
|
// show or hide their entry points.
|
|
// GET /api/v1/metadata/ai/status
|
|
func (h *MetadataAIHandler) HandleStatus(w http.ResponseWriter, r *http.Request) {
|
|
writeJSON(w, http.StatusOK, map[string]any{
|
|
"enabled": h.service.Enabled(),
|
|
"on_view": h.service.OnViewMode(),
|
|
})
|
|
}
|
|
|
|
// WriteMetadataAIDisabledStatus answers the status probe with a clean negative
|
|
// when no metadata AI handler is wired.
|
|
func WriteMetadataAIDisabledStatus(w http.ResponseWriter, _ *http.Request) {
|
|
writeJSON(w, http.StatusOK, map[string]any{"enabled": false, "on_view": "off"})
|
|
}
|
|
|
|
type translateDescriptionRequest struct {
|
|
// TargetLanguage echoes the detail response's pending_translation_language.
|
|
TargetLanguage string `json:"target_language"`
|
|
}
|
|
|
|
// HandleTranslateOnView is the viewer-facing on-demand description
|
|
// translation: any profile that can access the item may request its
|
|
// descriptions in the language the detail response reported missing. Gated by
|
|
// metadata_ai.on_view; duplicate viewers collapse onto one job and recently
|
|
// failed targets are not retried (cooldown in the service).
|
|
// POST /api/v1/items/{id}/translate-description
|
|
func (h *MetadataAIHandler) HandleTranslateOnView(w http.ResponseWriter, r *http.Request) {
|
|
contentID := chi.URLParam(r, "id")
|
|
var req translateDescriptionRequest
|
|
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
|
writeError(w, http.StatusBadRequest, "invalid_request", "Invalid request body")
|
|
return
|
|
}
|
|
if req.TargetLanguage == "" {
|
|
writeError(w, http.StatusBadRequest, "bad_request", "target_language is required")
|
|
return
|
|
}
|
|
|
|
scope, ok := access.GetScope(r.Context())
|
|
if !ok || h.ItemAccess == nil {
|
|
writeError(w, http.StatusForbidden, "forbidden", "Viewer access is required")
|
|
return
|
|
}
|
|
filter := catalog.AccessFilter{
|
|
AllowedLibraryIDs: scope.AllowedLibraryIDs,
|
|
DisabledLibraryIDs: scope.DisabledLibraryIDs,
|
|
MaxContentRating: scope.MaxContentRating,
|
|
UserID: scope.UserID,
|
|
ProfileID: scope.ProfileID,
|
|
}
|
|
if err := h.ItemAccess.EnsureAccessible(r.Context(), contentID, filter); err != nil {
|
|
if errors.Is(err, catalog.ErrItemNotFound) {
|
|
writeError(w, http.StatusNotFound, "not_found", "Item not found")
|
|
return
|
|
}
|
|
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to authorize item")
|
|
return
|
|
}
|
|
|
|
var requestedBy *int
|
|
if userID := apimw.GetUserID(r.Context()); userID != 0 {
|
|
requestedBy = &userID
|
|
}
|
|
|
|
job, err := h.service.RequestOnView(r.Context(), contentID, req.TargetLanguage, requestedBy)
|
|
if err != nil {
|
|
switch {
|
|
case errors.Is(err, translation.ErrNotConfigured):
|
|
writeError(w, http.StatusServiceUnavailable, "not_configured",
|
|
"On-view translation is not enabled on this server")
|
|
case errors.Is(err, translation.ErrInvalidRequest):
|
|
writeError(w, http.StatusBadRequest, "bad_request", err.Error())
|
|
default:
|
|
slog.Error("failed to request on-view translation",
|
|
"content_id", contentID, "error", err)
|
|
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to start translation")
|
|
}
|
|
return
|
|
}
|
|
|
|
writeJSON(w, http.StatusAccepted, map[string]any{"job": job})
|
|
}
|
|
|
|
type translateMetadataRequest struct {
|
|
TargetLanguage string `json:"target_language"`
|
|
IncludeChildren *bool `json:"include_children"` // default true
|
|
Force bool `json:"force"`
|
|
}
|
|
|
|
// HandleTranslate enqueues a translation job for an item.
|
|
// POST /api/v1/admin/items/{id}/metadata-translation
|
|
func (h *MetadataAIHandler) HandleTranslate(w http.ResponseWriter, r *http.Request) {
|
|
contentID := chi.URLParam(r, "id")
|
|
var req translateMetadataRequest
|
|
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
|
|
writeError(w, http.StatusBadRequest, "invalid_request", "Invalid request body")
|
|
return
|
|
}
|
|
if req.TargetLanguage == "" {
|
|
writeError(w, http.StatusBadRequest, "bad_request", "target_language is required")
|
|
return
|
|
}
|
|
includeChildren := true
|
|
if req.IncludeChildren != nil {
|
|
includeChildren = *req.IncludeChildren
|
|
}
|
|
|
|
var requestedBy *int
|
|
if userID := apimw.GetUserID(r.Context()); userID != 0 {
|
|
requestedBy = &userID
|
|
}
|
|
|
|
job, err := h.service.Enqueue(r.Context(), translation.JobRequest{
|
|
TargetKind: translation.TargetItem,
|
|
ContentID: contentID,
|
|
TargetLanguage: req.TargetLanguage,
|
|
IncludeChildren: includeChildren,
|
|
Force: req.Force,
|
|
RequestedBy: requestedBy,
|
|
})
|
|
if err != nil {
|
|
switch {
|
|
case errors.Is(err, translation.ErrNotConfigured):
|
|
writeError(w, http.StatusServiceUnavailable, "not_configured",
|
|
"Metadata AI translation is not configured on this server")
|
|
case errors.Is(err, translation.ErrInvalidRequest):
|
|
writeError(w, http.StatusBadRequest, "bad_request", err.Error())
|
|
default:
|
|
slog.Error("failed to enqueue metadata translation",
|
|
"content_id", contentID, "error", err)
|
|
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to start translation")
|
|
}
|
|
return
|
|
}
|
|
|
|
writeJSON(w, http.StatusAccepted, map[string]any{"job": job})
|
|
}
|
|
|
|
// HandleListJobs lists recent translation jobs for an item; the metadata
|
|
// editor polls this for progress.
|
|
// GET /api/v1/admin/items/{id}/metadata-translation/jobs
|
|
func (h *MetadataAIHandler) HandleListJobs(w http.ResponseWriter, r *http.Request) {
|
|
jobs, err := h.service.ListJobs(r.Context(), chi.URLParam(r, "id"))
|
|
if err != nil {
|
|
writeError(w, http.StatusInternalServerError, "list_error", "Failed to list jobs")
|
|
return
|
|
}
|
|
writeJSON(w, http.StatusOK, map[string]any{"jobs": jobs})
|
|
}
|
|
|
|
// HandleCancelJob cancels a job belonging to the item in the URL.
|
|
// POST /api/v1/admin/items/{id}/metadata-translation/jobs/{job_id}/cancel
|
|
func (h *MetadataAIHandler) HandleCancelJob(w http.ResponseWriter, r *http.Request) {
|
|
jobID, err := strconv.ParseInt(chi.URLParam(r, "job_id"), 10, 64)
|
|
if err != nil {
|
|
writeError(w, http.StatusBadRequest, "invalid_id", "Invalid job ID")
|
|
return
|
|
}
|
|
job, err := h.service.GetJob(r.Context(), jobID)
|
|
if err != nil {
|
|
if errors.Is(err, translation.ErrJobNotFound) {
|
|
writeError(w, http.StatusNotFound, "not_found", "Job not found")
|
|
return
|
|
}
|
|
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to load job")
|
|
return
|
|
}
|
|
// The curation guard authorized {id}; the job must belong to it.
|
|
if job.ContentID != chi.URLParam(r, "id") {
|
|
writeError(w, http.StatusNotFound, "not_found", "Job not found")
|
|
return
|
|
}
|
|
if err := h.service.Cancel(r.Context(), jobID); err != nil {
|
|
writeError(w, http.StatusInternalServerError, "internal_error", "Failed to cancel job")
|
|
return
|
|
}
|
|
w.WriteHeader(http.StatusNoContent)
|
|
}
|