* docs: define ebook architecture matching audiobooks * docs: plan ebook audiobook-parity implementation * feat: add ebook scanner parser foundation * fix: harden ebook scanner foundation * fix: handle ebook isbn labels * fix: guard ebook subtree scans * feat: scan ebook libraries in core * fix: preserve ebook scan people credits * fix: refresh ebook scan metadata safely * feat: persist ebook series membership * test: cover ebook series persistence decisions * fix: address ebook scanner PR review * docs: clarify ebook foundation PR scope * feat: add ebook metadata enricher * fix: harden ebook poster cache * feat: wire ebook metadata sync task * feat: expose ebook library metadata setup * feat: add ebook catalog scope support * feat: add ebook detail view * feat: label ebook file versions by format * feat: use file-size copy for downloads * feat: use file language in download dialog * test: cover ebook detail authors and downloads * fix: drop narrator credits from ebook scanner merges * fix: align ebook collection filters with book media * fix: drop asin provider ids from ebook enrichment * fix: force ebook people refresh for stale narrators * chore: omit ebook planning docs from branch * feat: add ebook detail related content * feat: add ebook reader file entrypoint * feat: render ebooks with foliate reader * feat: persist ebook reader progress * feat: add ebook reader controls * feat: extract ebook pdf metadata * feat: favor scanner isbn during ebook enrichment * feat: extract fbz ebook metadata * feat: count cbz ebook pages * feat: show ebook file page counts * feat: show ebook download summaries * feat: switch ebook reader files * feat: prefer epub for ebook read action * feat: surface ebook reader progress * feat: sync ebook reader progress cache * feat: hide ebook read action for unsupported files * feat: filter ebook reader file selector * fix: serve fbz ebook archives with reader mime type * fix: detect fbz ebooks from compound filename * fix: authorize fbz ebooks from compound filename * fix: scope ebook catalog facets * fix: reject narrator queries for ebooks * fix: build ebook recommendation text from authors * fix: include ebooks in embedding eligibility * fix: include ebooks in recommendation media mix * fix: include ebooks in recently added recommendations * feat: include ebook progress in recommendation signals * feat: include ebooks in continue watching sections * feat: include ebooks in catalog progress metrics * fix: read ebook isbn from epub metadata * fix: filter ebook asin provider aliases * fix: fall back from unsupported ebook reader files * fix: sort ebook catalogs by reader progress * fix: filter ebook catalogs by reader progress * fix: include ebooks in last watched catalog filters * feat: reflect ebook reader progress in item user state * feat: share ebook progress state across item surfaces * feat: report ebook scan progress * fix: include ebook activity in recommendations * fix: expose ebook reader progress on item detail * fix: support ebook subtree scans * fix: honor profile header for ebook item progress * fix: add ebook library default sections * fix: route ebook continue cards to reader * fix: hide watched toggle for ebooks * fix: route ebook watch tonight cards to reader * fix: route ebook hero actions to reader * fix: detect archive ebook reader formats by filename * feat: cache embedded ebook covers during scan * fix: encode ebook hero reader links * fix: persist non-epub ebook reader progress * fix: scope narrator catalog badges to audiobooks * fix: merge ebook reader progress during item repair * fix: label ebook progress filters as read * fix: show ebook related rails as book covers * fix: remove txt ebook reader support * fix: reject txt ebook reader files * fix: label ebook advanced filters as read * fix: label ebook personalized sorts as read * fix: remove plain text reader loader path * test: cover ebook unread catalog rules * fix: preserve ebook reader library context * fix: link ebook genres with library scope * fix: encode related rail item links * fix: encode catalog card item links * fix: encode hero and continue item links * fix: encode watch tonight item links * fix: encode recommendation and search item links * test: cover ebook scan format set * fix: label ebook search results clearly * fix: make global search prompt media neutral * fix: encode catalog read API ids * fix: encode item API ids * fix: include ebook reader vendor in docker build * fix: make ebook reader build clean * fix: clean ebook embedded descriptions * docs: plan ebook reader shell parity * feat: add ebook reader shell controls * fix: widen ebook scrolled reader flow * fix: remove scrolled reader content width cap * docs: plan ebook reader full parity * feat: persist ebook reader config * feat: add ebook annotations and bookmarks * feat: add ebook reader tools and aids * feat: add ebook advanced reader settings * fix: keep ebook reader panel in viewport * fix: use foliate sizing units for ebook scroll flow * fix: keep ebook settings controls readable * fix: simplify ebook reader settings controls * feat(ebooks): extract local covers during scan (#98) * feat(ebooks): extract local covers during scan * fix(ebooks): read nullable poster paths during cover scan * fix(catalog): coalesce nullable media artwork fields * fix(ebooks): group sibling formats by book identity * fix(ebooks): tolerate legacy ebook metadata encodings * fix(ebooks): decode PDF hex metadata strings * fix(ebooks): harden local cover extraction and format grouping Address review findings on the local cover scan: - Restrict generic sidecar covers (cover.jpg, folder.png, ...) to single-book directories, always accept images named after the book file, and apply exactly one cover per reconcile with sidecar taking precedence over the embedded cover. - Replace the read-then-write poster update with an atomic conditional UPDATE (ItemRepository.SetLocalPoster) so provider/admin artwork is never clobbered by concurrent writers, and refresh locally owned posters when the extracted cover bytes change (thumbhash compare). - Preserve UTF-8 PDF Info strings (including a UTF-8 BOM) instead of forcing everything through Windows-1252; the cp1252 fallback now only applies to non-UTF-8 bytes. - Select EPUB covers by manifest media-type with properties="cover-image" outranking the EPUB2 meta name="cover" id, so XHTML cover pages no longer shadow the real image. - Order CBZ pages naturally (2.jpg before 10.jpg, ch2/ before ch10/) when picking the cover page, via a single O(n) min-scan. - Bump the ebook content group key scheme to version 2 and reprocess rows written under older versions so pre-existing libraries gain sibling-format grouping instead of accumulating duplicates. - Group different formats only (a same-format sibling with colliding sparse metadata stays a separate item) and stop a joining sibling's embedded metadata from overwriting a provider-matched item. - Decode any IANA-labelled OPF/FB2 XML charset (windows-1251, koi8-r, shift_jis, ...) via x/net/html/charset, and wire the charset reader into FB2 parsing which previously had none. - Strip the full .fb2.zip double extension from filename-derived titles and group keys. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * feat(ebooks): add reader profiles and ruler (#99) * feat(ebooks): extract local covers during scan * fix(ebooks): read nullable poster paths during cover scan * fix(catalog): coalesce nullable media artwork fields * fix(ebooks): group sibling formats by book identity * fix(ebooks): tolerate legacy ebook metadata encodings * fix(ebooks): decode PDF hex metadata strings * feat(ebooks): add reader profiles and ruler * fix(ebooks): address reader ruler and profile review findings - skip renderer setStyles/render when computed styles and attributes are unchanged, so ruler position updates no longer re-style the book view - drag the ruler via a local draft that commits on release, with the surface rect cached at pointer-down - migrate font values persisted before the generic stacks (Inter, Georgia, Merriweather, legacy serif) so the font select never renders blank, with a Custom fallback option for unknown values - make the ruler band click-through and move dragging to a dedicated keyboard-accessible slider handle so links and text selection keep working under the band - share font stacks between options and profiles via READER_FONT_STACKS - surface the active reading profile, move presets to the top of the settings panel, and drop the redundant profile button aria-labels Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ebooks): resolve prefer-const lint error in readest document lib `pnpm run lint` failed on the branch because `direction` is never reassigned in getDirection; split the destructure so only the reassigned `writingMode` stays mutable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Quick <31828688+Quick104@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * Merge branch 'main' into work/ebooks-reader-base Brings the ebook integration branch up to date with main (audiobook library redesign, continue-watching rework and card affordances, quic-go bump, jellycompat fixes). Conflict resolutions favor main's generalized mechanisms and register ebooks with them: - media scope validation goes through IsValidMediaScope (now including "ebook" alongside main's "video" group scope), in Go and in the web filter/search types - continue-watching uses main's typed rails; reading-type sections pull resume points from ebook_reader_progress and the ebook library default section is wired to ContinueTypeConfig(ContinueTypeReading) - item_repo keeps main's derived select-list machinery (itemColumnExpr) and both poster accessors (GetPoster/SetLocalPoster for ebook covers, GetPosterPath for audiobook covers) - web cards/hero/watch-tonight adopt main's buildMediaPlayHref helpers, which now route ebooks to /reader/ebook and encode content ids; ebook affordances (BookOpen icon, Read verb, percent-read subtitle) carry over onto main's reworked components - LibraryForm ebook support ported into main's refactored useLibraryForm/libraryTypes modules Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(docker): copy foliate-js vendor into Dockerfile.dev frontend stage foliate-js is a file:vendor/foliate-js dependency, so pnpm install needs the vendor directory before the lockfile install layer. The production Dockerfile already copies it; the dev image was missed, breaking make dev-deploy with ENOENT on /app/web/vendor/foliate-js. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(ebooks): render Continue Reading sections as upright poster cards All-ebook continue sections previously fell through to the horizontal 16:9 wide card; include ebooks in the poster-variant check so book covers render in their natural 2:3 framing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ui): stop related-rail highlight ring clipping on detail pages Move the current-item ring onto the cover artwork with a themed ring-offset color (matching the sidebar profile highlight) and give the scroll container top headroom so the ring is not cut off by overflow-x-auto. Applies to both ebook and audiobook detail rails. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scanner): harden ebook scanning against data loss and bad metadata - Reconcile missing ebook files like video/audio, with real per-root walk failure tracking (failed/unmounted roots are excluded from deletion), symlinked-root support via the shared logical walker, and the empty-root cleanup allowance before any destructive reconciliation. - Create ebook items as 'pending' so enrichment can promote them to 'matched' (backfill migration included), and protect matched items from re-scan clobbering: title/year skipped, people/series fill-empty only. - PDF metadata: scan head + tail windows (non-linearized PDFs keep the Info dict at the end), require proper key delimiters, head values win. - Cap plain .fb2 reads like .fbz entries; drop .md as an ebook format. - gofmt internal/scanner/audiobook.go (pre-existing drift). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ebooks): make enrichment failures non-terminal with dedicated backoff state - Provider errors now record a failure (capped retries) instead of stamping last_refreshed, which permanently excluded items after transient outages. - Unconfigured metadata chains and the scan-window membership race skip the item without stamping or burning a retry. - Failure tracking moves to a new ebook_enrichment_state table, decoupling it from media_items.refresh_failures (shared with metadata refresh debt). - Preserve non-author people credits when persisting enrichment results. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(catalog): gate ebook progress on hidden history and centralize threshold - Apply user_history_hidden_items gating (video semantics) to the ebook watched/in-progress filters, progress sort plan, and Continue Reading. - Continue Reading pages past dismissed items via the shared collector and dedupes items across pages (also fixes the video path's latent exposure). - Centralize the 0.9 finished threshold as models.EbookFinishedProgressThreshold with a single SQL-interpolated mirror in catalog. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(recommendations): correct watcher counting and wire ebook taste signals - itemWatchersQuery dedupes to distinct (watcher, item) rows so one binge-watcher can no longer satisfy minWatchers; the eligibility floor now counts distinct accounts rather than profiles. - Hidden-history gating on GetEbookReaderProgressForUser (signal reader). - Ebook reading produces canonical implicit taste signals (weighted like the equivalent movie progress ratio); ebooks join taste-seed candidates. - Stale GetRecentlyAddedItems doc comment corrected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(api): harden ebook reader endpoints and serve a Content-Security-Policy - Serve a CSP on all SPA HTML responses: blob/srcdoc book iframes inherit it, so script-src 'self' 'wasm-unsafe-eval' blocks script execution from malicious book content (sandbox alone is defeated by the WebKit allow-scripts requirement). Threat model documented on the constant. - X-Content-Type-Options: nosniff on frontend, jellycompat, and ebook file responses; MIME resolution can no longer fall through to octet-stream for an admitted ebook file. - Annotation PATCH: presence-aware field semantics (absent keeps, present sets/clears), invariant re-validation on the merged row, and an atomic SELECT ... FOR UPDATE read-merge-write. - Request size caps (413) on progress/config/annotation writes; Content-Disposition via mime.FormatMediaType; hidden-history gating in the shared ebook progress lister; FK-cascade indexes for reader tables. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(api): native read-state endpoints for ebooks - POST/DELETE /watched/{id} accepts ebook content IDs: mark read upserts progress 1.0 preserving the reader's file/location (or picks the preferred reader file for never-opened books); mark unread mirrors video unwatch semantics and deletes the progress row. - /history/remove accepts ebooks: hides via user_history_hidden_items without touching the reading position (hidden != unread; next reading activity resurfaces the book, mirroring video re-watch). - Access-filter checks match the video branch; shared logic lives in ebook_read_state.go. Sort metrics/user-state thresholds use the shared constant; profile-header fallback deduplicated. Clients: response is {type: "ebook", affected_count: 1, played: bool}; the existing watched SSE event fires. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(web): harden the ebook reader UI - Open-flow race: cancellation checked after every await with full stale-run teardown (no wrong-file progress saves, no leaked views/blob URLs); book.destroy() on cleanup. - Progress: monotonic stale-response guard; visibilitychange flush uses the refresh-capable client, pagehide uses keepalive; per-book cross-format progress documented as deliberate. - Settings: side effects out of the setState updater; local edits no longer clobbered by late server config; pending saves flushed on unmount/pagehide. - TTS: generation token so Stop actually stops (Chromium/Firefox synthetic events); Media Session uninstalled on unmount. - External book links: http(s) only, opened with noopener,noreferrer. - apiBlob 512 MiB guard with a user-facing error; fraction bookmarks navigable; search-result key collisions fixed; dead e-ink code removed; getLibrarySortRelevanceScope deduplicated; md format dropped. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(web): mark read/unread affordances for ebooks - Item detail gets a Mark Read/Unread button; card menus drop the ebook gate and share type-aware labels/toasts (also dedupes audiobook wording). - Watched-state invalidation includes the reader progress query key so the Continue button and percent refresh after toggling. - Continue Reading dismiss copy for ebooks; dismissal path now URL-encodes item IDs (ebook content IDs can contain reserved characters). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: record the PR #124 review and hardening pass Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: rxwatcher <rxwatcher@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
349 lines
12 KiB
JavaScript
349 lines
12 KiB
JavaScript
const normalizeWhitespace = str => str ? str
|
|
.replace(/[\t\n\f\r ]+/g, ' ')
|
|
.replace(/^[\t\n\f\r ]+/, '')
|
|
.replace(/[\t\n\f\r ]+$/, '') : ''
|
|
const getElementText = el => normalizeWhitespace(el?.textContent)
|
|
|
|
const NS = {
|
|
XLINK: 'http://www.w3.org/1999/xlink',
|
|
EPUB: 'http://www.idpf.org/2007/ops',
|
|
}
|
|
|
|
const MIME = {
|
|
XML: 'application/xml',
|
|
XHTML: 'application/xhtml+xml',
|
|
}
|
|
|
|
const STYLE = {
|
|
'strong': ['strong', 'self'],
|
|
'emphasis': ['em', 'self'],
|
|
'style': ['span', 'self'],
|
|
'a': 'anchor',
|
|
'strikethrough': ['s', 'self'],
|
|
'sub': ['sub', 'self'],
|
|
'sup': ['sup', 'self'],
|
|
'code': ['code', 'self'],
|
|
'image': 'image',
|
|
}
|
|
|
|
const TABLE = {
|
|
'tr': ['tr', {
|
|
'th': ['th', STYLE, ['colspan', 'rowspan', 'align', 'valign']],
|
|
'td': ['td', STYLE, ['colspan', 'rowspan', 'align', 'valign']],
|
|
}, ['align']],
|
|
}
|
|
|
|
const POEM = {
|
|
'epigraph': ['blockquote'],
|
|
'subtitle': ['h2', STYLE],
|
|
'text-author': ['p', STYLE],
|
|
'date': ['p', STYLE],
|
|
'stanza': ['div', 'self'],
|
|
'v': ['div', STYLE],
|
|
}
|
|
|
|
const SECTION = {
|
|
'title': ['header', {
|
|
'p': ['h1', STYLE],
|
|
'empty-line': ['br'],
|
|
}],
|
|
'epigraph': ['blockquote', 'self'],
|
|
'image': 'image',
|
|
'annotation': ['aside'],
|
|
'section': ['section', 'self'],
|
|
'p': ['p', STYLE],
|
|
'poem': ['blockquote', POEM],
|
|
'subtitle': ['h2', STYLE],
|
|
'cite': ['blockquote', 'self'],
|
|
'empty-line': ['br'],
|
|
'table': ['table', TABLE],
|
|
'text-author': ['p', STYLE],
|
|
}
|
|
POEM['epigraph'].push(SECTION)
|
|
|
|
const BODY = {
|
|
'image': 'image',
|
|
'title': ['section', {
|
|
'p': ['h1', STYLE],
|
|
'empty-line': ['br'],
|
|
}],
|
|
'epigraph': ['section', SECTION],
|
|
'section': ['section', SECTION],
|
|
}
|
|
|
|
class FB2Converter {
|
|
constructor(fb2) {
|
|
this.fb2 = fb2
|
|
this.doc = document.implementation.createDocument(NS.XHTML, 'html')
|
|
// use this instead of `getElementById` to allow images like
|
|
// `<image l:href="#img1.jpg" id="img1.jpg" />`
|
|
this.bins = new Map(Array.from(this.fb2.getElementsByTagName('binary'),
|
|
el => [el.id, el]))
|
|
}
|
|
getImageSrc(el) {
|
|
const href = el.getAttributeNS(NS.XLINK, 'href')
|
|
if (!href) return 'data:,'
|
|
const [, id] = href.split('#')
|
|
if (!id) return href
|
|
const bin = this.bins.get(id)
|
|
return bin
|
|
? `data:${bin.getAttribute('content-type')};base64,${bin.textContent}`
|
|
: href
|
|
}
|
|
image(node) {
|
|
const el = this.doc.createElement('img')
|
|
el.alt = node.getAttribute('alt')
|
|
el.title = node.getAttribute('title')
|
|
el.setAttribute('src', this.getImageSrc(node))
|
|
return el
|
|
}
|
|
anchor(node) {
|
|
const el = this.convert(node, { 'a': ['a', STYLE] })
|
|
el.setAttribute('href', node.getAttributeNS(NS.XLINK, 'href'))
|
|
if (node.getAttribute('type') === 'note')
|
|
el.setAttributeNS(NS.EPUB, 'epub:type', 'noteref')
|
|
return el
|
|
}
|
|
convert(node, def) {
|
|
// not an element; return text content
|
|
if (node.nodeType === 3) return this.doc.createTextNode(node.textContent)
|
|
if (node.nodeType === 4) return this.doc.createCDATASection(node.textContent)
|
|
if (node.nodeType === 8) return this.doc.createComment(node.textContent)
|
|
|
|
const d = def?.[node.nodeName]
|
|
if (!d) return null
|
|
if (typeof d === 'string') return this[d](node)
|
|
|
|
const [name, opts, attrs] = d
|
|
const el = this.doc.createElement(name)
|
|
|
|
// copy the ID, and set class name from original element name
|
|
if (node.id) el.id = node.id
|
|
el.classList.add(node.nodeName)
|
|
|
|
// copy attributes
|
|
if (Array.isArray(attrs)) for (const attr of attrs) {
|
|
const value = node.getAttribute(attr)
|
|
if (value) el.setAttribute(attr, value)
|
|
}
|
|
|
|
// process child elements recursively
|
|
const childDef = opts === 'self' ? def : opts
|
|
let child = node.firstChild
|
|
while (child) {
|
|
const childEl = this.convert(child, childDef)
|
|
if (childEl) el.append(childEl)
|
|
child = child.nextSibling
|
|
}
|
|
return el
|
|
}
|
|
}
|
|
|
|
const parseXML = async blob => {
|
|
const buffer = await blob.arrayBuffer()
|
|
const str = new TextDecoder('utf-8').decode(buffer)
|
|
const parser = new DOMParser()
|
|
const doc = parser.parseFromString(str, MIME.XML)
|
|
const encoding = doc.xmlEncoding
|
|
// `Document.xmlEncoding` is deprecated, and already removed in Firefox
|
|
// so parse the XML declaration manually
|
|
|| str.match(/^<\?xml\s+version\s*=\s*["']1.\d+"\s+encoding\s*=\s*["']([A-Za-z0-9._-]*)["']/)?.[1]
|
|
if (encoding && encoding.toLowerCase() !== 'utf-8') {
|
|
const str = new TextDecoder(encoding).decode(buffer)
|
|
return parser.parseFromString(str, MIME.XML)
|
|
}
|
|
return doc
|
|
}
|
|
|
|
const style = URL.createObjectURL(new Blob([`
|
|
@namespace epub "http://www.idpf.org/2007/ops";
|
|
body > img, section > img {
|
|
display: block;
|
|
margin: auto;
|
|
}
|
|
.title h1 {
|
|
text-align: center;
|
|
}
|
|
body > section > .title, body.notesBodyType > .title {
|
|
margin: 3em 0;
|
|
}
|
|
body.notesBodyType > section .title h1 {
|
|
text-align: start;
|
|
}
|
|
body.notesBodyType > section .title {
|
|
margin: 1em 0;
|
|
}
|
|
p {
|
|
text-indent: 1em;
|
|
margin: 0;
|
|
}
|
|
:not(p) + p, p:first-child {
|
|
text-indent: 0;
|
|
}
|
|
.stanza {
|
|
text-indent: 0;
|
|
margin: 1em 0;
|
|
}
|
|
.text-author, .date {
|
|
text-align: end;
|
|
}
|
|
.text-author:before {
|
|
content: "—";
|
|
}
|
|
table {
|
|
border-collapse: collapse;
|
|
}
|
|
td, th {
|
|
padding: .25em;
|
|
}
|
|
a[epub|type~="noteref"] {
|
|
font-size: .75em;
|
|
vertical-align: super;
|
|
}
|
|
body:not(.notesBodyType) > .title, body:not(.notesBodyType) > .epigraph {
|
|
margin: 3em 0;
|
|
}
|
|
`], { type: 'text/css' }))
|
|
|
|
const template = html => `<?xml version="1.0" encoding="utf-8"?>
|
|
<html xmlns="http://www.w3.org/1999/xhtml">
|
|
<head><link href="${style}" rel="stylesheet" type="text/css"/></head>
|
|
<body>${html}</body>
|
|
</html>`
|
|
|
|
// name of custom ID attribute for TOC items
|
|
const dataID = 'data-foliate-id'
|
|
|
|
export const makeFB2 = async blob => {
|
|
const book = {}
|
|
const doc = await parseXML(blob)
|
|
const converter = new FB2Converter(doc)
|
|
|
|
const $ = x => doc.querySelector(x)
|
|
const $$ = x => [...doc.querySelectorAll(x)]
|
|
const getPerson = el => {
|
|
const nick = getElementText(el.querySelector('nickname'))
|
|
if (nick) return nick
|
|
const first = getElementText(el.querySelector('first-name'))
|
|
const middle = getElementText(el.querySelector('middle-name'))
|
|
const last = getElementText(el.querySelector('last-name'))
|
|
const name = [first, middle, last].filter(x => x).join(' ')
|
|
const sortAs = last
|
|
? [last, [first, middle].filter(x => x).join(' ')].join(', ')
|
|
: null
|
|
return { name, sortAs }
|
|
}
|
|
const getDate = el => el?.getAttribute('value') ?? getElementText(el)
|
|
const annotation = $('title-info annotation')
|
|
book.metadata = {
|
|
title: getElementText($('title-info book-title')),
|
|
identifier: getElementText($('document-info id')),
|
|
language: getElementText($('title-info lang')),
|
|
author: $$('title-info author').map(getPerson),
|
|
translator: $$('title-info translator').map(getPerson),
|
|
contributor: $$('document-info author').map(getPerson)
|
|
// techincially the program probably shouldn't get the `bkp` role
|
|
// but it has been so used by calibre, so ¯\_(ツ)_/¯
|
|
.concat($$('document-info program-used').map(getElementText))
|
|
.map(x => Object.assign(typeof x === 'string' ? { name: x } : x,
|
|
{ role: 'bkp' })),
|
|
publisher: getElementText($('publish-info publisher')),
|
|
published: getDate($('title-info date')),
|
|
modified: getDate($('document-info date')),
|
|
description: annotation ? converter.convert(annotation,
|
|
{ annotation: ['div', SECTION] }).innerHTML : null,
|
|
subject: $$('title-info genre').map(getElementText),
|
|
}
|
|
if ($('coverpage image')) {
|
|
const src = converter.getImageSrc($('coverpage image'))
|
|
book.getCover = () => fetch(src).then(res => res.blob())
|
|
} else book.getCover = () => null
|
|
|
|
// get convert each body
|
|
const bodyData = Array.from(doc.querySelectorAll('body'), body => {
|
|
const converted = converter.convert(body, { body: ['body', BODY] })
|
|
return [Array.from(converted.children, el => {
|
|
// get list of IDs in the section
|
|
const ids = [el, ...el.querySelectorAll('[id]')].map(el => el.id)
|
|
return { el, ids }
|
|
}), converted]
|
|
})
|
|
|
|
const urls = []
|
|
const sectionData = bodyData[0][0]
|
|
// make a separate section for each section in the first body
|
|
.map(({ el, ids }, id) => {
|
|
// set up titles for TOC
|
|
const titles = Array.from(
|
|
el.querySelectorAll(':scope > section > .title'),
|
|
(el, index) => {
|
|
el.setAttribute(dataID, index)
|
|
const section = el.closest('section')
|
|
const size = new TextEncoder().encode(section.innerHTML).length
|
|
- Array.from(section.querySelectorAll('[src]'))
|
|
.reduce((sum, el) => sum + (el.getAttribute('src')?.length ?? 0), 0)
|
|
return { title: getElementText(el), index, size, href: `${id}#${index}` }
|
|
})
|
|
return { ids, titles, el }
|
|
})
|
|
// for additional bodies, only make one section for each body
|
|
.concat(bodyData.slice(1).map(([sections, body]) => {
|
|
const ids = sections.map(s => s.ids).flat()
|
|
body.classList.add('notesBodyType')
|
|
return { ids, el: body, linear: 'no' }
|
|
}))
|
|
.map(({ ids, titles, el, linear }) => {
|
|
const str = template(el.outerHTML)
|
|
const blob = new Blob([str], { type: MIME.XHTML })
|
|
const url = URL.createObjectURL(blob)
|
|
urls.push(url)
|
|
const title = normalizeWhitespace(
|
|
el.querySelector('.title, .subtitle, p')?.textContent
|
|
?? (el.classList.contains('title') ? el.textContent : ''))
|
|
return {
|
|
ids, title, titles, load: () => url,
|
|
createDocument: () => new DOMParser().parseFromString(str, MIME.XHTML),
|
|
// doo't count image data as it'd skew the size too much
|
|
size: blob.size - Array.from(el.querySelectorAll('[src]'),
|
|
el => el.getAttribute('src')?.length ?? 0)
|
|
.reduce((a, b) => a + b, 0),
|
|
linear,
|
|
}
|
|
})
|
|
|
|
const idMap = new Map()
|
|
book.sections = sectionData.map((section, index) => {
|
|
const { ids, load, createDocument, size, linear, titles } = section
|
|
for (const id of ids) if (id) idMap.set(id, index)
|
|
return { id: index, load, createDocument, size, linear, subitems: titles }
|
|
})
|
|
|
|
book.toc = sectionData.map(({ title, titles }, index) => {
|
|
const id = index.toString()
|
|
return {
|
|
label: title,
|
|
href: id,
|
|
subitems: titles?.length ? titles.map(({ title, index }) => ({
|
|
label: title,
|
|
href: `${id}#${index}`,
|
|
})) : null,
|
|
}
|
|
}).filter(item => item)
|
|
|
|
book.resolveHref = href => {
|
|
const [a, b] = href.split('#')
|
|
return a
|
|
// the link is from the TOC
|
|
? { index: Number(a), anchor: doc => doc.querySelector(`[${dataID}="${b}"]`) }
|
|
// link from within the page
|
|
: { index: idMap.get(b), anchor: doc => doc.getElementById(b) }
|
|
}
|
|
book.splitTOCHref = href => href?.split('#')?.map(x => Number(x)) ?? []
|
|
book.getTOCFragment = (doc, id) => doc.querySelector(`[${dataID}="${id}"]`)
|
|
|
|
book.destroy = () => {
|
|
for (const url of urls) URL.revokeObjectURL(url)
|
|
}
|
|
return book
|
|
}
|