source-resolution.md

Book Agent 0.1.0 · 本版随附原文,按章节提供导览;完整原文可在文末展开。文内本机路径属于示例,请替换为你的实际路径。

本版本其他文档与许可
# Original source resolution

`SourceResolver` reads the package's original PDF or EPUB. It never rewrites the Markdown, database, source fields, or publication. The cache and supplemental mappings go to the independent runtime directory. Cached names include the original file SHA-256; a changed file size or modification time triggers a fresh fingerprint and invalidates parsed EPUB content.

Windows input and runtime paths use the package's shared extended-path helper before containment comparisons. This preserves cache separation even when one caller uses a regular drive path and another uses `\\?\` syntax. Synthetic tests cover mixed spelling and PDF read/render beyond 260 characters.

`locate(hit, hints)` first checks an existing source locator against its declared source hash, publication type, page bounds or EPUB manifest/anchor. A matching declared source hash verifies that locator's file provenance; without one, exact native text can verify the alignment. Bounds checks alone leave `locator_verified=false`. When a text-bearing page disagrees with an unverified locator, a bounded text search follows. A source version conflict remains explicit and suppresses verification.

Text may come from the hit itself, a bounded Markdown character range under `source.char_start/char_end`, or a Markdown line range. The resolver then tries native text, a partial exact snippet, heading/neighbor probes, and conservative text similarity. Partial, heading and fuzzy matches are candidates with `locator_verified=false`. A similarity value describes text matching, never confidence in an answer. Multiple matches retain their candidates and clear verification; `source_locator` stays null until a unique range is available.

PDF output separates three concepts:

| Field | Meaning |
| --- | --- |
| `pdf_page_index` | Zero based index in the file |
| `pdf_page_number` | One based file page number |
| `printed_page_label` | PDF page label, or null when unavailable |

A PDF label is publication metadata. It does not independently prove the printed number drawn on the page. If no label is present, cite “PDF file page N”. `read_source` reports `text_origin=pdf_native_text`; an image-only page has no native text and an explicit limitation. There is no automatic OCR.

PDF search defaults to at most 40 pages and accepts `hints.pdf_page_indices` for a specific bounded range. Later pages require an existing locator or supplied page hint. `hints.text_probes` and the hit's `heading_path` can narrow alignment. This mechanism does not silently realign an entire book. Supplemental results are written as individual atomic JSON files under runtime `source_map/`.

EPUB inspection follows `META-INF/container.xml`, OPF manifest and spine. Outputs preserve spine order, member XHTML href, fragment, chapter title, image src/href, alt, caption and nearby text. EPUB citations use chapter and href/anchor, with no invented stable page number. OPF-relative and complete member hrefs are accepted. External resources, archive escapes, absolute members, links and duplicate members are blocked; nothing is extracted or downloaded. XML declarations with a DOCTYPE or ENTITY are rejected. Malformed/non-XML HTML and complex EPUB layouts may require a future format adapter.

`read_epub_image(locator, image_href)` only reads an image declared in the EPUB image manifest and referenced by the requested chapter. Pillow decodes raster formats into a bounded PNG. SVG and unsupported raster formats return a clear error. Metadata may expose a blocked image reference with `available=false`; it never fetches it.

The source tests use synthetic publications. They exercise Chinese/space paths, original-file preservation, page labels, image-only pages, ambiguous text, stale locator hashes, EPUB anchors/images and path escapes. They verify location and image generation, not visual understanding.

原文 SHA-256:74b3e1e61eb553eaddfe104cc6f8265def3f54e836f2e24b2da07a943aceb5b6