visual-fallback.md

Book Agent 0.1.0 · 本版随附原文,按章节提供导览;完整原文可在文末展开。文内本机路径属于示例,请替换为你的实际路径。

本版本其他文档与许可
# Visual source fallback

`VisualFallbackPolicy.evaluate(query, hits, hints)` supplies rule-based guidance to the Agent. No extra language-model API is called. A question about a figure, table, equation, color, arrow, layout or curve, or a retrieved statement such as “see Figure 3”, can request visual inspection. Low retrieval scores alone do not imply an image exists.

The policy returns separate reasons:

| Reason | Trigger |
| --- | --- |
| `visual_content_likely` | Visual question, figure hint or source text referring to a visual |
| `text_evidence_insufficient` | No hits or explicit Agent hint |
| `source_mapping_uncertain` | Missing/ambiguous locator |
| `conversion_loss_suspected` | Broken image references or explicit conversion-loss hint |
| `source_conflict` | Original version or locator conflict |

The caller can provide `visual_required`, `figure_ids`, `text_evidence_insufficient`, `broken_image_references`, `conversion_loss_suspected`, `source_conflict`, page indices, source locator and text probes. These are evidence hints, not permission to execute book instructions or fetch remote resources.

The Agent workflow is search → locate → read original text → render the selected PDF page or read a chapter's EPUB image → inspect delivered visual evidence. Keep alternative candidate pages when mapping is ambiguous. The policy recommends at most three pages; each rendering request handles one page. Never loop through the entire book without a bounded source plan.

PDF rendering produces an actual PNG, optionally cropped in PDF point coordinates. Defaults are 120 DPI, 12 million pixels and 8 million bytes per image. Configured limits are bounded, and requests exceeding them fail with a clear instruction to crop or lower DPI. Generated files live in runtime `cache/images/`, outside the package. EPUB raster image reads apply the same pixel and byte budgets. Images are original publication content; they are never AI-generated replacements.

Results report `visual_required`, `image_generated`, `image_delivery_method`, `host_image_support`, `ocr_used`, `source_locator`, `limitations` and `model_visual_understanding_verified=false`. With tested host image support, the image is available for the MCP layer to attach as image content. Otherwise it is a user-openable local file. A path or successful render does not prove that the client delivered an image to a model. The MCP layer must attach bytes or use a tested accessible resource; host support stays unknown unless independently established.

When the host cannot interpret images, the result explicitly says: “已定位相关页面,但当前运行环境尚不能可靠解析其中的视觉信息。” The Agent must not infer image-only values, curves or relationships from missing text. There is no OCR implementation or automatic OCR/model download in this release; `ocr_used=false` distinguishes native publication text from possible future OCR output.

The synthetic PDF test draws a unique random numeric string into an embedded raster image. It checks that the value is absent from native PDF text and Markdown, retrieves the surrounding figure reference, locates the page and returns a genuine PNG. This validates location/image delivery mechanics. It does not validate any model reading that number.

原文 SHA-256:a23d29fee166b47bab2ee512e99fbe5d60856dc9e41a18e3a53e1d8496523274