sqlite-adapters.md

Book Agent 0.1.0 · 本版随附原文,按章节提供导览;完整原文可在文末展开。文内本机路径属于示例,请替换为你的实际路径。

本版本其他文档与许可
# Existing SQLite adapters

`PrebuiltVectorStoreAdapter` is the single database interface used by Core. It exposes `inspect`, `validate`, `get_record`, `iter_records`, `get_source_metadata`, `search` and `close`. Records contain their existing ID, full text and source metadata; returned dense hits expose the original metric score.

Connections use URI `mode=ro`, connection-local `query_only=ON`, and `trusted_schema=OFF`. Extension loading is never enabled. `immutable=1` is not used. Existing WAL, SHM or rollback-journal sidecars cause `sqlite_snapshot_incomplete` before opening. The adapter will not checkpoint, repair, add indexes or change journaling. Obtain a complete clean snapshot upstream when this diagnostic occurs.

The supported pipeline contract is `bge-m3-lossless-v1`, based on the real example inspection: `metadata(key,value)` plus `chunks(ordinal,char_start,char_end,text,tokens,text_sha256,embedding)`, with `ordinal` as primary key. It requires exact BGE-M3 1024D little-endian float32 normalized metadata and uses the reviewed script's query normalization and dot score semantics. Metadata conflicts, checksum failures, offset gaps, invalid vectors or full-text reconstruction mismatches fail validation.

A second documented adapter supports `book-agent-explicit-v1`. This is an opt-in interchange format for new integrations; existing packages are never rewritten into it. Example vector manifest:

```json
{
  "version": "book-agent-explicit-v1",
  "model": "original-index-model",
  "dimensions": 2,
  "embedding_format": "json",
  "normalization": "l2",
  "metric": "dot",
  "sqlite_schema": {
    "table": "passages",
    "id": "identifier",
    "text": "body",
    "vector": "values_json",
    "source_json": "source",
    "source_columns": {}
  }
}
```

Column identifiers are quoted and checked against actual schema; record ID must be a declared primary key. `source_columns` maps normalized metadata fields to actual column names; `source_json` identifies a JSON-object column. No table or BLOB format is guessed. Unknown versions return `unknown_sqlite_schema` with observed table names and an adapter requirement. Virtual record tables are refused with an extension diagnostic.

Supported vector encodings: JSON numeric arrays; declared little- or big-endian float32 and float64 BLOBs (`little-endian float32`, `float32-le` and corresponding float64/big-endian spellings). The decoder checks dimensions, byte lengths, finite values and nonzero magnitude. Compressed, quantized and extension-specific formats require their own documented decoder; they are explicitly unsupported. Pickle and executable deserializers are never used.

Supported metrics are dot, cosine and Euclidean L2. Dot/cosine scores sort higher first; L2 distances sort lower first. Core must preserve score meaning and direction when fusing candidates. Sequential scans reuse existing vectors and keep a bounded candidate heap; the configurable default memory budget is 64 MiB, with explicit errors if records or retained evidence exceed it. No replacement database or new document vectors are created.

Source metadata absence is a warning, not an invented page number. The example provides exact Markdown offsets and hashes but no PDF page mapping. SourceResolver may separately produce a confidence-labelled match to an original page. Lexical indexes, alignment caches and runtime metadata belong in the runtime directory.

Tests cover the real schema and audited script score parity, a different table/column schema, JSON and float64/BLOB decoding, declared big-endian float32, NaN/Inf/zero/wrong dimensions, unknown encodings, missing source metadata, unsupported virtual tables, WAL dependencies and immutable input hashes/directories.

原文 SHA-256:b1edc7f9d2ee6efb926374c8af39bb208504c03509794ceb2546b95a01b65c97