query-runtime.md

Book Agent 0.1.0 · 本版随附原文,按章节提供导览;完整原文可在文末展开。文内本机路径属于示例,请替换为你的实际路径。

本版本其他文档与许可
# Query runtime inheritance

Book Agent uses the document vectors already supplied by the book. It encodes only the current question after resolving the original retrieval-space contract. There is no document embedding API, index rebuild command or model-choice wizard. Equal dimensions alone do not establish a compatible space.

Three optional execution modes

The default is `none`. Fulltext retrieval and original-source reading remain available without weights or API credentials.

| Mode | Question encoding | Requirements |
| --- | --- | --- |
| `none` | Unavailable; Core returns a lexical fallback | None |
| `offline` | Pinned original BGE-M3 on CPU | Verified local weight directory and optional `local-query` dependencies |
| `online` | Original inherited hosted endpoint | Explicit remote authorization, exact endpoint allowlist and an environment-variable credential reference |

```shell
book-agent configure BOOK_ID --vector-search none
book-agent configure BOOK_ID --vector-search offline --query-model-path ./BAAI-bge-m3
book-agent configure BOOK_ID --vector-search online --allow-remote-query --allow-query-endpoint https://api.siliconflow.cn/v1/embeddings --query-auth-env SILICONFLOW_API_KEY
```

The library additionally accepts `auto`: a configured local path is tried first, and a local failure can fall through to the hosted encoder only when remote requests are separately authorized. `offline` never contacts an encoder service. Merely setting a path or credential does not activate vectors when execution is `none`.

返回章节目录

Resolving the existing space

Resolution compares archive/vector manifests, SQLite metadata, validated Markdown reconstruction, the statically inspected search script and the audited format adapter. Disagreements fail explicitly. Incomplete query contracts remain `unverified`. Changing the packaged search script invalidates its automatic audited compatibility. Packaged scripts are never imported or executed.

For the real `bge-m3-lossless-v1` package, the audited producer uses SiliconFlow's exact `BAAI/bge-m3` hosted model at `https://api.siliconflow.cn/v1/embeddings`, raw question text without instruction or prefix, little-endian float32 document vectors and L2 normalization. The scorer inherits dot-product semantics. `status: verified` and `compatibility: compatible_pipeline_runtime` refer to that audited contract; they do not certify live credentials, publisher identity, service availability or natural-language relevance. The original provider revision is `provider-managed-unpinned`.

返回章节目录

Installing the original model for offline questions

Install the optional Python runtime dependencies and verify an existing offline directory:

```shell
python -m pip install "prebuilt-book-agent[local-query]"
book-agent install-query-model BOOK_ID --model-dir ./BAAI-bge-m3
book-agent configure BOOK_ID --vector-search offline --query-model-path ./BAAI-bge-m3
```

An explicit `--download` on `install-query-model` obtains only the fixed public upstream artifacts if the target directory does not exist. Installation without that flag never downloads. An existing directory must already contain every exact artifact; partial or unrelated content is not overwritten. Search never downloads or repairs weights. Missing weights or dependencies produce a structured reason and lexical fallback, while strict search raises that reason.

The directory contains six pinned upstream files and a small `book-agent-query-contract.json` receipt. The receipt records relative names, byte sizes, SHA-256 values, model identity, revisions and recipe. It contains no machine paths or credentials and can move with the files. Verify or register the directory again at its actual location on another machine. Share one directory among books with this same verified contract; no duplicate model copy is needed in each package, Skill or runtime. Keep it outside read-only book packages.

The standard portable executable supplies lightweight fulltext and HTTP query support. Local neural questions use the Python runtime with the `local-query` extra and this separate weight directory. Weights are not embedded in the standard executable.

返回章节目录

Fixed artifacts and recipe

Model: [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3). Tokenizer and configuration revision: `5617a9f61b028005a4858fdac845db406aefb181`. Safetensors revision: `9a0624b896d81da7492a910ffa53731274b6cf3d`.

| File | Bytes | SHA-256 |
| --- | ---: | --- |
| `config.json` | 687 | `26159e7ad065073448460117eb24b7a4572f6f4e78eadff65dc0a11c052449fa` |
| `tokenizer_config.json` | 444 | `a62b2b6784f990259fddef5f16388693a8043be4f69179e6a5257eeb3f9abac4` |
| `special_tokens_map.json` | 964 | `8c785abebea9ae3257b61681b4e6fd8365ceafde980c21970d001e834cf10835` |
| `tokenizer.json` | 17,098,108 | `21106b6d7dab2952c1d496fb21d5dc9db75c28ed361a05f5020bbba27810dd08` |
| `sentencepiece.bpe.model` | 5,069,051 | `cfc8146abe2a0488e9e2a0c56de7952f7c11ab059eca145a0a727afce0db2865` |
| `model.safetensors` | 2,271,064,456 | `993b2248881724788dcab8c644a91dfd63584b6e5604ff2037cb5541e1e38e7e` |

The local recipe is `bge-m3-dense-cls-l2-float32-v1`: one raw string, original fast XLM-Roberta tokenizer, special tokens included, `XLMRobertaModel` in CPU float32 evaluation mode, CLS from `last_hidden_state[0,0,:]`, and L2 normalization to a finite 1024-dimensional representation. This follows the [upstream BGE-M3 encoder](https://github.com/FlagOpen/FlagEmbedding/blob/master/FlagEmbedding/inference/embedder/encoder_only/m3.py). Local loading forces `local_files_only=True`, `trust_remote_code=False`, `use_safetensors=True` and eager attention. No Pickle checkpoint or remote code is loaded. Generic local model declarations are not executable adapters.

The process caches at most one local query adapter and serializes its queries. The adapter retains one official eager encoder layer at a time in CPU float32 RAM, and reads only the current question's embedding rows. It closes Safetensors mappings before computation; no full model or second weight copy is created. This follows the upstream [partial tensor loading API](https://huggingface.co/docs/safetensors/main/index). Each query rereads the encoder layers, so disk speed can limit latency. Default inference uses two CPU threads and a budget of 512 tokenizer tokens including special tokens. Runtime configuration can set `local_cpu_threads` from 1 to 4 and `local_max_tokens` from 1 to 8192. The public Core question limit remains 8,000 characters; the lower-level encoder also enforces its own 16,000-character ceiling. Oversized questions are rejected before forward, without truncation. The 2.27 GB checkpoint is needed on disk; longer token budgets increase activation memory and CPU time.

返回章节目录

Offline MCP deadlines

Cold encoding includes dependency loading and weight checks. For offline MCP, set the server's per-tool deadline to 600 seconds:

```shell
book-agent --data-dir ./runtime-data mcp BOOK_ID --timeout 600
```

Codex's host tool deadline is separate. Set `tool_timeout_sec = 600` on the corresponding MCP server entry and include `"--timeout", "600"` in its arguments. Codex supports separate `startup_timeout_sec` and `tool_timeout_sec` fields in each `[mcp_servers.<name>]` entry. [Official OpenAI MCP configuration](https://learn.chatgpt.com/docs/extend/mcp). For example, replace the executable/data paths and `BOOK_ID` with the installed offline Python runtime and its registered book:

```toml
[mcp_servers.book_agent_offline]
command = "F:/Tools/BookAgent/query-runtime/Scripts/python.exe"
args = ["-B", "-m", "book_agent", "--data-dir", "F:/BookAgentData", "mcp", "BOOK_ID", "--timeout", "600"]
startup_timeout_sec = 90
tool_timeout_sec = 600
```

When updating an entry generated by `install-host`, retain that entry's name and actual paths. Both the server and host deadline must allow the slow tool call. A CLI search in another process does not warm this MCP process; the lightweight portable EXE is not the optional offline Python runtime.

返回章节目录

Conversion and numeric compatibility limits

The original pinned model revision exposes `pytorch_model.bin`. This implementation uses the public Hugging Face [Safetensors conversion PR #130](https://huggingface.co/BAAI/bge-m3/discussions/130), whose fixed commit is a direct child of that original revision. The conversion was published by the Hugging Face conversion bot and remains unmerged. Tokenizer/configuration artifacts stay pinned to the original revision. The original unsafe binary is never downloaded or loaded by Book Agent.

Artifact identity, tokenizer identity and CLS/L2 recipe are checked. Numerical parity between this local checkpoint and the original **unpinned hosted encoder** has not been independently established. Responses expose `provider_numeric_parity_verified: false` and the local-use warning. Using the same published model and recipe provides a documented inheritance path, while it does not prove provider-specific numerical equivalence. [Upstream model metadata](https://huggingface.co/BAAI/bge-m3/blob/5617a9f61b028005a4858fdac845db406aefb181/README.md) declares MIT; the separate weight bundle includes the upstream source and license information.

返回章节目录

Authorized online questions

```json
{
  "execution": "online",
  "allow_remote_requests": true,
  "allowed_endpoints": ["https://api.siliconflow.cn/v1/embeddings"],
  "auth_env": "SILICONFLOW_API_KEY",
  "timeout_seconds": 30
}
```

Set the secret in the process environment. JSON configuration and book metadata only name the environment variable. Model and endpoint overrides cannot silently replace the inherited space. A manifest URL is not authorization: remote use defaults to false, exact endpoint allowlisting is required, redirects are refused and diagnostics omit credentials. Only the current question is sent to the query encoder. Remote reranking separately requires authorization and sends candidate excerpts. The Harness's own language model network use is separate.

An explicit paired hosted encoder can declare `query_runtime` in a `book-agent-explicit-v1` manifest:

```json
{
  "query_runtime": {
    "provider": "openai_compatible",
    "query_model": "existing-query-encoder",
    "paired_document_model": "original-index-model",
    "revision": "fixed-encoder-revision",
    "compatibility": "declared_paired_encoder",
    "dimensions": 1024,
    "normalization": "l2",
    "metric": "dot",
    "query_prefix": "query: ",
    "query_instruction": "",
    "endpoint": "https://approved.example/v1/embeddings",
    "auth_env": "EXISTING_QUERY_KEY"
  }
}
```

This adapter supports explicitly paired document/query encoders. It requires a complete revision, normalization, paired-space declaration and validated source identity. Runtime endpoint authorization is still required. A generic `sentence_transformers` declaration remains unavailable until an audited safe execution adapter exists.

返回章节目录

Validation evidence

The latest complete release regression finished with **276 passed, 2 skipped in 161.38 seconds** (`.work/tests-release-final.log` and `.xml`). It includes 14 actual tiny-model tests and 29 installer/bootstrap unit cases. Separate **actual default Skill bootstrap installed 46 wheels** offline in **73.259 seconds** (`.work/acceptance/bootstrap-real-install.json`). The **dedicated `release/install-offline.py` installed 67 packages** in an independent CPython 3.13 venv using no-index local hashed wheels, with pip check reporting no broken requirements, in **233.145 seconds** (`.work/offline-install-real-final.summary.json` and `.log`). The optional `local-query` bootstrap route had its 67-wheel plan verified but was not actually installed through bootstrap.

Both receiving-runtime probes loaded Book Agent from their own installed venv, returned three default lexical hits, verified original-source location, produced a Codex host dry-run, checked default 46/optional 67-wheel bootstrap plans, blocked network connections and preserved the original archive without generating document vectors. Default probe time was **16.747 seconds** (`.work/acceptance/base-bootstrap-receiving-runtime.json`). Optional probe time was **245.664 seconds** (`.work/acceptance/optional-receiving-runtime.json`), including cold imports of PyTorch `2.14.1+cpu`, Transformers `4.57.6`, Safetensors `0.8.0` and SentencePiece `0.2.2`, CPU tensor math and plan hash verification. That total is not a search latency measurement; no full model forward was repeated in either receiving-runtime probe. The complete-model evidence below uses the earlier verified Python runtime. Installation/probe reports captured the same code/dependency artifacts before this documentation update; final manifest identity will be regenerated with the documents.

Mock tests check public-artifact integrity receipts, relocation, CLS rather than mean pooling, raw inputs, one cached model adapter, thread limits, no silent truncation, offline loader flags, model-space rejection, tamper detection and remote authorization. They do not execute a real model. The optional actual CPU integration test runs only when `BOOK_AGENT_QUERY_MODEL_TEST_PATH` points to explicitly installed verified weights; it does not download. It checks a finite 1024-dimensional unit vector from a natural-language question. Real SQLite search with that question must be recorded separately from stored-record self-vector scorer tests.

On 2026-10-03, **14 actual tiny-model tests passed in 131.04 seconds**, with no skipped cases. PyTorch `2.14.1+cpu`, Transformers `4.57.6` and Safetensors `0.8.0` executed the model math. Four inputs (unpadded, padded, explicit positions and 3D mask) matched the official full eager XLM-R encoder's final hidden states and CLS/L2 output at `rtol=1e-5`, `atol=1e-6`. Further cases checked embedding-row slicing, mappings closed before computation, one resident CPU layer, corrupt checkpoints and unsupported input/configuration rejection. Evidence is `.work/query-streaming-resumed-threads.xml` and its accompanying log/summary. This test does not load the delivered full checkpoint.

The separate **actual BGE-M3 CPU acceptance passed** using the pinned delivered Safetensors directory and original SQLite document vectors, with socket/urllib network entry points blocked. It produced 1024 finite dimensions with L2 norm `0.9999999338`; retrieval reported `ready / encoded / offline`, used semantic search and returned record IDs `1`, `2`, `3`. Source backtrace located physical PDF page 2 (`page_index=1`). The archive's SHA/size/mtime and all 29 input files were unchanged; document embeddings generated: `0`. Evidence: `.work/acceptance/offline-real-query.json`.

In that local run, cold encoding took **33.609 seconds**; the second Core search took **0.780 seconds for the same question after model/dependency warmup**, with question encoding repeated. The process caches the tokenizer and streaming model adapter, not query vectors; each forward reads the encoder layers again. The complete **88.695-second** run included weight registration/hash checks. This second-search timing is not generalized to other questions. The report does not record process peak memory; these timings are one machine's observation, not a performance or memory guarantee. Hosted-provider numerical parity and actual local reranking remain unverified. Separate native Codex CLI/Python MCP question answering passed with lexical fallback; it did not exercise the offline neural encoder or independently validate automatic Skill triggering.

A second actual blocked-network run, `.work/acceptance/offline-query-instrumented.json`, also passed. It counted **one real forward for the same-question search and one for a fresh-question search**. After model/dependency warmup, Core search took **0.973 seconds** for the same question and **0.869 seconds** for the new question; measured forward times were **0.621 / 0.854 seconds**. The new question used semantic retrieval. Initial encoding took **155.887 seconds**, and the complete run including weight verification took **256.183 seconds**. The archive and 29 input files remained unchanged, and no document embeddings were generated. Both cold measurements include dependency loading and checks before encoding; their differing times show unstable startup in these local runs, not a universal latency range or guarantee.

Service tests use injected transport to check model inheritance, prefixes, one-query input, missing credentials, timeouts and invalid responses. They are contract tests, not live SiliconFlow validation. A live service acceptance needs a separately authorized endpoint and available credentials. Supplied-vector search remains possible offline, with representation provenance reported as `user-supplied / unverified` unless independently established.

离线模式的新 `install-host` 安装会自动为 MCP 设置 600 秒工具期限,并同步 Codex 的 `tool_timeout_sec`;其他模式保持原默认参数。已有旧离线配置若出现所有权冲突,先用 `uninstall-host BOOK_ID --host codex` 安全卸载本项目拥有且未改变的配置,再重新安装。手工接入其他宿主时,仍需在宿主中配置匹配的工具期限。

返回章节目录

查看完整原文(逐字保留)
# Query runtime inheritance

Book Agent uses the document vectors already supplied by the book. It encodes only the current question after resolving the original retrieval-space contract. There is no document embedding API, index rebuild command or model-choice wizard. Equal dimensions alone do not establish a compatible space.

## Three optional execution modes

The default is `none`. Fulltext retrieval and original-source reading remain available without weights or API credentials.

| Mode | Question encoding | Requirements |
| --- | --- | --- |
| `none` | Unavailable; Core returns a lexical fallback | None |
| `offline` | Pinned original BGE-M3 on CPU | Verified local weight directory and optional `local-query` dependencies |
| `online` | Original inherited hosted endpoint | Explicit remote authorization, exact endpoint allowlist and an environment-variable credential reference |

```shell
book-agent configure BOOK_ID --vector-search none
book-agent configure BOOK_ID --vector-search offline --query-model-path ./BAAI-bge-m3
book-agent configure BOOK_ID --vector-search online --allow-remote-query --allow-query-endpoint https://api.siliconflow.cn/v1/embeddings --query-auth-env SILICONFLOW_API_KEY
```

The library additionally accepts `auto`: a configured local path is tried first, and a local failure can fall through to the hosted encoder only when remote requests are separately authorized. `offline` never contacts an encoder service. Merely setting a path or credential does not activate vectors when execution is `none`.

## Resolving the existing space

Resolution compares archive/vector manifests, SQLite metadata, validated Markdown reconstruction, the statically inspected search script and the audited format adapter. Disagreements fail explicitly. Incomplete query contracts remain `unverified`. Changing the packaged search script invalidates its automatic audited compatibility. Packaged scripts are never imported or executed.

For the real `bge-m3-lossless-v1` package, the audited producer uses SiliconFlow's exact `BAAI/bge-m3` hosted model at `https://api.siliconflow.cn/v1/embeddings`, raw question text without instruction or prefix, little-endian float32 document vectors and L2 normalization. The scorer inherits dot-product semantics. `status: verified` and `compatibility: compatible_pipeline_runtime` refer to that audited contract; they do not certify live credentials, publisher identity, service availability or natural-language relevance. The original provider revision is `provider-managed-unpinned`.

## Installing the original model for offline questions

Install the optional Python runtime dependencies and verify an existing offline directory:

```shell
python -m pip install "prebuilt-book-agent[local-query]"
book-agent install-query-model BOOK_ID --model-dir ./BAAI-bge-m3
book-agent configure BOOK_ID --vector-search offline --query-model-path ./BAAI-bge-m3
```

An explicit `--download` on `install-query-model` obtains only the fixed public upstream artifacts if the target directory does not exist. Installation without that flag never downloads. An existing directory must already contain every exact artifact; partial or unrelated content is not overwritten. Search never downloads or repairs weights. Missing weights or dependencies produce a structured reason and lexical fallback, while strict search raises that reason.

The directory contains six pinned upstream files and a small `book-agent-query-contract.json` receipt. The receipt records relative names, byte sizes, SHA-256 values, model identity, revisions and recipe. It contains no machine paths or credentials and can move with the files. Verify or register the directory again at its actual location on another machine. Share one directory among books with this same verified contract; no duplicate model copy is needed in each package, Skill or runtime. Keep it outside read-only book packages.

The standard portable executable supplies lightweight fulltext and HTTP query support. Local neural questions use the Python runtime with the `local-query` extra and this separate weight directory. Weights are not embedded in the standard executable.

### Fixed artifacts and recipe

Model: [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3). Tokenizer and configuration revision: `5617a9f61b028005a4858fdac845db406aefb181`. Safetensors revision: `9a0624b896d81da7492a910ffa53731274b6cf3d`.

| File | Bytes | SHA-256 |
| --- | ---: | --- |
| `config.json` | 687 | `26159e7ad065073448460117eb24b7a4572f6f4e78eadff65dc0a11c052449fa` |
| `tokenizer_config.json` | 444 | `a62b2b6784f990259fddef5f16388693a8043be4f69179e6a5257eeb3f9abac4` |
| `special_tokens_map.json` | 964 | `8c785abebea9ae3257b61681b4e6fd8365ceafde980c21970d001e834cf10835` |
| `tokenizer.json` | 17,098,108 | `21106b6d7dab2952c1d496fb21d5dc9db75c28ed361a05f5020bbba27810dd08` |
| `sentencepiece.bpe.model` | 5,069,051 | `cfc8146abe2a0488e9e2a0c56de7952f7c11ab059eca145a0a727afce0db2865` |
| `model.safetensors` | 2,271,064,456 | `993b2248881724788dcab8c644a91dfd63584b6e5604ff2037cb5541e1e38e7e` |

The local recipe is `bge-m3-dense-cls-l2-float32-v1`: one raw string, original fast XLM-Roberta tokenizer, special tokens included, `XLMRobertaModel` in CPU float32 evaluation mode, CLS from `last_hidden_state[0,0,:]`, and L2 normalization to a finite 1024-dimensional representation. This follows the [upstream BGE-M3 encoder](https://github.com/FlagOpen/FlagEmbedding/blob/master/FlagEmbedding/inference/embedder/encoder_only/m3.py). Local loading forces `local_files_only=True`, `trust_remote_code=False`, `use_safetensors=True` and eager attention. No Pickle checkpoint or remote code is loaded. Generic local model declarations are not executable adapters.

The process caches at most one local query adapter and serializes its queries. The adapter retains one official eager encoder layer at a time in CPU float32 RAM, and reads only the current question's embedding rows. It closes Safetensors mappings before computation; no full model or second weight copy is created. This follows the upstream [partial tensor loading API](https://huggingface.co/docs/safetensors/main/index). Each query rereads the encoder layers, so disk speed can limit latency. Default inference uses two CPU threads and a budget of 512 tokenizer tokens including special tokens. Runtime configuration can set `local_cpu_threads` from 1 to 4 and `local_max_tokens` from 1 to 8192. The public Core question limit remains 8,000 characters; the lower-level encoder also enforces its own 16,000-character ceiling. Oversized questions are rejected before forward, without truncation. The 2.27 GB checkpoint is needed on disk; longer token budgets increase activation memory and CPU time.

### Offline MCP deadlines

Cold encoding includes dependency loading and weight checks. For offline MCP, set the server's per-tool deadline to 600 seconds:

```shell
book-agent --data-dir ./runtime-data mcp BOOK_ID --timeout 600
```

Codex's host tool deadline is separate. Set `tool_timeout_sec = 600` on the corresponding MCP server entry and include `"--timeout", "600"` in its arguments. Codex supports separate `startup_timeout_sec` and `tool_timeout_sec` fields in each `[mcp_servers.<name>]` entry. [Official OpenAI MCP configuration](https://learn.chatgpt.com/docs/extend/mcp). For example, replace the executable/data paths and `BOOK_ID` with the installed offline Python runtime and its registered book:

```toml
[mcp_servers.book_agent_offline]
command = "F:/Tools/BookAgent/query-runtime/Scripts/python.exe"
args = ["-B", "-m", "book_agent", "--data-dir", "F:/BookAgentData", "mcp", "BOOK_ID", "--timeout", "600"]
startup_timeout_sec = 90
tool_timeout_sec = 600
```

When updating an entry generated by `install-host`, retain that entry's name and actual paths. Both the server and host deadline must allow the slow tool call. A CLI search in another process does not warm this MCP process; the lightweight portable EXE is not the optional offline Python runtime.

### Conversion and numeric compatibility limits

The original pinned model revision exposes `pytorch_model.bin`. This implementation uses the public Hugging Face [Safetensors conversion PR #130](https://huggingface.co/BAAI/bge-m3/discussions/130), whose fixed commit is a direct child of that original revision. The conversion was published by the Hugging Face conversion bot and remains unmerged. Tokenizer/configuration artifacts stay pinned to the original revision. The original unsafe binary is never downloaded or loaded by Book Agent.

Artifact identity, tokenizer identity and CLS/L2 recipe are checked. Numerical parity between this local checkpoint and the original **unpinned hosted encoder** has not been independently established. Responses expose `provider_numeric_parity_verified: false` and the local-use warning. Using the same published model and recipe provides a documented inheritance path, while it does not prove provider-specific numerical equivalence. [Upstream model metadata](https://huggingface.co/BAAI/bge-m3/blob/5617a9f61b028005a4858fdac845db406aefb181/README.md) declares MIT; the separate weight bundle includes the upstream source and license information.

## Authorized online questions

```json
{
  "execution": "online",
  "allow_remote_requests": true,
  "allowed_endpoints": ["https://api.siliconflow.cn/v1/embeddings"],
  "auth_env": "SILICONFLOW_API_KEY",
  "timeout_seconds": 30
}
```

Set the secret in the process environment. JSON configuration and book metadata only name the environment variable. Model and endpoint overrides cannot silently replace the inherited space. A manifest URL is not authorization: remote use defaults to false, exact endpoint allowlisting is required, redirects are refused and diagnostics omit credentials. Only the current question is sent to the query encoder. Remote reranking separately requires authorization and sends candidate excerpts. The Harness's own language model network use is separate.

An explicit paired hosted encoder can declare `query_runtime` in a `book-agent-explicit-v1` manifest:

```json
{
  "query_runtime": {
    "provider": "openai_compatible",
    "query_model": "existing-query-encoder",
    "paired_document_model": "original-index-model",
    "revision": "fixed-encoder-revision",
    "compatibility": "declared_paired_encoder",
    "dimensions": 1024,
    "normalization": "l2",
    "metric": "dot",
    "query_prefix": "query: ",
    "query_instruction": "",
    "endpoint": "https://approved.example/v1/embeddings",
    "auth_env": "EXISTING_QUERY_KEY"
  }
}
```

This adapter supports explicitly paired document/query encoders. It requires a complete revision, normalization, paired-space declaration and validated source identity. Runtime endpoint authorization is still required. A generic `sentence_transformers` declaration remains unavailable until an audited safe execution adapter exists.

## Validation evidence

The latest complete release regression finished with **276 passed, 2 skipped in 161.38 seconds** (`.work/tests-release-final.log` and `.xml`). It includes 14 actual tiny-model tests and 29 installer/bootstrap unit cases. Separate **actual default Skill bootstrap installed 46 wheels** offline in **73.259 seconds** (`.work/acceptance/bootstrap-real-install.json`). The **dedicated `release/install-offline.py` installed 67 packages** in an independent CPython 3.13 venv using no-index local hashed wheels, with pip check reporting no broken requirements, in **233.145 seconds** (`.work/offline-install-real-final.summary.json` and `.log`). The optional `local-query` bootstrap route had its 67-wheel plan verified but was not actually installed through bootstrap.

Both receiving-runtime probes loaded Book Agent from their own installed venv, returned three default lexical hits, verified original-source location, produced a Codex host dry-run, checked default 46/optional 67-wheel bootstrap plans, blocked network connections and preserved the original archive without generating document vectors. Default probe time was **16.747 seconds** (`.work/acceptance/base-bootstrap-receiving-runtime.json`). Optional probe time was **245.664 seconds** (`.work/acceptance/optional-receiving-runtime.json`), including cold imports of PyTorch `2.14.1+cpu`, Transformers `4.57.6`, Safetensors `0.8.0` and SentencePiece `0.2.2`, CPU tensor math and plan hash verification. That total is not a search latency measurement; no full model forward was repeated in either receiving-runtime probe. The complete-model evidence below uses the earlier verified Python runtime. Installation/probe reports captured the same code/dependency artifacts before this documentation update; final manifest identity will be regenerated with the documents.

Mock tests check public-artifact integrity receipts, relocation, CLS rather than mean pooling, raw inputs, one cached model adapter, thread limits, no silent truncation, offline loader flags, model-space rejection, tamper detection and remote authorization. They do not execute a real model. The optional actual CPU integration test runs only when `BOOK_AGENT_QUERY_MODEL_TEST_PATH` points to explicitly installed verified weights; it does not download. It checks a finite 1024-dimensional unit vector from a natural-language question. Real SQLite search with that question must be recorded separately from stored-record self-vector scorer tests.

On 2026-10-03, **14 actual tiny-model tests passed in 131.04 seconds**, with no skipped cases. PyTorch `2.14.1+cpu`, Transformers `4.57.6` and Safetensors `0.8.0` executed the model math. Four inputs (unpadded, padded, explicit positions and 3D mask) matched the official full eager XLM-R encoder's final hidden states and CLS/L2 output at `rtol=1e-5`, `atol=1e-6`. Further cases checked embedding-row slicing, mappings closed before computation, one resident CPU layer, corrupt checkpoints and unsupported input/configuration rejection. Evidence is `.work/query-streaming-resumed-threads.xml` and its accompanying log/summary. This test does not load the delivered full checkpoint.

The separate **actual BGE-M3 CPU acceptance passed** using the pinned delivered Safetensors directory and original SQLite document vectors, with socket/urllib network entry points blocked. It produced 1024 finite dimensions with L2 norm `0.9999999338`; retrieval reported `ready / encoded / offline`, used semantic search and returned record IDs `1`, `2`, `3`. Source backtrace located physical PDF page 2 (`page_index=1`). The archive's SHA/size/mtime and all 29 input files were unchanged; document embeddings generated: `0`. Evidence: `.work/acceptance/offline-real-query.json`.

In that local run, cold encoding took **33.609 seconds**; the second Core search took **0.780 seconds for the same question after model/dependency warmup**, with question encoding repeated. The process caches the tokenizer and streaming model adapter, not query vectors; each forward reads the encoder layers again. The complete **88.695-second** run included weight registration/hash checks. This second-search timing is not generalized to other questions. The report does not record process peak memory; these timings are one machine's observation, not a performance or memory guarantee. Hosted-provider numerical parity and actual local reranking remain unverified. Separate native Codex CLI/Python MCP question answering passed with lexical fallback; it did not exercise the offline neural encoder or independently validate automatic Skill triggering.

A second actual blocked-network run, `.work/acceptance/offline-query-instrumented.json`, also passed. It counted **one real forward for the same-question search and one for a fresh-question search**. After model/dependency warmup, Core search took **0.973 seconds** for the same question and **0.869 seconds** for the new question; measured forward times were **0.621 / 0.854 seconds**. The new question used semantic retrieval. Initial encoding took **155.887 seconds**, and the complete run including weight verification took **256.183 seconds**. The archive and 29 input files remained unchanged, and no document embeddings were generated. Both cold measurements include dependency loading and checks before encoding; their differing times show unstable startup in these local runs, not a universal latency range or guarantee.

Service tests use injected transport to check model inheritance, prefixes, one-query input, missing credentials, timeouts and invalid responses. They are contract tests, not live SiliconFlow validation. A live service acceptance needs a separately authorized endpoint and available credentials. Supplied-vector search remains possible offline, with representation provenance reported as `user-supplied / unverified` unless independently established.

离线模式的新 `install-host` 安装会自动为 MCP 设置 600 秒工具期限,并同步 Codex 的 `tool_timeout_sec`;其他模式保持原默认参数。已有旧离线配置若出现所有权冲突,先用 `uninstall-host BOOK_ID --host codex` 安全卸载本项目拥有且未改变的配置,再重新安装。手工接入其他宿主时,仍需在宿主中配置匹配的工具期限。

原文 SHA-256:01d217d6ce05a80c9140dd7460562caa9272d32531edca59e7e0139ab02d79ef