AIOKFRAGAgentsArchitecture

Connecting OKF to RAG: source versions, chunks and query contracts

··3 min read

An index can still split a table after its source has been moved into an OKF bundle. Headings, links and versions make information available; preserving it during ingestion and using it during retrieval require implementation.

This site maintains knowledge/index.md as an entry point for agents reading project conventions. That demonstrates file navigation, not a deployed vector service or a measured retrieval improvement. The following is a proposed contract for connecting such sources to a retrieval system.

Define what each result must carry

For a hypothetical order-retention rule, the path might be policies/orders.md, the section ID retention and the environment prod. Also record the source revision, content hash, index version and effective date. A path locates a document; a revision identifies the particular edition. These are proposed application fields, not claims about mandatory OKF fields.

json
{
  "source_id": "policies/orders.md#retention",
  "source_revision": "2",
  "environment": "prod",
  "effective_from": "2026-09-01",
  "content_hash": "<computed from source bytes>",
  "index_version": "retention-demo-v2",
  "parent_id": "policies/orders.md"
}

The chunker must preserve structure

Keep section paths, table headers and scope in retrievable data. If a table exceeds the segment limit, split rows while retaining column names, units and the parent ID. Definitions spanning segments need a way to retrieve their parent. A fixed character-count split can lose meaning even in carefully organized sources.

Explicit links also do not implement traversal. Define allowed edges, a hop limit, cycle handling and a missing-target result. Apply the same authorization to source access and retrieval; do not retrieve restricted content and expect the model to ignore it.

Switch versions instead of mixing policies

When the source moves from revision 1 to 2, build a new index version and verify additions, changes and deletions. Switch the active index only after checks pass. If they fail, retain the previous version and disclose its date. This is a proposed design, not a report of a service already tested on this site.

Check at least four cases: the applicable production rule, isolation of test rules, absence of deleted documents and an unknown answer when evidence is missing. The acceptance condition is use of the correct source version and scope, not simply production of an answer.

Choose retrieval by the information available

With a known document ID, fetch it directly and check its version. For a vague description, use full-text or vector search to find candidates and enforce environment, effective-date and access conditions. Expand relationships through IDs or links when necessary. This is not a universal rule to traverse OKF before every vector search; choose based on available identifiers and observed retrieval failures.

The answer should expose its source ID, revision and cited passage. If only expired or incompatible rules are available, report that gap instead of drawing a conclusion. A file format cannot guarantee correctness; these checks make failures traceable.

Use one case to check the boundaries

For the production-retention question, the failure might be a test document ranking higher or an old production revision remaining active. The first calls for an environment filter; the second calls for index-version checks. Rewording the prompt does not fix either ingestion problem. The companion article supplies a model-free ranking and filtering example; use it to check the fields before adding real embeddings and answer evaluation.

The bundle maintainer owns source content and links. Indexing owns segmentation, versions and deletion. Retrieval owns constraints and fetching. Generation owns citations and handling missing evidence. Explicit ownership identifies which step has fallen behind a source update.

RAG 檢索範例 / Retrieval failure example本站 knowledge 入口 / This site’s knowledge index