---
title: Documentation Retrieval
description: Reference for published Markdown source links and the documentation retrieval chain - providers, multi-site MkDocs fallback configuration, location resolution rules, the AI Search filter contract, and the evaluation harness.
---

# Documentation Retrieval

How `saif docs`, the Forge MCP tools, and Smithy's `search_docs`/`get_doc` find and fetch documentation pages. Design rationale for the tiered chain lives in [ADR 0011](decisions/0011-shared-platform-sdk-and-tiered-documentation-search.md); this page documents what the chain does today.

---

## Published Markdown source

Forge's rendered documentation pages expose a **Download Markdown source** action. Readers can download the source or copy the link for a consumer to use as its first retrieval source, before extracting content from HTML. The existing action uses `page.file.src_uri | url` in `docs/overrides/partials/actions.html` to link to the page's source path relative to the site; `docs/hooks/copy_markdown.py` copies the `.md` files into the published output.

For example, `/forge/reference/documentation-retrieval/` links to `/forge/reference/documentation-retrieval.md`. A page authored as `reference/index.md` instead links to `/forge/reference/index.md`, not `/forge/reference.md`. The hook also creates the latter sibling copy for non-root `index.md` pages unless a real source file already occupies that path. The site root's source remains `/forge/index.md`; the hook creates no sibling copy for it.

The MkDocs provider does **not** discover this action in HTML. It independently builds its first fetch URL from the configured [`RawMarkdownUrlPattern`](#mkdocs-fallback-configuration), which defaults to `{base}/{path}.md`, using the site's base URL and resolved document path. For ordinary pages this reaches the same source URL as the action; for directory index pages it relies on the sibling copies instead. Consumers following the action's link retain the actual `index.md` source path without deriving it from the rendered URL. The provider's existing [HTML fallback](#rules-within-a-site) still applies when its raw Markdown fetch fails.

---

## 🔗 Provider chain

`DocumentationService` (`src/dotnet/SAIF.Platform.Sdk/Documentation/DocumentationService.cs`) holds an ordered list of `IDocumentationProvider` instances, sorted ascending by `Order`. For both search and retrieval the first provider returning a non-empty result wins. A provider that throws is logged and skipped, and the chain continues; only a cancellation requested by the caller propagates.

| Host | Provider | `Name` | `Order` | Search | Fetch |
| ---- | -------- | ------ | ------- | ------ | ----- |
| Smithy | `AiSearchDocumentationProvider` | `ai_search` | 0 | ✅ Semantic, cross-repo | ❌ Always returns `null` |
| Smithy | `MkDocsDocumentationProvider` | `mkdocs` | 100 | ✅ Lexical | ✅ |
| CLI | `SmithyRemoteDocumentationProvider` | `smithy_remote` | 0 | ✅ Brokered through Smithy MCP | ✅ Brokered through Smithy MCP |
| CLI | `MkDocsDocumentationProvider` | `mkdocs` | 100 | ✅ Lexical | ✅ |

Registration:

- `AddForgeSdk` registers `MkDocsDocumentationProvider` for every host.
- Smithy adds `AddAiSearchDocumentation` on top, which also requires a `SearchIndexClient` for the private AI Search endpoint.
- The CLI adds `SmithyRemoteDocumentationProvider`, which calls Smithy's `search_docs`/`get_doc` over the bearer-authenticated MCP bridge. Any failure (unconfigured, unreachable, unauthenticated, timeout) yields an empty result and falls through. A failed connection is remembered for a 5-minute cooldown so later calls fall through without re-incurring the connect timeout.

Because `ai_search` is search-only, full-page retrieval on Smithy always comes from the MkDocs provider.

### Provider attribution

`DocumentationService` tags the current activity with three attribution tags. Each is set to the winning provider's `Name`, or to the sentinel `none` when no provider produced a result.

| Tag | Set by | Use it for |
| --- | ------ | ---------- |
| `docs.search.provider` | `SearchAsync` only | Which provider answered the search. |
| `docs.fetch.provider` | `GetDocumentAsync` only | Which provider served the full page. |
| `docs.provider` | Both | Continuity with dashboards and queries that predate the split. |

Prefer the split tags. Search and fetch usually run inside the same ambient activity, because an agent searches and then immediately fetches one of the returned locations. `docs.provider` is written by both operations, so the fetch value overwrites the search value and the activity ends up reporting only the last writer. `docs.search.provider` and `docs.fetch.provider` never collide, so an activity can show `ai_search` for discovery and `mkdocs` for retrieval at the same time - which is the normal shape on Smithy, where `ai_search` is search-only.

Tags ride `Activity.Current` where one exists. The same value is always written as a structured log property at Debug, so a host without tracing still gets attribution.

!!! warning "Attribution cannot distinguish Smithy's backends"
    From the CLI the winning provider reports `smithy_remote` whether Smithy answered from Azure AI Search or fell back to its own MkDocs index. Treat CLI-side provider attribution as coarse routing information, not as a backend breakdown.

---

## 🩺 Diagnostics vocabulary

Resolution outcomes use a fixed string vocabulary so diagnostics stay comparable across hosts and runs.

| Constant set | Values |
| ------------ | ------ |
| `MkDocsResolutionReasons` | `resolved`, `not_found`, `unknown_site`, `ambiguous`, `site_unavailable`, `fetch_failed` |
| `MkDocsResolutionRules` | `exact`, `variant`, `canonical-leaf`, `none` |

Each `MkDocsLocationResolution` carries the requested location, the reason code, the rule that fired, the resolved site and location when it resolved, ambiguity candidates in `site/location` form, and a human-readable `Detail`. Log message wording is not part of this contract; the reason and rule strings are.

---

## ⚙️ MkDocs fallback configuration

Bound from the `Forge:Documentation` configuration section into `ForgeDocumentationOptions`.

| Key | Default | Description |
| --- | ------- | ----------- |
| `Origin` | `https://docs.saif.com` | Origin hosting every platform docs site. |
| `DefaultSite` | `forge` | Site tried first for a location with no site segment. |
| `Sites` | `forge`, `cloud-foundations`, `iac-azure-modules`, `iac-okta-modules`, `iac-aws-modules` | Sites published under `Origin`. |
| `SearchIndexPath` | `/search/search_index.json` | Search index path relative to a site's base URL. |
| `CacheFileNameTemplate` | `{site}-docs-search-index.json` | Per-site cache file name; `{site}` is the placeholder. When omitted, the site name is prepended, so `docs-index.json` becomes `forge-docs-index.json` for Forge. |
| `CacheDurationHours` | `1` | How long a disk-cached index is treated as fresh on initial load. Does not expire the in-memory index. |
| `RawMarkdownUrlPattern` | `{base}/{path}.md` | Raw markdown URL pattern. `null` disables raw markdown and always converts HTML to markdown. |

Derived values and behavior:

- A site's base URL is `{Origin}/{site}`. There are no per-site base URLs.
- `DefaultSite` is inserted into the site list when configuration omits it, so it is always served.
- Site names are de-duplicated case-insensitively, blank entries are dropped, and a `CacheDurationHours` of zero or less falls back to one hour.
- Each site has its own index, its own cache file under `%LOCALAPPDATA%\saif-cli\cache\`, and its own load gate. Nothing loads at startup; indexes load lazily on first use, and cross-site search loads them concurrently.
- A site whose index fails to load (unpublished, unreachable, malformed, or empty) is logged at Debug and marked failed for the process lifetime, then skipped. Other sites are unaffected.
- When a network fetch fails, a stale cache file is used as a fallback if one exists.
- Characters invalid in a file name are replaced with `-` in the site-name substitution when building the cache path; the template itself is not sanitized.
- A query of exactly `*` lists pages instead of term-scoring them, with anchor entries (`page/#section`) excluded so a page appears once. `saif docs list` depends on this. A query that merely contains `*`, such as `* widget`, is scored normally. A wildcard listing is not an exhaustive page inventory: `DocumentationService` caps every search at 25 results, so `saif docs list` and a `*` search from any caller return at most 25 pages.
- Because every page gets the same score under a wildcard listing, per-site results are interleaved round-robin across loaded sites before the global result limit is applied, so a large default-site index (Forge has thousands of pages) cannot fill the entire limit and hide every other site.

### Adding a site

Add the site name to `Sites`. Nothing else is required, provided the site publishes an index at `{Origin}/{name}/search/search_index.json`. A listed site that does not serve an index costs one round trip per process and is then skipped.

### Current publishing status

| Site | Index at `{Origin}/{site}/search/search_index.json` |
| ---- | --------------------------------------------------- |
| `forge` | ✅ Serves |
| `iac-azure-modules` | ✅ Serves |
| `iac-okta-modules` | ✅ Serves |
| `cloud-foundations` | ❌ 404 - root-publishes to `docs.saif.com/search/search_index.json` |
| `iac-aws-modules` | ❌ 404 - not yet published |

Both are expected to move to the standard layout, so they are listed and skipped rather than special-cased. Until then the fallback covers three sites.

---

## 🧭 Location resolution

`MkDocsSearchService.ResolveLocationAsync` accepts both bare paths (Forge pages, whose `forge/` prefix upstream providers strip) and site-qualified paths.

### Search result locations

**A location returned by search can be passed straight back to `get_doc` or `saif docs get`.** This is a location-resolution contract, not a guarantee of page availability, freshness, or agreement between the indexed and fetched revisions.

| Result from | `Location` shape | Example |
| ----------- | ---------------- | ------- |
| The default site (`forge`) | Bare path, unchanged from the index | `reference/documentation-retrieval` |
| Any other site | `{site}/{path}` | `iac-azure-modules/reference/naming` |

A non-default site's results are qualified because resolution tries `DefaultSite` first: a bare path that also exists on `forge` would fetch the wrong page, and a bare path matching several sites would come back `ambiguous`. Qualifying at search time removes the collision instead of asking the caller to guess. A site's root page qualifies to the bare site name.

`Url` is always built from the unqualified path, so it keeps pointing at the published page. Use `Location` to fetch and `Url` to link.

!!! note "Default-site results are unchanged"
    Qualification applies only to non-default sites. Forge results are byte-identical to what search returned before, so existing callers and saved locations keep working.

### Site selection

1. Normalize: trim surrounding slashes and drop any `#fragment`.
2. If the leading segment names a configured site, resolve the remainder against that site. If that misses, the unstripped path is retried against the same site, which covers a site-named top-level directory inside the site itself. If that site's index cannot load, the result is `site_unavailable`.
3. Otherwise resolve against `DefaultSite` first. A `resolved` or `ambiguous` outcome there is returned immediately. If the default site's index cannot load, this step is skipped and remembered so step 5 can distinguish an outage from a genuine miss.
4. Otherwise try every remaining loaded site, keeping both `resolved` and `ambiguous` outcomes. Exactly one outcome is returned as it stands, so ambiguity found inside a single non-default site is reported rather than discarded. More than one outcome is reported `ambiguous` across sites, never picked silently; the candidate list merges each resolved page as `site/location` with the candidates any ambiguous site already qualified.
5. No match: `unknown_site` when the path contains a `/`, its leading segment contains a hyphen, and that segment is not a top-level directory in any loaded site. Otherwise `site_unavailable` if the default site could not load in step 3, since a bare path was never actually checked against it and a miss on the fallback sites alone cannot confirm it is missing. Anything else is `not_found`.

### Rules within a site

Applied in order; matching is case-insensitive.

| Rule | Matches | Ambiguity handling |
| ---- | ------- | ------------------ |
| `exact` | The normalized path is an indexed canonical page. | n/a |
| `variant` | The path after stripping a trailing `.md`, `.html`, or `.htm`, stripping a trailing `/index`, and mapping a bare `index` to the site root. | More than one match returns `ambiguous` with candidates. |
| `canonical-leaf` | The final segment matches exactly one canonical page in that site. | More than one canonical page sharing the leaf returns `not_found` rather than a guess. |

All three lookups are built over canonical pages: anchor entries (`page/#section`) in the search index normalize onto the page that owns them, and the first entry for a page wins.

Leaf uniqueness is checked within each site, not against one combined index. A bare path that misses on the default site can resolve by leaf on another site; competing matches on multiple non-default sites return `ambiguous`. A default-site match still wins without consulting other sites. An `exact` hit is silent; `variant` and `canonical-leaf` resolutions log at Information with the requested location's length, site, resolved location, and rule. The raw requested location can carry a token or signed URL, so these events log only its length.

Content is fetched as raw markdown from `RawMarkdownUrlPattern` first, falling back to the published HTML converted to markdown when raw content is unavailable, empty, or actually HTML. If neither source supplies non-whitespace content, retrieval returns `fetch_failed`, not a document containing only generated front matter. On a successful fetch the returned document carries the **resolved** location, URL, and site, alongside the originally requested location, and the markdown is prefixed with `source`, `site`, and `retrieved` front matter.

### Limits

!!! warning "What resolution cannot recover"
    - **Stale index, removed page.** Resolution reads only the in-memory index. A path that still exists in the index resolves `exact`, then fails at fetch time with `fetch_failed`.
    - **Stale index, page still served.** A page can return HTTP 200 and non-empty content from a different revision than the search result. The current chain does not detect that skew.
    - **Renamed leaf.** A rename that changes a page's final segment is unrecoverable by leaf matching.
    - **Unknown site vs. wrong path.** Distinguishing an unconfigured site qualifier from a plainly wrong content path is a heuristic (slash present, hyphenated leading segment, not a known top-level directory). A single-word leading segment is always treated as content.
    - **Default-site short circuit.** A bare path that resolves on the default site is returned without consulting other sites, so a same-leaf page elsewhere is never reported as ambiguous.
    - **Directory restructuring, not just a rename.** This consolidation moved pages between top-level directories (for example `guides/development/index` to `build/index`). Each moved page carries `moved_from` front matter recording its prior paths, and the docs build turns those into HTML redirects on both docs.saif.com/forge/ and forge.saif.com, so a browser following an old URL lands on the new page. Retrieval does not use that mapping: `ResolveLocationAsync` does not consult `moved_from`, and the `.md` twin at an old path is not redirected. A caller passing a pre-move location gets `not_found` by design and should search again to find the current location.
    - **Troubleshooting ID redirects.** The docs build also redirects `troubleshoot/<SAIFTRBL ID>/` to each article, but `ResolveLocationAsync` does not resolve those paths for the same reason it ignores `moved_from`: redirect stubs are not in the search index. Search for the ID instead; the [catalog](../troubleshoot/catalog.md) and each article's frontmatter carry it.
    - **No section-level fetch.** `get_doc` returns the whole page and drops any `#fragment` during normalization, so it cannot return a single H2 section. [#1173](https://github.com/saif-corp/forge/issues/1173) declined section fetch because articles are short and single-symptom, and the AI Search provider is search-only.

### Staleness and open revision-verification criterion

There is no enforced staleness bound. The indexer schedule affects AI Search freshness, while `CacheDurationHours` controls only whether MkDocs accepts a disk cache on initial load. `MkDocsSearchService` keeps a loaded index for the process lifetime and reuses a stale cache file after a network failure, so a long-lived host can serve arbitrarily old results.

Search results and fetched pages carry no comparable source-commit metadata. HTTP 200, non-empty content, and a successful fetch therefore establish availability only, not freshness or revision agreement. The evaluation harness's `commitSha` identifies the measured CLI build, not the documentation revision.

**The same-source-commit verification and detectable-skew acceptance criterion in [#1177](https://github.com/saif-corp/forge/issues/1177) remains open, and client-side revision correlation is not the planned next step.** Producer-side source-identifier metadata is tracked in [saif-corp/cloud-foundations#160](https://github.com/saif-corp/cloud-foundations/issues/160). That work stands on its own for build attribution and orphaned-document detection rather than gating this criterion, so treat it as related scope, not a blocker. This change implements neither Forge-side correlation nor explicit skew detection. Periodic index refresh could reduce staleness without new metadata, but it would not prove revision agreement or bound skew to the indexer schedule alone: refresh intervals, publishing delays, and failed refreshes also matter. Periodic refresh is also unimplemented here.

---

## 🔍 AI Search filter contract

`AiSearchDocumentationOptions` binds from `Forge:Documentation:AiSearch`.

| Key | Default | Description |
| --- | ------- | ----------- |
| `IndexName` | `saif-docs` | Index over platform docs, populated by cloud-foundations. |
| `SemanticConfigurationName` | `docs-semantic` | Semantic configuration used for ranking. |
| `PlatformDocsBaseUrl` | `https://docs.saif.com` | Base URL that blob-storage paths are rewritten to. |
| `MkDocsSitePrefix` | `forge` | Leading path segment stripped from `Location` so the MkDocs fallback can resolve it. Empty disables stripping. |
| `SearchFilter` | *(empty - no filter)* | OData `$filter` sent with every search request. Empty or whitespace sends no filter. Smithy configures `format eq '.html'`; see [Code default vs. shipped value](#code-default-vs-shipped-value). |

The `Default` column is the C# property initializer in `AiSearchDocumentationOptions`, which is what any host gets before configuration binds over it.

When `SearchFilter` holds a non-whitespace value, it is assigned to the request-level `SearchOptions.Filter`. It is never applied after results return. Azure AI Search evaluates `$filter` **before** semantic reranking, so a request-level filter shapes the candidate set that gets reranked. A client-side post-filter would leave the Markdown twin of each page consuming result slots before the trim happens.

### Code default vs. shipped value

Filtering is off in code and on in Smithy. Those are two different things and both are deliberate.

| | Value | Source |
| --- | ----- | ------ |
| Code default | *(empty)* - no filter, every indexed document is a candidate | `SearchFilter` initializer in `src/smithy/src/Smithy/Documentation/AiSearchDocumentationOptions.cs` |
| Shipped in Smithy | `format eq '.html'` - drops the Markdown twin of each indexed page | `Forge:Documentation:AiSearch:SearchFilter` in `src/smithy/src/Smithy/appsettings.json` |

The split is the opt-in boundary. The filter only behaves as measured against an index whose documents actually carry a populated `format` field, which requires the `saif-docs` indexer reset. Defaulting to empty means a host pointed at an un-reset index keeps working unchanged, and enabling the filter stays an explicit, per-host decision made once that host's index has been reset.

!!! warning "Standing up another host"
    Filtering is **not** on by default. If you configure a host other than Smithy against the `saif-docs` index, you get no filter until you set `Forge:Documentation:AiSearch:SearchFilter` yourself, and search can return `.md` twins alongside rendered pages. Variant resolution now makes these `.md` locations fetchable when the canonical page exists, but duplicates still consume ranking slots. Confirm `.html` coverage on the index you are pointing at, then enable the filter to remove those duplicate candidates.

### Measured effect of the filter

Measured against platform dev after the `saif-docs` indexer reset, with the same CLI binary either side so only server configuration changed:

| Metric | Filter off | Filter on |
| ------ | ---------- | --------- |
| hit@1 | 1/15 (6.7%) | **8/15 (53.3%)** |
| hit@5 | 7/15 (46.7%) | **11/15 (73.3%)** |
| fetch success | 15/15 (100%) | 13/15 (86.7%) |

Discovery improves sharply because the Markdown twin of a page no longer outranks the rendered page it duplicates. The filter-off hit@1 of 6.7% understates real behaviour: most top results were the `.md` twin of an acceptable location, and location comparison deliberately does not strip `.md`. Scored with the suffix stripped, filter-off hit@1 would be 7/15.

Fetch success drops by two cases, and neither is caused by the filter. Both are index-hygiene problems that the filter merely promotes into the top slot: one stale root-published blob and one phantom entry with no page behind it ([cloud-foundations#158](https://github.com/saif-corp/cloud-foundations/issues/158)).

!!! warning "The filter depends on the indexer reset"
    `format` is null on every document indexed before the `saif-docs` indexer reset, so `format eq '.html'` against an un-reset corpus makes search return almost nothing. The reset has been run and coverage confirmed indirectly by the measurement above - a filtered search returning results at all proves `format` is populated. Direct verification from a developer machine is still not possible: the AI Search service has public network access disabled ([cloud-foundations#77](https://github.com/saif-corp/cloud-foundations/issues/77)) and `listAdminKeys` is denied by RBAC.

    Any environment whose index has not been reset must leave `SearchFilter` empty until it has.

### Changing or disabling the filter

To change the filter, for example to narrow it further:

1. Set `Forge:Documentation:AiSearch:SearchFilter` in Smithy's configuration to the new OData expression.
2. Re-run the [evaluation harness](#evaluation-harness) with a distinct label and compare aggregates against the previous run.

To disable filtering entirely, for example when pointing a host at an index that has not been reset:

1. Set `Forge:Documentation:AiSearch:SearchFilter` to an empty string. Any whitespace-only value has the same effect. Removing an override restores the value from lower-priority configuration, which in Smithy's shipped `appsettings.json` is `format eq '.html'`; only removing the key from all configuration sources reaches the empty code default.
2. Expect duplicate candidates to consume ranking slots again; the filter-off column above records the measured discovery effect. `saif docs get` now resolves `.md` twins through the variant rule when the canonical page exists, so disabling the filter does not itself make those locations unfetchable.

No code change is required either way. The provider captures `IOptions<AiSearchDocumentationOptions>.Value` in its constructor, so a configuration change takes effect when the host restarts, not on the next request.

### Result previews

The preview shown for each AI Search result is a query-centred snippet cut from the document's `content` field. The indexed `description` is used only when that snippet comes back empty, with the same whitespace normalization and length limit.

The order matters. On rendered MkDocs pages `description` is the site-wide meta description, identical on every page, so preferring it renders every result with the same marketing sentence. That stayed hidden while Markdown twins dominated results, because `.md` blobs carry no `description` at all - enabling the `.html` filter would have surfaced it on every result at once. The two changes are coupled and ship together.

Snippet selection splits queries on whitespace, including tabs and line breaks, skips query terms shorter than three characters whenever a longer term is available, and scores candidate windows by how many distinct query terms they contain, preferring the earliest window on a tie. It considers at most the first 20 occurrences of each retained term, so it does not exhaustively search every possible window. Rendered HTML carries navigation chrome, so without the length rule a stop word such as `to` anchors the snippet on `Skip to content` instead of the prose the reader asked for.

Text is also whitespace-collapsed before the window is cut. Indexed HTML extracts arrive with long runs of newlines and indentation between fragments of real text, so an uncollapsed 200-character window can spend its whole budget on layout and surface only a handful of words. Collapsing raised the average preview from 136 to 206 visible characters across the evaluation set and removed every near-empty preview.

#### Preview output contract

`previewText` is normalized, not verbatim source. Consumers that render, compare, or diff it should expect:

- **Every run of whitespace collapses to exactly one U+0020 space.** The test is `char.IsWhiteSpace`, so this covers Unicode whitespace generally - tabs, `\r`, `\n`, form feeds, non-breaking and other Unicode spaces - not just ASCII blanks.
- **Source line breaks are not preserved.** A paragraph break, a list, and a code block all flatten into one continuous line. There is no way to recover the original layout from the preview.
- **The text is trimmed at both ends.** Leading and trailing whitespace is dropped rather than collapsed to a space.
- **The character budget uses .NET string length (UTF-16 code units).** `maxLength` (200 by default) applies after collapsing whitespace, not to grapheme clusters or display width. A preview contains at most 200 code units of normalized source text before ellipses.
- **Ellipses sit outside the budget.** A leading `...` is added when the window does not start at the beginning of the document, and a trailing `...` when it does not reach the end, so a truncated preview is up to 206 characters long.

The normalization is deliberate and stable. Do not treat a preview that differs from the source page's whitespace as a defect, and do not compare `previewText` byte-for-byte against page content.

!!! note "Previews inherit whatever the index stores as content"
    `content` holds the full text extract of a rendered page, including the navigation tree and the on-page table of contents, so a snippet can open on link labels before reaching prose. Pages with substantial body text recover within the window; thin index pages, whose extract is mostly navigation, do not. Improving that further belongs in the indexer, which would need to extract the page body rather than the whole document. No client-side snippet heuristic can recover prose from a document that contains very little.

### Release-notes ranking

Release-notes de-emphasis is the same knob — `format eq '.html' and section ne 'release-notes'` — and is deliberately deferred pending evaluation data. It is an editorial judgement, not a wiring gap. A second, independent mechanism also exists: the MkDocs lexical scorer reads an optional per-page `boost`, emitted by the Material search plugin from a page's `search.boost` front matter, and multiplies the term score by it. Values that are NaN or negative are clamped to zero.

---

## 🧪 Evaluation harness

`tools/docs-eval/Invoke-DocsRetrievalEval.ps1` replays `tools/docs-eval/eval-set.json` (14 cases, v2) through the CLI and records a JSON run.

```powershell
# Measure the CLI on PATH and write a labelled run record
.\tools\docs-eval\Invoke-DocsRetrievalEval.ps1 -Label pre-fix

# Validate the harness offline, without a CLI
.\tools\docs-eval\Invoke-DocsRetrievalEval.ps1 -SelfTest
```

### Parameters

| Parameter | Default | Description |
| --------- | ------- | ----------- |
| `-EvalSetPath` | `eval-set.json` beside the script | Fixture set to replay. |
| `-OutputPath` | `runs\<yyyyMMdd-HHmmss>-<label>.json` beside the script | Run record destination; the parent directory is created if missing. |
| `-Label` | `unlabeled` | Free-form tag stamped into the record, for example `pre-fix` and `post-fix`. |
| `-SaifPath` | `saif` | Executable to measure. Point it at a specific build to measure that tree. |
| `-CommitSha` | Derived; see below | Commit of the build being measured. When omitted the harness derives it from a SHA embedded in the CLI's informational version, and otherwise records `null` with a warning. Supply it when the version carries no SHA, even for a binary inside this checkout: its path cannot prove which commit built it. |
| `-SearchLimit` | `10` | Results requested per search, range 5-50. The shared documentation service caps the effective request at 25. hit@5 always scores the first five results, so a value below 5 is rejected rather than recorded as a mislabeled metric. |
| `-PassThru` | off | Also emit the run record on the pipeline. |
| `-SelfTest` | off | Run parsing, scoring, aggregation, the search-limit guard, and commit-attribution assertions against synthetic JSON and exit. Its own parameter set, so it takes no other parameters, and it runs before the search-limit guard. |

### What it measures

1. Runs `saif docs search <query> --limit <SearchLimit> --format json`.
2. Scores discovery: `hitAt1` (top result is in `acceptableLocations`) and `hitAt5` (any of the first five), plus `hitRank`.
3. Runs `saif docs get <location> --format json` on **the location search actually returned**, not the fixture's expected location. Fetching a known-good fixture path would bypass the search-returns-an-unfetchable-location failure mode the harness exists to catch.
4. Scores fetch success: exit code 0, parseable JSON, non-empty `content`. This does not verify documentation freshness or source-commit agreement.
5. Records the search snippet raw as `previewText` alongside `expectedPreviewTerms`. **Preview quality is never scored**; compare it by hand.

Location comparison trims whitespace and slashes and is case-insensitive. It does not strip a `.md` suffix or an `/index` segment, because whether those shapes resolve is part of what is being measured; the eval set lists them as separate acceptable locations.

### Run record

| Field | Contents |
| ----- | -------- |
| `schemaVersion`, `label`, `timestamp` | Run identity. `schemaVersion` is `2`. |
| `commitSha`, `commitShaSource` | Commit of the **measured build**, and how it was attributed: `parameter` (explicit `-CommitSha`), `cli-version` (SHA parsed from the CLI's informational version), or `unknown` (`commitSha` is `null`). The executable's path never determines its commit. |
| `harnessCommitSha` | HEAD of the checkout **running the harness**, or `unknown` when git cannot report it. Always the harness, never the measured build. |
| `cliVersion`, `cliPath`, `resolvedCliPath` | What was measured: the reported version, the `-SaifPath` value as passed, and the executable it resolved to. |
| `evalSetPath`, `evalSetVersion`, `searchLimit`, `hitAtNDepth` | Harness inputs. |
| `knownLimitations` | Carried into every run, including the `smithy_remote` attribution limitation. |
| `aggregate` | `caseCount`, `hitAt1Count`/`hitAt1Rate`, `hitAt5Count`/`hitAt5Rate`, `fetchSuccessCount`/`fetchSuccessRate`. |
| `cases` | Per-case record: locations, hit flags and rank, search and fetch exit codes and errors, provider fields, raw preview. |

!!! warning "Commit attribution in `schemaVersion` 1 records"
    Version 1 stamped `commitSha` with the harness checkout's HEAD unconditionally, whatever binary `-SaifPath` resolved to. Every run recorded during this change is a version 1 record, and several measured a CLI published outside this checkout, so their `commitSha` identifies the harness rather than the build. Read those values as harness provenance only. The measurements themselves stand: the build behind the pdev figures was verified through the pipeline rather than through the run record. Version 2 records separate the two fields and refuse to guess.

!!! note "Read the aggregate, not the exit code"
    A measurement run exits 0 even when cases fail. Every failure - a non-zero CLI exit, an empty result set, unparseable JSON - is recorded on the case and the run continues. `-SelfTest` is the exception: it exits 1 when an assertion fails.

Run records under `tools/docs-eval/runs/` are git-ignored and are not committed. Eval-set case fixtures name their failure mode: `framing-prose`, `stale-location`, `cross-site`, `guide-over-release-notes`, and `release-notes-intended`.

---

## 📈 Production signals

The harness only runs on demand against a fixed 15-case set, so it cannot tell you that retrieval degraded last Tuesday. Smithy emits the subset of these signals that can be observed without ground truth, under the `Smithy` meter and activity source, collected by SAIF service defaults over OTLP.

| Signal | Kind | Dimensions | Answers |
| ------ | ---- | ---------- | ------- |
| `smithy.docs.searches` | Counter | `smithy.docs.source`, `smithy.docs.outcome` (`ok`, `empty`) | How often does search match nothing at all? |
| `smithy.docs.fetches` | Counter | `smithy.docs.source`, `smithy.docs.outcome` (`ok`, `not_found`) | How often does a fetch return no non-empty document? Includes direct requests, not just locations returned by search. |
| `smithy.docs.result_count` | Span tag | - | Result-set size distribution for a given search. |
| `smithy.docs.min_preview_chars` | Span tag | - | Visible, whitespace-collapsed length of the thinnest snippet in a result set. |

`smithy.docs.outcome` was added as an extra dimension on the existing `smithy.docs.searches` counter rather than as a separate metric, so queries that group only by `smithy.docs.source` keep aggregating exactly as they did before.

The zero-result event in `DocumentationService` records query length and term count at Information, not raw query text. The search Debug events in `ForgeTools.SearchDocsCoreAsync` and `AiSearchDocumentationProvider.SearchAsync` also use only those safe query counts; MCP `search_docs` uses that shared logging without separately logging the query. The agent's documentation retrieval span also omits the raw query. Documentation queries are free text that can carry pasted tokens or signed URLs, so these application-owned instrumentation points do not emit raw queries even when a host enables Debug logging. This is not a guarantee about framework-level telemetry or other host instrumentation.

### What fetch failure detects

Before the variant-resolution fix, search-side metrics looked healthy throughout the failure this page documents. Search returned results, including `.md` twins the old fetch path could not resolve. `smithy.docs.fetches` with `smithy.docs.outcome` provides a production availability signal, but it includes all fetches and does not correlate them with prior search hits. The harness measures a different population: the top search result from each fixture. Its fetch-success rate moved from 13.3% to 93.3% once resolution was fixed. Those are historical measurements, not the current behavior of `.md` locations.

A sustained `not_found` rate flags an availability problem such as a missing page, broken index, or changed publishing layout. A stale index or page that still produces a successful fetch records `ok`; this counter cannot detect freshness or revision skew. Those checks remain part of the [open revision-verification criterion](#staleness-and-open-revision-verification-criterion).

`smithy.docs.min_preview_chars` exists for the same reason. Three separate preview defects shipped without any signal, because a search that returns the correct page with an unusable snippet still counts as a successful search. The thinnest snippet in a result set is a cheap proxy that catches a whole-result-set regression.

!!! warning "hit@1 and hit@5 cannot be emitted in production"
    Both require knowing which page *should* have won, which only the eval set supplies. There is no runtime substitute. Treat the harness as the ranking-quality instrument and these metrics as the availability-and-plausibility instrument; they answer different questions and neither replaces the other.

---

## 📚 Resources

- [ADR 0011 - Shared platform SDK and tiered documentation search](decisions/0011-shared-platform-sdk-and-tiered-documentation-search.md)
- [Smithy MCP tools](smithy.md)
