Documentation Retrieval¶
How saif docs, the Forge MCP tools, and Smithy's search_docs/get_doc find and fetch documentation pages. Design rationale for the tiered chain lives in ADR 0011; this page documents what the chain does today.
Published Markdown source¶
Forge's rendered documentation pages expose a Download Markdown source action. Readers can download the source or copy the link for a consumer to use as its first retrieval source, before extracting content from HTML. The existing action uses page.file.src_uri | url in docs/overrides/partials/actions.html to link to the page's source path relative to the site; docs/hooks/copy_markdown.py copies the .md files into the published output.
For example, /forge/reference/documentation-retrieval/ links to /forge/reference/documentation-retrieval.md. A page authored as reference/index.md instead links to /forge/reference/index.md, not /forge/reference.md. The hook also creates the latter sibling copy for non-root index.md pages unless a real source file already occupies that path. The site root's source remains /forge/index.md; the hook creates no sibling copy for it.
The MkDocs provider does not discover this action in HTML. It independently builds its first fetch URL from the configured RawMarkdownUrlPattern, which defaults to {base}/{path}.md, using the site's base URL and resolved document path. For ordinary pages this reaches the same source URL as the action; for directory index pages it relies on the sibling copies instead. Consumers following the action's link retain the actual index.md source path without deriving it from the rendered URL. The provider's existing HTML fallback still applies when its raw Markdown fetch fails.
๐ Provider chain¶
DocumentationService (src/dotnet/SAIF.Platform.Sdk/Documentation/DocumentationService.cs) holds an ordered list of IDocumentationProvider instances, sorted ascending by Order. For both search and retrieval the first provider returning a non-empty result wins. A provider that throws is logged and skipped, and the chain continues; only a cancellation requested by the caller propagates.
| Host | Provider | Name |
Order |
Search | Fetch |
|---|---|---|---|---|---|
| Smithy | AiSearchDocumentationProvider |
ai_search |
0 | โ Semantic, cross-repo | โ Always returns null |
| Smithy | MkDocsDocumentationProvider |
mkdocs |
100 | โ Lexical | โ |
| CLI | SmithyRemoteDocumentationProvider |
smithy_remote |
0 | โ Brokered through Smithy MCP | โ Brokered through Smithy MCP |
| CLI | MkDocsDocumentationProvider |
mkdocs |
100 | โ Lexical | โ |
Registration:
AddForgeSdkregistersMkDocsDocumentationProviderfor every host.- Smithy adds
AddAiSearchDocumentationon top, which also requires aSearchIndexClientfor the private AI Search endpoint. - The CLI adds
SmithyRemoteDocumentationProvider, which calls Smithy'ssearch_docs/get_docover the bearer-authenticated MCP bridge. Any failure (unconfigured, unreachable, unauthenticated, timeout) yields an empty result and falls through. A failed connection is remembered for a 5-minute cooldown so later calls fall through without re-incurring the connect timeout.
Because ai_search is search-only, full-page retrieval on Smithy always comes from the MkDocs provider.
Provider attribution¶
DocumentationService tags the current activity with three attribution tags. Each is set to the winning provider's Name, or to the sentinel none when no provider produced a result.
| Tag | Set by | Use it for |
|---|---|---|
docs.search.provider |
SearchAsync only |
Which provider answered the search. |
docs.fetch.provider |
GetDocumentAsync only |
Which provider served the full page. |
docs.provider |
Both | Continuity with dashboards and queries that predate the split. |
Prefer the split tags. Search and fetch usually run inside the same ambient activity, because an agent searches and then immediately fetches one of the returned locations. docs.provider is written by both operations, so the fetch value overwrites the search value and the activity ends up reporting only the last writer. docs.search.provider and docs.fetch.provider never collide, so an activity can show ai_search for discovery and mkdocs for retrieval at the same time - which is the normal shape on Smithy, where ai_search is search-only.
Tags ride Activity.Current where one exists. The same value is always written as a structured log property at Debug, so a host without tracing still gets attribution.
Attribution cannot distinguish Smithy's backends
From the CLI the winning provider reports smithy_remote whether Smithy answered from Azure AI Search or fell back to its own MkDocs index. Treat CLI-side provider attribution as coarse routing information, not as a backend breakdown.
๐ฉบ Diagnostics vocabulary¶
Resolution outcomes use a fixed string vocabulary so diagnostics stay comparable across hosts and runs.
| Constant set | Values |
|---|---|
MkDocsResolutionReasons |
resolved, not_found, unknown_site, ambiguous, site_unavailable, fetch_failed |
MkDocsResolutionRules |
exact, variant, canonical-leaf, none |
Each MkDocsLocationResolution carries the requested location, the reason code, the rule that fired, the resolved site and location when it resolved, ambiguity candidates in site/location form, and a human-readable Detail. Log message wording is not part of this contract; the reason and rule strings are.
โ๏ธ MkDocs fallback configuration¶
Bound from the Forge:Documentation configuration section into ForgeDocumentationOptions.
| Key | Default | Description |
|---|---|---|
Origin |
https://docs.saif.com |
Origin hosting every platform docs site. |
DefaultSite |
forge |
Site tried first for a location with no site segment. |
Sites |
forge, cloud-foundations, iac-azure-modules, iac-okta-modules, iac-aws-modules |
Sites published under Origin. |
SearchIndexPath |
/search/search_index.json |
Search index path relative to a site's base URL. |
CacheFileNameTemplate |
{site}-docs-search-index.json |
Per-site cache file name; {site} is the placeholder. When omitted, the site name is prepended, so docs-index.json becomes forge-docs-index.json for Forge. |
CacheDurationHours |
1 |
How long a disk-cached index is treated as fresh on initial load. Does not expire the in-memory index. |
RawMarkdownUrlPattern |
{base}/{path}.md |
Raw markdown URL pattern. null disables raw markdown and always converts HTML to markdown. |
Derived values and behavior:
- A site's base URL is
{Origin}/{site}. There are no per-site base URLs. DefaultSiteis inserted into the site list when configuration omits it, so it is always served.- Site names are de-duplicated case-insensitively, blank entries are dropped, and a
CacheDurationHoursof zero or less falls back to one hour. - Each site has its own index, its own cache file under
%LOCALAPPDATA%\saif-cli\cache\, and its own load gate. Nothing loads at startup; indexes load lazily on first use, and cross-site search loads them concurrently. - A site whose index fails to load (unpublished, unreachable, malformed, or empty) is logged at Debug and marked failed for the process lifetime, then skipped. Other sites are unaffected.
- When a network fetch fails, a stale cache file is used as a fallback if one exists.
- Characters invalid in a file name are replaced with
-in the site-name substitution when building the cache path; the template itself is not sanitized. - A query of exactly
*lists pages instead of term-scoring them, with anchor entries (page/#section) excluded so a page appears once.saif docs listdepends on this. A query that merely contains*, such as* widget, is scored normally. A wildcard listing is not an exhaustive page inventory:DocumentationServicecaps every search at 25 results, sosaif docs listand a*search from any caller return at most 25 pages. - Because every page gets the same score under a wildcard listing, per-site results are interleaved round-robin across loaded sites before the global result limit is applied, so a large default-site index (Forge has thousands of pages) cannot fill the entire limit and hide every other site.
Adding a site¶
Add the site name to Sites. Nothing else is required, provided the site publishes an index at {Origin}/{name}/search/search_index.json. A listed site that does not serve an index costs one round trip per process and is then skipped.
Current publishing status¶
| Site | Index at {Origin}/{site}/search/search_index.json |
|---|---|
forge |
โ Serves |
iac-azure-modules |
โ Serves |
iac-okta-modules |
โ Serves |
cloud-foundations |
โ 404 - root-publishes to docs.saif.com/search/search_index.json |
iac-aws-modules |
โ 404 - not yet published |
Both are expected to move to the standard layout, so they are listed and skipped rather than special-cased. Until then the fallback covers three sites.
๐งญ Location resolution¶
MkDocsSearchService.ResolveLocationAsync accepts both bare paths (Forge pages, whose forge/ prefix upstream providers strip) and site-qualified paths.
Search result locations¶
A location returned by search can be passed straight back to get_doc or saif docs get. This is a location-resolution contract, not a guarantee of page availability, freshness, or agreement between the indexed and fetched revisions.
| Result from | Location shape |
Example |
|---|---|---|
The default site (forge) |
Bare path, unchanged from the index | reference/documentation-retrieval |
| Any other site | {site}/{path} |
iac-azure-modules/reference/naming |
A non-default site's results are qualified because resolution tries DefaultSite first: a bare path that also exists on forge would fetch the wrong page, and a bare path matching several sites would come back ambiguous. Qualifying at search time removes the collision instead of asking the caller to guess. A site's root page qualifies to the bare site name.
Url is always built from the unqualified path, so it keeps pointing at the published page. Use Location to fetch and Url to link.
Default-site results are unchanged
Qualification applies only to non-default sites. Forge results are byte-identical to what search returned before, so existing callers and saved locations keep working.
Site selection¶
- Normalize: trim surrounding slashes and drop any
#fragment. - If the leading segment names a configured site, resolve the remainder against that site. If that misses, the unstripped path is retried against the same site, which covers a site-named top-level directory inside the site itself. If that site's index cannot load, the result is
site_unavailable. - Otherwise resolve against
DefaultSitefirst. Aresolvedorambiguousoutcome there is returned immediately. If the default site's index cannot load, this step is skipped and remembered so step 5 can distinguish an outage from a genuine miss. - Otherwise try every remaining loaded site, keeping both
resolvedandambiguousoutcomes. Exactly one outcome is returned as it stands, so ambiguity found inside a single non-default site is reported rather than discarded. More than one outcome is reportedambiguousacross sites, never picked silently; the candidate list merges each resolved page assite/locationwith the candidates any ambiguous site already qualified. - No match:
unknown_sitewhen the path contains a/, its leading segment contains a hyphen, and that segment is not a top-level directory in any loaded site. Otherwisesite_unavailableif the default site could not load in step 3, since a bare path was never actually checked against it and a miss on the fallback sites alone cannot confirm it is missing. Anything else isnot_found.
Rules within a site¶
Applied in order; matching is case-insensitive.
| Rule | Matches | Ambiguity handling |
|---|---|---|
exact |
The normalized path is an indexed canonical page. | n/a |
variant |
The path after stripping a trailing .md, .html, or .htm, stripping a trailing /index, and mapping a bare index to the site root. |
More than one match returns ambiguous with candidates. |
canonical-leaf |
The final segment matches exactly one canonical page in that site. | More than one canonical page sharing the leaf returns not_found rather than a guess. |
All three lookups are built over canonical pages: anchor entries (page/#section) in the search index normalize onto the page that owns them, and the first entry for a page wins.
Leaf uniqueness is checked within each site, not against one combined index. A bare path that misses on the default site can resolve by leaf on another site; competing matches on multiple non-default sites return ambiguous. A default-site match still wins without consulting other sites. An exact hit is silent; variant and canonical-leaf resolutions log at Information with the requested location's length, site, resolved location, and rule. The raw requested location can carry a token or signed URL, so these events log only its length.
Content is fetched as raw markdown from RawMarkdownUrlPattern first, falling back to the published HTML converted to markdown when raw content is unavailable, empty, or actually HTML. If neither source supplies non-whitespace content, retrieval returns fetch_failed, not a document containing only generated front matter. On a successful fetch the returned document carries the resolved location, URL, and site, alongside the originally requested location, and the markdown is prefixed with source, site, and retrieved front matter.
Limits¶
What resolution cannot recover
- Stale index, removed page. Resolution reads only the in-memory index. A path that still exists in the index resolves
exact, then fails at fetch time withfetch_failed. - Stale index, page still served. A page can return HTTP 200 and non-empty content from a different revision than the search result. The current chain does not detect that skew.
- Renamed leaf. A rename that changes a page's final segment is unrecoverable by leaf matching.
- Unknown site vs. wrong path. Distinguishing an unconfigured site qualifier from a plainly wrong content path is a heuristic (slash present, hyphenated leading segment, not a known top-level directory). A single-word leading segment is always treated as content.
- Default-site short circuit. A bare path that resolves on the default site is returned without consulting other sites, so a same-leaf page elsewhere is never reported as ambiguous.
- Directory restructuring, not just a rename. This consolidation moved pages between top-level directories (for example
guides/development/indextobuild/index). Each moved page carriesmoved_fromfront matter recording its prior paths, and the docs build turns those into HTML redirects on both docs.saif.com/forge/ and forge.saif.com, so a browser following an old URL lands on the new page. Retrieval does not use that mapping:ResolveLocationAsyncdoes not consultmoved_from, and the.mdtwin at an old path is not redirected. A caller passing a pre-move location getsnot_foundby design and should search again to find the current location. - Troubleshooting ID redirects. The docs build also redirects
troubleshoot/<SAIFTRBL ID>/to each article, butResolveLocationAsyncdoes not resolve those paths for the same reason it ignoresmoved_from: redirect stubs are not in the search index. Search for the ID instead; the catalog and each article's frontmatter carry it. - No section-level fetch.
get_docreturns the whole page and drops any#fragmentduring normalization, so it cannot return a single H2 section. #1173 declined section fetch because articles are short and single-symptom, and the AI Search provider is search-only.
Staleness and open revision-verification criterion¶
There is no enforced staleness bound. The indexer schedule affects AI Search freshness, while CacheDurationHours controls only whether MkDocs accepts a disk cache on initial load. MkDocsSearchService keeps a loaded index for the process lifetime and reuses a stale cache file after a network failure, so a long-lived host can serve arbitrarily old results.
Search results and fetched pages carry no comparable source-commit metadata. HTTP 200, non-empty content, and a successful fetch therefore establish availability only, not freshness or revision agreement. The evaluation harness's commitSha identifies the measured CLI build, not the documentation revision.
The same-source-commit verification and detectable-skew acceptance criterion in #1177 remains open, and client-side revision correlation is not the planned next step. Producer-side source-identifier metadata is tracked in saif-corp/cloud-foundations#160. That work stands on its own for build attribution and orphaned-document detection rather than gating this criterion, so treat it as related scope, not a blocker. This change implements neither Forge-side correlation nor explicit skew detection. Periodic index refresh could reduce staleness without new metadata, but it would not prove revision agreement or bound skew to the indexer schedule alone: refresh intervals, publishing delays, and failed refreshes also matter. Periodic refresh is also unimplemented here.
๐ AI Search filter contract¶
AiSearchDocumentationOptions binds from Forge:Documentation:AiSearch.
| Key | Default | Description |
|---|---|---|
IndexName |
saif-docs |
Index over platform docs, populated by cloud-foundations. |
SemanticConfigurationName |
docs-semantic |
Semantic configuration used for ranking. |
PlatformDocsBaseUrl |
https://docs.saif.com |
Base URL that blob-storage paths are rewritten to. |
MkDocsSitePrefix |
forge |
Leading path segment stripped from Location so the MkDocs fallback can resolve it. Empty disables stripping. |
SearchFilter |
(empty - no filter) | OData $filter sent with every search request. Empty or whitespace sends no filter. Smithy configures format eq '.html'; see Code default vs. shipped value. |
The Default column is the C# property initializer in AiSearchDocumentationOptions, which is what any host gets before configuration binds over it.
When SearchFilter holds a non-whitespace value, it is assigned to the request-level SearchOptions.Filter. It is never applied after results return. Azure AI Search evaluates $filter before semantic reranking, so a request-level filter shapes the candidate set that gets reranked. A client-side post-filter would leave the Markdown twin of each page consuming result slots before the trim happens.
Code default vs. shipped value¶
Filtering is off in code and on in Smithy. Those are two different things and both are deliberate.
| Value | Source | |
|---|---|---|
| Code default | (empty) - no filter, every indexed document is a candidate | SearchFilter initializer in src/smithy/src/Smithy/Documentation/AiSearchDocumentationOptions.cs |
| Shipped in Smithy | format eq '.html' - drops the Markdown twin of each indexed page |
Forge:Documentation:AiSearch:SearchFilter in src/smithy/src/Smithy/appsettings.json |
The split is the opt-in boundary. The filter only behaves as measured against an index whose documents actually carry a populated format field, which requires the saif-docs indexer reset. Defaulting to empty means a host pointed at an un-reset index keeps working unchanged, and enabling the filter stays an explicit, per-host decision made once that host's index has been reset.
Standing up another host
Filtering is not on by default. If you configure a host other than Smithy against the saif-docs index, you get no filter until you set Forge:Documentation:AiSearch:SearchFilter yourself, and search can return .md twins alongside rendered pages. Variant resolution now makes these .md locations fetchable when the canonical page exists, but duplicates still consume ranking slots. Confirm .html coverage on the index you are pointing at, then enable the filter to remove those duplicate candidates.
Measured effect of the filter¶
Measured against platform dev after the saif-docs indexer reset, with the same CLI binary either side so only server configuration changed:
| Metric | Filter off | Filter on |
|---|---|---|
| hit@1 | 1/15 (6.7%) | 8/15 (53.3%) |
| hit@5 | 7/15 (46.7%) | 11/15 (73.3%) |
| fetch success | 15/15 (100%) | 13/15 (86.7%) |
Discovery improves sharply because the Markdown twin of a page no longer outranks the rendered page it duplicates. The filter-off hit@1 of 6.7% understates real behaviour: most top results were the .md twin of an acceptable location, and location comparison deliberately does not strip .md. Scored with the suffix stripped, filter-off hit@1 would be 7/15.
Fetch success drops by two cases, and neither is caused by the filter. Both are index-hygiene problems that the filter merely promotes into the top slot: one stale root-published blob and one phantom entry with no page behind it (cloud-foundations#158).
The filter depends on the indexer reset
format is null on every document indexed before the saif-docs indexer reset, so format eq '.html' against an un-reset corpus makes search return almost nothing. The reset has been run and coverage confirmed indirectly by the measurement above - a filtered search returning results at all proves format is populated. Direct verification from a developer machine is still not possible: the AI Search service has public network access disabled (cloud-foundations#77) and listAdminKeys is denied by RBAC.
Any environment whose index has not been reset must leave SearchFilter empty until it has.
Changing or disabling the filter¶
To change the filter, for example to narrow it further:
- Set
Forge:Documentation:AiSearch:SearchFilterin Smithy's configuration to the new OData expression. - Re-run the evaluation harness with a distinct label and compare aggregates against the previous run.
To disable filtering entirely, for example when pointing a host at an index that has not been reset:
- Set
Forge:Documentation:AiSearch:SearchFilterto an empty string. Any whitespace-only value has the same effect. Removing an override restores the value from lower-priority configuration, which in Smithy's shippedappsettings.jsonisformat eq '.html'; only removing the key from all configuration sources reaches the empty code default. - Expect duplicate candidates to consume ranking slots again; the filter-off column above records the measured discovery effect.
saif docs getnow resolves.mdtwins through the variant rule when the canonical page exists, so disabling the filter does not itself make those locations unfetchable.
No code change is required either way. The provider captures IOptions<AiSearchDocumentationOptions>.Value in its constructor, so a configuration change takes effect when the host restarts, not on the next request.
Result previews¶
The preview shown for each AI Search result is a query-centred snippet cut from the document's content field. The indexed description is used only when that snippet comes back empty, with the same whitespace normalization and length limit.
The order matters. On rendered MkDocs pages description is the site-wide meta description, identical on every page, so preferring it renders every result with the same marketing sentence. That stayed hidden while Markdown twins dominated results, because .md blobs carry no description at all - enabling the .html filter would have surfaced it on every result at once. The two changes are coupled and ship together.
Snippet selection splits queries on whitespace, including tabs and line breaks, skips query terms shorter than three characters whenever a longer term is available, and scores candidate windows by how many distinct query terms they contain, preferring the earliest window on a tie. It considers at most the first 20 occurrences of each retained term, so it does not exhaustively search every possible window. Rendered HTML carries navigation chrome, so without the length rule a stop word such as to anchors the snippet on Skip to content instead of the prose the reader asked for.
Text is also whitespace-collapsed before the window is cut. Indexed HTML extracts arrive with long runs of newlines and indentation between fragments of real text, so an uncollapsed 200-character window can spend its whole budget on layout and surface only a handful of words. Collapsing raised the average preview from 136 to 206 visible characters across the evaluation set and removed every near-empty preview.
Preview output contract¶
previewText is normalized, not verbatim source. Consumers that render, compare, or diff it should expect:
- Every run of whitespace collapses to exactly one U+0020 space. The test is
char.IsWhiteSpace, so this covers Unicode whitespace generally - tabs,\r,\n, form feeds, non-breaking and other Unicode spaces - not just ASCII blanks. - Source line breaks are not preserved. A paragraph break, a list, and a code block all flatten into one continuous line. There is no way to recover the original layout from the preview.
- The text is trimmed at both ends. Leading and trailing whitespace is dropped rather than collapsed to a space.
- The character budget uses .NET string length (UTF-16 code units).
maxLength(200 by default) applies after collapsing whitespace, not to grapheme clusters or display width. A preview contains at most 200 code units of normalized source text before ellipses. - Ellipses sit outside the budget. A leading
...is added when the window does not start at the beginning of the document, and a trailing...when it does not reach the end, so a truncated preview is up to 206 characters long.
The normalization is deliberate and stable. Do not treat a preview that differs from the source page's whitespace as a defect, and do not compare previewText byte-for-byte against page content.
Previews inherit whatever the index stores as content
content holds the full text extract of a rendered page, including the navigation tree and the on-page table of contents, so a snippet can open on link labels before reaching prose. Pages with substantial body text recover within the window; thin index pages, whose extract is mostly navigation, do not. Improving that further belongs in the indexer, which would need to extract the page body rather than the whole document. No client-side snippet heuristic can recover prose from a document that contains very little.
Release-notes ranking¶
Release-notes de-emphasis is the same knob โ format eq '.html' and section ne 'release-notes' โ and is deliberately deferred pending evaluation data. It is an editorial judgement, not a wiring gap. A second, independent mechanism also exists: the MkDocs lexical scorer reads an optional per-page boost, emitted by the Material search plugin from a page's search.boost front matter, and multiplies the term score by it. Values that are NaN or negative are clamped to zero.
๐งช Evaluation harness¶
tools/docs-eval/Invoke-DocsRetrievalEval.ps1 replays tools/docs-eval/eval-set.json (14 cases, v2) through the CLI and records a JSON run.
# Measure the CLI on PATH and write a labelled run record
.\tools\docs-eval\Invoke-DocsRetrievalEval.ps1 -Label pre-fix
# Validate the harness offline, without a CLI
.\tools\docs-eval\Invoke-DocsRetrievalEval.ps1 -SelfTest
Parameters¶
| Parameter | Default | Description |
|---|---|---|
-EvalSetPath |
eval-set.json beside the script |
Fixture set to replay. |
-OutputPath |
runs\<yyyyMMdd-HHmmss>-<label>.json beside the script |
Run record destination; the parent directory is created if missing. |
-Label |
unlabeled |
Free-form tag stamped into the record, for example pre-fix and post-fix. |
-SaifPath |
saif |
Executable to measure. Point it at a specific build to measure that tree. |
-CommitSha |
Derived; see below | Commit of the build being measured. When omitted the harness derives it from a SHA embedded in the CLI's informational version, and otherwise records null with a warning. Supply it when the version carries no SHA, even for a binary inside this checkout: its path cannot prove which commit built it. |
-SearchLimit |
10 |
Results requested per search, range 5-50. The shared documentation service caps the effective request at 25. hit@5 always scores the first five results, so a value below 5 is rejected rather than recorded as a mislabeled metric. |
-PassThru |
off | Also emit the run record on the pipeline. |
-SelfTest |
off | Run parsing, scoring, aggregation, the search-limit guard, and commit-attribution assertions against synthetic JSON and exit. Its own parameter set, so it takes no other parameters, and it runs before the search-limit guard. |
What it measures¶
- Runs
saif docs search <query> --limit <SearchLimit> --format json. - Scores discovery:
hitAt1(top result is inacceptableLocations) andhitAt5(any of the first five), plushitRank. - Runs
saif docs get <location> --format jsonon the location search actually returned, not the fixture's expected location. Fetching a known-good fixture path would bypass the search-returns-an-unfetchable-location failure mode the harness exists to catch. - Scores fetch success: exit code 0, parseable JSON, non-empty
content. This does not verify documentation freshness or source-commit agreement. - Records the search snippet raw as
previewTextalongsideexpectedPreviewTerms. Preview quality is never scored; compare it by hand.
Location comparison trims whitespace and slashes and is case-insensitive. It does not strip a .md suffix or an /index segment, because whether those shapes resolve is part of what is being measured; the eval set lists them as separate acceptable locations.
Run record¶
| Field | Contents |
|---|---|
schemaVersion, label, timestamp |
Run identity. schemaVersion is 2. |
commitSha, commitShaSource |
Commit of the measured build, and how it was attributed: parameter (explicit -CommitSha), cli-version (SHA parsed from the CLI's informational version), or unknown (commitSha is null). The executable's path never determines its commit. |
harnessCommitSha |
HEAD of the checkout running the harness, or unknown when git cannot report it. Always the harness, never the measured build. |
cliVersion, cliPath, resolvedCliPath |
What was measured: the reported version, the -SaifPath value as passed, and the executable it resolved to. |
evalSetPath, evalSetVersion, searchLimit, hitAtNDepth |
Harness inputs. |
knownLimitations |
Carried into every run, including the smithy_remote attribution limitation. |
aggregate |
caseCount, hitAt1Count/hitAt1Rate, hitAt5Count/hitAt5Rate, fetchSuccessCount/fetchSuccessRate. |
cases |
Per-case record: locations, hit flags and rank, search and fetch exit codes and errors, provider fields, raw preview. |
Commit attribution in schemaVersion 1 records
Version 1 stamped commitSha with the harness checkout's HEAD unconditionally, whatever binary -SaifPath resolved to. Every run recorded during this change is a version 1 record, and several measured a CLI published outside this checkout, so their commitSha identifies the harness rather than the build. Read those values as harness provenance only. The measurements themselves stand: the build behind the pdev figures was verified through the pipeline rather than through the run record. Version 2 records separate the two fields and refuse to guess.
Read the aggregate, not the exit code
A measurement run exits 0 even when cases fail. Every failure - a non-zero CLI exit, an empty result set, unparseable JSON - is recorded on the case and the run continues. -SelfTest is the exception: it exits 1 when an assertion fails.
Run records under tools/docs-eval/runs/ are git-ignored and are not committed. Eval-set case fixtures name their failure mode: framing-prose, stale-location, cross-site, guide-over-release-notes, and release-notes-intended.
๐ Production signals¶
The harness only runs on demand against a fixed 15-case set, so it cannot tell you that retrieval degraded last Tuesday. Smithy emits the subset of these signals that can be observed without ground truth, under the Smithy meter and activity source, collected by SAIF service defaults over OTLP.
| Signal | Kind | Dimensions | Answers |
|---|---|---|---|
smithy.docs.searches |
Counter | smithy.docs.source, smithy.docs.outcome (ok, empty) |
How often does search match nothing at all? |
smithy.docs.fetches |
Counter | smithy.docs.source, smithy.docs.outcome (ok, not_found) |
How often does a fetch return no non-empty document? Includes direct requests, not just locations returned by search. |
smithy.docs.result_count |
Span tag | - | Result-set size distribution for a given search. |
smithy.docs.min_preview_chars |
Span tag | - | Visible, whitespace-collapsed length of the thinnest snippet in a result set. |
smithy.docs.outcome was added as an extra dimension on the existing smithy.docs.searches counter rather than as a separate metric, so queries that group only by smithy.docs.source keep aggregating exactly as they did before.
The zero-result event in DocumentationService records query length and term count at Information, not raw query text. The search Debug events in ForgeTools.SearchDocsCoreAsync and AiSearchDocumentationProvider.SearchAsync also use only those safe query counts; MCP search_docs uses that shared logging without separately logging the query. The agent's documentation retrieval span also omits the raw query. Documentation queries are free text that can carry pasted tokens or signed URLs, so these application-owned instrumentation points do not emit raw queries even when a host enables Debug logging. This is not a guarantee about framework-level telemetry or other host instrumentation.
What fetch failure detects¶
Before the variant-resolution fix, search-side metrics looked healthy throughout the failure this page documents. Search returned results, including .md twins the old fetch path could not resolve. smithy.docs.fetches with smithy.docs.outcome provides a production availability signal, but it includes all fetches and does not correlate them with prior search hits. The harness measures a different population: the top search result from each fixture. Its fetch-success rate moved from 13.3% to 93.3% once resolution was fixed. Those are historical measurements, not the current behavior of .md locations.
A sustained not_found rate flags an availability problem such as a missing page, broken index, or changed publishing layout. A stale index or page that still produces a successful fetch records ok; this counter cannot detect freshness or revision skew. Those checks remain part of the open revision-verification criterion.
smithy.docs.min_preview_chars exists for the same reason. Three separate preview defects shipped without any signal, because a search that returns the correct page with an unusable snippet still counts as a successful search. The thinnest snippet in a result set is a cheap proxy that catches a whole-result-set regression.
hit@1 and hit@5 cannot be emitted in production
Both require knowing which page should have won, which only the eval set supplies. There is no runtime substitute. Treat the harness as the ranking-quality instrument and these metrics as the availability-and-plausibility instrument; they answer different questions and neither replaces the other.