Skip to content

Smithy

Smithy is Forge's AI developer assistant. Its MCP server exposes Smithy's platform agent over the Model Context Protocol and also exposes deterministic documentation tools backed by Forge's shared documentation service.

This page documents Smithy's public MCP tool surface:

  • ask
  • search_docs
  • get_doc

For the design rationale behind explicit conversation IDs and MCP structuredContent, see ADR 0014.


ask

Ask Smithy about the Forge developer platform, Azure DevOps pipelines, builds, work items, repositories, or developer workflows. A single platform agent handles the whole turn, including any tool calls it makes, rather than routing to separate specialist agents.

By default Smithy also has live Azure DevOps access: it can query work items, look up builds, list pull requests and pipelines, and search code across projects. Those calls run under the caller's on-behalf-of identity, so results are scoped to what that user can already see; Smithy has no elevated or shared Azure DevOps credentials.

Parameters

Parameter Required Type Description
question Yes string Your question for Smithy.
conversationId No string Optional stable conversation identifier. Omit it for a one-shot question, or reuse the same value to continue a prior thread.

conversationId constraints

Constraint Value
Minimum length 1 character
Maximum length 128 characters
Allowed characters Letters, numbers, -, _
Idle expiration 30 days of inactivity

If conversationId is present but invalid, the tool returns an MCP error.

One-shot and multi-turn behavior

  • One-shot: Omit conversationId. Smithy answers the question and does not read or write conversation history.
  • Multi-turn: Reuse the same conversationId on later calls. Smithy only loads and persists history when the request includes both:
  • a caller-supplied conversationId
  • an authenticated Entra object identifier (oid)
  • Unauthenticated calls: If the caller is not authenticated, Smithy treats the call as one-shot even when conversationId is supplied.
  • Expired or missing conversations: If a supplied conversationId does not resolve to an existing conversation, Smithy starts a new one and returns a note in the structured response.

Response

ask returns the answer in two forms:

  1. Plain text MCP tool content, so text-only clients continue to work unchanged
  2. MCP structuredContent with the envelope below
{
  "answer": "Use the forge MCP tools for read-only CLI lookups before shelling out to saif.",
  "conversationId": "my-stable-id",
  "note": null
}
Field Type Description
answer string Smithy's plain-text answer.
conversationId string or null The active conversation ID when Smithy is actually using persisted conversation state. It is null for one-shot calls and unauthenticated calls.
note string or null Present only when Smithy could not load the requested conversation, such as when it was not found or expired after 30 days of inactivity.

Examples

One-shot call:

{
  "question": "How do I inspect Azure DevOps build logs from Forge?"
}

Multi-turn call:

{
  "question": "Continue from the last answer and show me the exact tool names.",
  "conversationId": "build-log-help"
}

Timeout behavior

ask enforces a wall-clock time budget for the whole turn (default 90 seconds, covering conversation load/history bounding, tool calls, and model round trips, not just the model call). If the turn does not finish inside that budget, the tool returns an MCP tool error instead of leaving the connection hanging:

{
  "error": "timeout",
  "timeoutSeconds": 90,
  "message": "Smithy could not answer within 90 seconds. The turn (conversation storage, history bounding, tool calls, and the model response) did not complete within that time budget. Retry in a minute, or use the deterministic 'search_docs' and 'get_doc' tools directly."
}

This is returned with IsError: true on the MCP result. The budget covers every step of the turn, not just the model call, so a timeout does not by itself identify which dependency was slow. Callers should treat it as a transient latency signal: retry the ask call after a short delay, or fall back to search_docs/get_doc for deterministic lookups that do not depend on the rest of the turn.

Content and iteration limits

Beyond the wall-clock timeout, ask has two other internal limits that shape long or multi-tool answers:

Limit Default Behavior
Documentation content cap 12,000 characters (~3,000 tokens) per page When the agent loop fetches a documentation page internally while answering, content beyond this cap is truncated and a notice is appended. This keeps one oversized page from dominating the turn's token budget, since a tool result is resent on every subsequent round trip.
Tool-call round trips 12 per turn ask stops the agent loop after this many tool-call round trips, independent of the time budget. A turn that is still making progress but needs more round trips than this will end early rather than continuing indefinitely.

Both limits are separate failure modes from the timeout above: a turn can hit either one well before the 90-second budget expires. If an answer looks incomplete or references only part of a long page, or a complex multi-tool question ends abruptly, use search_docs/get_doc directly instead of ask. search_docs returns ranked snippets rather than page content, and get_doc returns the full, uncapped page. Uncapped retrieval applies to resolvable Forge doc pages: the AI Search provider is search-only, so a hit in another platform repository has no full-page fallback and get_doc reports it as not available.

History bounding for multi-turn conversations

Multi-turn conversations (see One-shot and multi-turn behavior) accumulate stored history across turns. To keep that history within the model's context window, ask bounds it deterministically before each turn:

Setting Default Behavior
History budget 76,800 tokens, or 50 stored turns Once loaded history exceeds either limit, the oldest complete user/assistant exchanges are dropped first, one exchange at a time, until both limits are satisfied.
Newest exchange Always kept The most recent exchange is always retained, even if it alone exceeds the token budget, so the model always has its immediately preceding context.

Dropped turns are discarded outright, not summarized: there is no LLM call in this path, and no rolling summary is carried forward. Callers relying on exact recall of very old turns in a long-running conversation should account for this — once an exchange ages out of the window, it is gone. Token counts are estimated (a character-based heuristic), not exact.

These limits apply to the prior history supplied to the model, not to the persisted record and not to the whole prompt. History is bounded when it is loaded, and the newly completed exchange is appended afterward, so a saved conversation can hold 52 turns at the default limit and exceed the token budget by that last exchange. It is bounded again on the next load, so this does not compound. The prompt itself is also larger than the history budget, since the system prompt, the current question, and tool definitions are added on top. One-shot calls do not accumulate history and are unaffected.


search_docs

Searches the Forge platform documentation for relevant content. Use this tool to find documentation about Forge features, guides, and reference material. After finding relevant results, use get_doc to retrieve the full content of a specific document.

Parameters

Parameter Required Type Description
query Yes string Search query, such as terraform modules, authentication, or migration guide.
maxResults No integer Maximum number of results to return. Default 10, maximum 25.

Response

search_docs returns a JSON array. Each array item has this shape:

[
  {
    "Title": "Calling Downstream APIs",
    "Path": "build/apis/calling-apis",
    "Url": "https://docs.saif.com/forge/build/apis/calling-apis/",
    "Snippet": "...clients are generated with Kiota, and every request carries the downstream scope you configured. Register the client...",
    "Score": 0.92
  }
]
Field Type Description
Title string Document title.
Path string Document location path to pass to get_doc.
Url string Public documentation URL, which identifies the site that served the result.
Snippet string Query-centred excerpt cut from the document body. See Snippet behavior.
Score number Search relevance score.

Results can come from more than one documentation site

Retrieval covers configured platform docs sites (forge, iac-azure-modules, iac-okta-modules, and others), not just Forge. Pass Path to get_doc unchanged. A result from the default Forge site returns a bare path such as build/apis/calling-apis; a result from another site keeps a site-qualified path such as iac-azure-modules/release-notes/4.2.0. This preserves routing, not a guarantee that the page is available or agrees with the indexed revision. Do not strip the leading site segment: a bare path resolves against Forge first, so stripping it can fetch the wrong page or report ambiguity. Url identifies the published page. Site attribution rides in Path and Url rather than a separate field. See Documentation retrieval for the full resolution contract and availability limits.

Snippet behavior

Snippet is a window of roughly 200 characters cut from the matched document's body and centred on the query, with leading and trailing ... where the text was truncated. It is not a fixed prefix of the page, so two results for the same page under different queries return different excerpts.

How the window is chosen:

  • Query terms shorter than 3 characters are skipped whenever a longer term is available, so a snippet is not anchored on a stop word such as to.
  • Candidate windows are scored by how many distinct query terms they cover, and the earliest window wins a tie.
  • Whitespace runs are collapsed to single spaces and the text is trimmed before the window is cut, so a snippet does not preserve the source's line breaks or indentation.
  • If no query term matches, the snippet falls back to the leading characters of the document.

On the Smithy AI Search path the indexed description is used only when the body-derived snippet comes back empty. It is a fallback, not the preferred source, because on rendered MkDocs pages description holds the site-wide meta description and is identical on every page.

Thin pages can still preview as navigation text

The index stores the full text extract of a rendered page, including the navigation tree and the on-page table of contents. Pages with substantial body text recover inside the window, but a thin index page whose extract is mostly navigation can preview on link labels. Treat a snippet as a relevance hint and call get_doc for the actual content.


get_doc

Gets the full content of a specific Forge documentation page. Use this after search_docs and pass the location path from search results.

Parameters

Parameter Required Type Description
location Yes string Document location path, such as learn/install-saif-cli or reference/version-compatibility.

Response

Successful lookups return a JSON object with this shape:

{
  "Location": "reference/version-compatibility",
  "Url": "https://docs.saif.com/forge/reference/version-compatibility/",
  "Content": "# Version Compatibility\n..."
}

If the document is not found, get_doc returns:

{
  "NotFound": true,
  "Location": "reference/missing-page",
  "Message": "Document not found at location 'reference/missing-page'. Use 'search_docs' to find available documents."
}
Field Type Description
Location string On success, the resolved location the content was actually retrieved from, relative to the site that served it, which can differ from the location argument. On a not-found response, the location the caller asked for.
Url string Public documentation URL of the resolved page, which identifies the site that served it.
Content string Full markdown content of the document, prefixed with source, site, and retrieved front matter. Never truncated.
NotFound boolean Present only on not-found responses.
Message string Present only on not-found responses.

Resolved locations

location is resolved before the page is fetched, so a request can succeed against a page whose path is not exactly what you passed. Two input shapes are accepted, and both are what search_docs hands you:

  • Bare path, such as reference/version-compatibility. Resolved against the default Forge site first, then against the remaining configured sites.
  • Site-qualified path, such as iac-azure-modules/release-notes/4.2.0. Routed straight to that site. This is what search_docs returns for every non-default site, so pass it back verbatim.

Within each site, an exact indexed path wins first, followed by variants that strip a .md, .html, .htm, or /index suffix, then a unique final-segment match. Leaf uniqueness is checked per site. Bare paths that miss on Forge can resolve by leaf on another configured site; competing matches on multiple non-default sites return ambiguity. A default-site match wins without consulting other sites. Multiple pages sharing a variant key return ambiguity, while multiple pages sharing only a leaf return no match from that site's leaf lookup.

The returned Location is the resolved path relative to the site that served it, so a site-qualified request comes back without its leading site segment. Use Url to confirm which site answered. Compare Location against the path portion of what you sent when you need to know whether a normalization happened; the requested location is not returned as a separate field.

Not every search hit is fetchable

Smithy's AI Search provider is search-only, so full-page retrieval always comes from the MkDocs index. A search hit in another platform repository that the MkDocs index does not cover returns the not-found shape. See Documentation retrieval for the provider chain and the limits of resolution.