A pi-ai route had to name an installed catalog provider, served that catalog's models verbatim, and could override only the endpoint. An OpenAI-compatible gateway, a self-hosted server, or a model newer than the pinned pi-ai release was therefore unreachable, and a stale context window could not be corrected without upgrading the package. A route is now a declaration whose defaults come from the installed catalog. `catalog.ts` merges that catalog under the profile's own model entries, `provider.ts` builds the pi-ai Provider (reusing the catalog provider when the route keeps its protocol, so implementations this package cannot reconstruct keep working), and the adapter serves every operation from one `createModels()` collection. That also retires the `@earendil-works/pi-ai/compat` import, which pi-ai documents as a temporary entry point it deletes with its ModelManager migration. Credentials stay on the harness seam: the resolved key rides the request as pi-ai's highest-priority auth override, so `Models` holds no credential store and a named-but-missing reference still fails loud instead of falling back to an unrelated ambient key. A model's configured maxTokens now reaches the seam as defaultMaxTokens.
18 KiB
@deepseek-ai/dsh-llm-pi-ai
English | 中文
Generic multi-provider adapter for the harness LLM seam backed by @earendil-works/pi-ai. One plugin instance owns a dict of provider profiles keyed by route; every request selects a profile with GenerateOptions.provider and resolves GenerateOptions.model against that route's configured catalog. A route naming an installed pi-ai provider inherits its endpoint, wire protocol, and model catalog as defaults and overrides them field by field; a route pi-ai does not ship is declared outright, so an OpenAI-compatible gateway, a self-hosted server, or a provider newer than the installed catalog is configuration rather than a code change.
The package root exposes the Cordis plugin contract, PiAiAdapter, and supportedProtocols(); profile resolution, catalog materialization, provider construction, replay conversion, and stream conversion remain package-internal.
Config
Configure credentials, the model catalog, and deployment-specific transport settings per provider, keyed by the provider route itself. Prefer apiKeyEnv — a credential reference resolved per request — over a literal apiKey, so no secret enters this file. Omitting both is what leaves the route unauthenticated, which for an installed catalog route means pi-ai's provider-native ambient discovery; a configured reference that resolves to nothing fails the request with MISSING_CREDENTIAL instead, because falling through would authenticate with whatever unrelated key the environment happens to hold. One credential serves every model on its route.
- id: llm
name: '@deepseek-ai/dsh-llm-pi-ai'
config:
providers:
# Catalog route: endpoint, protocol, and models all come from pi-ai.
openai:
apiKeyEnv: OPENAI_API_KEY
baseURL: https://proxy.example.com:8443
reasoning: high
retryPolicy:
mode: normal
maxRetries: 3
backoff:
initialDelayMs: 500
maxDelayMs: 10000
jitterRatio: 0.1
# Catalog route with its catalog narrowed to one model and that model's
# capacity corrected; every unset field still comes from the catalog.
anthropic:
apiKeyEnv: ANTHROPIC_API_KEY
streamIdleTimeoutMs: 300000
models:
- id: claude-sonnet-4-5
contextWindow: 200000
# Hand-declared route: pi-ai ships nothing under this key, so the profile
# supplies the whole provider.
acme-gateway:
displayName: Acme Gateway
apiKeyEnv: ACME_GATEWAY_API_KEY
api: openai-completions
baseURL: https://gateway.acme.example/v1
models:
- id: acme-large
name: Acme Large
contextWindow: 65536
maxTokens: 4096
The dict shape makes duplicate routes unrepresentable, and the pre-release array shape (with per-profile provider fields) fails load with migration directions. providers may also be empty or omitted entirely: the adapter then mounts dormant — zero routes, no extra catalog entries — and registers routes the moment the llm-pi-ai: settings section supplies profiles, dropping them again when it empties. Dormant or not, the plugin declares every installed catalog provider in the configurable-provider directory (ctx.llm.listConfigurableProviders(), settings path providers.<provider>), joined with every route the current profiles declare, so configuration surfaces can offer the full catalog before any route exists and can still address a hand-declared one. Which adapters exist is composition; which providers run can be entirely the user's settings document. Registration with ctx.llm is atomic: a collision with any provider route already owned by another adapter fails plugin loading without registering the remaining routes. Model ids are not lifecycle config; a model the route does not configure fails before any provider request with LlmError('UNKNOWN_MODEL').
Catalog resolution
A profile's models list replaces the route's installed catalog rather than extending it; omitting it (or leaving it empty) serves that catalog unchanged. Each entry defaults its unset fields from the installed model of the same id, so narrowing a catalog route to two models, correcting one capacity, or adding a model newer than the installed catalog are all one-line edits. Only the fields the harness consumes are configurable — id, name, contextWindow, maxTokens, and reasoning; pricing and input modalities have no harness consumer and ride the installed entry or are absent, while reasoning-level spellings and OpenAI-compatibility quirks have no configuration surface at all because restating them cannot be validated.
Resolution fails loud, naming the offending route and model, when a route cannot be served: a model the installed catalog does not describe needs an explicit contextWindow and maxTokens, and a route the catalog does not ship needs api, baseURL, and a non-empty models list. api accepts the protocols in supportedProtocols() — pi-ai's own streaming API set — and is only needed when the catalog cannot supply one: a model absent from the catalog inherits the protocol its shipped siblings agree on, so adding a model to a single-protocol catalog route restates nothing.
baseURL sets the endpoint of every model on the route, so private proxies such as https://proxy.example.com:8443 remain supported; a catalog route that omits it keeps each catalog model's own endpoint. Naming api on a catalog route repoints the whole route at that protocol, which is how a deployment moves a provider between, say, Responses and Chat Completions.
Dynamic configuration (settings + credentials)
The adapter reads its profiles through a thunk once per operation instead of freezing them at construction. The plugin registers the llm-pi-ai namespace on the optional ctx.settings seam with this same Config schema and its cordis.yml entry as the composition base, and because providers is a dict, the base and the user's llm-pi-ai: settings section merge per provider: a user can add a route, override one field of a composition route, or point a route at another proxy, all effective on the next request with no restart. Without a mounted settings service the entry config alone drives the adapter, unchanged.
Credentials resolve per stream call: a non-empty literal apiKey wins, then apiKeyEnv through the optional ctx.credentials seam ($DSH_HOME/.env under the live environment; exactly that variable without a mounted seam). A profile naming no credential at all — and only that case — defers to pi-ai's ambient discovery. The route set and each route's captured retry policy are the registration-level facts: when either changes, the plugin replaces its registration atomically (same adapter instance, candidate set validated first), so a route another adapter already owns leaves the previous routes serving and reverting to a working configuration re-applies. Provider key order never counts as a change. A live settings snapshot naming an unknown provider (or failing any other resolver bound) keeps the last good profiles and logs the failure; the entry config itself still fails plugin load.
The adapter exposes each configured route's models through ctx.llm.listModels(provider). This is provider-neutral selector metadata read from the same pi-ai Models collection the request path uses, so discovery does not create a second model registry. ctx.llm.resolveModelInfo(provider, model) performs that exact descriptor lookup once and returns its identity, context window, configured output cap, and selectable thinking levels, keeping authoritative metadata on the route-owning adapter rather than its consumers. A model's maxTokens becomes the seam's defaultMaxTokens, so a request that names no output cap carries the configured one.
The reasoning.efforts list is pi-ai's ordered getSupportedThinkingLevels(model) result without filtering or normalization, including off and the model-specific availability of xhigh or max. The Harness exposes each canonical pi-ai level as an opaque ID; provider/model wire spellings remain inside pi-ai's thinkingLevelMap. A non-reasoning model therefore exposes pi-ai's off choice. The profile reasoning value, including off, is the deployment default when configured; omitting it preserves the provider default. Per-request GenerateOptions.reasoningEffort takes precedence, and any explicit value absent from the exact model capability fails with UNSUPPORTED_REASONING_EFFORT before network I/O instead of being clamped. pi-ai's common stream options represent off by omitting reasoning.
Supported profile fields are apiKey, apiKeyEnv, displayName, api, baseURL, models, headers, reasoning, thinkingBudgets, cacheRetention, transport, timeoutMs, websocketConnectTimeoutMs, streamIdleTimeoutMs, and retryPolicy. Each profile's optional retry policy is captured with that provider route; omission uses bounded normal defaults. The stream-idle interval is a positive finite Node timer delay, defaults to five minutes, and covers only an outstanding provider read, not consumer think time. Harness app attribution wins a conflicting configured header name.
The adapter forces pi-ai's SDK maxRetries to zero so one stream() call makes one provider request. The removed profile fields maxRetries and maxRetryDelayMs fail load instead of silently multiplying or hiding the separately composed agent-level retry budget. Idle expiry aborts the SDK's stable request signal and surfaces TIMEOUT; an earlier caller abort remains ABORTED.
Provider/model routing and replay
Each resolved route contributes one pi-ai Provider to the adapter's createModels() collection, and requests reach the provider through Models.streamSimple(). A catalog route that keeps its catalog protocol reuses the installed provider with its model list replaced, because that provider owns API implementations this package cannot reconstruct — Bedrock loads its Smithy module through a separate entry point — so rebuilding it from parts would silently narrow which providers work. Every other route is built by createProvider() over the protocol table behind supportedProtocols(), whose entries are the same factories pi-ai's own provider factories use.
Credentials never enter that collection. The harness resolves a route's key through its own seam before the request reaches pi-ai and passes it as the request's apiKey option, which pi-ai treats as the highest-priority auth override; Models therefore holds no credential store, and the harness keeps its fail-loud reference semantics. A route naming no credential resolves as configured-but-keyless and leaves the requirement to the protocol, which is where it actually lives.
The selected model descriptor supplies the protocol implementation. This includes native API differences such as OpenAI models whose descriptor uses the Responses API rather than Chat Completions; the harness adapter does not hardcode endpoint selection by model name.
Successful assistant responses store a versioned, lossless-JSON replay state beside their durable provider/model provenance. At request time, LlmService passes replay state only when the historical provider route and target provider route are currently owned by this same PiAiAdapter instance. The adapter validates the state and restores pi-ai response ids and provider signatures even when the target provider or model changes; pi-ai then decides which metadata its target API can reuse. History without replay state is translated as foreign provider-neutral content and never impersonates a native pi-ai response.
If a listener rewrites assembled assistant content, the loop drops replay state before logging the message because its provider metadata no longer describes the content. Invalid versions, malformed metadata, provenance provider/model mismatches, and content/block mismatches fail explicitly with LlmError('INVALID_REPLAY_STATE').
Vocabulary differences
- pi-ai tool-call arguments are parsed objects; the harness stores raw JSON strings. The adapter parses input and re-stringifies output.
- pi-ai reports failures as in-stream error events; these map to
finish {kind:'error'|'aborted', failure}chunks. Provider-specific error text distinguishes terminalQUOTAfrom transientRATE_LIMIT, while text and usage signals evaluated against the resolved model's context window normalize overflow toCONTEXT_WINDOW_EXCEEDED. A terminalstopwhose message carries no content blocks maps to afinish {kind:'error'}with codeEMPTY_RESPONSE(retried by default policy) instead of a successful empty message. - pi-ai folds reasoning tokens into output usage; there is no separate reasoning count to map.
- pi-ai's
offthinking level crosses the Harness capability seam unchanged and becomes an omitted pi-ai commonreasoningoption at dispatch. GenerateOptions.stopis rejected withUNSUPPORTED_OPTIONbecause pi-ai's common streaming surface cannot guarantee it across providers.
App attribution
Every request carries the shared attribution header from dsh-llm's attributionHeaders(), merged through pi-ai's headers stream option. Provider-specific app-attribution headers are not synthesized. See dsh-llm § App attribution.
Dependency weight
pi-ai installs several provider SDKs and lazy-loads the one selected by the catalog model. The dependency weight is isolated to this opt-in adapter package.
Model Experience
Provider request through pi-ai
What the model sees
The selected catalog model receives GenerateOptions.system, history, tools, and sampling fields supported by pi-ai's common streaming API. This package adds no prompt prose. Provider-native replay metadata is restored only when the adapter validates it for the historical content.
Token effect
Provider tokenization governs exact input. Conversion adds no model-visible text; replay metadata may let a native API reuse provider-side state.
KV Cache effect
Conversion preserves logical request order without adding text, while the selected provider's serialization and replay state determine reuse. Changing adapter instance, provider, model, or any upstream request token may prevent reuse from the first difference.
Provider response
What the model sees
pi-ai events become harness reasoning, text, tool-call, usage, and finish chunks. Parsed tool arguments cross the harness boundary as raw JSON strings.
Token effect
Generated content affects later inputs only after the loop records it. pi-ai folds reasoning tokens into output usage when the provider does not report them separately.
KV Cache effect
Recorded response content appends to the next request and does not invalidate its earlier reusable prefix. Unrecorded transport metadata and usage accounting do not affect cache identity.
Known Limitations and Deferred Work
- Settings can add or override routes, not remove composition routes — the user layer merges over the composition
base, so deleting acordis.yml-provided provider is a composition change;replaceon the namespace only resets the user layer. headerscan carry a credential the redactor never sees — the profile'sheadersdict is plain strings, soAuthorizationorapi-keyset there is returned verbatim by a redacteddescribe()and rendered by any configuration UI. Store credentials asapiKeyEnvreferences; making the dict write-only is deferred with the rest of the wire-boundary work.- Model discovery is configuration, not a provider query — the route's catalog is whatever
settings.yamlsays; nothing fetches a provider's/modelsendpoint, so a model list is only as current as its last edit. A one-shot discovery action that offers a provider's live list for the user to adopt belongs to the configuration surface and is deferred with it. - One wire protocol per route —
apiapplies to the whole route, so a mixed-protocol catalog route (an OpenAI-style catalog spanning Responses and Chat Completions) cannot host a model of the other protocol, and adding a model such a route does not describe requires namingapiand moving every model onto it. Splitting the provider across two route keys is the workaround. - An unauthenticated route depends on its protocol — naming no credential resolves the route as configured-but-keyless, but pi-ai's OpenAI-compatible implementation still requires an API key or an
Authorizationheader, so a keyless local server needs a placeholderapiKeyor anAuthorizationentry inheaders. GenerateOptions.stopis unsupported — pi-ai's common stream options cannot guarantee stop-sequence behavior across providers, so the adapter rejects the field.- In-history
systemmessages use pi-ai's common context conversion — provider-specific placement follows pi-ai rather than a harness-owned wire override. - Provider HTTP status is unavailable — pi-ai error events do not expose a stable HTTP status across providers; failures expose only stable harness error codes.
- Retry policy is provider-owned, not an SDK retry — each provider profile may configure nested
retryPolicy, whichdsh-llm-retryexecutes at the agent failed-step seam; pi-ai SDK retries stay disabled so durable agent steps andllm/retryevents own every visible attempt, and directctx.llm.stream()calls remain single-attempt.