fix(llm): answer a catalog route's models from pi-ai's own registry
Clicking "fetch available models" on a built-in provider went to the network. That is the wrong source: pi-ai's registry is the authoritative list for its own providers, and it carries the context windows and output caps a `GET /models` listing does not disclose. Asking api.deepseek.com what DeepSeek serves is both slower and worse, and against an endpoint that answers a different shape it failed outright. Interrogation is still keyed by settings namespace — the provider being added has no route — but the request may now name the route it is editing. An adapter that already describes that route answers from what it knows, needs no endpoint at all, and never touches the network; only a route the catalog does not describe reaches the wire, and one naming no endpoint is told to set one or enter its models by hand. `ConfigurableProviderView` gained `supportsDiscovery` so a surface offers the action where a namespace can answer instead of hardcoding an adapter family. Three narrower corrections ride along. Discovery no longer claims Azure or Codex: Azure authenticates with an `api-key` header and an `api-version` query despite its OpenAI lineage, and Codex uses OAuth, so both reported an authentication failure as a provider with no models. Cancellation during the body read escaped as the raw abort reason rather than a coded ABORTED. And the schema comment claiming the probe key is never logged overstated it: the host neither stores nor returns it, but it rides the client's outgoing envelope like every other secret-bearing payload, and redacting that tap is a configuration-plane-wide change.
This commit is contained in:
@@ -1,12 +1,17 @@
|
||||
/**
|
||||
* One-shot interrogation of a provider endpoint's model listing, serving the
|
||||
* configuration surface's "fetch available models" action.
|
||||
* Answering "which models can this provider serve?" for the configuration
|
||||
* surface's "fetch available models" action.
|
||||
*
|
||||
* This is deliberately *not* a catalog refresh. Nothing here is stored: the
|
||||
* request carries a draft the user is still editing — an endpoint and a
|
||||
* credential neither of which may exist in `settings.yaml` yet — and the reply
|
||||
* is candidate metadata the surface offers for adoption. `settings.yaml`
|
||||
* remains the only thing that decides what a route serves.
|
||||
* A route the installed pi-ai catalog ships is answered **from that catalog**,
|
||||
* with no network call at all: pi-ai's registry is the authoritative list for
|
||||
* its own providers, and it carries the capacities a listing endpoint would
|
||||
* not disclose. Only a route the catalog does not describe — a gateway, a
|
||||
* self-hosted server — is interrogated over the wire.
|
||||
*
|
||||
* Neither path is a catalog refresh. Nothing here is stored: the request
|
||||
* carries a draft the user is still editing, and the reply is candidate
|
||||
* metadata the surface offers for adoption. `settings.yaml` remains the only
|
||||
* thing that decides what a route serves.
|
||||
*
|
||||
* Only OpenAI-compatible protocols are interrogated. Their listing is the one
|
||||
* shape a gateway, a self-hosted server, and the official endpoints all agree
|
||||
@@ -20,16 +25,17 @@
|
||||
import { LlmError } from '@deepseek-ai/dsh-llm'
|
||||
import type { LlmDiscoveredModel, LlmModelDiscoveryRequest } from '@deepseek-ai/dsh-llm'
|
||||
import { attributionHeaders } from '@deepseek-ai/dsh-llm'
|
||||
import { catalogModels } from './catalog.ts'
|
||||
|
||||
/**
|
||||
* Protocols whose model listing this module can read. Every entry speaks
|
||||
* OpenAI's `GET /models` shape; pi-ai's other protocols are absent because a
|
||||
* wrong guess at their response shape would be reported as an empty provider
|
||||
* rather than as the gap it is.
|
||||
* Protocols whose model listing this module can read: the two that speak
|
||||
* OpenAI's `GET /models` shape with bearer auth. Azure is absent despite its
|
||||
* OpenAI lineage — it authenticates with an `api-key` header and requires an
|
||||
* `api-version` query — and Codex authenticates through OAuth; guessing at
|
||||
* either would report an authentication failure as a provider with no models.
|
||||
* pi-ai's remaining protocols are absent for the same reason.
|
||||
*/
|
||||
const LISTABLE_PROTOCOLS: ReadonlySet<string> = new Set([
|
||||
'azure-openai-responses',
|
||||
'openai-codex-responses',
|
||||
'openai-completions',
|
||||
'openai-responses',
|
||||
])
|
||||
@@ -165,6 +171,26 @@ function readListing(body: unknown): LlmDiscoveredModel[] {
|
||||
export async function discoverModels(
|
||||
request: LlmModelDiscoveryRequest,
|
||||
): Promise<readonly LlmDiscoveredModel[]> {
|
||||
// A catalog route already has its answer, and a better one: the installed
|
||||
// entries carry context windows and output caps no listing endpoint reports.
|
||||
if (request.provider !== undefined) {
|
||||
const installed = catalogModels(request.provider)
|
||||
if (installed.size > 0) {
|
||||
return [...installed.values()].map(model => ({
|
||||
id: model.id,
|
||||
name: model.name,
|
||||
contextWindow: model.contextWindow,
|
||||
maxTokens: model.maxTokens,
|
||||
}))
|
||||
}
|
||||
}
|
||||
if (request.baseURL === undefined || request.baseURL.length === 0) {
|
||||
throw new LlmError(
|
||||
`pi-ai ships no catalog for provider "${request.provider ?? ''}", so its models can only come from its`
|
||||
+ " endpoint; set a baseURL, or enter this provider's models by hand",
|
||||
'DISCOVERY_FAILED',
|
||||
)
|
||||
}
|
||||
const api = request.api ?? 'openai-completions'
|
||||
if (!LISTABLE_PROTOCOLS.has(api)) {
|
||||
throw new LlmError(
|
||||
@@ -196,7 +222,18 @@ export async function discoverModels(
|
||||
'DISCOVERY_FAILED',
|
||||
)
|
||||
}
|
||||
const text = await readBounded(response, url)
|
||||
let text: string
|
||||
try {
|
||||
text = await readBounded(response, url)
|
||||
} catch (error: unknown) {
|
||||
// Cancellation during the body read rejects with the abort reason, which
|
||||
// may be any value; the caller gets the same coded failure it would have
|
||||
// for a cancellation before the request went out.
|
||||
if (request.signal?.aborted) {
|
||||
throw new LlmError('model discovery aborted by caller', 'ABORTED', { cause: error })
|
||||
}
|
||||
throw error
|
||||
}
|
||||
let body: unknown
|
||||
try {
|
||||
body = JSON.parse(text)
|
||||
|
||||
@@ -4,6 +4,8 @@ import { afterEach, describe, expect, it } from 'vitest'
|
||||
import { Context } from 'cordis'
|
||||
import LlmService, { userAgent } from '@deepseek-ai/dsh-llm'
|
||||
import * as LlmPiAi from '@deepseek-ai/dsh-llm-pi-ai'
|
||||
import { getBuiltinModels } from '@earendil-works/pi-ai/providers/all'
|
||||
import { discoverModels } from '../src/discovery.ts'
|
||||
|
||||
const servers: Server[] = []
|
||||
|
||||
@@ -25,6 +27,7 @@ async function listingServer(behavior: {
|
||||
status?: number
|
||||
body?: string
|
||||
chunks?: string[]
|
||||
holdOpenMs?: number
|
||||
}): Promise<ListingServer> {
|
||||
const paths: string[] = []
|
||||
const headers: IncomingMessage['headers'][] = []
|
||||
@@ -35,7 +38,10 @@ async function listingServer(behavior: {
|
||||
// No declared length: the ceiling has to hold on what is read.
|
||||
response.writeHead(behavior.status ?? 200, { 'content-type': 'application/json' })
|
||||
for (const chunk of behavior.chunks) response.write(chunk)
|
||||
response.end()
|
||||
if (behavior.holdOpenMs === undefined) { response.end(); return }
|
||||
// Left open so a caller's cancellation lands while the body is still
|
||||
// being read rather than after it completed.
|
||||
setTimeout(() => { response.end() }, behavior.holdOpenMs)
|
||||
return
|
||||
}
|
||||
const body = behavior.body ?? '{}'
|
||||
@@ -60,6 +66,39 @@ async function harness(): Promise<Context> {
|
||||
return ctx
|
||||
}
|
||||
|
||||
describe('catalog-route model discovery', () => {
|
||||
it('answers from the installed registry, with capacities and no network call', async () => {
|
||||
const server = await listingServer({ body: JSON.stringify({ data: [{ id: 'from-the-endpoint' }] }) })
|
||||
const ctx = await harness()
|
||||
|
||||
const models = await ctx.llm.discoverModels('llm-pi-ai', { provider: 'deepseek', baseURL: server.url })
|
||||
|
||||
// pi-ai's own registry is the authority for its own providers, and it
|
||||
// carries what a listing endpoint would not disclose.
|
||||
expect(models.map(model => model.id).sort())
|
||||
.toEqual(getBuiltinModels('deepseek').map(model => model.id).sort())
|
||||
expect(models.every(model => (model.contextWindow ?? 0) > 0 && (model.maxTokens ?? 0) > 0)).toBe(true)
|
||||
expect(server.paths).toEqual([])
|
||||
})
|
||||
|
||||
it('needs no endpoint for a route the catalog describes', async () => {
|
||||
const ctx = await harness()
|
||||
await expect(ctx.llm.discoverModels('llm-pi-ai', { provider: 'deepseek' })).resolves.not.toHaveLength(0)
|
||||
})
|
||||
|
||||
it('says where a route the catalog does not describe must get its models', async () => {
|
||||
const ctx = await harness()
|
||||
await expect(ctx.llm.discoverModels('llm-pi-ai', { provider: 'acme-gateway' }))
|
||||
.rejects.toThrow(/ships no catalog for provider "acme-gateway".*set a baseURL/s)
|
||||
// A form that cleared the field says the same thing as one that never had it.
|
||||
await expect(ctx.llm.discoverModels('llm-pi-ai', { provider: 'acme-gateway', baseURL: '' }))
|
||||
.rejects.toThrow(/set a baseURL/)
|
||||
// The seam refuses a request naming neither, so the module's own guard for
|
||||
// that shape is only reachable by calling it directly.
|
||||
await expect(discoverModels({})).rejects.toThrow(/set a baseURL/)
|
||||
})
|
||||
})
|
||||
|
||||
describe('draft-provider model discovery', () => {
|
||||
it('reads an OpenAI-compatible listing and keeps the capacities it discloses', async () => {
|
||||
const server = await listingServer({
|
||||
@@ -171,12 +210,27 @@ describe('draft-provider model discovery', () => {
|
||||
.rejects.toMatchObject({ code: 'DISCOVERY_FAILED' })
|
||||
})
|
||||
|
||||
it('says which protocols it cannot interrogate rather than guessing a shape', async () => {
|
||||
it.each(['anthropic-messages', 'azure-openai-responses', 'openai-codex-responses', 'google-generative-ai'])(
|
||||
'says it cannot interrogate %s rather than guessing a shape',
|
||||
async (api) => {
|
||||
// Azure authenticates with an `api-key` header and an `api-version`
|
||||
// query despite its OpenAI lineage, and Codex uses OAuth; guessing at
|
||||
// either would report an auth failure as a provider with no models.
|
||||
const ctx = await harness()
|
||||
await expect(ctx.llm.discoverModels('llm-pi-ai', { baseURL: 'https://gateway.example/v1', api }))
|
||||
.rejects.toMatchObject({ code: 'DISCOVERY_UNSUPPORTED' })
|
||||
},
|
||||
)
|
||||
|
||||
it('reports cancellation during the body read as an abort, not a raw reason', async () => {
|
||||
const ctx = await harness()
|
||||
await expect(ctx.llm.discoverModels('llm-pi-ai', {
|
||||
baseURL: 'https://gateway.example/v1',
|
||||
api: 'anthropic-messages',
|
||||
})).rejects.toMatchObject({ code: 'DISCOVERY_UNSUPPORTED' })
|
||||
const controller = new AbortController()
|
||||
// Chunked, so the headers arrive and the cancellation lands mid-body.
|
||||
const slow = await listingServer({ chunks: ['{"data":[', '{"id":"a"}'], holdOpenMs: 400 })
|
||||
const probe = ctx.llm.discoverModels('llm-pi-ai', { baseURL: slow.url, signal: controller.signal })
|
||||
setTimeout(() => { controller.abort('test cancellation') }, 40)
|
||||
|
||||
await expect(probe).rejects.toMatchObject({ code: 'ABORTED' })
|
||||
})
|
||||
|
||||
it('honors caller cancellation', async () => {
|
||||
|
||||
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write packages/llm/llm/README.md
|
||||
README.md: 60cc94b6375030955136b4efaf969b69bca2530a
|
||||
README.zh.md: 5b24a1e311c37d13dc4f287e5ae4b57efe00e4e6
|
||||
README.md: 3a7ec1e8daa33d825fadc15e6481781da48571c4
|
||||
README.zh.md: d5a60a574a7947de83c44df85ce71e16b54be9f4
|
||||
|
||||
@@ -26,7 +26,7 @@ An adapter registry plus a single streaming call surface, interceptable via a wa
|
||||
|
||||
`LlmService` preserves errors from final adapter selection, synchronous dispatch, iterator construction, and iteration, and binds their provenance to the exact stream handle returned for that model call. `isLlmAdapterFailure(stream, value)` reports only errors from that call's final adapter boundary; `llmFailureOf(stream, value)` returns the adjacent immutable `LlmFailure`; `llmRetryPolicyOf(stream)` returns the immutable policy of the exact registration selected at that boundary, even if the route is later disposed or replaced. A call that never reaches a final adapter has no serving policy. Nested model calls, `llm/stream` middleware, and downstream consumer failures remain unclassified for the outer call. Classification never replaces or mutates the adapter's original coded `Error`.
|
||||
|
||||
Interrogating an endpoint is configuration-time work over a *draft*, which is why it is keyed by settings namespace rather than by provider route: the provider a surface is adding does not exist yet, so there is no route to name. The request carries the endpoint, the protocol, and a credential the harness uses for that one interrogation and never stores — nothing here reads or writes settings or credentials, and the reply is candidate metadata a surface may offer for adoption, never a registered catalog. `LlmDiscoveredModel` makes every field but `id` optional because most provider listings disclose an id and nothing else; a surface adopting one still owes the capacities its adapter requires. Duplicate and unusable ids are dropped, an unserved namespace fails with `NO_DISCOVERY`, and an empty namespace or endpoint fails with `INVALID_DISCOVERY`.
|
||||
Interrogating an endpoint is configuration-time work over a *draft*, which is why it is keyed by settings namespace rather than by provider route: the provider a surface is adding does not exist yet, so there is no route to name. The request may still *name* a route it is editing, and an adapter that already describes that route should answer from its own knowledge — better metadata, no network call — which is why `baseURL` is optional and one of the two is required. The request otherwise carries the endpoint, the protocol, and a credential the harness uses for that one interrogation and never stores — nothing here reads or writes settings or credentials, and the reply is candidate metadata a surface may offer for adoption, never a registered catalog. `LlmDiscoveredModel` makes every field but `id` optional because most provider listings disclose an id and nothing else; a surface adopting one still owes the capacities its adapter requires. Duplicate and unusable ids are dropped, an unserved namespace fails with `NO_DISCOVERY`, and a request naming neither a route nor an endpoint fails with `INVALID_DISCOVERY`.
|
||||
|
||||
Provider and model metadata is a discovery surface, not a routing whitelist. `registerAdapter()` still owns provider exclusivity and captures the adapter's retry policy for each route, while an adapter may accept model ids absent from `listModels()`; consumers must not reject a request because its model is unlisted. Returned selector metadata is detached and invalid or duplicate adapter entries fail with `INVALID_ADAPTER` or `INVALID_CATALOG`.
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@
|
||||
|
||||
`LlmService` 保留来自最终适配器选择、同步 dispatch、iterator 构造与迭代的错误,并将其溯源绑定到该次模型调用返回的精确流句柄。`isLlmAdapterFailure(stream, value)` 只报告该调用最终适配器边界的错误;`llmFailureOf(stream, value)` 返回关联的不可变 `LlmFailure`;`llmRetryPolicyOf(stream)` 返回在该边界选中的确切注册所对应的不可变策略,即使之后释放或替换路由也不变。未到达最终适配器的调用没有服务策略。嵌套模型调用、`llm/stream` middleware 和下游消费方失败对外层调用仍未分类。分类绝不替换或更改适配器原有的带代码 `Error`。
|
||||
|
||||
询问端点属于配置期针对**草稿**的操作,因此以 settings namespace 而非提供方路由为键:界面正在新增的提供方还不存在,也就没有路由可点名。请求携带端点、协议,以及一条 harness 只用于这一次询问、绝不存储的凭据——这里既不读也不写 settings 与 credentials,回复是界面可供用户采纳的候选元数据,而不是已注册的 catalog。`LlmDiscoveredModel` 除 `id` 外每个字段都是可选的,因为大多数提供方列表只公布 id;采纳其中一条的界面仍要补上其适配器所需的容量。重复与不可用的 id 会被丢弃,无人服务的 namespace 以 `NO_DISCOVERY` 失败,空 namespace 或空端点以 `INVALID_DISCOVERY` 失败。
|
||||
询问端点属于配置期针对**草稿**的操作,因此以 settings namespace 而非提供方路由为键:界面正在新增的提供方还不存在,也就没有路由可点名。但请求仍可**点名**它正在编辑的路由,而已经描述该路由的适配器应当用自己的知识作答——元数据更好,且无需联网——这正是 `baseURL` 可选、两者必居其一的原因。除此之外,请求携带端点、协议,以及一条 harness 只用于这一次询问、绝不存储的凭据——这里既不读也不写 settings 与 credentials,回复是界面可供用户采纳的候选元数据,而不是已注册的 catalog。`LlmDiscoveredModel` 除 `id` 外每个字段都是可选的,因为大多数提供方列表只公布 id;采纳其中一条的界面仍要补上其适配器所需的容量。重复与不可用的 id 会被丢弃,无人服务的 namespace 以 `NO_DISCOVERY` 失败,既不点名路由也不给端点的请求以 `INVALID_DISCOVERY` 失败。
|
||||
|
||||
提供方与模型元数据是发现接口,不是路由白名单。`registerAdapter()` 仍拥有提供方排他性,并为每条路由捕获适配器的重试策略;适配器则可以接受 `listModels()` 中不存在的模型 id,消费方禁止因模型未列出而拒绝请求。返回的 selector 元数据与输入脱离,无效或重复适配器配置项会以 `INVALID_ADAPTER` 或 `INVALID_CATALOG` 失败。
|
||||
|
||||
|
||||
@@ -517,8 +517,10 @@ export class LlmService extends Service {
|
||||
if (discover === undefined) {
|
||||
throw new LlmError(`no model discovery is registered for "${settingsNs}"`, 'NO_DISCOVERY')
|
||||
}
|
||||
if (request.baseURL.length === 0) {
|
||||
throw new LlmError('model discovery needs a non-empty baseURL', 'INVALID_DISCOVERY')
|
||||
// One of the two identifies what to describe: a route the adapter knows, or
|
||||
// an endpoint to ask. Neither leaves nothing to answer about.
|
||||
if ((request.provider ?? '').length === 0 && (request.baseURL ?? '').length === 0) {
|
||||
throw new LlmError('model discovery needs a provider route or a baseURL', 'INVALID_DISCOVERY')
|
||||
}
|
||||
const discovered = await discover(request)
|
||||
const seen = new Set<string>()
|
||||
|
||||
@@ -146,8 +146,18 @@ export interface LlmConfigurableProvider {
|
||||
* route: a provider being added has no route to name.
|
||||
*/
|
||||
export interface LlmModelDiscoveryRequest {
|
||||
/** Endpoint to interrogate. */
|
||||
baseURL: string
|
||||
/**
|
||||
* Route the draft is editing, when it edits an existing one. A route whose
|
||||
* adapter already knows its models answers from that knowledge instead of
|
||||
* asking the endpoint — the adapter's own registry is the better answer, and
|
||||
* it costs no network call.
|
||||
*/
|
||||
provider?: string
|
||||
/**
|
||||
* Endpoint to interrogate. Optional because a route the adapter already
|
||||
* describes needs none; a route it does not must supply one.
|
||||
*/
|
||||
baseURL?: string
|
||||
/** Wire protocol the endpoint speaks, when the draft names one. */
|
||||
api?: string
|
||||
/** Credential for this interrogation alone; the harness never stores it. */
|
||||
|
||||
@@ -255,5 +255,11 @@ describe('model discovery registry', () => {
|
||||
.rejects.toMatchObject({ code: 'NO_DISCOVERY' })
|
||||
await expect(ctx.llm.discoverModels('llm-example', { baseURL: '' }))
|
||||
.rejects.toMatchObject({ code: 'INVALID_DISCOVERY' })
|
||||
await expect(ctx.llm.discoverModels('llm-example', { provider: '', baseURL: '' }))
|
||||
.rejects.toMatchObject({ code: 'INVALID_DISCOVERY' })
|
||||
await expect(ctx.llm.discoverModels('llm-example', {}))
|
||||
.rejects.toMatchObject({ code: 'INVALID_DISCOVERY' })
|
||||
// Naming a route alone is enough: the adapter may know it without an endpoint.
|
||||
await expect(ctx.llm.discoverModels('llm-example', { provider: 'known-route' })).resolves.toEqual([])
|
||||
})
|
||||
})
|
||||
|
||||
Reference in New Issue
Block a user