Resolve compaction policy per routed model

This commit is contained in:
Yichen Jiang
2026-07-20 15:34:00 +08:00
parent b877ede82d
commit cfa180c127
54 changed files with 1210 additions and 319 deletions

View File

@@ -544,7 +544,7 @@ Waterfall around every streaming model call (retry, replay, routing). Bound to t
- `options` — the full request. A LOOP-built request arrives deep-frozen (mutation throws): its content is a pure function of the session log (the reconstructability Agent Note), so listeners read it, never rewrite it. A hand-built one-shot (compaction summarize) is the caller's own object and stays mutable here.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L43)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L50)
## session/*

View File

@@ -6,7 +6,7 @@
The abstract `llm` service: an adapter registry plus a streaming model-call surface, interceptable via the `llm/stream` waterfall.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L97)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L118)
### ctx.llm.registerAdapter(providers, adapter)
@@ -29,7 +29,7 @@ Register an adapter for the given provider routes. Throws `LlmError` with code `
**Returns** the disposer that unregisters all of them.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L112)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L133)
### ctx.llm.listProviders()
@@ -45,7 +45,7 @@ Describe provider routes with a registered adapter.
**Returns** detached provider metadata in registration order.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L143)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L164)
### ctx.llm.listModels(provider)
@@ -65,7 +65,30 @@ Discover models advertised by one registered provider. Catalog membership is adv
**Returns** detached model metadata in adapter-preferred order.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L153)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L174)
### ctx.llm.resolveModelContext(provider, model)
```ts website-api
/**
* Resolve context capacity from the adapter that owns one exact route.
* This query is independent of the advisory model catalog: an unlisted model
* may return metadata, while `undefined` never rejects later routing.
* @param provider - registered provider route to inspect.
* @param model - exact model id passed to the adapter.
* @returns detached context metadata, or `undefined` when the adapter has none.
*/
async resolveModelContext( provider: string, model: string, ): Promise<LlmModelContext | undefined>
```
Resolve context capacity from the adapter that owns one exact route. This query is independent of the advisory model catalog: an unlisted model may return metadata, while `undefined` never rejects later routing.
- `provider` — registered provider route to inspect.
- `model` — exact model id passed to the adapter.
**Returns** detached context metadata, or `undefined` when the adapter has none.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L209)
### ctx.llm.stream(options)
@@ -91,4 +114,4 @@ Stream one model call as raw chunks (token-level deltas). Throws `LlmError` with
**Returns** the chunk stream, possibly wrapped by `llm/stream` listeners.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L264)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/llm/src/index.ts#L308)

View File

@@ -6,18 +6,7 @@
Replay owner for one service-wide estimator and isolated per-session folds.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L106)
### ctx.tokenMeter.contextWindow
```ts website-api
/** Provider context-window capacity used by pressure consumers. */
readonly contextWindow: number
```
Provider context-window capacity used by pressure consumers.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L112)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L82)
### ctx.tokenMeter.measure(session, requestHeader?)
@@ -50,7 +39,7 @@ Provider usage is reused only when the latest successful call's canonical reques
**Returns** a detached deeply immutable pressure and surface measurement.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L143)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L114)
### ctx.tokenMeter.estimateMessage(message)
@@ -69,4 +58,4 @@ Heuristically price one model-visible message.
**Returns** content and role-framing tokens under the fixed service heuristic.
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L181)
[Source](https://github.com/deepseek-harness/deepseek-harness/blob/master/packages/llm/token-meter/src/index.ts#L152)