diff --git a/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.md b/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.md deleted file mode 100644 index f0a0bdb8ea..0000000000 --- a/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.md +++ /dev/null @@ -1,39 +0,0 @@ -# Agent Note: Host-owned Web session metrics - -Status: implemented - -English | [中文](2026-07-28-host-owned-web-session-metrics.zh.md) - -## Problem - -A Web stats line derived from the currently loaded conversation nodes is window-dependent under pagination. Compaction can replace visible content without preserving historical usage, while the selected model does not prove that a request used its route or capacity. Cache-write tokens also risk being folded into a cache-hit formula whose denominator has different semantics. - -## Decision - -The Host owns one session-level metrics projection. It incrementally folds the complete durable event log, keys settled usage by `(turn, step)`, and replaces an earlier usage record for the same key instead of double-counting chunk and message forms. Uncached input, output, cache reads, and cache writes remain four disjoint cumulative buckets. Compaction can change the current prompt surface without erasing historical usage. - -Current context pressure is the point-in-time `tokenMeter.measure(session).totalTokens`. Capacity instead belongs to the latest model request attempt observed by the current live mux connection. `LlmService.prepareCall()` retains the context metadata obtained by the exact lookup that also validates reasoning/defaults. After the final provider/model is fixed and the outer `llm/stream` call returns a handle, the loop publishes one contained `agent/model-request` notification. This boundary observes an attempt, not proof of provider I/O: preparation or a synchronous outer waterfall failure emits nothing, while short-circuit handles and later lazy adapter construction, iteration failure, or abort still count. - -The tail `session.history` response carries durable usage and pressure, while older pages omit them. Live changes use `session/metrics` mux frames. Both forms carry a durable-log revision and a projection revision; the client accepts only nondecreasing revisions and preserves metrics across older-page prepend. - -ApiProxy forwards each notification as a distinct `session/model-request` frame only to mux connections already open when dispatch occurs. It never places the frame in `session.history` or a subscription baseline. The client keeps durable `metrics` and transient `modelRequestContextWindow` as separate snapshot fields, replaces or explicitly clears the capacity on the next observed request, and clears both fields on `session/subscribed`; reconnect, restore, and a new subscription therefore start unknown until another request is observed. - -The Web stats line joins the durable projection and live capacity only at presentation. It renders uncached input, output, and cache reads separately, computes cache hit as `cacheRead / (uncachedInput + cacheRead)`, and shows current context as a percentage only when the current connection observed a capacity. Cache writes never enter that percentage. Visible nodes continue to supply only turn and step counts. - -## Alternatives considered - -**Fold the loaded node window in React.** This cannot survive pagination or compaction and duplicates durable-log semantics in a presentation package. - -**Send usage only with raw assistant events.** Reconnect and older-page stitching would still need the client to reconstruct a full-log aggregate, and duplicate usage forms would need protocol-specific repair there. - -**Reuse one total-token field for cache hit.** Cache reads, cache writes, and uncached input represent distinct provider accounting buckets; combining them would make the displayed rate misleading. - -**Query the selected route before dispatch.** Selection may never produce a request, and a second metadata lookup can race the registration-bound lookup that actually validates and dispatches the call. - -**Persist or replay the latest request capacity.** That would make a former request look current on reconnect or restore even though the new connection observed no request. The denominator is deliberately live and opportunistic. - -## Consequences - -Token totals remain stable across pagination, replay, compaction, and browser reconnect. The client stores a small detached durable projection plus one connection-local denominator instead of scanning the conversation window, and the status row remains readable for large histories through compact number formatting. - -The Host performs one incremental log fold per session and schedules durable projection updates only for usage, request-header, or surface-changing events; text and reasoning deltas do not publish metrics. A new connection omits the percentage until it observes a request with context metadata. A later request without metadata clears the denominator, while deployments without a token meter still retain the durable counters and label context unavailable instead of fabricating pressure. diff --git a/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.zh.md b/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.zh.md deleted file mode 100644 index 4535fd8a4b..0000000000 --- a/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.zh.md +++ /dev/null @@ -1,39 +0,0 @@ -# Agent Note: Host 拥有的 Web 会话指标 - -Status: implemented - -[English](2026-07-28-host-owned-web-session-metrics.md) | 中文 - -## 问题 - -Web 统计行若根据当前加载的会话节点推导指标,其结果会随分页窗口变化。压缩(compaction)可以替换可见内容,却无法保留历史用量;所选模型也不能证明某次请求实际采用了该模型的路由或容量。缓存写入 token 还可能被计入缓存命中率公式,而该公式的分母具有不同语义。 - -## 决策 - -Host 拥有一项会话级指标投影。它以增量方式归并完整的持久事件日志,按 `(turn, step)` 标识已结算用量;同一标识再次出现时,会替换较早的用量记录,而不会重复统计分片和消息两种形态。未缓存输入、输出、缓存读取与缓存写入保持为四个彼此独立的累计计数项。压缩可以改变当前提示词表层,但不会抹除历史用量。 - -当前上下文压力是即时的 `tokenMeter.measure(session).totalTokens`。容量则属于当前实时 mux 连接观察到的最新模型请求尝试。`LlmService.prepareCall()` 会保留同一次精确查询取得的上下文元数据,该查询也负责校验推理设置与默认值。最终提供方/模型确定且外层 `llm/stream` 调用返回句柄后,循环会发布一条 `agent/model-request` 通知,并收容该通知的失败。这个边界观察到的是一次尝试,并不能证明提供方 I/O 已开始:准备阶段或外层 waterfall(瀑布式事件)的同步失败不会发出通知,而短路句柄以及之后的惰性适配器构造、迭代失败或中止仍会计入。 - -`session.history` 尾页响应携带持久用量与压力,较早页面则省略这两项。实时变更使用 `session/metrics` mux 帧。两种形式都携带持久日志修订号和投影修订号;客户端只接受不减小的修订号,并在向前加载较早页面时保留指标。 - -ApiProxy 只把每条通知作为独立的 `session/model-request` 帧转发给分派发生时已经打开的 mux 连接。它绝不会把该帧放入 `session.history` 或订阅基线。客户端把持久 `metrics` 与临时 `modelRequestContextWindow` 保存在彼此独立的快照字段中,在观察到下一次请求时替换或显式清除容量,并在收到 `session/subscribed` 时清除这两个字段;因此,重连、恢复和新订阅都会从未知容量开始,直到观察到另一次请求。 - -Web 统计行只在展示时结合持久投影与实时容量。它分别呈现未缓存输入、输出与缓存读取,通过 `cacheRead / (uncachedInput + cacheRead)` 计算缓存命中率,并且只有当前连接观察到容量时,才把当前上下文显示为该容量的百分比。缓存写入绝不计入缓存命中率。可见节点仍然只提供轮次和步骤计数。 - -## 备选方案 - -**在 React 中归并已加载的节点窗口。** 此方案无法跨越分页或压缩保留数据,还会在展示包中重复实现持久日志语义。 - -**只随原始 assistant 事件发送用量。** 重连和较早页面拼接仍会要求客户端重建完整日志聚合,而且重复的用量形态需要在客户端按协议专门修复。 - -**为缓存命中率复用单一的 token 总数字段。** 缓存读取、缓存写入与未缓存输入是提供方记账中的不同计数项;将它们合并会使显示的比率产生误导。 - -**在分派前查询所选路由。** 选择操作可能永远不会产生请求;第二次元数据查询还可能与实际校验并分派调用的、绑定注册项的查询发生竞态。 - -**持久化或回放最新请求的容量。** 即使新连接没有观察到任何请求,这也会让先前请求在重连或恢复后显得仍然有效。该分母刻意只采用实时且恰好可得的数据。 - -## 后果 - -token 总量在分页、回放、压缩和浏览器重连期间保持稳定。客户端存储一项小型、脱耦的持久投影与一个连接本地分母,无需扫描会话窗口;状态行采用紧凑数字格式,因此在较长的历史记录中仍然清晰易读。 - -Host 为每个会话执行一次增量日志归并,仅为用量事件、请求头事件或表层变更事件调度持久投影更新;文本与推理(reasoning)增量不会发布指标。新连接在观察到带上下文元数据的请求之前不会显示百分比。后续不带元数据的请求会清除该分母;未部署 token 计量器时,系统仍保留持久计数器,并把上下文标示为不可用,而不会虚构压力值。 diff --git a/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.i18n.yaml b/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.i18n.yaml similarity index 52% rename from .agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.i18n.yaml rename to .agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.i18n.yaml index 8c72912377..6aa9498b9c 100644 --- a/.agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.i18n.yaml +++ b/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.i18n.yaml @@ -1,6 +1,6 @@ # Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each # side as of the last confirmed-consistent state. Both languages carry equal authority; # after editing either side, bring the other along and re-record with: -# pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-28-host-owned-web-session-metrics.md -2026-07-28-host-owned-web-session-metrics.md: f0a0bdb8ea4c1ba1f983d016c3cfde9c67a0424d -2026-07-28-host-owned-web-session-metrics.zh.md: 4535fd8a4b760aa031c11ab766b6be583b9b740d +# pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.md +2026-07-29-projected-token-usage-and-request-context.md: 3ac8c29f7752828f2d833359293c7b1c513d75ca +2026-07-29-projected-token-usage-and-request-context.zh.md: 02b03d1516a39b7727b214b63a9039623b69c56b diff --git a/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.md b/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.md new file mode 100644 index 0000000000..3ac8c29f77 --- /dev/null +++ b/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.md @@ -0,0 +1,45 @@ +# Agent Note: Projected token usage and request context + +Status: implemented + +English | [中文](2026-07-29-projected-token-usage-and-request-context.zh.md) + +## Problem + +A Web stats line derived from the currently loaded conversation nodes is window-dependent under pagination. Compaction can replace visible content without preserving historical usage. Conversely, context occupancy describes one real request boundary: a selected model is only an intention, and combining token pressure from one moment with capacity resolved for another route creates a false percentage. + +These two values therefore have different lifetimes. Provider-reported billing is durable, replayable session state. Request pressure and registration-bound capacity are an opportunistic live observation that must disappear across a connection generation. + +## Decision + +`@deepseek-ai/dsh-token-meter` registers the generic `tokenUsage` session projection when `ctx.sessionProjections` is present. The projection folds the complete durable log into uncached input, output, cache-read, and cache-write buckets. An `assistant/chunk` usage sample survives a later failed request; an `assistant/message` usage value replaces the earlier value for the same `(turn, step)` instead of being counted twice. Reasoning tokens remain an output subdivision and are not added again. Compaction and surface replacement do not erase earlier billing. + +The projection uses the standard projection lifecycle and wire path. History tail baselines, `session/projection` live frames, higher-seq-wins client storage, JSON checkpoints, cache recovery, and unit unload all remain generic. There is no token-specific history field, mux frame, projector, revision counter, or client fence. + +`LlmService.prepareCall()` retains context metadata from the exact lookup that also validates reasoning and captures the adapter registration. After the outer stream call returns its handle and before iteration begins, AgentLoop emits one contained `agent/model-request` notification. Preparation or a synchronous outer waterfall failure emits nothing; short-circuit handles and later iterator construction, iteration, or abort failures still count as an observed request attempt. + +ApiProxy handles that notification synchronously. It reads `tokenMeter.measure(agent.session).totalTokens` once when the optional service is present and combines the result with the same prepared call's registration-bound `contextWindow`. It broadcasts one atomic `session/model-request` frame containing the route, turn, step, and whichever of `contextTokens` and `contextWindow` are available. Measurement failure omits only the numerator. The frame goes only to mux connections already open at that instant; history, subscription baselines, reconnect, and restore never replay it. + +The client stores the complete latest request frame as `ConversationSnapshot.modelRequest`. Every later frame replaces the entire snapshot, so omitted fields clear earlier values. `SessionManager` temporarily holds a pre-instantiation frame, while a new subscription generation, disconnect, or session removal clears both resident and pending values. Model selection alone does not change this snapshot. + +The Web `StatsLine` reads `tokenUsage` through the standard `useProjection` hook and reads request telemetry plus visible nodes through `useSession`. It renders uncached input, output, and cache reads separately, computes cache hit as `cacheRead / (uncachedInput + cacheRead)`, and shows context occupancy only when one request snapshot contains both numerator and capacity. Visible nodes continue to supply only turn and step counts. The existing inline text UI is retained; the model selector gains no circle or other accessory. + +## Alternatives considered + +**A custom session metrics history field and mux frame.** This duplicated the generic projection protocol, cache, recovery, and seq fencing while coupling durable billing to transient request pressure. + +**Fold the loaded node window in React.** This cannot survive pagination or compaction and makes a presentation package reconstruct log semantics. + +**Publish usage only with final assistant messages.** A request that reports a usage chunk and then fails would lose provider billing. + +**Query capacity from the selected model.** Selection may never produce a request, and a second metadata lookup can disagree with the registration-bound lookup used by the actual call. + +**Persist or replay the latest request snapshot.** A request from a prior connection would appear current after restore even though no new request was observed. + +**Add a context circle beside the model selector.** That placement suggests selected-model state. The existing stats line expresses the request-scoped semantics without introducing a duplicate UI or data path. + +## Consequences + +Token totals stay stable across pagination, compaction, replay, and reconnect because they are ordinary durable projection state. Context occupancy is deliberately unknown after reconnect until a new real request is observed. Deployments without token-meter or without model capacity still publish the request route and clear stale optional fields instead of fabricating a percentage. + +ApiProxy performs one synchronous optional measurement and one frame conversion per observed request. It owns no per-session metrics cache or refresh queue. The browser keeps one generic projection value plus one small connection-local request snapshot, and streaming text deltas do not force the stats line to recompute. diff --git a/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.zh.md b/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.zh.md new file mode 100644 index 0000000000..02b03d1516 --- /dev/null +++ b/.agents/notes/implemented/architecture/2026-07-29-projected-token-usage-and-request-context.zh.md @@ -0,0 +1,45 @@ +# Agent Note:token 用量投影与请求上下文 + +Status: implemented + +[English](2026-07-29-projected-token-usage-and-request-context.md) | 中文 + +## 问题 + +Web 统计行若根据当前已加载的会话节点推导,其结果会在分页时依赖当前窗口。压缩(compaction)可以替换可见内容,却不保留历史用量。另一方面,上下文占用率描述的是一个真实请求边界:所选模型只代表意图;若把某一时刻的 token 压力与另一路由解析出的容量组合,就会产生虚假百分比。 + +因此,这两个值具有不同的生命周期。提供方报告的计费用量属于持久、可回放的会话状态。请求压力和与注册项绑定的容量则是恰好可得的实时观测,跨连接代次时必须消失。 + +## 决策 + +当 `ctx.sessionProjections` 存在时,`@deepseek-ai/dsh-token-meter` 会注册通用的 `tokenUsage` 会话投影。该投影将完整持久日志归并为未缓存输入、输出、缓存读取和缓存写入四类计数项。即使后续请求失败,`assistant/chunk` 用量样本仍会保留;同一 `(turn, step)` 的 `assistant/message` 用量值会替换先前值,不会重复计数。推理(reasoning)token 仍是输出的细分项,不会再次累加。压缩和表层替换不会抹除先前的计费用量。 + +该投影使用标准的投影生命周期与协议路径。历史尾页基线、`session/projection` 实时帧、seq 高者胜的客户端存储、JSON 检查点、缓存恢复和单元卸载均保持通用机制。系统没有任何 token 专用的历史字段、mux 帧、投影器、修订计数器或客户端 seq 防护机制。 + +`LlmService.prepareCall()` 会保留精确查询得到的上下文元数据;同一次查询还会校验推理强度并捕获适配器注册项。外层流调用返回句柄后、开始迭代前,AgentLoop 会发出一条失败受收容的 `agent/model-request` 通知。准备阶段或外层 waterfall(瀑布式事件)的同步失败不会发出通知;短路句柄以及之后的迭代器构造失败、迭代失败或中止仍算作一次已观测的请求尝试。 + +ApiProxy 会同步处理该通知。当可选服务存在时,它会读取一次 `tokenMeter.measure(agent.session).totalTokens`,并将结果与同一准备完成调用中绑定注册项的 `contextWindow` 合并。它会广播一个原子 `session/model-request` 帧,其中包含路由、轮次、步骤,以及可用的 `contextTokens`/`contextWindow` 字段。测量失败时只省略分子。该帧只发送给当时已经打开的 mux 连接;历史记录、订阅基线、重连和恢复都绝不回放该帧。 + +客户端将最新的完整请求帧存储为 `ConversationSnapshot.modelRequest`。每个后续帧都会替换整个快照,因此省略字段会清除先前值。`SessionManager` 会临时保存一个实例化前帧;新订阅代次、断开连接或移除会话时,则会同时清除常驻值和待处理值。仅选择模型不会改变该快照。 + +Web `StatsLine` 通过标准 `useProjection` 钩子读取 `tokenUsage`,并通过 `useSession` 读取请求观测数据与可见节点。它分别显示未缓存输入、输出与缓存读取,通过 `cacheRead / (uncachedInput + cacheRead)` 计算缓存命中率,并且只有同一份请求快照同时包含分子与容量时才显示上下文占用率。可见节点仍只提供轮次和步骤计数。系统保留现有的行内文本 UI;模型选择器不增加圆环或其他附属控件。 + +## 备选方案 + +**自定义会话指标历史字段和 mux 帧。** 这会重复实现通用投影协议、缓存、恢复和 seq 防护机制,并将持久计费用量与临时请求压力耦合。 + +**在 React 中归并已加载的节点窗口。** 此方案无法跨分页或压缩保留数据,还会迫使展示包重建日志语义。 + +**仅随最终 assistant 消息发布用量。** 如果请求报告一个用量分片后失败,就会丢失提供方计费用量。 + +**从所选模型查询容量。** 选择操作可能永远不会产生请求;第二次元数据查询还可能与实际调用所用的、绑定注册项的查询不一致。 + +**持久化或回放最新请求快照。** 即使没有观察到任何新请求,先前连接的请求也会在恢复后显得仍是当前请求。 + +**在模型选择器旁增加上下文圆环。** 该位置会让人以为这是所选模型的状态。现有统计行可以表达按请求作用域的语义,无需引入重复的 UI 或数据路径。 + +## 后果 + +token 总量在分页、压缩、回放和重连期间保持稳定,因为它们属于普通的持久投影状态。重连后,上下文占用率会刻意保持未知,直到系统观察到新的真实请求。未部署 token-meter 或模型不提供容量时,系统仍会发布请求路由,并清除陈旧的可选字段,而不会虚构百分比。 + +ApiProxy 会为每次已观测请求执行一次可选的同步测量和一次帧转换。它不拥有任何逐会话指标缓存或刷新队列。浏览器只保留一个通用投影值和一个小型连接本地请求快照;流式文本增量不会迫使统计行重新计算。 diff --git a/apps/web/tests/question-composer.e2e.ts b/apps/web/tests/question-composer.e2e.ts index 6ecdb683b3..8689a74fc2 100644 --- a/apps/web/tests/question-composer.e2e.ts +++ b/apps/web/tests/question-composer.e2e.ts @@ -104,7 +104,7 @@ describe('web e2e: resident question composer round trip', () => { const inner = child.getBoundingClientRect() return Math.max(box.top - inner.top, inner.bottom - box.bottom) }))) - const list = rows[0]?.parentElement ?? null + const list = card.querySelector('[data-question-scroll]') return { rows: rows.length, spill: Math.max(...spill), diff --git a/apps/web/tests/snapshots/code-mode-round/ui.expected.md b/apps/web/tests/snapshots/code-mode-round/ui.expected.md index 4847924619..dccd3fd7dc 100644 --- a/apps/web/tests/snapshots/code-mode-round/ui.expected.md +++ b/apps/web/tests/snapshots/code-mode-round/ui.expected.md @@ -12,6 +12,7 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - 'button "Think The user wants me to write a single `run_code` program that:"': - img - img @@ -28,7 +29,7 @@ - img - text: Think The program ran successfully. Let me now reply DONE as instructed. - paragraph: DONE -- text: cache hit 52% · 17,490 tokens · 1 turns · 2 steps +- text: 8.3k uncached input · 252 output · 9k cache read · cache hit 52% · context 7% of 128k · 1 turns · 2 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/cordis-tool-round/ui.expected.md b/apps/web/tests/snapshots/cordis-tool-round/ui.expected.md index e5e5626be3..b9c1ea7b2d 100644 --- a/apps/web/tests/snapshots/cordis-tool-round/ui.expected.md +++ b/apps/web/tests/snapshots/cordis-tool-round/ui.expected.md @@ -12,6 +12,7 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - button "Think The user wants me to:": - img - img @@ -42,7 +43,7 @@ - img - text: Think All three calls succeeded. I should now reply exactly "CORDIS_UI_DONE" and stop. - paragraph: CORDIS_UI_DONE -- text: cache hit 77% · 66,813 tokens · 1 turns · 4 steps +- text: 15.3k uncached input · 312 output · 51.2k cache read · cache hit 77% · context 13% of 128k · 1 turns · 4 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/lifecycle-chrome/reloaded.expected.md b/apps/web/tests/snapshots/lifecycle-chrome/reloaded.expected.md index 33d1f7e6bf..1a0b740ce6 100644 --- a/apps/web/tests/snapshots/lifecycle-chrome/reloaded.expected.md +++ b/apps/web/tests/snapshots/lifecycle-chrome/reloaded.expected.md @@ -12,12 +12,13 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - button "Think The user wants me to reply with a single word. Let me comply.": - img - img - text: Think The user wants me to reply with a single word. Let me comply. - paragraph: LIGHTHOUSE -- text: cache hit 99% · 7,810 tokens · 1 turns · 1 steps +- text: 109 uncached input · 21 output · 7.7k cache read · cache hit 99% · context unknown · 1 turns · 1 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/live-interactions/cancel.expected.md b/apps/web/tests/snapshots/live-interactions/cancel.expected.md index 3d092b17ec..bf87981ea2 100644 --- a/apps/web/tests/snapshots/live-interactions/cancel.expected.md +++ b/apps/web/tests/snapshots/live-interactions/cancel.expected.md @@ -12,8 +12,9 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - paragraph: partial -- text: 已停止 0 tokens · 1 turns · 1 steps +- text: 已停止 0 uncached input · 0 output · 0 cache read · context 4% of 128k · 1 turns · 1 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/live-interactions/error-auth.expected.md b/apps/web/tests/snapshots/live-interactions/error-auth.expected.md index 5272bcf2d1..9fe1195192 100644 --- a/apps/web/tests/snapshots/live-interactions/error-auth.expected.md +++ b/apps/web/tests/snapshots/live-interactions/error-auth.expected.md @@ -12,6 +12,8 @@ - img - button "编辑": - img +- button "▸ 上下文注入" +- text: 0 uncached input · 0 output · 0 cache read · context 4% of 128k · 0 turns · 0 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/live-interactions/retry.expected.md b/apps/web/tests/snapshots/live-interactions/retry.expected.md index 5935872557..6ad6d27487 100644 --- a/apps/web/tests/snapshots/live-interactions/retry.expected.md +++ b/apps/web/tests/snapshots/live-interactions/retry.expected.md @@ -12,12 +12,13 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - button "Think The user is asking for a one-sentence description of event sourcing. This is a straightforward knowledge question that doesn't require any skill loading or tool calls.": - img - img - text: Think The user is asking for a one-sentence description of event sourcing. This is a straightforward knowledge question that doesn't require any skill loading or tool calls. - paragraph: Event sourcing is a pattern where all changes to an application's state are stored as an immutable, append-only sequence of events, rather than persisting only the current state, enabling full auditability, temporal queries, and event-driven architectures. -- text: cache hit 99% · 7,869 tokens · 1 turns · 1 steps +- text: 110 uncached input · 79 output · 7.7k cache read · cache hit 99% · context 4% of 128k · 1 turns · 1 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/question-composer/answered.expected.md b/apps/web/tests/snapshots/question-composer/answered.expected.md index 91ff2cdf88..a243c38c48 100644 --- a/apps/web/tests/snapshots/question-composer/answered.expected.md +++ b/apps/web/tests/snapshots/question-composer/answered.expected.md @@ -12,6 +12,7 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - button "Think The user wants me to use the ask_user_question tool with specific parameters. Let me do exactly that.": - img - img @@ -25,7 +26,7 @@ - img - text: Think The user answered "Blue". I should now reply with the single word DONE and stop. - paragraph: DONE -- text: cache hit 95% · 8,769 tokens · 1 turns · 2 steps +- text: 397 uncached input · 180 output · 8.2k cache read · cache hit 95% · context 4% of 128k · 1 turns · 2 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/apps/web/tests/snapshots/seeded-history/ui.expected.md b/apps/web/tests/snapshots/seeded-history/ui.expected.md index 4f9181f702..ac9c92dbfd 100644 --- a/apps/web/tests/snapshots/seeded-history/ui.expected.md +++ b/apps/web/tests/snapshots/seeded-history/ui.expected.md @@ -1,32 +1,41 @@ - banner: - navigation "Session hierarchy": - button "Use the read tool twice" [disabled] - - text: · 1 turns - tablist: - tab "Chat" [selected] - tab "Trajectory" - tab "Waterfall" - text: "Use the read tool twice in one assistant message: read a.txt and b.txt. Then reply with the single word DONE and stop." +- button "复制": + - img +- button "在新对话中分支": + - img +- button "编辑": + - img - button "Think The user wants me to read a.txt and b.txt, then reply with \"DONE\". Let me do both reads in parallel.": + - img - img - text: Think The user wants me to read a.txt and b.txt, then reply with "DONE". Let me do both reads in parallel. -- button: - - img -- text: Read a.txt -- button: - - img -- text: Read b.txt +- img +- text: Read +- button "a.txt" +- img +- text: Read +- button "b.txt" - button "Think Both files have been read. a.txt contains \"alpha\" and b.txt contains \"beta\". I'll now reply with DONE as instructed.": + - img - img - text: Think Both files have been read. a.txt contains "alpha" and b.txt contains "beta". I'll now reply with DONE as instructed. - paragraph: DONE -- text: cache hit 98% · 15,962 tokens · 1 turns · 2 steps +- text: 339 uncached input · 135 output · 15.5k cache read · cache hit 98% · context unknown · 1 turns · 2 steps - textbox "Message the agent" - button "Add attachment": - img +- text: Danger Full Access - combobox "Access mode": - - option "Read-only" [selected] - - option "Read-write" + - option "Read Only" + - option "Workspace Write" + - option "Danger Full Access" [selected] - button "选择模型,当前 deepseek-v4-flash": - text: deepseek-v4-flash - img diff --git a/apps/web/tests/snapshots/steering/mid-steer.expected.md b/apps/web/tests/snapshots/steering/mid-steer.expected.md index 8d33ea6283..bbfae558c9 100644 --- a/apps/web/tests/snapshots/steering/mid-steer.expected.md +++ b/apps/web/tests/snapshots/steering/mid-steer.expected.md @@ -12,6 +12,7 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - button "Think The user wants me to use the ask_user_question tool to ask them a specific question with the given parameters. Let me do exactly that.": - img - img @@ -19,9 +20,7 @@ - button: - img - img -- text: "Tool call ask_user_question · {\"questions\": [{\"id\": \"checkpoint\", \"question\": \"Ready to continue?\", \"header\": \"Checkpoint\", \"options\": [{\"label\": \"Yes\"}, {\"label\": \"No\"}]}]} 等待回答(1 题)" -- button "▸ 问题内容" -- text: cache hit 98% · 7,946 tokens · 1 turns · 1 steps +- text: "Tool call ask_user_question · {\"questions\": [{\"id\": \"checkpoint\", \"question\": \"Ready to continue?\", \"header\": \"Checkpoint\", \"options\": [{\"label\": \"Yes\"}, {\"label\": \"No\"}]}]} 151 uncached input · 115 output · 7.7k cache read · cache hit 98% · context 4% of 128k · 1 turns · 1 steps" - region "Ready to continue?": - text: Checkpoint - heading "Ready to continue?" [level=2] diff --git a/apps/web/tests/snapshots/steering/settled.expected.md b/apps/web/tests/snapshots/steering/settled.expected.md index f08fc518e8..f6ed68511a 100644 --- a/apps/web/tests/snapshots/steering/settled.expected.md +++ b/apps/web/tests/snapshots/steering/settled.expected.md @@ -12,6 +12,7 @@ - img - button "编辑": - img +- button "▸ 上下文注入" - button "Think The user wants me to use the ask_user_question tool to ask them a specific question with the given parameters. Let me do exactly that.": - img - img @@ -25,7 +26,7 @@ - img - text: Think The user selected "Yes" and wants me to include the word "BANANA" in my final reply. Let me acknowledge their answer. - paragraph: Great, let's move forward. BANANA! -- text: cache hit 98% · 15,967 tokens · 1 turns · 2 steps +- text: 323 uncached input · 156 output · 15.5k cache read · cache hit 98% · context 6% of 128k · 1 turns · 2 steps - textbox "Message the agent" - button "Add attachment": - img diff --git a/docs/config-catalog.md b/docs/config-catalog.md index 3c0543e4ca..1026e1478b 100644 --- a/docs/config-catalog.md +++ b/docs/config-catalog.md @@ -1548,7 +1548,7 @@ Source: [`packages/context/time-context/src/index.ts:20`](../packages/context/ti export type TokenMeterConfig = Record ``` -Source: [`packages/llm/token-meter/src/types.ts:10`](../packages/llm/token-meter/src/types.ts) +Source: [`packages/llm/token-meter/src/types.ts:12`](../packages/llm/token-meter/src/types.ts) ## `@deepseek-ai/dsh-tool-bash` diff --git a/docs/cordis-catalog/services.md b/docs/cordis-catalog/services.md index 31574e0eb4..3ab05fc0a1 100644 --- a/docs/cordis-catalog/services.md +++ b/docs/cordis-catalog/services.md @@ -2031,7 +2031,7 @@ estimateMessage(message: Message): number Types: [EpochHeader](../core-data-structures/session.md) · [Message](../core-data-structures/core.md) · [Session](../core-data-structures/session.md) · [TokenMeasurement](../core-data-structures/token-meter.md) -Source: [`packages/llm/token-meter/src/index.ts:82`](../../packages/llm/token-meter/src/index.ts) +Source: [`packages/llm/token-meter/src/index.ts:85`](../../packages/llm/token-meter/src/index.ts) ## `ctx.toolResultPrune` — `ToolResultPruneService` diff --git a/docs/event-producer-consumer.md b/docs/event-producer-consumer.md index 1c36786cce..391059ec2a 100644 --- a/docs/event-producer-consumer.md +++ b/docs/event-producer-consumer.md @@ -9,7 +9,7 @@ This matrix shows which packages dispatch each harness-owned event and which pac | --- | --- | --- | --- | --- | | `agent-loop/config-start-failed` | `emit` | [`packages/core/agent-loop/src/index.ts:148`](../packages/core/agent-loop/src/index.ts) | [`agent-loop`](../packages/core/agent-loop) (`events.dispatch`) | [`tui`](../packages/ui/tui) | | `agent/cancel-requested` | `emit` | [`packages/core/agent/src/types.ts:296`](../packages/core/agent/src/types.ts) | [`agent-loop`](../packages/core/agent-loop) (`emitAgentEvent`) | [`goal-session`](../packages/goal/goal-session) | -| `agent/created` | `emit` | [`packages/core/agent/src/types.ts:228`](../packages/core/agent/src/types.ts) | [`agent`](../packages/core/agent) (`events.dispatch`) | `apiproxy`, [`goal-session`](../packages/goal/goal-session), [`tui`](../packages/ui/tui) | +| `agent/created` | `emit` | [`packages/core/agent/src/types.ts:228`](../packages/core/agent/src/types.ts) | [`agent`](../packages/core/agent) (`events.dispatch`) | [`goal-session`](../packages/goal/goal-session), [`tui`](../packages/ui/tui) | | `agent/disposed` | `emit` | [`packages/core/agent/src/types.ts:237`](../packages/core/agent/src/types.ts) | [`agent`](../packages/core/agent) (`events.dispatch`) | [`agent-loop`](../packages/core/agent-loop), [`goal-session`](../packages/goal/goal-session), [`tui`](../packages/ui/tui) | | `agent/error` | `emit` | [`packages/core/agent/src/types.ts:424`](../packages/core/agent/src/types.ts) | [`agent-loop`](../packages/core/agent-loop) (`emitAgentEvent`) | `apiproxy`, [`goal-session`](../packages/goal/goal-session), [`session-telemetry`](../packages/telemetry/session-telemetry), [`tui`](../packages/ui/tui) | | `agent/inbox/dequeue` | `emit` | [`packages/core/agent/src/types.ts:269`](../packages/core/agent/src/types.ts) | [`agent-loop`](../packages/core/agent-loop) (`emitAgentEvent`) | [`agent`](../packages/core/agent), `apiproxy`, [`tui`](../packages/ui/tui) | diff --git a/packages/client/connection/src/client/api.ts b/packages/client/connection/src/client/api.ts index 05ad4325ff..a773269257 100644 --- a/packages/client/connection/src/client/api.ts +++ b/packages/client/connection/src/client/api.ts @@ -12,7 +12,7 @@ export type { WorkspaceApi, WorkspaceId, WorkspaceView, CommandsApi, CommandDescriptor, SkillsApi, SkillEntry, ModelCatalogFailure, ModelCatalogModel, ModelProviderGroup, ModelReasoning, - ModelReasoningEffort, ModelTarget, SessionMetrics, SessionModels, SessionProjectionsBlock, + ModelReasoningEffort, ModelRequestTelemetry, ModelTarget, SessionModels, SessionProjectionsBlock, GoalsApi, GoalRef, } from '@deepseek-ai/dsh-host-apiproxy/api' export type { ToolCallView, ToolResultView } from '@deepseek-ai/dsh-tools/presentation' diff --git a/packages/client/connection/src/client/fixture.ts b/packages/client/connection/src/client/fixture.ts index cab616aabd..0843341191 100644 --- a/packages/client/connection/src/client/fixture.ts +++ b/packages/client/connection/src/client/fixture.ts @@ -15,6 +15,7 @@ import type { AssistantMessage, ContentBlock, MessageSource, + TokenUsage, ToolResultMessage, UserMessage, } from '@deepseek-ai/dsh-llm' @@ -103,6 +104,16 @@ function sid(id: string): SessionId { return id as SessionId } +/** Deterministic provider billing attached to fixture assistant messages. */ +function fixtureUsage(turn: number, step: number): TokenUsage { + return { + inputTokens: 20 + turn % 5, + outputTokens: 8 + step, + cacheReadTokens: turn === 0 ? 0 : 80, + cacheWriteTokens: turn % 10 === 0 ? 4 : 0, + } +} + /** fx-alpha history script: 60 turns (~130+ messages -> 3 pages at PAGE_MESSAGES=50), * mixing reasoning blocks / tool call+result / steering / context. */ function buildAlphaLog(): SessionEvent[] { @@ -110,7 +121,17 @@ function buildAlphaLog(): SessionEvent[] { let time = Date.now() - 3_600_000 const push = (e: Record): number => { const seq = events.length - events.push({ seq, time: (time += 800), ...e }) + const data = e['data'] as Record | undefined + const authored = e['type'] === 'assistant/message' && data !== undefined + ? { + ...e, + data: { + ...data, + usage: fixtureUsage(data['turn'] as number, data['step'] as number), + }, + } + : e + events.push({ seq, time: (time += 800), ...authored }) return seq } for (let turn = 0; turn < 60; turn++) { @@ -369,6 +390,60 @@ function permissionSelectOf( } } +interface FixtureTokenUsageProjection { + uncachedInputTokens: number + outputTokens: number + cacheReadTokens: number + cacheWriteTokens: number +} + +/** Fixture parallel of token-meter's last-sample-replacing usage projection. */ +function tokenUsageOf(log: readonly SessionEvent[]): FixtureTokenUsageProjection { + const totals: FixtureTokenUsageProjection = { + uncachedInputTokens: 0, + outputTokens: 0, + cacheReadTokens: 0, + cacheWriteTokens: 0, + } + let last: { + turn: number + step: number + buckets: FixtureTokenUsageProjection + } | null = null + for (const event of log) { + const item = event as unknown as { + type: string + data: { + turn?: number + step?: number + usage?: TokenUsage + chunk?: { type?: string; usage?: TokenUsage } + } + } + const usage = item.type === 'assistant/chunk' && item.data.chunk?.type === 'usage' + ? item.data.chunk.usage + : item.type === 'assistant/message' + ? item.data.usage + : undefined + if (usage === undefined || item.data.turn === undefined || item.data.step === undefined) continue + const buckets: FixtureTokenUsageProjection = { + uncachedInputTokens: usage.inputTokens, + outputTokens: usage.outputTokens, + cacheReadTokens: usage.cacheReadTokens ?? 0, + cacheWriteTokens: usage.cacheWriteTokens ?? 0, + } + const previous = last?.turn === item.data.turn && last.step === item.data.step + ? last.buckets + : undefined + totals.uncachedInputTokens += buckets.uncachedInputTokens - (previous?.uncachedInputTokens ?? 0) + totals.outputTokens += buckets.outputTokens - (previous?.outputTokens ?? 0) + totals.cacheReadTokens += buckets.cacheReadTokens - (previous?.cacheReadTokens ?? 0) + totals.cacheWriteTokens += buckets.cacheWriteTokens - (previous?.cacheWriteTokens ?? 0) + last = { turn: item.data.turn, step: item.data.step, buckets } + } + return totals +} + function projectionValuesOf(log: readonly SessionEvent[]): Record { const values: Record = {} const titleEvent = log.findLast(item => (item as { type: string }).type === 'session/title') @@ -383,12 +458,28 @@ function projectionValuesOf(log: readonly SessionEvent[]): Record[] { const type = (event as { type: string }).type + if ( + (type === 'assistant/chunk' + && (event as unknown as { data: { chunk?: { type?: string } } }).data.chunk?.type === 'usage') + || (type === 'assistant/message' + && (event as unknown as { data: { usage?: TokenUsage } }).data.usage !== undefined) + ) { + return [{ + type: 'session/projection', + sessionId: id, + key: 'tokenUsage', + value: tokenUsageOf(log), + seq: event.seq, + }] + } if (type === 'session/title') { const values = projectionValuesOf(log) /* v8 ignore next -- the advancing title event is in the log, so the key is present. */ @@ -853,7 +944,16 @@ export function createFixtureApi(options: FixtureOptions = {}): ApiProxy { replays.delete(id) const done = pieces.slice(0, i).join('') append(id, { type: 'assistant/chunk', data: { turn, step, chunk: { type: 'block-end', index: 0, block: { type: 'text', text: done } } } }) - append(id, { type: 'assistant/message', surfaceOp: 'append', data: { turn, step, message: assistantMessage(text(aborted ? `${done}(已中断)` : done)) } }) + append(id, { + type: 'assistant/message', + surfaceOp: 'append', + data: { + turn, + step, + message: assistantMessage(text(aborted ? `${done}(已中断)` : done)), + usage: fixtureUsage(turn, step), + }, + }) append(id, { type: 'step/end', data: { turn, step } }) append(id, { type: 'turn/end', data: { turn, reason: { kind: aborted ? 'cancelled' : 'completed' } } }) setRunning(id, false) @@ -1036,6 +1136,19 @@ export function createFixtureApi(options: FixtureOptions = {}): ApiProxy { append(id, { type: 'plan/mode', data: { active: plan.wanted } }) } append(id, { type: 'user/message', surfaceOp: 'append', data: userMessage(content) }) + const target = modelTargets.get(id) ?? { provider: 'deepseek', model: 'deepseek-v4-flash' } + const usage = tokenUsageOf(logOf(id)) + emitMux({ + type: 'session/model-request', + sessionId: id, + turn, + step: 0, + provider: target.provider, + model: target.model, + contextTokens: usage.uncachedInputTokens + usage.outputTokens + + usage.cacheReadTokens + usage.cacheWriteTokens, + contextWindow: 128_000, + }) startReply( id, turn, diff --git a/packages/client/connection/src/client/index.ts b/packages/client/connection/src/client/index.ts index 29cb02c6fa..4097a23036 100644 --- a/packages/client/connection/src/client/index.ts +++ b/packages/client/connection/src/client/index.ts @@ -17,7 +17,7 @@ export type { ToolCallView, ToolResultView, WorkspaceApi, WorkspaceId, WorkspaceView, CommandsApi, CommandDescriptor, SkillsApi, SkillEntry, ModelCatalogFailure, ModelCatalogModel, ModelProviderGroup, ModelReasoning, - ModelReasoningEffort, ModelTarget, SessionMetrics, SessionModels, SessionProjectionsBlock, + ModelReasoningEffort, ModelRequestTelemetry, ModelTarget, SessionModels, SessionProjectionsBlock, RpcRequest, RpcResponse, RpcResult, RpcError, RpcErrorCode, ClientRequest, ServerResponse, ServerRequest, ClientResponse, RpcMessage, RpcReceipt, IApiClient, SessionId, SessionEvent, ContentBlock, StreamChunk, diff --git a/packages/client/connection/tests/fixture.spec.ts b/packages/client/connection/tests/fixture.spec.ts index 09a3efecd3..d25a4974ca 100644 --- a/packages/client/connection/tests/fixture.spec.ts +++ b/packages/client/connection/tests/fixture.spec.ts @@ -84,6 +84,12 @@ describe('createFixtureApi', () => { }, plan: { active: false, pending: false }, goal: null, + tokenUsage: { + uncachedInputTokens: 0, + outputTokens: 0, + cacheReadTokens: 0, + cacheWriteTokens: 0, + }, } }, }) }) @@ -188,6 +194,20 @@ describe('createFixtureApi', () => { expect(types).toContain('assistant/chunk') expect(types).toContain('assistant/message') expect(types.at(-1)).toBe('turn/end') + expect(frames).toContainEqual({ + type: 'session/model-request', + sessionId: id, + turn: 0, + step: 0, + provider: 'deepseek', + model: 'deepseek-v4-flash', + contextTokens: 0, + contextWindow: 128_000, + }) + expect(frames.some(frame => + frame.type === 'session/projection' + && frame.key === 'tokenUsage' + && (frame.value as { outputTokens?: number }).outputTokens === 8)).toBe(true) const finalize = frames.find((f): f is Extract => f.type === 'session/event' && f.event.type === 'assistant/message') expect(JSON.stringify(finalize?.event.data)).toContain('(已中断)') // Idle cancel: no replay in flight, must not explode; running flips false. @@ -219,7 +239,7 @@ describe('createFixtureApi', () => { const envelopes: RpcRequest[] = [] for await (const envelope of api.events.mux(req({}), abort.signal)) { envelopes.push(envelope) - if (envelopes.length >= 8) abort.abort() + if (envelopes.length >= 9) abort.abort() } return envelopes } @@ -227,16 +247,18 @@ describe('createFixtureApi', () => { const second = await openOnce() expect(first[0]?.payload).toMatchObject({ type: 'session/subscribed', sessionId: 'fx-alpha' }) expect((first[0]?.payload as { lastSeq: number }).lastSeq).toBeGreaterThan(0) - // Projection baseline frames follow the subscribed frame (title + todos + permissions + plan + goal units). + // Projection baseline frames follow subscribed (domain units + token usage). expect(first[1]?.payload).toMatchObject({ type: 'session/projection', sessionId: 'fx-alpha', key: 'title', value: 'Fixture 历史会话' }) expect(first[2]?.payload).toMatchObject({ type: 'session/projection', sessionId: 'fx-alpha', key: 'todos' }) expect(first[3]?.payload).toMatchObject({ type: 'session/projection', sessionId: 'fx-alpha', key: 'permissions' }) expect(first[4]?.payload).toMatchObject({ type: 'session/projection', sessionId: 'fx-alpha', key: 'plan', value: { active: false, pending: false } }) expect(first[5]?.payload).toMatchObject({ type: 'session/projection', sessionId: 'fx-alpha', key: 'goal', value: null }) - expect(first[6]?.payload).toMatchObject({ type: 'approval/requested', toolName: 'dangerous_tool' }) - expect(second[6]?.rpcId).toBe(first[6]?.rpcId) // stable rpcId across replays (host replay semantics) - expect(first[7]?.payload).toMatchObject({ type: 'question/requested', sessionId: 'fx-alpha' }) - expect(second[7]?.rpcId).toBe(first[7]?.rpcId) + expect(first[6]?.payload).toMatchObject({ type: 'session/projection', sessionId: 'fx-alpha', key: 'tokenUsage' }) + expect(first[7]?.payload).toMatchObject({ type: 'approval/requested', toolName: 'dangerous_tool' }) + expect(second[7]?.rpcId).toBe(first[7]?.rpcId) // stable rpcId across replays (host replay semantics) + expect(first[8]?.payload).toMatchObject({ type: 'question/requested', sessionId: 'fx-alpha' }) + expect(second[8]?.rpcId).toBe(first[8]?.rpcId) + expect(first.some(envelope => envelope.payload.type === 'session/model-request')).toBe(false) }) it('steer with no replay in flight falls through to a fresh queued turn; non-text blocks stringify empty', async () => { diff --git a/packages/client/runtime/README.i18n.yaml b/packages/client/runtime/README.i18n.yaml index e74dacded2..9f33a7e7ae 100644 --- a/packages/client/runtime/README.i18n.yaml +++ b/packages/client/runtime/README.i18n.yaml @@ -2,5 +2,5 @@ # side as of the last confirmed-consistent state. Both languages carry equal authority; # after editing either side, bring the other along and re-record with: # pnpm run verify-translation-pairing --write packages/client/runtime/README.md -README.md: 879730733f3d89c3962d8c54bcfd53795a980049 -README.zh.md: 1b1d7f03d1f5f03c054dfeaa790a9f6f91e0dca2 +README.md: ba9a7d455e8a193f23884411eb1928a10f21ddd0 +README.zh.md: 873fefca48585efed010589917f2c63653b08e5b diff --git a/packages/client/runtime/README.md b/packages/client/runtime/README.md index 879730733f..ba9a7d455e 100644 --- a/packages/client/runtime/README.md +++ b/packages/client/runtime/README.md @@ -2,7 +2,7 @@ English | [中文](README.zh.md) -Client cordis boot and React-free object services: SlotsService wraps SlotCore and supplies renderer data sources; SessionsService owns Session objects, list/scope/history state; WorkspacesService depends on SessionsService and owns Workspace objects, list/actions, default-target derivation, and the New Session blank-reuse entry (`connectWorkspace`). The runtime fans the shared Host stream into both managers. Client sessions are always Host-born (Session+Agent+cwd in one `session.create`); the client holds no pre-entity session state — a session's Agent scope (the client mirror of host dsh-scope, keyed by the shared agent/session id) is born when its row enters the list mirror and dies with the prune. Contract: api-contracts v3 §4. Each `Session` holds a generic `ProjectionValueStore` seeded from the history-tail `projections` block and updated by `session/projection` frames under higher-seq-wins; domain keys (including `todos` and `title`) are read via `projections.faceOf` / `useProjection`, not via `ConversationSnapshot`. Durable `ConversationSnapshot.metrics` instead comes from the separate history-tail value and live `session/metrics` frames because point-in-time token-meter pressure can advance at the same durable log revision; only nondecreasing log and projection revisions are accepted. `ConversationSnapshot.modelRequestContextWindow` separately retains capacity from the latest `session/model-request` observed on the current mux connection. A later request replaces or clears that value, while `session/subscribed` clears both metrics ordering and capacity; reconnect, restore, and a new subscription therefore show no percentage until another request is observed. Missing metrics remain `null` rather than being inferred from the visible node window. +Client cordis boot and React-free object services: SlotsService wraps SlotCore and supplies renderer data sources; SessionsService owns Session objects, list/scope/history state; WorkspacesService depends on SessionsService and owns Workspace objects, list/actions, default-target derivation, and the New Session blank-reuse entry (`connectWorkspace`). The runtime fans the shared Host stream into both managers. Client sessions are always Host-born (Session+Agent+cwd in one `session.create`); the client holds no pre-entity session state — a session's Agent scope (the client mirror of host dsh-scope, keyed by the shared agent/session id) is born when its row enters the list mirror and dies with the prune. Contract: api-contracts v3 §4. Each `Session` holds a generic `ProjectionValueStore` seeded from the history-tail `projections` block and updated by `session/projection` frames under higher-seq-wins; domain keys (including `todos`, `title`, and `tokenUsage`) are read via `projections.faceOf` / `useProjection`, not via `ConversationSnapshot`. `ConversationSnapshot.modelRequest` separately retains the complete latest `session/model-request` observed on the current mux connection. Each frame replaces the whole snapshot, so omitted numerator or capacity fields clear an earlier value. `SessionManager` buffers one pre-instantiation snapshot, while `session/subscribed`, disconnect, and removal clear resident and pending values; reconnect, restore, and a new subscription therefore show no context percentage until another request is observed. Model selection alone does not alter request telemetry. ## Workspace and Session lists diff --git a/packages/client/runtime/README.zh.md b/packages/client/runtime/README.zh.md index 1b1d7f03d1..873fefca48 100644 --- a/packages/client/runtime/README.zh.md +++ b/packages/client/runtime/README.zh.md @@ -2,7 +2,7 @@ [English](README.md) | 中文 -客户端 cordis 启动与不依赖 React 的对象服务:SlotsService 包装 SlotCore 并提供 renderer 数据源;SessionsService 拥有 Session 对象、列表/scope/history 状态;WorkspacesService 依赖 SessionsService,拥有 Workspace 对象、列表/操作、默认目标派生,以及 New Session 空会话复用入口(`connectWorkspace`)。运行时把共享 Host 流分发给两个 manager。客户端 Session 一律由 Host 出生(一次 `session.create` 同瞬产出 Session+Agent+cwd);客户端不持有任何实体化之前的会话状态——Agent scope(host dsh-scope 的客户端镜像,以 agent/session 共用 id 为键)在会话行进入列表镜像时出生,随 prune 死亡。契约:api-contracts v3 §4。每个 `Session` 持有一个通用的 `ProjectionValueStore`,由历史尾页的 `projections` 块播种,并经 `session/projection` 帧按 seq 高者胜更新;领域键(含 `todos` 与 `title`)经 `projections.faceOf`/`useProjection` 读取,不经 `ConversationSnapshot`。持久的 `ConversationSnapshot.metrics` 则来自独立的 history 尾页值与实时 `session/metrics` 帧,因为即时 token-meter 压力可以在相同持久日志修订号上继续变化;客户端只接受日志修订号与投影修订号均不减小的数据。`ConversationSnapshot.modelRequestContextWindow` 另行保留当前 mux 连接观察到的最新 `session/model-request` 容量。后续请求会替换或清除该值,`session/subscribed` 则同时清除指标顺序状态与容量;因此,重连、恢复和新订阅都不会显示百分比,直到观察到另一次请求。缺失的 metrics 保持为 `null`,而不是根据可见节点窗口推断。 +客户端 cordis 启动与不依赖 React 的对象服务:SlotsService 包装 SlotCore 并提供 renderer 数据源;SessionsService 拥有 Session 对象、列表/scope/history 状态;WorkspacesService 依赖 SessionsService,拥有 Workspace 对象、列表/操作、默认目标派生,以及 New Session 空会话复用入口(`connectWorkspace`)。运行时把共享 Host 流分发给两个 manager。客户端 Session 一律由 Host 出生(一次 `session.create` 同瞬产出 Session+Agent+cwd);客户端不持有任何实体化之前的会话状态——Agent scope(host dsh-scope 的客户端镜像,以 agent/session 共用 id 为键)在会话行进入列表镜像时出生,随 prune 死亡。契约:api-contracts v3 §4。每个 `Session` 持有一个通用的 `ProjectionValueStore`,由历史尾页的 `projections` 块播种,并经 `session/projection` 帧按 seq 高者胜更新;领域键(含 `todos`、`title` 与 `tokenUsage`)经 `projections.faceOf`/`useProjection` 读取,不经 `ConversationSnapshot`。`ConversationSnapshot.modelRequest` 另行保留当前 mux 连接观察到的最新完整 `session/model-request`。每个帧都会替换整个快照,因此分子或容量字段一旦缺失,就会清除先前值。`SessionManager` 会缓冲一个实例化前快照;`session/subscribed`、断开连接和移除会话则会清除常驻值与待处理值;因此,重连、恢复和新订阅都不会显示上下文百分比,直到观察到另一次请求。仅选择模型不会改变请求观测数据。 ## Workspace 与 Session 列表 diff --git a/packages/client/runtime/src/client/sessions/conversation.ts b/packages/client/runtime/src/client/sessions/conversation.ts index 78ec9e2cd8..f211b209fd 100644 --- a/packages/client/runtime/src/client/sessions/conversation.ts +++ b/packages/client/runtime/src/client/sessions/conversation.ts @@ -6,7 +6,7 @@ import type { CommandId } from '@deepseek-ai/dsh-commands/brand' import type { ContentBlock } from '@deepseek-ai/dsh-llm/types' import type { - RpcError, SessionId, SessionMetrics, ToolCallView, ToolResultView, + ModelRequestTelemetry, RpcError, SessionId, ToolCallView, ToolResultView, } from '@deepseek-ai/dsh-client-connection/client' import type { PendingInteraction } from './pending.ts' @@ -267,16 +267,6 @@ export interface ConversationSnapshot { */ blank: boolean lastAgentError: string | null - /** - * Host-owned cumulative usage/current pressure. Independent of `nodes` - * pagination; null until a tail response or live metrics frame supplies a - * current durable value. - */ - metrics: SessionMetrics | null - /** - * Capacity from the latest model-request attempt observed on this mux - * generation. Absent before the first such request, after a request whose - * registration exposes no capacity, and after `session/subscribed`. - */ - modelRequestContextWindow?: number + /** Latest atomic model-request snapshot on this mux generation. */ + modelRequest: ModelRequestTelemetry | null } diff --git a/packages/client/runtime/src/client/sessions/manager.ts b/packages/client/runtime/src/client/sessions/manager.ts index ff22ca92db..d53c9c4e94 100644 --- a/packages/client/runtime/src/client/sessions/manager.ts +++ b/packages/client/runtime/src/client/sessions/manager.ts @@ -2,7 +2,10 @@ // dispatch entry + list state, constructed and held by SessionsService (one per client runtime). // List data never enters zustand; React connects via subscribe/getListSnapshot. -import type { IApiClient, HostFrame, MuxFrame, RpcError, RpcRequest, RpcResult, SessionId, SessionSummary, WorkspaceId } from '@deepseek-ai/dsh-client-connection/client' +import type { + HostFrame, IApiClient, ModelRequestTelemetry, MuxFrame, RpcError, RpcRequest, + RpcResult, SessionId, SessionSummary, WorkspaceId, +} from '@deepseek-ai/dsh-client-connection/client' // Value import from the inline-safe wire layer (not the connection plugin): // plugin-to-plugin value imports are a bundle purity error. import { transportError } from '@deepseek-ai/dsh-host-apiproxy/api' @@ -58,11 +61,11 @@ export class SessionManager { * frames are low-frequency; overflow drops oldest) and dropped on session-removed (audit S7). */ private readonly pendingBuffers = new Map[]>() /** - * Latest model capacity observed for an uninstantiated session on the + * Latest request telemetry observed for an uninstantiated session on the * current mux generation. Unlike durable history, this transient frame * cannot be backfilled when get() lazily creates the Session. */ - private readonly modelRequestContextWindows = new Map() + private readonly modelRequests = new Map() /** Outstanding approval questions per session, keyed by approvalId (idempotent under mux-open * replays of the same requested frame). Manager-owned rather than read off Session instances * because the sidebar must light up for sessions never instantiated. Cleared per connection @@ -171,7 +174,7 @@ export class SessionManager { } private createSession(sessionId: SessionId): Session { - const modelRequestContextWindow = this.modelRequestContextWindows.get(sessionId) + const modelRequest = this.modelRequests.get(sessionId) return new Session(sessionId, this.api, { // The sender's local first-send flip mirrors into the list row so the // session surfaces (lists filter on blank) before any host frame lands. @@ -179,7 +182,7 @@ export class SessionManager { this.recordMutation({ kind: 'engaged', sessionId: engaged.sessionId }) }, projections: this.projectionStore(sessionId), - ...(modelRequestContextWindow === undefined ? {} : { modelRequestContextWindow }), + ...(modelRequest === undefined ? {} : { modelRequest }), }) } @@ -355,13 +358,13 @@ export class SessionManager { return } if (frame.type === 'session/model-request') { - // Transient and non-replayable: retain the latest capacity until lazy - // instantiation. An absent value explicitly clears an earlier one. - if (frame.contextWindow === undefined) this.modelRequestContextWindows.delete(frame.sessionId) - else this.modelRequestContextWindows.set(frame.sessionId, frame.contextWindow) + // Transient and non-replayable: retain the whole latest request until + // lazy instantiation. Missing fields replace rather than inherit. + const { type: _type, sessionId, ...modelRequest } = frame + this.modelRequests.set(sessionId, modelRequest) } if (frame.type === 'session/subscribed') { - this.modelRequestContextWindows.delete(frame.sessionId) + this.modelRequests.delete(frame.sessionId) // Rows past the host's durable baseline rode state a restart lost; drop // them so last-wins cannot pin a phantom value over recomputed truth. this.projectionStores.get(frame.sessionId)?.truncate(frame.lastSeq) @@ -440,7 +443,7 @@ export class SessionManager { this.recordMutation({ kind: 'remove', sessionId: frame.sessionId }) this.sessions.get(frame.sessionId)?.handleRemoved() // instance survives (resident-instance rule), only flagged in the snapshot this.pendingBuffers.delete(frame.sessionId) // a removed session's buffered frames must not replay on a future instantiation - this.modelRequestContextWindows.delete(frame.sessionId) // connection-local request capacity dies with the Host session + this.modelRequests.delete(frame.sessionId) // connection-local request telemetry dies with the Host session this.waitingApprovals.delete(frame.sessionId) // a removed session cannot wait on anyone this.projectionStores.delete(frame.sessionId) // removed sessions drop their projection rows with the instance return @@ -481,7 +484,7 @@ export class SessionManager { if (kept.length === 0) this.pendingBuffers.delete(sessionId) else this.pendingBuffers.set(sessionId, kept) } - this.modelRequestContextWindows.clear() + this.modelRequests.clear() for (const session of this.sessions.values()) session.handleReconnecting() } diff --git a/packages/client/runtime/src/client/sessions/session.ts b/packages/client/runtime/src/client/sessions/session.ts index 8b99940009..765c3f5fa1 100644 --- a/packages/client/runtime/src/client/sessions/session.ts +++ b/packages/client/runtime/src/client/sessions/session.ts @@ -5,7 +5,7 @@ import type { ContentBlock } from '@deepseek-ai/dsh-llm/types' import type { SessionEvent } from '@deepseek-ai/dsh-session/types' import type { HistoryEntry, IApiClient, MuxFrame, RpcError, RpcId, RpcResult, - SessionId, SessionMetrics, ToolEventView, + ModelRequestTelemetry, SessionId, ToolEventView, } from '@deepseek-ai/dsh-client-connection/client' // Value import from the inline-safe wire layer (not the connection plugin): // plugin-to-plugin value imports are a bundle purity error. @@ -43,8 +43,8 @@ export interface SessionOptions { * private store (bare object-layer construction). */ projections?: ProjectionValueStore - /** Model capacity already observed on this mux generation before lazy construction. */ - modelRequestContextWindow?: number + /** Request telemetry already observed on this mux generation before lazy construction. */ + modelRequest?: ModelRequestTelemetry } /** Queue-row preview cap: the dock renders one line, the full content never leaves the host mirror. */ @@ -110,10 +110,8 @@ export class Session implements SessionFace { private queueCache: { rev: number; value: QueuedMessage[] } | null = null private frozenRev = 0 private nodesCache: { folded: readonly ConversationNode[]; frozenRev: number; value: readonly ConversationNode[] } | null = null - /** Host-owned durable usage/current-pressure projection. */ - private metrics: SessionMetrics | null = null - /** Latest capacity observed on this mux connection, independent of durable metrics arrival. */ - private contextWindow: number | undefined + /** Latest atomic request snapshot observed on this mux connection. */ + private modelRequest: ModelRequestTelemetry | null /** `run_code` sub-dispatches by parent callId (window-derived, like openCalls). Appends * copy-on-write the per-parent array so published snapshot references never mutate. */ private codeDispatches = new Map() @@ -175,7 +173,7 @@ export class Session implements SessionFace { private readonly options: SessionOptions = {}, ) { this.projections = options.projections ?? new ProjectionValueStore() - this.contextWindow = options.modelRequestContextWindow + this.modelRequest = options.modelRequest ?? null this.snapshotCache = this.buildSnapshot() } @@ -331,10 +329,10 @@ export class Session implements SessionFace { * in-flight open first — its history request rode the dead connection and must not settle * the fresh generation into 'error' (audit S4). */ async resync(): Promise { - // Queue, metrics, and request capacity are NOT cleared here: onConnected + // Queue and request telemetry are NOT cleared here: onConnected // (which drives resync) races the mux frames — fresh-generation state may - // have landed already, and the host never resends it. session/subscribed - // owns the generation reset before the queue snapshot and metrics frames. + // have landed already, and the host never resends request telemetry. + // session/subscribed owns the reset before the queue snapshot. if (this.openState === 'cold') return // never opened: no window to rebuild (doOpen flips to 'loading' synchronously, so cold implies no in-flight open) this.openGeneration++ this.openPromise = null @@ -415,24 +413,22 @@ export class Session implements SessionFace { this.queueRev++ changed = true } - if (this.contextWindow !== undefined) { - this.contextWindow = undefined - changed = true - } - if (this.metrics !== null) { - this.metrics = null + if (this.modelRequest !== null) { + this.modelRequest = null changed = true } if (changed) this.notifier.markDirty() return } - case 'session/metrics': { - this.installMetrics(frame.metrics) - return - } case 'session/model-request': { - if (this.contextWindow === frame.contextWindow) return - this.contextWindow = frame.contextWindow + const { + type: _type, + sessionId: _sessionId, + ...modelRequest + } = frame + // Whole-frame replacement is load-bearing: an omitted numerator or + // capacity clears that field from the preceding request. + this.modelRequest = modelRequest this.notifier.markDirty() return } @@ -508,17 +504,16 @@ export class Session implements SessionFace { /** Connection-loss boundary: clear values that are not replayed before the next stream starts. */ handleReconnecting(): void { this.openGeneration++ - if (this.metrics === null && this.contextWindow === undefined) return - this.metrics = null - this.contextWindow = undefined + if (this.modelRequest === null) return + this.modelRequest = null this.notifier.markDirty() } - /** host/session-removed relay: flag the resident snapshot and clear connection-local capacity. */ + /** host/session-removed relay: flag the resident snapshot and clear request telemetry. */ handleRemoved(): void { - const changed = !this.removed || this.contextWindow !== undefined + const changed = !this.removed || this.modelRequest !== null this.removed = true - this.contextWindow = undefined + this.modelRequest = null if (changed) this.notifier.markDirty() } @@ -567,7 +562,6 @@ export class Session implements SessionFace { result.value.events, result.value.hasMore, result.value.projections, - result.value.metrics, ) // Gap detection (§D.3-4): baseline past the window tail and liveBuffer did not cover it -> pull the tail page once more. const tailSeq = this.windowTailSeq() @@ -579,7 +573,6 @@ export class Session implements SessionFace { result.value.events, result.value.hasMore, result.value.projections, - result.value.metrics, ) } } @@ -606,7 +599,6 @@ export class Session implements SessionFace { entries: HistoryEntry[], hasMore: boolean, projections: ProjectionsBaseline | undefined, - metrics: SessionMetrics | undefined, ): void { this.events = entries.map(e => e.event) this.views = entries.map(e => e.view) @@ -615,7 +607,6 @@ export class Session implements SessionFace { this.foldAdapter.reset(this.events, this.baseSeq, this.views) this.rebuildDerivedFromWindow() if (projections !== undefined) this.projections.seed(projections) - if (metrics !== undefined) this.installMetrics(metrics) const buffered = this.liveBuffer this.liveBuffer = [] for (const item of buffered) this.appendLive(item.event, item.view) @@ -668,7 +659,6 @@ export class Session implements SessionFace { result.value.events, result.value.hasMore, result.value.projections, - result.value.metrics, ) } } catch (error) { @@ -854,20 +844,6 @@ export class Session implements SessionFace { return tail === undefined ? null : tail.seq } - /** Install a metrics snapshot unless a newer durable or publication revision already landed. */ - private installMetrics(metrics: SessionMetrics): void { - const current = this.metrics - if ( - current !== null - && ( - metrics.logRevision < current.logRevision - || metrics.projectionRevision < current.projectionRevision - ) - ) return - this.metrics = metrics - this.notifier.markDirty() - } - private buildSnapshot(): ConversationSnapshot { const { nodes: folded, degraded } = this.foldAdapter.nodes() // Frozen interrupted nodes ride fractional seqs: a stable merge keeps them in flow order. @@ -920,10 +896,7 @@ export class Session implements SessionFace { promptError: this.promptError, blank: this.blankBit, lastAgentError: this.lastAgentError, - metrics: this.metrics, - ...(this.contextWindow === undefined - ? {} - : { modelRequestContextWindow: this.contextWindow }), + modelRequest: this.modelRequest, } } } diff --git a/packages/client/runtime/tests/client-apply.spec.ts b/packages/client/runtime/tests/client-apply.spec.ts index 7399c46f6b..1e2937b464 100644 --- a/packages/client/runtime/tests/client-apply.spec.ts +++ b/packages/client/runtime/tests/client-apply.spec.ts @@ -102,7 +102,7 @@ describe('runtime client apply', () => { expect(bench.api.callsOf('session.create')).toHaveLength(1) }) - it('clears connection-local Session state on disconnect but not connected', async () => { + it('clears connection-local request telemetry on reconnect but not connected', async () => { const bench = await mount() const sessions = bench.ctx.get('sessions') as SessionsService bench.sinks?.onHostEnvelope?.({ @@ -112,21 +112,8 @@ describe('runtime client apply', () => { await Promise.resolve() const session = sessions.binding('s-state' as never)?.session if (session === undefined) throw new Error('session binding missing') - const currentMetrics = { - projectionRevision: 4, - logRevision: 10, - uncachedInputTokens: 10, - outputTokens: 4, - cacheReadTokens: 90, - cacheWriteTokens: 3, - contextTokens: 35, - } bench.sinks?.onMuxEnvelope?.({ - rpcId: 'metrics' as never, - payload: { type: 'session/metrics', sessionId: 's-state', metrics: currentMetrics } as never, - }) - bench.sinks?.onMuxEnvelope?.({ - rpcId: 'capacity' as never, + rpcId: 'request' as never, payload: { type: 'session/model-request', sessionId: 's-state', @@ -134,19 +121,20 @@ describe('runtime client apply', () => { step: 1, provider: 'test', model: 'alpha', + contextTokens: 32_000, contextWindow: 128_000, } as never, }) bench.sinks?.onConnected?.() - expect(session.getSnapshot()).toMatchObject({ - metrics: currentMetrics, - modelRequestContextWindow: 128_000, + expect(session.getSnapshot().modelRequest).toMatchObject({ + model: 'alpha', + contextTokens: 32_000, + contextWindow: 128_000, }) - bench.sinks?.onDisconnected?.() - expect(session.getSnapshot().metrics).toBeNull() - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() + bench.sinks?.onStateChange?.('reconnecting') + expect(session.getSnapshot().modelRequest).toBeNull() }) it('stops the stream loop when the plugin fiber unloads', async () => { diff --git a/packages/client/runtime/tests/fake-api.ts b/packages/client/runtime/tests/fake-api.ts index 829417e44b..0654446c5b 100644 --- a/packages/client/runtime/tests/fake-api.ts +++ b/packages/client/runtime/tests/fake-api.ts @@ -4,7 +4,7 @@ import type { CommandId } from '@deepseek-ai/dsh-commands/brand' import type { ClientResponse, CommandDescriptor, HostFrame, IApiClient, ModelTarget, MuxFrame, - RpcError, RpcReceipt, RpcRequest, RpcResponse, SessionId, SessionMetrics, SessionModels, + RpcError, RpcReceipt, RpcRequest, RpcResponse, SessionId, SessionModels, SessionProjectionsBlock, SkillEntry, WorkspaceId, WorkspaceView, } from '@deepseek-ai/dsh-client-connection/client' @@ -69,7 +69,6 @@ export class FakeApiClient implements IApiClient { events: never[] hasMore: boolean projections?: SessionProjectionsBlock - metrics?: SessionMetrics }>> = () => Promise.resolve(ok({ events: [], hasMore: false })) diff --git a/packages/client/runtime/tests/manager.spec.ts b/packages/client/runtime/tests/manager.spec.ts index 62f4cc32d1..d5a4237fc1 100644 --- a/packages/client/runtime/tests/manager.spec.ts +++ b/packages/client/runtime/tests/manager.spec.ts @@ -4,7 +4,7 @@ */ import { describe, expect, it, vi } from 'vitest' -import type { SessionId, SessionMetrics } from '@deepseek-ai/dsh-client-connection/client' +import type { SessionId } from '@deepseek-ai/dsh-client-connection/client' import { SessionManager } from '../src/client/sessions/manager.ts' import { FakeApiClient, deferred, err, ok } from './fake-api.ts' import { entries, plainTurn } from './event-script.ts' @@ -41,7 +41,7 @@ describe('instances', () => { expect(manager.get(S2).getSnapshot().pending).toEqual([]) }) - it('retains the latest transient model capacity until lazy instantiation', () => { + it('retains the latest transient request snapshot until lazy instantiation', () => { const api = new FakeApiClient() const manager = new SessionManager(api) manager.handleMuxEnvelope({ @@ -53,6 +53,7 @@ describe('instances', () => { step: 1, provider: 'test', model: 'alpha', + contextTokens: 12_000, contextWindow: 128_000, }, }) @@ -65,14 +66,22 @@ describe('instances', () => { step: 2, provider: 'test', model: 'beta', + contextTokens: 32_000, contextWindow: 256_000, }, }) - expect(manager.get(S1).getSnapshot().modelRequestContextWindow).toBe(256_000) + expect(manager.get(S1).getSnapshot().modelRequest).toEqual({ + turn: 1, + step: 2, + provider: 'test', + model: 'beta', + contextTokens: 32_000, + contextWindow: 256_000, + }) }) - it('retains explicit capacity clearing before lazy instantiation', () => { + it('retains whole-frame replacement before lazy instantiation', () => { const api = new FakeApiClient() const manager = new SessionManager(api) manager.handleMuxEnvelope({ @@ -99,10 +108,15 @@ describe('instances', () => { }, }) - expect(manager.get(S1).getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(manager.get(S1).getSnapshot().modelRequest).toEqual({ + turn: 1, + step: 2, + provider: 'test', + model: 'unknown-capacity', + }) }) - it('clears retained capacity on subscribed and resident capacity on removal', () => { + it('clears retained request telemetry on subscribed and removal', () => { const api = new FakeApiClient() const manager = new SessionManager(api) manager.handleMuxEnvelope({ @@ -122,7 +136,7 @@ describe('instances', () => { payload: { type: 'session/subscribed', sessionId: S1, lastSeq: 0 }, }) const session = manager.get(S1) - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(session.getSnapshot().modelRequest).toBeNull() manager.handleMuxEnvelope({ rpcId: 'request-after-subscribe' as never, @@ -136,12 +150,12 @@ describe('instances', () => { contextWindow: 256_000, }, }) - expect(session.getSnapshot().modelRequestContextWindow).toBe(256_000) + expect(session.getSnapshot().modelRequest?.contextWindow).toBe(256_000) manager.handleHostEnvelope({ rpcId: 'removed' as never, payload: { type: 'host/session-removed', sessionId: S1 }, }) - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(session.getSnapshot().modelRequest).toBeNull() manager.handleMuxEnvelope({ rpcId: 'request-before-lazy-removal' as never, @@ -159,28 +173,15 @@ describe('instances', () => { rpcId: 'lazy-removed' as never, payload: { type: 'host/session-removed', sessionId: S2 }, }) - expect(manager.get(S2).getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(manager.get(S2).getSnapshot().modelRequest).toBeNull() }) - it('clears resident metrics and capacity plus lazy capacity before reconnect', () => { + it('clears resident and lazy request telemetry on disconnect', () => { const api = new FakeApiClient() const manager = new SessionManager(api) const session = manager.get(S1) - const currentMetrics: SessionMetrics = { - projectionRevision: 4, - logRevision: 10, - uncachedInputTokens: 10, - outputTokens: 4, - cacheReadTokens: 90, - cacheWriteTokens: 3, - contextTokens: 35, - } manager.handleMuxEnvelope({ - rpcId: 'metrics' as never, - payload: { type: 'session/metrics', sessionId: S1, metrics: currentMetrics }, - }) - manager.handleMuxEnvelope({ - rpcId: 'resident-capacity' as never, + rpcId: 'resident-request' as never, payload: { type: 'session/model-request', sessionId: S1, @@ -188,11 +189,12 @@ describe('instances', () => { step: 1, provider: 'test', model: 'resident', + contextTokens: 35, contextWindow: 128_000, }, }) manager.handleMuxEnvelope({ - rpcId: 'lazy-capacity' as never, + rpcId: 'lazy-request' as never, payload: { type: 'session/model-request', sessionId: S2, @@ -200,19 +202,20 @@ describe('instances', () => { step: 1, provider: 'test', model: 'lazy', + contextTokens: 70, contextWindow: 256_000, }, }) - expect(session.getSnapshot()).toMatchObject({ - metrics: currentMetrics, - modelRequestContextWindow: 128_000, + expect(session.getSnapshot().modelRequest).toMatchObject({ + model: 'resident', + contextTokens: 35, + contextWindow: 128_000, }) - manager.handleReconnecting() + manager.handleDisconnected() - expect(session.getSnapshot().metrics).toBeNull() - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() - expect(manager.get(S2).getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(session.getSnapshot().modelRequest).toBeNull() + expect(manager.get(S2).getSnapshot().modelRequest).toBeNull() }) it('caps the pending buffer at 32 keeping the newest, and drops it on session-removed', () => { diff --git a/packages/client/runtime/tests/session.spec.ts b/packages/client/runtime/tests/session.spec.ts index 827e9f853e..3994a691cb 100644 --- a/packages/client/runtime/tests/session.spec.ts +++ b/packages/client/runtime/tests/session.spec.ts @@ -10,7 +10,7 @@ import { describe, expect, it, vi } from 'vitest' import { Context } from 'cordis' import type { SessionEvent } from '@deepseek-ai/dsh-session/types' import type { - SessionId, SessionMetrics, SessionProjectionsBlock, + SessionId, SessionProjectionsBlock, } from '@deepseek-ai/dsh-client-connection/client' import { Session } from '../src/client/sessions/session.ts' import { FakeApiClient, deferred, err, ok } from './fake-api.ts' @@ -29,34 +29,15 @@ function histResponse( events: SessionEvent[], hasMore = false, projections?: SessionProjectionsBlock, - metrics?: SessionMetrics, ) { // history now returns HistoryEntry[] ({event, view?}); these tests are view-less. return Promise.resolve(ok({ events: entries(events) as never[], hasMore, ...projections === undefined ? {} : { projections }, - ...metrics === undefined ? {} : { metrics }, })) } -function metrics( - projectionRevision: number, - logRevision: number, - over: Partial = {}, -): SessionMetrics { - return { - projectionRevision, - logRevision, - uncachedInputTokens: 10, - outputTokens: 4, - cacheReadTokens: 90, - cacheWriteTokens: 3, - contextTokens: 35, - ...over, - } -} - describe('open', () => { it('installs the tail page: cold → loading → open with window and nodes in place', async () => { const { api, session } = makeSession() @@ -70,19 +51,7 @@ describe('open', () => { expect(snapshot.openState).toBe('open') expect(snapshot.hasMore).toBe(true) expect(snapshot.nodes.map(n => n.kind)).toEqual(['user', 'assistant']) - expect(snapshot.metrics).toBeNull() - }) - - it('installs full-log metrics independently of older history pages', async () => { - const { api, session } = makeSession() - const tailMetrics = metrics(4, 106) - api.onHistory = () => histResponse(plainTurn(100, 3, '问', '答'), true, undefined, tailMetrics) - await session.open() - expect(session.getSnapshot().metrics).toBe(tailMetrics) - - api.onHistory = () => histResponse(plainTurn(94, 2, '旧问', '旧答')) - await session.loadOlder() - expect(session.getSnapshot().metrics).toBe(tailMetrics) + expect(snapshot.modelRequest).toBeNull() }) it('is idempotent: concurrent opens share one history call, reopening when open is a no-op', async () => { @@ -147,17 +116,8 @@ describe('live event path', () => { expect(session.getSnapshot().nodes).toEqual(before.nodes) }) - it('keeps live capacity separate from durable metrics, replaces or clears it on requests, and resets at subscription', async () => { + it('replaces the whole request snapshot, clears omitted fields, and resets at subscription', async () => { const { session } = await opened() - const current = metrics(8, 10) - session.handleMuxEnvelope('m1' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: current, - }) - expect(session.getSnapshot().metrics).toBe(current) - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() - session.handleMuxEnvelope('request-1' as never, { type: 'session/model-request', sessionId: SID, @@ -165,32 +125,17 @@ describe('live event path', () => { step: 1, provider: 'test', model: 'alpha', + contextTokens: 32_000, contextWindow: 128_000, }) - expect(session.getSnapshot().metrics).toBe(current) - expect(session.getSnapshot().modelRequestContextWindow).toBe(128_000) - - session.handleMuxEnvelope('m2' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: metrics(9, 9, { uncachedInputTokens: 1 }), + expect(session.getSnapshot().modelRequest).toEqual({ + turn: 1, + step: 1, + provider: 'test', + model: 'alpha', + contextTokens: 32_000, + contextWindow: 128_000, }) - session.handleMuxEnvelope('m3' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: metrics(7, 11, { uncachedInputTokens: 2 }), - }) - expect(session.getSnapshot().metrics).toBe(current) - expect(session.getSnapshot().modelRequestContextWindow).toBe(128_000) - - const ordinaryUpdate = metrics(9, 11, { contextTokens: 40 }) - session.handleMuxEnvelope('m4' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: ordinaryUpdate, - }) - expect(session.getSnapshot().metrics).toBe(ordinaryUpdate) - expect(session.getSnapshot().modelRequestContextWindow).toBe(128_000) session.handleMuxEnvelope('request-2' as never, { type: 'session/model-request', @@ -200,16 +145,19 @@ describe('live event path', () => { provider: 'test', model: 'without-capacity', }) - expect(session.getSnapshot().metrics).toEqual(ordinaryUpdate) - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(session.getSnapshot().modelRequest).toEqual({ + turn: 2, + step: 1, + provider: 'test', + model: 'without-capacity', + }) session.handleMuxEnvelope('sub' as never, { type: 'session/subscribed', sessionId: SID, lastSeq: 5, }) - expect(session.getSnapshot().metrics).toBeNull() - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(session.getSnapshot().modelRequest).toBeNull() session.handleMuxEnvelope('request-3' as never, { type: 'session/model-request', sessionId: SID, @@ -217,19 +165,17 @@ describe('live event path', () => { step: 1, provider: 'test', model: 'beta', + contextTokens: 20, contextWindow: 256_000, }) - const nextGeneration = metrics(0, 10, { contextTokens: 20 }) - session.handleMuxEnvelope('m5' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: nextGeneration, + expect(session.getSnapshot().modelRequest).toMatchObject({ + turn: 3, + contextTokens: 20, + contextWindow: 256_000, }) - expect(session.getSnapshot().metrics).toBe(nextGeneration) - expect(session.getSnapshot().modelRequestContextWindow).toBe(256_000) }) - it('publishes a subscribed reset when capacity arrived before durable metrics', async () => { + it('publishes a subscribed reset when request telemetry arrived first', async () => { const { session } = await opened() session.handleMuxEnvelope('request' as never, { type: 'session/model-request', @@ -238,17 +184,17 @@ describe('live event path', () => { step: 1, provider: 'test', model: 'alpha', + contextTokens: 8_000, contextWindow: 128_000, }) - expect(session.getSnapshot().metrics).toBeNull() - expect(session.getSnapshot().modelRequestContextWindow).toBe(128_000) + expect(session.getSnapshot().modelRequest?.contextWindow).toBe(128_000) session.handleMuxEnvelope('sub' as never, { type: 'session/subscribed', sessionId: SID, lastSeq: 5, }) - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() + expect(session.getSnapshot().modelRequest).toBeNull() }) it('materializes a command node from live lifecycle frames and reproduces it from a history window', async () => { @@ -876,72 +822,61 @@ describe('remaining branches', () => { }) describe('resync', () => { - it('fences pre-disconnect history behind a fresh mux metrics baseline', async () => { + it('clears request telemetry on reconnect and drops a stale in-flight history response', async () => { const { api, session } = makeSession() const stale = deferred>>() api.onHistory = () => stale.promise const opening = session.open() - const oldLiveMetrics = metrics(8, 10) - session.handleMuxEnvelope('old-metrics' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: oldLiveMetrics, - }) - session.handleMuxEnvelope('old-capacity' as never, { + session.handleMuxEnvelope('old-request' as never, { type: 'session/model-request', sessionId: SID, turn: 1, step: 1, provider: 'test', model: 'old', + contextTokens: 20, contextWindow: 128_000, }) session.handleReconnecting() - expect(session.getSnapshot().metrics).toBeNull() - expect(session.getSnapshot().modelRequestContextWindow).toBeUndefined() - const freshMetrics = metrics(0, 1, { contextTokens: 20 }) - session.handleMuxEnvelope('fresh-metrics' as never, { - type: 'session/metrics', - sessionId: SID, - metrics: freshMetrics, - }) + expect(session.getSnapshot().modelRequest).toBeNull() stale.resolve(ok({ events: entries(plainTurn(0, 0, '旧问', '旧答')) as never[], hasMore: false, - metrics: metrics(99, 99, { contextTokens: 999 }), })) await opening expect(session.getSnapshot().nodes).toEqual([]) - expect(session.getSnapshot().metrics).toBe(freshMetrics) + expect(session.getSnapshot().modelRequest).toBeNull() api.onHistory = () => histResponse(plainTurn(6, 1, '新问', '新答')) await session.resync() expect(session.getSnapshot().openState).toBe('open') expect(session.getSnapshot().nodes.map(node => node.seq)).toEqual([7, 9]) - expect(session.getSnapshot().metrics).toBe(freshMetrics) + expect(session.getSnapshot().modelRequest).toBeNull() }) - it('preserves fresh-generation metrics that arrive before a failing history refresh', async () => { + it('preserves a fresh-generation request snapshot when history resync fails', async () => { const { api, session } = makeSession() - const oldMetrics = metrics(8, 10) - api.onHistory = () => histResponse(plainTurn(0, 0, 'a', 'b'), false, undefined, oldMetrics) + api.onHistory = () => histResponse(plainTurn(0, 0, 'a', 'b')) await session.open() - expect(session.getSnapshot().metrics).toBe(oldMetrics) session.handleMuxEnvelope('sub' as never, { type: 'session/subscribed', sessionId: SID, lastSeq: 5, }) - expect(session.getSnapshot().metrics).toBeNull() + expect(session.getSnapshot().modelRequest).toBeNull() - const freshMetrics = metrics(0, 10, { contextTokens: 20 }) - session.handleMuxEnvelope('fresh-metrics' as never, { - type: 'session/metrics', + session.handleMuxEnvelope('fresh-request' as never, { + type: 'session/model-request', sessionId: SID, - metrics: freshMetrics, + turn: 2, + step: 1, + provider: 'test', + model: 'fresh', + contextTokens: 20, + contextWindow: 256_000, }) api.onHistory = () => Promise.resolve(err({ code: 'internal', @@ -953,7 +888,14 @@ describe('resync', () => { expect(session.getSnapshot()).toMatchObject({ openState: 'error', - metrics: freshMetrics, + modelRequest: { + turn: 2, + step: 1, + provider: 'test', + model: 'fresh', + contextTokens: 20, + contextWindow: 256_000, + }, }) }) diff --git a/packages/client/test-runtime/src/fixtures.ts b/packages/client/test-runtime/src/fixtures.ts index 4219d233e2..7c9d74aae7 100644 --- a/packages/client/test-runtime/src/fixtures.ts +++ b/packages/client/test-runtime/src/fixtures.ts @@ -62,6 +62,7 @@ export function conversationSnapshot(sessionId: SessionId): ConversationSnapshot promptError: null, blank: false, lastAgentError: null, + modelRequest: null, } } diff --git a/packages/client/ui-conversation/README.i18n.yaml b/packages/client/ui-conversation/README.i18n.yaml index 6508a18b98..388b2044b8 100644 --- a/packages/client/ui-conversation/README.i18n.yaml +++ b/packages/client/ui-conversation/README.i18n.yaml @@ -2,5 +2,5 @@ # side as of the last confirmed-consistent state. Both languages carry equal authority; # after editing either side, bring the other along and re-record with: # pnpm run verify-translation-pairing --write packages/client/ui-conversation/README.md -README.md: b74127ffe11df40fcad0c97e2fc6896ec79ad9d1 -README.zh.md: f6dfc3ca61ea80758c9f64ca98f034b0832ee89e +README.md: c60a38ec26d4bca15ce23e51dbaedf3ef36022e4 +README.zh.md: 65c5554a78960fc4f86a6a370772c95429c648c4 diff --git a/packages/client/ui-conversation/README.md b/packages/client/ui-conversation/README.md index b74127ffe1..c60a38ec26 100644 --- a/packages/client/ui-conversation/README.md +++ b/packages/client/ui-conversation/README.md @@ -20,7 +20,7 @@ Per-session UI state for selection and the active view lives in the declared cha The composer bar declares session-scoped single seats for `'conversation.input.plan'` (right of the local access-mode control) and `'conversation.input.model'` (immediately before the pending indicator and send/stop button), plus list slots for overlay, dock, left, and right input extensions. Feature packages own each control and its state; ui-conversation supplies placement, the `locked` owner prop, and the standard slot shares. While the `plan` projection's effective target is plan mode, InputBar swaps its textarea placeholder to the plan-task wording (a host-folded value read through the standard-kit `useProjection`; owner-supplied placeholders win). The resident no-session shell uses `DisabledInputBar` and therefore dispatches no session-scoped control seats. -The chat stats line reads durable token counters/current pressure from `ConversationSnapshot.metrics` and joins them only at presentation with the separate connection-local `modelRequestContextWindow`; visible nodes supply only the existing turn/step counts. It renders uncached input, output, and cache reads as separate compact values, computes cache hit as `cacheRead / (uncachedInput + cacheRead)` without cache writes, and shows context occupancy only after the current mux connection observes a model request with capacity. Before that request, after reconnect/restore/new subscription, or after a request without capacity, the percentage is omitted and context is labeled unknown rather than queried ahead or reconstructed from history. +The chat stats line reads full-log billing from the generic `tokenUsage` projection and joins it only at presentation with the connection-local atomic `ConversationSnapshot.modelRequest`; visible nodes supply only the existing turn/step counts. It renders uncached input, output, and cache reads as separate compact values, computes cache hit as `cacheRead / (uncachedInput + cacheRead)` without cache writes, and shows context occupancy only when the same observed request snapshot contains both `contextTokens` and `contextWindow`. Before that request, after reconnect/restore/new subscription, or after a request missing either field, context is labeled unknown rather than queried from the selected model or reconstructed from history. The existing inline stats row remains the sole context UI; the model selector has no circle or accessory. `src/client/` is organized for the future package split: `contract/` is the sole inter-domain shared face (`slots.ts` slot declarations + composed slot props including the tool-row contract, `views.ts` shared primitives, `tool-call-model.ts`); the `skeleton/`, `chat/`, and `toolviews/` (sample registrants) domain directories import contract files and never each other; `apply.ts` is the only assembly point allowed to import all three domains. The `/client` export surface is the contract only — `apply`/`inject`, the two service classes, and the `contract/` type families; implementation components (skeleton, chat rows) and the store factory stay internal and reach the page exclusively through apply's slot registrations (tests take them via the `./src/*` subpath). diff --git a/packages/client/ui-conversation/README.zh.md b/packages/client/ui-conversation/README.zh.md index f6dfc3ca61..65c5554a78 100644 --- a/packages/client/ui-conversation/README.zh.md +++ b/packages/client/ui-conversation/README.zh.md @@ -20,7 +20,7 @@ todo 两个面就是在该形状上的两个注册项,都是普通注册方插 输入栏为 `'conversation.input.plan'`(位于本地 access 模式控件右侧)和 `'conversation.input.model'`(渲染在 pending 指示器与发送/停止按钮之前)声明会话作用域的单实例 seat,并为 overlay、dock、left 和 right 输入扩展声明列表 slot。各功能包拥有相应控件及其状态;ui-conversation 提供放置位置、`locked` owner prop 和标准 slot share。当 `plan` 投影的有效目标为 plan mode 时,InputBar 将文本框 placeholder 切换为 plan 任务措辞(经标准套件 `useProjection` 读取的 host 折叠值;owner 提供的 placeholder 优先)。常驻无会话壳使用 `DisabledInputBar`,因此不会分发任何会话作用域的控件 seat。 -聊天统计行从 `ConversationSnapshot.metrics` 读取持久的 token 计数/当前压力,并且只在展示时把它们与独立的连接本地 `modelRequestContextWindow` 结合;可见节点仅提供既有的轮次和步骤计数。它以相互独立的紧凑值显示未缓存输入、输出与缓存读取,通过 `cacheRead / (uncachedInput + cacheRead)` 计算缓存命中率而不计入缓存写入,并且只有当前 mux 连接观察到带容量的模型请求后才显示上下文占用率。在该请求之前、重连/恢复/新订阅之后,或在请求不带容量之后,系统都会省略百分比,并把上下文标为「未知」,而不会提前查询或根据历史记录重建。 +聊天统计行从通用 `tokenUsage` 投影读取完整日志计费用量,并且只在展示时把它与连接本地的原子快照 `ConversationSnapshot.modelRequest` 结合;可见节点仅提供既有的轮次和步骤计数。它以相互独立的紧凑值显示未缓存输入、输出与缓存读取,通过 `cacheRead / (uncachedInput + cacheRead)` 计算缓存命中率而不计入缓存写入,并且只有同一份已观测请求快照同时包含 `contextTokens` 与 `contextWindow` 时才显示上下文占用率。在该请求之前、重连/恢复/新订阅之后,或在请求缺少任一字段之后,系统都会把上下文标为「未知」,而不会从所选模型查询或根据历史记录重建。现有的行内统计行仍是唯一的上下文 UI;模型选择器不增加圆环或附属控件。 `src/client/` 按未来的包拆分组织:`contract/` 是唯一的跨领域共享表层(`slots.ts` slot 声明 + 组合后的 slot props,包括工具行契约、`views.ts` 共享原语、`tool-call-model.ts`);`skeleton/`、`chat/` 和 `toolviews/`(示例注册方)领域目录只导入 contract 文件,彼此绝不导入;`apply.ts` 是唯一允许导入全部三个领域的组装点。`/client` 导出表层只包含契约:`apply`/`inject`、两个服务类和 `contract/` 类型家族;实现组件(骨架、聊天行)与 store factory 保持内部状态,只能通过 apply 的 slot 注册到达页面(测试通过 `./src/*` 子路径获取它们)。 diff --git a/packages/client/ui-conversation/package.json b/packages/client/ui-conversation/package.json index 42812dce24..88be12b5af 100644 --- a/packages/client/ui-conversation/package.json +++ b/packages/client/ui-conversation/package.json @@ -44,6 +44,7 @@ "@deepseek-ai/dsh-client-ui-slash": "^0.0.1", "@deepseek-ai/dsh-client-ui-slots": "^0.0.1", "@deepseek-ai/dsh-invariants": "^0.0.1", + "@deepseek-ai/dsh-token-meter": "^0.0.1", "cordis": "^4.0.0-rc.7", "react": "^18.2.0" }, @@ -52,6 +53,7 @@ "@deepseek-ai/dsh-plan-mode": "workspace:^", "@deepseek-ai/dsh-session-projection": "workspace:^", "@deepseek-ai/dsh-tool-todo": "workspace:^", + "@deepseek-ai/dsh-token-meter": "workspace:^", "@deepseek-ai/dsh-client-ui-layout": "workspace:^", "@deepseek-ai/dsh-client-ui-primitives": "workspace:^", "@deepseek-ai/dsh-client-ui-slash": "workspace:^", diff --git a/packages/client/ui-conversation/src/client/chat/ChatView.tsx b/packages/client/ui-conversation/src/client/chat/ChatView.tsx index 1cbb00a4ce..36a2cde7b9 100644 --- a/packages/client/ui-conversation/src/client/chat/ChatView.tsx +++ b/packages/client/ui-conversation/src/client/chat/ChatView.tsx @@ -221,7 +221,9 @@ function StreamingTail({ useSession, onGrow }: { * The chat view slot entry: pure component over the composed props (tool rows * render through the declared keyed hole's renderSlot share). */ -export function ChatView({ useSession, useSessions, useStore, renderSlot, sessionId, openFile, loadOlder }: ChatViewSlotProps) { +export function ChatView({ + useProjection, useSession, useSessions, useStore, renderSlot, sessionId, openFile, loadOlder, +}: ChatViewSlotProps) { const nodes = useSession(s => s.nodes) // Workspace root off the session list row: path summaries display relative to it. const cwd = useSessions(s => s.byId[sessionId]?.cwd) @@ -381,7 +383,7 @@ export function ChatView({ useSession, useSessions, useStore, renderSlot, sessio {running && } - + {!atBottom && (