feat: add economy/maximum presets, tool-lab and subagent-cursor extensions
Some checks failed
CI / windows node 24 / wine blocking (push) Has been skipped
CI / node 22.19 (push) Has been skipped
CI / node 26 (push) Has been skipped
CI / python 3.10 / keyless SDK (push) Has been skipped
CI / python runtime / release-shaped Linux x64 (push) Has been skipped
CI / wine apt cache (push) Successful in 7s
CI / serial / linux (push) Has been skipped
Deploy documentation / build (push) Failing after 1m25s
Deploy documentation / deploy (push) Has been skipped
Landlock Run / Matrix (push) Successful in 5s
Release (vendor) / Pack npm tarballs (push) Failing after 2m47s
Release (dsh) / Pack npm tarballs (push) Failing after 1m56s
Sandbox / sandbox e2e (landlock, ubuntu-24.04) (push) Failing after 1m57s
Sandbox / sandbox e2e (bwrap, ubuntu-latest) (push) Failing after 1m19s
Release (vendor) / Publish to npm (push) Has been skipped
Release (dsh) / Publish to npm (push) Has been skipped
CI / serial / windows (self-hosted standby) (push) Has been cancelled
CI / larger-runner-benchmark (16, linux, dsh-ubuntu-24-04-16core, typecheck) (push) Has been cancelled
Landlock Run / darwin (no platform package — degradation proof) (push) Has been cancelled
Landlock Run / ${{ matrix.platform }} (push) Has been cancelled
CI / node 24 / static (push) Has been cancelled
CI / node 24 / coverage (push) Has been cancelled
CI / node 24 / snapshots and artifacts (push) Has been cancelled
CI / windows node 24 / native complete (push) Has been cancelled
CI / serial / linux (self-hosted standby) (push) Has been cancelled
CI / serial / macos (push) Has been cancelled
CI / larger-runner-benchmark (16, windows, dsh-windows-2025-16core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (32, linux, dsh-ubuntu-24-04-32core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (32, windows, dsh-windows-2025-32core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (4, linux, dsh-ubuntu-24-04-4core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (4, windows, dsh-windows-2025-4core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (64, linux, dsh-ubuntu-24-04-64core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (64, windows, dsh-windows-2025-64core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (8, windows, dsh-windows-2025-8core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (96, linux, dsh-ubuntu-24-04-96core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (96, windows, dsh-windows-2025-96core, production-site) (push) Has been cancelled
CI / consolidated-runner-benchmark (16, linux, dsh-ubuntu-24-04-16core, 16) (push) Has been cancelled
CI / consolidated-runner-benchmark (16, windows, dsh-windows-2025-16core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (32, linux, dsh-ubuntu-24-04-32core, 32) (push) Has been cancelled
CI / consolidated-runner-benchmark (32, windows, dsh-windows-2025-32core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (4, linux, dsh-ubuntu-24-04-4core, 4) (push) Has been cancelled
CI / consolidated-runner-benchmark (4, windows, dsh-windows-2025-4core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (64, linux, dsh-ubuntu-24-04-64core, 32) (push) Has been cancelled
CI / consolidated-runner-benchmark (64, windows, dsh-windows-2025-64core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (8, linux, dsh-ubuntu-24-04-8core, 8) (push) Has been cancelled
CI / consolidated-runner-benchmark (8, windows, dsh-windows-2025-8core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (96, linux, dsh-ubuntu-24-04-96core, 32) (push) Has been cancelled
CI / consolidated-runner-benchmark (96, windows, dsh-windows-2025-96core, 2) (push) Has been cancelled
CI / all checks passed (push) Has been cancelled
Sandbox / sandbox e2e (seatbelt, macos-latest) (push) Has been cancelled
CI / larger-runner-benchmark (8, linux, dsh-ubuntu-24-04-8core, typecheck) (push) Has been cancelled
Sandbox / sandbox e2e (landlock, ubuntu-24.04-arm) (push) Has been cancelled
E2E (real DeepSeek API) / e2e (push) Failing after 1m24s
Some checks failed
CI / windows node 24 / wine blocking (push) Has been skipped
CI / node 22.19 (push) Has been skipped
CI / node 26 (push) Has been skipped
CI / python 3.10 / keyless SDK (push) Has been skipped
CI / python runtime / release-shaped Linux x64 (push) Has been skipped
CI / wine apt cache (push) Successful in 7s
CI / serial / linux (push) Has been skipped
Deploy documentation / build (push) Failing after 1m25s
Deploy documentation / deploy (push) Has been skipped
Landlock Run / Matrix (push) Successful in 5s
Release (vendor) / Pack npm tarballs (push) Failing after 2m47s
Release (dsh) / Pack npm tarballs (push) Failing after 1m56s
Sandbox / sandbox e2e (landlock, ubuntu-24.04) (push) Failing after 1m57s
Sandbox / sandbox e2e (bwrap, ubuntu-latest) (push) Failing after 1m19s
Release (vendor) / Publish to npm (push) Has been skipped
Release (dsh) / Publish to npm (push) Has been skipped
CI / serial / windows (self-hosted standby) (push) Has been cancelled
CI / larger-runner-benchmark (16, linux, dsh-ubuntu-24-04-16core, typecheck) (push) Has been cancelled
Landlock Run / darwin (no platform package — degradation proof) (push) Has been cancelled
Landlock Run / ${{ matrix.platform }} (push) Has been cancelled
CI / node 24 / static (push) Has been cancelled
CI / node 24 / coverage (push) Has been cancelled
CI / node 24 / snapshots and artifacts (push) Has been cancelled
CI / windows node 24 / native complete (push) Has been cancelled
CI / serial / linux (self-hosted standby) (push) Has been cancelled
CI / serial / macos (push) Has been cancelled
CI / larger-runner-benchmark (16, windows, dsh-windows-2025-16core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (32, linux, dsh-ubuntu-24-04-32core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (32, windows, dsh-windows-2025-32core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (4, linux, dsh-ubuntu-24-04-4core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (4, windows, dsh-windows-2025-4core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (64, linux, dsh-ubuntu-24-04-64core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (64, windows, dsh-windows-2025-64core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (8, windows, dsh-windows-2025-8core, production-site) (push) Has been cancelled
CI / larger-runner-benchmark (96, linux, dsh-ubuntu-24-04-96core, typecheck) (push) Has been cancelled
CI / larger-runner-benchmark (96, windows, dsh-windows-2025-96core, production-site) (push) Has been cancelled
CI / consolidated-runner-benchmark (16, linux, dsh-ubuntu-24-04-16core, 16) (push) Has been cancelled
CI / consolidated-runner-benchmark (16, windows, dsh-windows-2025-16core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (32, linux, dsh-ubuntu-24-04-32core, 32) (push) Has been cancelled
CI / consolidated-runner-benchmark (32, windows, dsh-windows-2025-32core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (4, linux, dsh-ubuntu-24-04-4core, 4) (push) Has been cancelled
CI / consolidated-runner-benchmark (4, windows, dsh-windows-2025-4core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (64, linux, dsh-ubuntu-24-04-64core, 32) (push) Has been cancelled
CI / consolidated-runner-benchmark (64, windows, dsh-windows-2025-64core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (8, linux, dsh-ubuntu-24-04-8core, 8) (push) Has been cancelled
CI / consolidated-runner-benchmark (8, windows, dsh-windows-2025-8core, 2) (push) Has been cancelled
CI / consolidated-runner-benchmark (96, linux, dsh-ubuntu-24-04-96core, 32) (push) Has been cancelled
CI / consolidated-runner-benchmark (96, windows, dsh-windows-2025-96core, 2) (push) Has been cancelled
CI / all checks passed (push) Has been cancelled
Sandbox / sandbox e2e (seatbelt, macos-latest) (push) Has been cancelled
CI / larger-runner-benchmark (8, linux, dsh-ubuntu-24-04-8core, typecheck) (push) Has been cancelled
Sandbox / sandbox e2e (landlock, ubuntu-24.04-arm) (push) Has been cancelled
E2E (real DeepSeek API) / e2e (push) Failing after 1m24s
- new economy and maximum agent presets with three-role pipeline skill - new packages/extensions/tool-lab (home-lab ComfyUI/Docling/Whishper tools) - new packages/subagent/subagent-cursor provider - openrouter balance UI with on-demand refresh - session projection context-seed boundary fold - regenerate docs catalogs; keep local searxng benchmark scripts
This commit is contained in:
@@ -0,0 +1,52 @@
|
||||
# Agent Note: Projection units fold against a header-derived context
|
||||
|
||||
Status: implemented
|
||||
|
||||
## Problem
|
||||
|
||||
A forked subagent's session log opens with a verbatim copy of its parent's log. `SessionHeader.seedLength` records that inherited prefix, and consumers that must attribute work to the session that actually did it already slice past it — `dsh-agent`'s inbox and `dsh-schedule`'s domain both do.
|
||||
|
||||
Projection units could not. `ProjectionDefinition.apply(state, event)` received only the event stream, so every unit folded the inherited prefix as if the child had produced it. For `openRouterCost` that made a child's value report its parent's spend: measured on a real fork with `seedLength` 109342, the child folded to $2.0401 over 575 priced steps while its own work was $0.1073 over 25. Summing a parent with its subagents — what the composer dock does to show delegated spend — therefore counted the parent once per forked child, inflating an $8.59 session to $12.79.
|
||||
|
||||
The log alone cannot supply the boundary. `session/end-seed` marks a constructor seed, but a resumed ordinary session appends one too, so the marker cannot distinguish a fork boundary from a resume boundary: the parent session in the measurement above carries four of them. Only the durable header separates the two.
|
||||
|
||||
## Decision
|
||||
|
||||
`ProjectionDefinition.apply` takes a third parameter, `ProjectionFoldContext` — per-session facts derived from the session's durable header, currently just `seedLength`. The registry supplies it on every call, so a unit reads session facts without ever touching a `Session`:
|
||||
|
||||
- `snapshot`, the eager drive, and the lazy cell build derive it from `session.header` through the exported `foldContextOf`.
|
||||
- `restore` takes it as a required fourth parameter, because a detached fold has no Session to derive it from. Its four callers each already hold the stored header: the projection cache's cold read (`tail.meta`, `whole.meta`), api-proxy's detached history baseline, and the subagent catalog's identity probe.
|
||||
|
||||
The parameter is required in the interface and unused by 13 of the 14 shipped units: a two-parameter implementation satisfies a three-parameter signature, so no other unit changed.
|
||||
|
||||
`openRouterCost` skips events below the boundary and bumps `stateVersion` to 4. It still reads `request/context` from the inherited prefix — those records carry no cost, and the route they establish is what attributes the child's first chunk-only step, which would otherwise fold unpriced.
|
||||
|
||||
`CostDock` sums the session plus its subagent subtree over the session list's `parentId`/`origin` rows. That sum is only sound because each value now describes its own session's work.
|
||||
|
||||
## What the context does not decide
|
||||
|
||||
`ProjectionFoldContext` is not a general escape hatch to session state. It carries header facts a pure fold legitimately needs and nothing derived from mutable session state, which would break the synchronous-fold and plain-JSON-state guarantees the persisted cache depends on. A unit describing the whole conversation rather than the session's own work — context pressure, the visible transcript — correctly ignores `seedLength` and folds the inherited prefix like any other event.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Reset the accumulated state on `session/end-seed`.** No framework change, and the marker is already durable. Rejected because a resumed ordinary session appends the same marker: resetting on it would zero a session's spend on every reopen. The marker cannot name which boundary it is.
|
||||
|
||||
**Track a second accumulator for "spend since the last marker" and let the client pick per `origin`.** Contained entirely in this plugin. Rejected because it is wrong for a resumed continuable subagent, whose own earlier segments fall before the last marker and would silently vanish from its total — trading a 20× over-count for a quiet under-count.
|
||||
|
||||
**Deduplicate by the `steps` map's `${turn}:${step}` keys across the lineage.** The inherited steps carry the parent's own keys, so a union would drop them. Rejected because sibling forks both continue from the same fork point and mint colliding keys for their own first steps, and a child seeded with nothing numbers from 0 exactly like its parent.
|
||||
|
||||
**Have the client stop adding forked children entirely.** Never inflated. Rejected because delegated spend is the figure the dock exists to show; hiding it answers the wrong question.
|
||||
|
||||
**Give `apply` the whole `Session`.** Simpler signature change. Rejected because it hands every unit a mutable, non-JSON object and an appendable log, inviting folds that read state outside the event they were given — the exact discipline the unit contract exists to enforce.
|
||||
|
||||
## Consequences
|
||||
|
||||
Every unit's value now means "this session's own work" or "the whole conversation" by explicit choice rather than by accident. The cost is one more parameter on the seam's central function and a required argument on `restore`, which is what makes the omission impossible to reintroduce silently: a caller with no header cannot compile.
|
||||
|
||||
The `stateVersion` bump discards persisted `openRouterCost` rows. Sessions whose rows are dropped and which are never reopened read as absent rather than refolded, because the cold read serves the cache and never folds a cold log — so historical subagent spend stays missing until something opens those sessions. `dsh-client-ui-openrouter-usage` records that gap under its known limitations; a cold-fold aggregate is a separate decision.
|
||||
|
||||
`tokenUsage` and `sessionStats` still fold a forked child's inherited prefix. That is now a deliberate, visible choice rather than an invisible one: whether those figures should describe own work or inherited history is a question for their owners, and the context they need is already in hand.
|
||||
|
||||
## Testing
|
||||
|
||||
`packages/session/session-projection/tests/registry.spec.ts` proves the context reaches `apply` on all three drive paths — eager drive, lazy cell build after events flowed, and `restore` with a caller-supplied context — through a unit that counts only its own events. `packages/llm/openrouter-usage/tests/projection.spec.ts` builds a real fork through `ctx.sessions.create(undefined, { seed, meta: { seedLength } })` and asserts the child excludes the inherited prefix while the parent's figure is untouched, plus that a chunk-only child step still prices from a route recorded in that prefix. `packages/client/ui-openrouter-usage/tests/cost-dock.client.spec.tsx` covers the subtree walk: nested chains, ordinary forks excluded, a descendant with no projection value skipped, and `parentId` cycle termination.
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-22-cursor-external-agent-provider.md
|
||||
2026-08-22-cursor-external-agent-provider.md: 1126ee883fecf56081ac44b00cb70a54f7a8cbe0
|
||||
2026-08-22-cursor-external-agent-provider.zh.md: 0565f29236c4319857444081d6e6f132f93064ee
|
||||
@@ -0,0 +1,57 @@
|
||||
# Agent Note: Cursor joins the external-agent providers, and all three become visible
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-08-22-cursor-external-agent-provider.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The subagent seam already had two external product providers — `codex` over the Codex app-server protocol and `claude-code` over the official Claude Agent SDK — but Cursor had none, and neither existing provider was reachable by a model in a shipped composition: every Agent Preset carried its tool row with `disabled: true`, and the committed base bundle mounted neither provider. The capability existed and nothing could use it.
|
||||
|
||||
Cursor is also the first of the three whose only headless prompt channel is a command-line argument. `cursor-agent --print` takes the task as a POSITIONAL argument, documents no `--` end-of-options separator, and reads no prompt from stdin. Codex sidesteps this entirely — its task travels inside JSON-RPC — and the Claude Code provider hands its prompt to an SDK. A Cursor provider has to decide what happens when model-authored text becomes argv.
|
||||
|
||||
## Decision
|
||||
|
||||
`@deepseek-ai/dsh-subagent-cursor` registers the fixed `cursor` provider over `cursor-agent --print --output-format stream-json`. It reuses the seam's out-of-process vocabulary (`NO_START_CAPABILITIES`, `resolveChildCwd`, `settleRunResult`, `subprocessRunHandle`) and owns only three product-specific things: the print-mode event decoder, the publication gate, and the argv admission rules.
|
||||
|
||||
**Publication gates on `system`/`init`.** Print mode has no handshake, so there is no "remote session exists" moment to wait for. The CLI's first event announces its own chat id after it has resolved credentials and model, which is the earliest point that proves a usable child exists — the structural equivalent of Codex returning an ephemeral thread. A successful `result` arriving with no prior `init` fails startup rather than resolving: a run the caller was never handed has nowhere to put an answer.
|
||||
|
||||
**Argv admission fails loud twice.** A task whose first character is `-` is rejected at admission, because the CLI would parse it as an option and nothing in the seam can escape it. A resolved Windows `.cmd` or `.bat` shim is rejected too: only `cmd.exe` can run one, and its command tail reparses — model-authored task text would become shell syntax. PATHEXT resolution prefers the `cursor-agent.exe` the native Windows installer provides, so the rejection names a fixable installation rather than blocking the platform.
|
||||
|
||||
**Cancellation settles the result, the seam stops the process.** There is no reply channel, so no protocol interrupt exists. The run's abort signal goes to the spawn spec, where the subprocess seam owns the termination escalation, and the same abort is raced into both the startup and result awaits so a cancelled run settles as `aborted` immediately instead of waiting out the grace period. Stdin is closed right after spawn, so a prompt the CLI still tries to read fails fast rather than stalling an unattended child forever.
|
||||
|
||||
**Only `result` completes a run.** The provider accepts `subtype: "success"` with `is_error: false` and a nonblank `result`. Print mode carries no machine-readable failure taxonomy, so every other ending is `error` and the provider produces neither `max-tokens` nor `refusal`. A malformed stdout line is a protocol failure, not something to skip — skipping it would hide a CLI version whose stream this contract cannot read. Unknown event kinds are ignored, because a newer CLI adding events must not break a contract that needs only three of them.
|
||||
|
||||
**All three providers move onto the shipped host plane, and their tool rows turn on.** The base bundle mounts `codex`, `claude-code`, and `cursor` once each, so a Profile enables one by removing `disabled` from its preset tool row instead of mounting a duplicate provider. The `code`, `cordis`, and `standard` presets carry enabled rows. The `economy` preset keeps all three `disabled: true`: its stated purpose is not to reach for external paid agents by default, and that reason survives this change.
|
||||
|
||||
Loading a provider starts no product process, so a deployment without a given CLI keeps an inert tool row whose call fails at call time. That is deliberate: the alternative — gating the row on a PATH probe at load — makes the model's roster depend on host state it cannot see, and turns a missing CLI into a silent absence instead of an answerable error.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Drive Cursor through the existing ACP provider.** `cursor-agent acp` speaks the Agent Client Protocol natively, so `dsh-subagent-acp` can drive it with configuration alone and no new package. Rejected as the primary path because it produces no `subagent_cursor` tool row, no product-specific config, and no product-specific stop-reason or failure mapping — Cursor would be reachable only by a deployment that hand-wrote an ACP row, which is the same invisibility this change exists to remove. The ACP path remains valid and is documented in the package README for a deployment that wants ACP's permission auto-answer policy or a longer-lived remote session.
|
||||
|
||||
**Pass the task through `cmd.exe` on Windows, as the Claude Code provider does for its `.cmd` shim.** That provider quotes the resolved executable into an environment variable that cmd expands once, and its remaining arguments are fixed SDK flags with no cmd metacharacters. Rejected here because the trick does not generalize to the payload: a quoted value containing `"` breaks out of its quoting, and an executable path cannot contain `"` while a model-authored task certainly can.
|
||||
|
||||
**Default `force` to true so a delegated child can actually edit.** Rejected: Cursor's own print-mode default only proposes changes, and the sibling providers are unattended-safe by the same logic — the Codex provider declines approvals and the Claude Code provider disables `AskUserQuestion`. A deployment that wants edits sets `force` explicitly, which is also where the decision is auditable.
|
||||
|
||||
**Pass `CURSOR_API_KEY` as `--api-key`.** The CLI accepts it. Rejected because argv is world-readable in a process listing; the credential goes through the provider's `env` config, which the subprocess seam already treats as a deliberate opt-in past its credential scrub.
|
||||
|
||||
**Add a `--stream-partial-output` delta path.** Rejected as unused surface: without it each `assistant` event is one complete message, which is exactly the seam's "last non-empty assistant message" selection rule with no accumulation state to own.
|
||||
|
||||
**Probe the CLI at load and skip the tool row when absent.** Rejected as above — it makes the model-visible roster a function of invisible host state, and a missing CLI is better reported as a failed call than as a tool that was never there.
|
||||
|
||||
## Consequences
|
||||
|
||||
Three external product agents are now visible to the model in three of the four shipped presets, which is the point and also the cost: each is a tool row the model can choose, and a deployment without the corresponding CLI pays one failed call to learn that. The `economy` preset is unchanged in behavior.
|
||||
|
||||
The committed base-bundle guard inverted. `packages/bundle/base/tests/base.spec.ts` previously asserted that the base layer mounts none of these providers and depends on none of them; it now asserts each is mounted exactly once and declared as a dependency. The old assertion encoded the opt-in-only policy this change replaces.
|
||||
|
||||
Cursor delegation has no continuation. The CLI supports resuming a chat by id, and the provider deliberately does not: the seam's one-shot contract is one process, one run, one result, and a resumable Cursor chat would need durable descriptor fields no consumer asks for yet.
|
||||
|
||||
## Testing
|
||||
|
||||
`packages/subagent/subagent-cursor/tests/subagent-cursor.spec.ts` drives the real event stream through a fake subprocess handle: the init gate, chunk-split and blank-line framing, last-non-empty-message selection with non-text blocks dropped, every unusable terminal result, a result without an announced session, malformed stdout, stream error, end of stream, and close. Its lifecycle cases cover the fixed argv with and without `force`/`trust`, stdin closure, immediate cancellation settling as `aborted` with partial output, post-publication exit and protocol failures flattened through the diagnostic sink, pre-spawn abort, the four startup rollback paths, an aggregate when rollback itself fails, run isolation, and the registered plugin's config, resolved executable, and warning text. The argv admission rules are tested per platform so the suite pins both outcomes on every host.
|
||||
|
||||
`packages/subagent/subagent-cursor/tests/loader-composition.e2e.ts` boots the public opt-in composition with an empty `PATH` and asserts the provider, the `subagent_cursor` schema, and the generic Job controls register while no run starts — loading must not probe or launch a CLI.
|
||||
|
||||
`packages/bundle/base/tests/base.spec.ts` pins the inverted policy: each of the three providers is mounted once and declared once.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Agent Note:Cursor 加入外部 agent 提供方,并让三者同时可见
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-08-22-cursor-external-agent-provider.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
subagent seam 早已拥有两个外部产品提供方——基于 Codex app-server 协议的 `codex`,以及基于官方 Claude Agent SDK 的 `claude-code`——但 Cursor 一个都没有,而且这两个既有提供方在任何已发布的组装中都无法被模型触及:每个 Agent Preset 都把其工具行设为 `disabled: true`,已提交的 base bundle 也没有挂载任何一个。能力存在,却没有任何东西能用它。
|
||||
|
||||
Cursor 还是三者中唯一一个仅以命令行参数作为无头提示词通道的产品。`cursor-agent --print` 按**位置**接收任务,未记载 `--` 选项终止符,也不从标准输入读取提示词。Codex 完全绕开了这个问题——它的任务在 JSON-RPC 内部传输;Claude Code 提供方则把提示词交给 SDK。而 Cursor 提供方必须决定:当模型撰写的文本变成 argv 时会发生什么。
|
||||
|
||||
## 决策
|
||||
|
||||
`@deepseek-ai/dsh-subagent-cursor` 基于 `cursor-agent --print --output-format stream-json` 注册固定的 `cursor` 提供方。它复用 seam 的进程外词汇(`NO_START_CAPABILITIES`、`resolveChildCwd`、`settleRunResult`、`subprocessRunHandle`),只拥有三件与产品相关的事:print 模式事件解码器、发布闸门,以及 argv 准入规则。
|
||||
|
||||
**发布以 `system`/`init` 为闸门。** print 模式没有握手,因此不存在可等待的“远端会话已就绪”时刻。该 CLI 的首个事件会在它解析出凭证与模型之后公布其自有会话 ID,这是能证明存在可用子级的最早时点——在结构上等同于 Codex 返回临时线程。若成功的 `result` 在没有先行 `init` 的情况下到达,启动会失败而非返回结果:调用方从未拿到的运行没有地方安放答案。
|
||||
|
||||
**argv 准入两处显式失败。** 首字符为 `-` 的任务在准入阶段被拒绝,因为该 CLI 会将其解析为选项,而 seam 中没有任何手段可以转义它。解析到 Windows `.cmd` 或 `.bat` 包装脚本同样被拒绝:只有 `cmd.exe` 能运行它,而其命令尾部会被重新解析——模型撰写的任务文本会变成 shell 语法。PATHEXT 解析会优先选择原生 Windows 安装程序提供的 `cursor-agent.exe`,因此这项拒绝指向的是一个可修复的安装问题,而不是封禁该平台。
|
||||
|
||||
**取消负责结束结果,seam 负责停止进程。** 由于没有回复通道,也就不存在协议层中断。本次运行的中止信号被交给 spawn 规格,由子进程 seam 拥有逐级终止机制;同一个中止信号同时被并入启动等待与结果等待的竞态,因此被取消的运行会立即判为 `aborted`,而不必等完宽限期。标准输入在 spawn 后立即关闭,因此该 CLI 若仍尝试读取提示词,会快速失败,而不是让无人值守的子级永远停滞。
|
||||
|
||||
**只有 `result` 能让运行完成。** 提供方只接受 `subtype: "success"` 且 `is_error: false` 并带非空白 `result` 的事件。print 模式不携带可供程序判读的失败分类,因此其他任何结束都是 `error`,且该提供方既不产生 `max-tokens` 也不产生 `refusal`。格式错误的标准输出行属于协议失败,而不是可以跳过的东西——跳过它会掩盖某个本约定无法读取其事件流的 CLI 版本。未知事件类别会被忽略,因为更新版 CLI 新增事件不应破坏一个只需要其中三种事件的约定。
|
||||
|
||||
**三个提供方一并进入已发布的宿主平面,其工具行随之开启。** base bundle 各挂载 `codex`、`claude-code`、`cursor` 一次,因此 Profile 启用其中之一的方式是从其 preset 工具行中删除 `disabled`,而不是挂载重复的提供方。`code`、`cordis` 与 `standard` preset 携带已启用的行。`economy` preset 三者均保持 `disabled: true`:它被明确设定为默认不动用外部付费 agent,而这条理由在本次变更后依然成立。
|
||||
|
||||
加载提供方不会启动任何产品进程,因此缺少某个 CLI 的部署会保留一个惰性工具行,其调用会在调用时失败。这是有意为之:另一种做法——在加载时以 PATH 探测为工具行设闸——会让模型的工具清单取决于它看不见的宿主状态,并把“缺少 CLI”从一个可回答的错误变成一次无声的缺席。
|
||||
|
||||
## 考虑过的替代方案
|
||||
|
||||
**通过既有 ACP 提供方驱动 Cursor。** `cursor-agent acp` 原生讲 Agent Client Protocol,因此 `dsh-subagent-acp` 仅凭配置即可驱动它,无需新包。作为主路径被否决,因为它不产生 `subagent_cursor` 工具行、没有产品专属配置,也没有产品专属的停止原因与失败映射——Cursor 将只能被手写 ACP 行的部署触及,而这正是本次变更要消除的那种不可见性。ACP 路径依然有效,并已在包 README 中记录,供需要 ACP 权限自动应答策略或更长生命周期远端会话的部署使用。
|
||||
|
||||
**在 Windows 上让任务穿过 `cmd.exe`,如 Claude Code 提供方处理其 `.cmd` 包装脚本那样。** 那个提供方把解析出的可执行文件加引号放进环境变量,由 cmd 展开一次,其余参数是不含 cmd 元字符的固定 SDK 标志。此处被否决,因为该技巧无法推广到载荷本身:含 `"` 的加引号值会突破引号,而可执行文件路径不可能含 `"`,模型撰写的任务却完全可能含有它。
|
||||
|
||||
**把 `force` 默认设为 true,好让被委托的子级真能编辑。** 被否决:Cursor 自身的 print 模式默认只提出改动,而同族提供方按同一逻辑保持无人值守安全——Codex 提供方拒绝审批,Claude Code 提供方禁用 `AskUserQuestion`。需要编辑的部署显式设置 `force`,那也正是该决策可被审计的位置。
|
||||
|
||||
**把 `CURSOR_API_KEY` 作为 `--api-key` 传入。** 该 CLI 接受它。被否决,因为 argv 在进程列表中对所有人可读;凭证走提供方的 `env` 配置,子进程 seam 已把该配置视为越过其凭证清除的一次有意选择。
|
||||
|
||||
**新增 `--stream-partial-output` 增量路径。** 作为无人使用的surface被否决:不启用它时,每个 `assistant` 事件就是一条完整消息,这恰好就是 seam 的“最后一条非空助手消息”选择规则,且无需拥有任何累积状态。
|
||||
|
||||
**在加载时探测 CLI,缺失则跳过工具行。** 同上被否决——那会让模型可见的工具清单成为不可见宿主状态的函数,而缺少 CLI 更适合报告为一次失败的调用,而非一个从未存在过的工具。
|
||||
|
||||
## 后果
|
||||
|
||||
三个外部产品 agent 现在在四个已发布 preset 中的三个里对模型可见,这既是目的也是代价:每一个都是模型可以选择的工具行,而缺少对应 CLI 的部署要用一次失败调用来得知这一点。`economy` preset 的行为未变。
|
||||
|
||||
已提交的 base bundle 闸门发生反转。`packages/bundle/base/tests/base.spec.ts` 此前断言 base 层不挂载这些提供方、也不依赖它们;现在它断言每一个都被挂载恰好一次并被声明为依赖。旧断言编码的正是本次变更所替换的“仅可选启用”政策。
|
||||
|
||||
Cursor 委托没有续接。该 CLI 支持按 ID 恢复对话,而提供方有意不做:seam 的 one-shot 约定是一个进程、一次运行、一个结果,可恢复的 Cursor 对话需要目前没有任何消费方要求的持久化描述符字段。
|
||||
|
||||
## 测试
|
||||
|
||||
`packages/subagent/subagent-cursor/tests/subagent-cursor.spec.ts` 通过伪造的子进程句柄驱动真实事件流:init 闸门、分块与空行的帧解析、丢弃非文本块的最后一条非空消息选择、每一种不可用的终止结果、没有公布会话的结果、格式错误的标准输出、流失败、流结束以及关闭。其生命周期用例覆盖带与不带 `force`/`trust` 的固定 argv、标准输入关闭、立即取消并携带部分输出判为 `aborted`、发布后退出与协议失败经诊断出口摊平、spawn 前中止、四条启动回滚路径、回滚自身失败时的聚合错误、运行间隔离,以及已注册插件的配置、解析出的可执行文件与告警文本。argv 准入规则按平台分别测试,因此该套件在任何宿主上都能钉住两种结果。
|
||||
|
||||
`packages/subagent/subagent-cursor/tests/loader-composition.e2e.ts` 以空 `PATH` 启动公开的可选组装,断言提供方、`subagent_cursor` schema 与通用 Job 控制工具均已注册且没有任何运行启动——加载不得探测或启动 CLI。
|
||||
|
||||
`packages/bundle/base/tests/base.spec.ts` 钉住反转后的政策:三个提供方各被挂载一次、各被声明一次。
|
||||
@@ -0,0 +1,35 @@
|
||||
# Agent Note: OpenRouter balance refreshes on click
|
||||
|
||||
Status: implemented
|
||||
|
||||
## Problem
|
||||
|
||||
The account balance badge polled `openRouterUsage.snapshot()` on a 60s interval and rendered whatever the host had last fetched. A user who had just topped up, or who had just watched a long run spend credits, had no way to ask for the current figure: the badge looked like a button, was not one, and the only way to move the number was to wait out the interval.
|
||||
|
||||
Widening the poll interval trades staleness for request volume in both directions and never answers "what is it right now".
|
||||
|
||||
## Decision
|
||||
|
||||
The gateway exposes a second Remote, `refresh()`, which fetches the account immediately and resolves with the resulting snapshot. Concurrent callers share one in-flight fetch, so repeated clicks cost one request, and the shared promise is released once it settles so a later click fetches again. A failed fetch resolves with the last-known snapshot rather than rejecting — the caller compares `updatedAt` if it wants to know whether the figure moved.
|
||||
|
||||
`refresh()` re-reads only the account, not the model pricing table. Pricing is a six-hour table whose staleness the badge does not express, and folding it in would make a click cost a second request for a figure the click is not about.
|
||||
|
||||
`BalanceBadge` calls `refresh` on click, marks itself `aria-busy` for the duration, and ignores further clicks until it settles. The poll continues underneath, so the badge stays current without clicks and a failed refresh is retried by the next tick.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Shorten the poll interval.** No new surface. Rejected because it multiplies requests for every user to serve the moment one user cares, and still cannot answer "right now".
|
||||
|
||||
**Have the click re-read the cached snapshot.** A one-line client change. Rejected because the cached value is exactly what the badge already shows; the click would appear to do something and do nothing.
|
||||
|
||||
**Push balance changes as a forwarded Remote event.** Instant and click-free. Rejected because the host learns of a change only by polling OpenRouter itself, so the event would carry the same staleness with more machinery. The badge's known limitations record that the figure is polled rather than pushed.
|
||||
|
||||
**Refresh pricing alongside the balance.** Rejected as above: a click about the balance should not pay for a table the user cannot see.
|
||||
|
||||
## Consequences
|
||||
|
||||
The badge is now honestly interactive — it looks like a button and behaves like one — at the cost of one more method on the gateway's Remote surface and a user-triggerable outbound request. The in-flight share bounds that: a held-down click is one fetch, not a stream of them.
|
||||
|
||||
## Testing
|
||||
|
||||
`packages/llm/openrouter-usage/tests/loader-composition.spec.ts` moves the mocked `/credits` figures between calls and asserts `refresh()` serves the moved balance, that two concurrent refreshes issue one fetch, and that a later refresh fetches again. `packages/client/ui-openrouter-usage/tests/balance-badge.client.spec.tsx` covers the click path: the fetched figure renders, a click while outstanding is ignored, and a failed refresh keeps the last-known figure and clears the busy state.
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-23-maximum-preset-three-role-pipeline.md
|
||||
2026-08-23-maximum-preset-three-role-pipeline.md: 1ce18e6c9dc68cea0d6bd0f60a5f579466ffc4a2
|
||||
2026-08-23-maximum-preset-three-role-pipeline.zh.md: 8aa375c60a104b5184dbafe4fc330edb793b6a6c
|
||||
@@ -0,0 +1,65 @@
|
||||
# Agent Note: the `maximum` agent preset and its three-role delivery pipeline
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-08-23-maximum-preset-three-role-pipeline.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The shipped preset roster covered a capability RANGE — `minimal` two tools, `standard` the usual coding agent, `code` the same catalog as one TypeScript program, `cordis` plus self-modification, `economy` the same catalog spent carefully — and its upper end was still short of what the deployment can compose. Persistent terminals, language-server queries, session history, durable time and tmux context, scheduled reminders, MCP servers, and the standalone editor each ship as a plugin that no shipped preset mounts, so the only way to get them was to author a preset by hand and rediscover which plane each row belongs to.
|
||||
|
||||
The roster also had no opinion about HOW a change gets made. Every preset offers `subagent` and `subagent_fork` as unshaped delegation: the model decides per call whether to delegate, what to say, and whether anyone checks the result. Specify, build, and verify are three different jobs with three different failure modes, and one agent doing all three in one context reviews its own work with its own assumptions in scope.
|
||||
|
||||
## Decision
|
||||
|
||||
`apps/cli/config/agent-presets/maximum/` is a sixth shipped preset (`order: 6`) that mounts every model-facing row this deployment can compose for one agent, and structures delegation as a fixed three-role pipeline.
|
||||
|
||||
Its catalog is `standard`'s plus: the six `terminal_*` tools over an entry-local PTY realm, `lsp` over an entry-local `lsp` realm, the five `session_*` read tools, `schedule_*`, `str_replace_editor`, the seven `cordis_*` tools, and `run_code` beside the native schemas (`tool-presentation` at `mode: both`). `time-context` and `tmux-context` inject per-step durable context. The three external product agents stay enabled, as they now are in `standard`.
|
||||
|
||||
The pipeline is three `tool-subagent` instances over the one `spawn` provider, distinguished only by `toolName` and the child `persona` the provider installs as a scoped section shadowing `deployment:persona`:
|
||||
|
||||
| Tool | Role | Reply sections |
|
||||
|---|---|---|
|
||||
| `subagent_architect` | reads the repository and specifies the change | GOAL / CONTEXT / PLAN / ACCEPTANCE / RISKS |
|
||||
| `subagent_implementer` | makes the repository satisfy that specification | CHANGES / VERIFICATION / DEVIATIONS |
|
||||
| `subagent_reviewer` | verifies the result against it and returns a verdict | VERDICT / EVIDENCE / DEFECTS |
|
||||
|
||||
All three are `backgroundMode: one-shot` — a stage's reply is the next stage's input, so the default must be the foreground call that returns it — and `maxDepth: 1`, which makes the roles leaves and keeps one pass's agent count equal to the stages run. The preset persona owns sequencing: architect, implementer, reviewer, with a FAIL verdict returning to the implementer and a third failed review going to the user instead of a fourth round. A `three-role-delivery-pipeline` skill ships in the preset's own `skills/` directory with the handoff rules, since a role child cannot see the parent's conversation and everything it needs must be pasted into its prompt.
|
||||
|
||||
The preset's `skill-filesystem` scans two custom roots: its own `skills/`, and the `cordis` preset's, so composition authoring is documented for the `cordis_*` tools it also carries. A copy landing beside no `cordis` directory scans a missing path, which is valid empty state for that provider.
|
||||
|
||||
### Rows that stay off, and why
|
||||
|
||||
Two rows ship `disabled` because what they need is a machine fact, not a deployment fact, and each carries the worked configuration a copy fills in; a third is absent because its name is already taken:
|
||||
|
||||
- **`lsp-stdio`** resolves every configured executable AT LOAD and rolls back every provider when one is missing, so enabling it in a shipped preset would fail the mount on any machine without that language server. The `lsp` service and `tool-lsp` are unconditional instead; without a provider the tool stays in the catalog and answers the structured `LSP_UNAVAILABLE`.
|
||||
- **`mcp-client`** binds one instance to one server — a command to spawn or a URL to reach.
|
||||
- **`tool-bash-persistent`** is absent rather than disabled: it registers under the name `bash`, which this preset's `tool-bash` already owns. `tool-terminal` supplies long-lived sessions under names that do not collide.
|
||||
|
||||
`web_fetch` stays off because the shipped host mounts no fetch provider; that provider defers SSRF protection and the model would choose the request target. Full-text session search stays off because the host mounts `session-query-sqlite` with `openAt: never` — `session_search` and `session_event_search` answer `SESSION_QUERY_SEARCH_DISABLED` while the three read and trace tools work, and changing that is a host patch, not a preset row.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Enforce the roles with `toolFilter` instead of personas.** A read-only architect and reviewer are exactly what `toolFilter` is for, and a persona cannot enforce anything. It is not usable from a shipped preset: the filter names GLOBAL tool names and `tools.restrict()` throws on an unknown one, while the shell tool is `bash` off Windows and `pwsh` on it — one list would be a startup failure on one platform. A deny list over the write tools would also be theater while the shell remains, since a child can write through it. The preset states the restriction in each role persona and the composition comment names the filter as the enforcement a deployment adds for its own platform.
|
||||
|
||||
**Drive the pipeline from a fixed `workflow` script.** `tool-workflow` runs deterministic orchestration, which is what a three-stage pipeline is. Its script comes from the model per call, not from the composition, and its own prompt guidance reserves it for explicit user requests for orchestration; `tool-ralph` takes a build-time script but fans out fresh identical rounds rather than distinct roles. Encoding the order in the persona keeps the stages in the transcript as ordinary tool calls the user can watch, interrupt, and read.
|
||||
|
||||
**One role tool with a `role` argument.** Child policy is fixed per `tool-subagent` instance — another persona means another instance — so a single tool could only pass the role in its prompt, leaving the role's rules as text the parent must remember to repeat rather than a section the provider installs.
|
||||
|
||||
**Add the capabilities to `standard` instead of a sixth preset.** `standard` is the default every new session gets, and this catalog is 53 tools plus a generated Code Mode SDK on every request. Keeping the maximal point separate leaves the default affordable and gives the roster an explicit capability-versus-cost axis, with `economy` at the other end.
|
||||
|
||||
**Mount `schedule` on the host plane, as the Web overlay does.** Its `agent/created` listener is scope-filtered, so a preset-mounted instance installs only on root agents joined to that preset — the property that makes it a legal preset row. Host mounting would hand the tools to every session including `minimal`, which is what the preset boundary exists to prevent.
|
||||
|
||||
## Consequences
|
||||
|
||||
The roster now spans from two tools to every tool, and the preset file is the readable inventory of what this deployment can compose for one agent — including the three rows that need machine-local configuration, each with the worked example a copy edits.
|
||||
|
||||
The cost is real and deliberate: every request carries the full native catalog and the Code Mode SDK, and one pipeline pass is three child agents with their own contexts. The composition header says so and names `standard` and `economy` as the cheaper points.
|
||||
|
||||
The three roles are advisory, not enforced. A reviewer child can write files; only its persona tells it not to. Enforcement requires the platform-specific `toolFilter` a deployment adds to its own copy.
|
||||
|
||||
`apps/cli` gains `dsh-tool-terminal` and `dsh-tool-session-query`: a preset's bare specifiers resolve from the host install's dependency surface, so a row no shipped composition mounted before needs the dependency added there.
|
||||
|
||||
## Testing
|
||||
|
||||
`apps/cli/tests/web-agent-presets.e2e.ts` mounts the preset on the real shipped Web composition and asserts the EXACT tool catalog — the assertion that catches a row registering into the wrong layer, which otherwise mounts cleanly and contributes nothing — plus that each role tool registers under its own name with foreground-by-default semantics. The Web authoring and selection goldens carry the new roster row.
|
||||
@@ -0,0 +1,65 @@
|
||||
# Agent Note:`maximum` agent preset 及其三角色交付流水线
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-08-23-maximum-preset-three-role-pipeline.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
随附的 preset 名册已经覆盖了一条能力**区间**——`minimal` 两个工具、`standard` 常规编码 agent、`code` 把同一份目录变成一个 TypeScript 程序、`cordis` 再加自我修改、`economy` 精打细算地花同一份目录——但它的上限仍然够不到本部署真正能组合出的东西。持久终端、语言服务器查询、会话历史、durable 时间与 tmux 上下文、定时提醒、MCP 服务器、独立编辑器,各自都以插件形式随附,却没有任何随附 preset 挂载它们;要用上就只能手写一个 preset,并重新弄清每一行属于哪个平面。
|
||||
|
||||
名册对**改动如何发生**也没有主张。每个 preset 都提供 `subagent` 和 `subagent_fork` 作为无形状的委派:是否委派、说什么、有没有人复核结果,都由模型逐次决定。规格、实现、验证是三件不同的工作,各有不同的失败方式;一个 agent 在同一个上下文里全做完,就是带着自己的假设复核自己的产出。
|
||||
|
||||
## 决定
|
||||
|
||||
`apps/cli/config/agent-presets/maximum/` 是第六个随附 preset(`order: 6`):它为一个 agent 挂载本部署能组合的全部面向模型的行,并把委派固定成三角色流水线。
|
||||
|
||||
它的目录是 `standard` 再加上:entry 本地 PTY realm 上的六个 `terminal_*` 工具、entry 本地 `lsp` realm 上的 `lsp`、五个 `session_*` 读取工具、`schedule_*`、`str_replace_editor`、七个 `cordis_*` 工具,以及与原生 schema 并存的 `run_code`(`tool-presentation` 取 `mode: both`)。`time-context` 与 `tmux-context` 逐步注入 durable 上下文。三个外部产品 agent 保持启用,与 `standard` 现在的状态一致。
|
||||
|
||||
流水线是同一个 `spawn` 提供方上的三个 `tool-subagent` 实例,彼此只由 `toolName` 和子 `persona` 区分——提供方会把该 persona 作为遮蔽 `deployment:persona` 的 scoped 段落安装到子 agent 上:
|
||||
|
||||
| 工具 | 角色 | 回复段落 |
|
||||
|---|---|---|
|
||||
| `subagent_architect` | 阅读仓库并写出改动规格 | GOAL / CONTEXT / PLAN / ACCEPTANCE / RISKS |
|
||||
| `subagent_implementer` | 让仓库满足该规格 | CHANGES / VERIFICATION / DEVIATIONS |
|
||||
| `subagent_reviewer` | 对照规格验证结果并给出结论 | VERDICT / EVIDENCE / DEFECTS |
|
||||
|
||||
三者都是 `backgroundMode: one-shot`——上一阶段的回复就是下一阶段的输入,所以默认必须是能把结果带回来的前台调用——并且 `maxDepth: 1`,这让角色成为叶子,一次流水线的 agent 数量恰好等于跑过的阶段数。排序由 preset persona 负责:架构师、实现者、审查者;FAIL 结论退回实现者,第三次审查失败则交给用户,而不是开始第四轮。preset 自带的 `skills/` 目录里随附一个 `three-role-delivery-pipeline` skill 说明交接规则,因为角色子 agent 看不到父会话的对话,它需要的一切都必须粘贴进它的提示词。
|
||||
|
||||
preset 的 `skill-filesystem` 扫描两个自定义根目录:它自己的 `skills/`,以及 `cordis` preset 的——这样它同时携带的 `cordis_*` 工具就有了组合编写文档。若某份拷贝落在没有 `cordis` 目录的位置,扫描到的就是一个不存在的路径,对该提供方而言是合法的空状态。
|
||||
|
||||
### 哪些行保持关闭,以及为什么
|
||||
|
||||
有两行以 `disabled` 随附,因为它们需要的是机器事实而非部署事实,并且各自带着拷贝时填写的示例配置;还有一行则因为名字已被占用而根本没有出现:
|
||||
|
||||
- **`lsp-stdio`** 会在**加载时**解析每个已配置的可执行文件,缺一个就回滚全部提供方;因此在随附 preset 里启用它,会让任何没装该语言服务器的机器挂载失败。作为替代,`lsp` 服务与 `tool-lsp` 无条件挂载;没有提供方时工具仍在目录中,并返回结构化的 `LSP_UNAVAILABLE`。
|
||||
- **`mcp-client`** 一个实例绑定一台服务器——要么是待启动的命令,要么是待访问的 URL。
|
||||
- **`tool-bash-persistent`** 是缺席而非禁用:它以 `bash` 之名注册,而这个名字已被本 preset 的 `tool-bash` 占用。长期存活的会话由 `tool-terminal` 以不冲突的名字提供。
|
||||
|
||||
`web_fetch` 保持关闭,因为随附宿主没有挂载 fetch 提供方:该提供方把 SSRF 防护留给别人,而请求目标将由模型选择。会话全文检索保持关闭,因为宿主挂载 `session-query-sqlite` 时取 `openAt: never`——`session_search` 与 `session_event_search` 会返回 `SESSION_QUERY_SEARCH_DISABLED`,三个读取与追踪工具照常可用;改变这一点是宿主 patch,不是 preset 里的一行。
|
||||
|
||||
## 备选方案
|
||||
|
||||
**用 `toolFilter` 而不是 persona 来强制角色。** 只读的架构师与审查者正是 `toolFilter` 的用途,而 persona 强制不了任何事。它在随附 preset 里不可用:过滤器写的是**全局**工具名,`tools.restrict()` 遇到未知名字会抛错,而 shell 工具在非 Windows 上是 `bash`、在 Windows 上是 `pwsh`——同一份名单必然在某个平台上启动失败。而且只要 shell 还在,针对写工具的 deny 名单也只是摆设,子 agent 照样能通过 shell 写入。preset 把限制写进各角色 persona,并在组合注释里指明:过滤器是部署方按自己平台补上的强制手段。
|
||||
|
||||
**用固定的 `workflow` 脚本驱动流水线。** `tool-workflow` 运行确定性编排,而三阶段流水线正是编排。但它的脚本由模型逐次撰写,而非来自组合;它自己的提示词指引也把它限定在用户明确要求编排时使用。`tool-ralph` 接受构建期固定脚本,但它扇出的是一轮轮相同的全新 agent,而不是彼此不同的角色。把顺序写进 persona,则让各阶段以普通工具调用的形式留在 transcript 里,用户可以看、可以打断、可以读。
|
||||
|
||||
**一个带 `role` 参数的角色工具。** 子 agent 策略是每个 `tool-subagent` 实例固定的——换一个 persona 就得换一个实例——所以单一工具只能把角色写在提示词里,角色规则于是退化成父 agent 必须记得每次重复的文本,而不是提供方安装的段落。
|
||||
|
||||
**把这些能力加进 `standard`,而不是新增第六个 preset。** `standard` 是每个新会话拿到的默认值,而这份目录是 53 个工具外加每次请求都要带上的 Code Mode SDK。把最大点单独放置,既让默认值保持可负担,也给名册一条明确的能力—成本轴,`economy` 在另一端。
|
||||
|
||||
**像 Web overlay 那样把 `schedule` 挂在宿主平面。** 它的 `agent/created` 监听是按 scope 过滤的,所以 preset 挂载的实例只会安装到加入该 preset 的根 agent 上——正是这一性质让它成为合法的 preset 行。挂在宿主上则会把这些工具发给包括 `minimal` 在内的每个会话,而这恰恰是 preset 边界要阻止的事。
|
||||
|
||||
## 影响
|
||||
|
||||
名册现在从两个工具一直延伸到全部工具,而这个 preset 文件本身就是一份可读的清单:本部署能为一个 agent 组合出什么,包括那三行需要机器本地配置的行,每行都带着拷贝后可直接编辑的示例。
|
||||
|
||||
代价真实且有意为之:每次请求都携带完整原生目录与 Code Mode SDK,一次流水线是三个各有上下文的子 agent。组合文件的头部注释说明了这一点,并指出 `standard` 与 `economy` 是更便宜的选择。
|
||||
|
||||
三个角色是约定而非强制。审查者子 agent 能写文件,只有它的 persona 让它别写。要强制,就需要部署方在自己的拷贝里加上与平台相符的 `toolFilter`。
|
||||
|
||||
`apps/cli` 新增了 `dsh-tool-terminal` 与 `dsh-tool-session-query` 依赖:preset 的裸 specifier 从宿主安装的依赖面解析,因此此前没有任何随附组合挂载过的行,需要在那里补上依赖。
|
||||
|
||||
## 测试
|
||||
|
||||
`apps/cli/tests/web-agent-presets.e2e.ts` 在真实的随附 Web 组合上挂载该 preset,断言**精确**的工具目录——这正是能抓住「某一行注册到了错误层级」的断言,否则它会干净地挂载却什么也不贡献——并断言每个角色工具以自己的名字注册、默认前台执行。Web 的编写与选择 golden 也带上了新的名册行。
|
||||
Reference in New Issue
Block a user