refactor(cli): exclude tmux context and source guard
This commit is contained in:
@@ -1,6 +0,0 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-27-tmux-location-context.md
|
||||
2026-07-27-tmux-location-context.md: b6f0b0cc85fa4808bd761f30bdbab8ad65e61717
|
||||
2026-07-27-tmux-location-context.zh.md: e79214e03296a87c622df91437385950c00eded6
|
||||
@@ -1,61 +0,0 @@
|
||||
# Agent Note: tmux-location context
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-27-tmux-location-context.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
An agent running inside tmux has no way to tell the model where it is: which session, window, and pane the process occupies, and how the window is laid out. A user directing several panes wants the model to orient itself to its own location so instructions like "the pane below" or "this window" resolve. The location must reach the model as durable, reconstructable context, not a system-prompt value rewritten in place, and must cost nothing when the location has not changed.
|
||||
|
||||
tmux exposes this without a daemon: `$TMUX_PANE` names the process's pane, and `tmux display-message -t "$TMUX_PANE" -p '<format>'` prints any pane/window/session field. The open question was how to observe it — pull on each preparation, or push from a tmux hook — and how to avoid a per-step token cost and hidden process-local state.
|
||||
|
||||
## Decision
|
||||
|
||||
`@deepseek-ai/dsh-tmux-context` is an opt-in function plugin in `packages/context/tmux-context/`, alongside the other bounded request-context enrichments that define neither a tool nor a service. Shipped examples do not mount it because tmux-location disclosure and its token cost are deployment policy.
|
||||
|
||||
**Pull on the first step of each turn, not a tmux push.** The plugin prepends an `agent/step` listener and acts only when `step === 1`. A pull model needs no background process, no hook installation in the user's tmux, and no teardown; it re-reads current state each turn so a moved, renamed, or re-laid-out pane is picked up naturally. Gating on the first step makes the reading per-turn: a location is stable within a turn, and re-querying every step would add cost without new information. A pane moved mid-turn is reflected on the next turn, which is the accepted tradeoff for the simpler design.
|
||||
|
||||
**Read through the `ctx.bash` seam, never raw `child_process`.** The listener runs the tmux/`ps` read commands through `ctx.bash`, so the deployment's sandbox and policy apply and the plugin owns no subprocess code. Absent `ctx.bash`, absent tmux env, a wrong field count, or an empty pane id each make the attempt a no-op, matching how `workspace-context` no-ops without an `fs` provider.
|
||||
|
||||
**Detect a real pane by tty, not by `$TMUX_PANE` alone.** `$TMUX_PANE` is inherited: a terminal launched from a tmux shell (a VS Code integrated terminal, a desktop launcher) carries `$TMUX`/`$TMUX_PANE` from that ancestor even though the process does not live in that pane, which otherwise injects a stale, wrong location. The command resolves this process's controlling terminal with `ps -o tty= -p <pid>` (the agent's own pid, passed in-process) and compares it to the pane's `#{pane_tty}`; fields are emitted only on a match. A genuine pane owns this process's tty; an inherited environment names some other pane's tty and reads as "not in tmux". Checking `$TMUX` instead does not help — it is inherited identically. This is the definitive discriminator and needs no allowlist of terminal emulators.
|
||||
|
||||
**Own location and layout only.** The queried fields are session name, window index/name, pane index/id, window/pane active flags, and `window_layout`. Pane and window pixel sizes are excluded (layout tree conveys structure; sizes are noisy and change on every terminal resize). Sibling-pane contents are never captured (`capture-pane`), keeping the reading small and avoiding scraping unrelated, possibly sensitive, output.
|
||||
|
||||
**Inject only on change, with optional interval floor.** When due, the plugin calls `agent.inject()` for one `user/message` with source `{ kind: 'plugin', plugin: 'tmux-context' }`. Change suppression compares the rendered state block (everything after the turn preamble line) against the latest injection of this source, found by scanning raw durable session events — so the schedule survives compaction and process resume without a process-local cache. The optional `refreshIntervalMs` (manually validated as a non-negative safe integer at plugin load) additionally suppresses injections within that window of the latest one.
|
||||
|
||||
### Text
|
||||
|
||||
```text
|
||||
tmux location (turn <turn>):
|
||||
session <session>, window <index> "<name>", pane <index> <pane-id>
|
||||
window active=<0|1>, pane active=<0|1>, layout <window-layout>
|
||||
```
|
||||
|
||||
The turn preamble is the volatile first line; the two-line state block below it is the unit compared for change suppression, so re-injection is driven by tmux state, not loop position.
|
||||
|
||||
### Durability and request reconstruction
|
||||
|
||||
Each reading is a normal surface node until compaction shadows it; the plugin contributes nothing to system-prompt assembly and `request/header` carries no tmux-context text. The reading records a preparation attempt, not a committed step: because the prepended listener runs first, its append may remain when a later `agent/step` listener cancels or fails the attempt, and the append-only log performs no rollback.
|
||||
|
||||
The published `./invariant` companion registers no runtime check: a reading is a per-turn snapshot of external tmux state, so the session holds no cross-event relation to validate, and scheduling and format stay pinned by the package's pipeline tests.
|
||||
|
||||
## Consequences
|
||||
|
||||
An agent booted inside tmux now receives its own session/window/pane location and window layout as durable, source-attributed context, updated per turn when the location changes. Deployments opt in through cordis.yml; the default spine and shipped examples stay silent. Outside a real tmux pane — including a terminal that merely inherited `$TMUX`/`$TMUX_PANE` — or without a `ctx.bash` executor, the plugin is inert with no error, so composing it is safe everywhere. Because the reading is one durable `user/message`, it survives compaction as ordinary history, contributes nothing to system-prompt assembly or request headers, and costs at most one two-line message per changed turn. The pull model adds one bash execution (through the sandboxed bash seam) on the first step of each turn that is due — internally a `ps` tty probe, a `tmux display-message` tty query, and the field query. Only the optional interval floor suppresses the query itself; an unchanged location is known only after querying, so it suppresses the injection alone.
|
||||
|
||||
## Testing
|
||||
|
||||
Unit tests pin: first-step injection and source/surface metadata; the `$TMUX_PANE`-keyed command including its `#{pane_tty}`-vs-`ps -o tty=` guard; step-gating; change suppression across turns and re-injection on a moved pane; positive-interval suppression and threshold; every no-op path (no bash, nonzero exit, wrong field count, empty pane id, aborted signal); prepended ordering before ordinary `agent/step` listeners; resilience to a corrupt prior reading (non-text block, single-line text); and config rejection of negative and non-integer intervals. Per-file coverage is 100%.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Push from a tmux hook / background watcher** — rejected: requires installing hooks in the user's tmux and a background process with teardown, to gain mid-step freshness that per-turn context does not need.
|
||||
- **Run every step** — rejected: location is stable within a turn; re-querying adds token cost without new information. Gating on `step === 1` yields per-turn readings.
|
||||
- **Raw `child_process`** — rejected: bypasses the sandbox/policy seam and hand-rolls subprocess code the `ctx.bash` executor already owns.
|
||||
- **Include pane/window pixel sizes** — rejected: sizes churn on every resize and add noise; the layout tree already conveys structure.
|
||||
- **Scrape sibling panes with `capture-pane`** — rejected: large, noisy, and privacy-sensitive; out of scope for "own location".
|
||||
- **Dynamic system-prompt section** — rejected: replacing a value erases the earlier readings behind prior reasoning and is not reconstructable; one durable attributed message records each location where it became visible.
|
||||
- **Trust `$TMUX_PANE` (or `$TMUX`) presence** — rejected: both are inherited by terminals launched from a tmux shell (VS Code integrated terminal), so a non-pane process injects a stale location. The pane `#{pane_tty}` vs. this process's controlling tty is the definitive check.
|
||||
- **Denylist known terminal emulators (e.g. `TERM_PROGRAM=vscode`)** — rejected: a partial, ever-growing list that still misses other launchers; the tty match is exact and launcher-agnostic.
|
||||
- **A runtime invariant validating each reading's turn, position, and format** — shipped initially, then removed: it re-derived the producer's own scheduling from the log and asserted a regex over text the same package had just rendered, so it restated `apply()` rather than checking an independent relation. Every failure it could report required an edit to this package, which its pipeline tests already catch. Reintroduce a companion check only for a relation the plugin does not itself compute — for example if readings gain cross-turn ordering or enclosure obligations that another package can violate.
|
||||
@@ -1,61 +0,0 @@
|
||||
# Agent Note:tmux 位置上下文
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-27-tmux-location-context.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
运行在 tmux 内的 agent 无法告诉模型自己身在何处:进程占据哪个 session、window、pane,以及 window 如何布局。当用户操作多个 pane 时,希望模型能对自身位置有所定位,从而让"下方的 pane""这个 window"之类的指令得以解析。位置必须以持久、可重建的上下文形式送达模型,而非在原地被改写的系统提示值,并且当位置未变化时不产生任何成本。
|
||||
|
||||
tmux 无需守护进程即可暴露这些信息:`$TMUX_PANE` 标识进程所在 pane,`tmux display-message -t "$TMUX_PANE" -p '<format>'` 可打印任意 pane/window/session 字段。待决问题在于如何观测——在每次准备时拉取,还是由 tmux hook 推送——以及如何避免逐步骤 token 成本与隐藏的进程内状态。
|
||||
|
||||
## 决策
|
||||
|
||||
`@deepseek-ai/dsh-tmux-context` 是位于 `packages/context/tmux-context/` 的可选启用型函数插件,与其他既不定义工具也不定义服务的有界请求上下文增强并列。随附示例不挂载它,因为 tmux 位置披露及其 token 成本属于部署策略。
|
||||
|
||||
**在每轮的第一个 step 拉取,而非 tmux 推送。** 插件前置注册一个 `agent/step` 监听器,仅在 `step === 1` 时动作。拉取模型无需后台进程、无需在用户的 tmux 中安装 hook、也无需清理;它每轮重新读取当前状态,因此被移动、改名或重新布局的 pane 都会被自然感知。以第一个 step 为门槛使读数按轮次生成:位置在一轮内是稳定的,逐步骤重复查询只会增加成本而不带来新信息。轮次中途移动的 pane 会在下一轮反映,这是换取更简单设计所接受的取舍。
|
||||
|
||||
**通过 `ctx.bash` seam 读取,绝不用裸 `child_process`。** 监听器通过 `ctx.bash` 运行 tmux/`ps` 只读命令,从而应用部署方的沙箱与策略,插件不拥有任何子进程代码。`ctx.bash` 缺失、tmux 环境缺失、字段数不符或 pane id 为空,都会使本次尝试成为空操作,与 `workspace-context` 在无 `fs` provider 时的空操作一致。
|
||||
|
||||
**以 tty 判定真实 pane,而非仅凭 `$TMUX_PANE`。** `$TMUX_PANE` 会被继承:从 tmux shell 启动的终端(VS Code 集成终端、桌面启动器)会从该祖先进程带上 `$TMUX`/`$TMUX_PANE`,即使进程并不位于那个 pane 中,否则就会注入一个陈旧且错误的位置。命令用 `ps -o tty= -p <pid>`(在进程内传入 agent 自身的 pid)解析本进程的控制终端,并与 pane 的 `#{pane_tty}` 比较;只有匹配时才输出字段。真正的 pane 拥有本进程的 tty;继承而来的环境指向的是另一个 pane 的 tty,因而被读作"不在 tmux 中"。改为检查 `$TMUX` 也无济于事——它同样会被继承。这是决定性的判别依据,且无需维护终端模拟器名单。
|
||||
|
||||
**仅自身位置与布局。** 查询字段为 session name、window index/name、pane index/id、window/pane 活动标志以及 `window_layout`。省略 pane 与 window 像素尺寸(布局树已传达结构;尺寸嘈杂且每次终端缩放都会变化)。从不采集相邻 pane 内容(`capture-pane`),使读数保持小巧,并避免抓取无关、可能敏感的输出。
|
||||
|
||||
**仅在变化时注入,并可选间隔下限。** 需要时,插件调用 `agent.inject()` 注入一条来源为 `{ kind: 'plugin', plugin: 'tmux-context' }` 的 `user/message`。变化抑制将渲染出的状态块(轮次前缀行之后的全部内容)与该来源的最近一次注入比较,后者通过扫描原始持久会话事件获得——因此调度可跨压缩与进程恢复存续,无需进程内缓存。可选的 `refreshIntervalMs`(在插件加载时手动校验为非负安全整数)会额外抑制距最近一次注入不足该窗口的注入。
|
||||
|
||||
### 文本
|
||||
|
||||
```text
|
||||
tmux location (turn <turn>):
|
||||
session <session>, window <index> "<name>", pane <index> <pane-id>
|
||||
window active=<0|1>, pane active=<0|1>, layout <window-layout>
|
||||
```
|
||||
|
||||
轮次前缀是易变的首行;其下的两行状态块才是变化抑制所比较的单元,因此重新注入由 tmux 状态驱动,而非循环位置。
|
||||
|
||||
### 持久性与请求重建
|
||||
|
||||
每条读数在被压缩遮蔽前都是普通表层节点;插件对系统提示装配毫无贡献,`request/header` 也不携带任何 tmux-context 文本。读数记录的是一次准备尝试,而非已提交的 step:由于前置监听器最先运行,当后续 `agent/step` 监听器取消或失败时其追加可能仍会保留,只追加的日志不做回滚。
|
||||
|
||||
发布的 `./invariant` 伴生插件不注册任何运行时检查:读数是外部 tmux 状态的按轮快照,会话中不存在需要校验的跨事件关系,调度与格式由本包的管线测试固定。
|
||||
|
||||
## 后果
|
||||
|
||||
启动于 tmux 内的 agent 现在会以持久、带来源标记的上下文收到自身的 session/window/pane 位置及 window 布局,并在位置变化时按轮次更新。部署方通过 cordis.yml 选择启用;默认 spine 与随附示例保持沉默。在真实 tmux pane 之外——包括仅继承了 `$TMUX`/`$TMUX_PANE` 的终端——或没有 `ctx.bash` 执行器时,插件保持惰性且不报错,因此在任何地方组合它都安全。由于读数是一条持久的 `user/message`,它作为普通历史经受压缩,对系统提示装配与请求头毫无贡献,且每个发生变化的轮次至多花费一条两行消息。拉取模型在每个到期轮次的第一个 step 增加一次 bash 执行(经沙箱化的 bash seam)——内部包含一次 `ps` tty 探测、一次 `tmux display-message` tty 查询和字段查询。只有可选的间隔下限会抑制查询本身;位置是否变化只有在查询之后才知道,因此它只抑制注入。
|
||||
|
||||
## 测试
|
||||
|
||||
单元测试固定了:首个 step 的注入及来源/表层元数据;以 `$TMUX_PANE` 为键的命令(含其 `#{pane_tty}` 与 `ps -o tty=` 的比对守卫);step 门槛;跨轮次的变化抑制与 pane 移动时的重新注入;正间隔抑制与阈值;每条空操作路径(无 bash、非零退出、字段数不符、pane id 为空、信号已取消);前置排序先于普通 `agent/step` 监听器;对损坏的历史读数(非文本块、单行文本)的容错;以及配置对负值与非整数间隔的拒绝。逐文件覆盖率为 100%。
|
||||
|
||||
## 考虑过的替代方案
|
||||
|
||||
- **由 tmux hook / 后台监视器推送**——否决:需要在用户的 tmux 中安装 hook,并引入带清理的后台进程,只为换取按轮次上下文并不需要的步内新鲜度。
|
||||
- **每个 step 都运行**——否决:位置在一轮内稳定;重复查询只增加 token 成本而无新信息。以 `step === 1` 为门槛得到按轮次读数。
|
||||
- **裸 `child_process`**——否决:绕过沙箱/策略 seam,并手写 `ctx.bash` 执行器已拥有的子进程代码。
|
||||
- **包含 pane/window 像素尺寸**——否决:尺寸每次缩放都变动、徒增噪声;布局树已传达结构。
|
||||
- **用 `capture-pane` 抓取相邻 pane**——否决:庞大、嘈杂且涉及隐私;超出"自身位置"范围。
|
||||
- **动态系统提示区块**——否决:替换某个值会抹去支撑先前推理的历史读数且不可重建;单条持久且带来源的消息在每个位置变得可见时予以记录。
|
||||
- **信任 `$TMUX_PANE`(或 `$TMUX`)存在即可**——否决:两者都会被从 tmux shell 启动的终端(VS Code 集成终端)继承,于是非 pane 进程会注入陈旧位置。pane 的 `#{pane_tty}` 与本进程控制终端的比对才是决定性检查。
|
||||
- **对已知终端模拟器设黑名单(如 `TERM_PROGRAM=vscode`)**——否决:名单不完整且会不断增长,仍会漏掉其他启动器;tty 比对精确且与启动器无关。
|
||||
- **用运行时 invariant 校验每条读数的轮次、位置与格式**——最初随包发布,随后移除:它从日志中重新推导生产者自身的调度,并对同一个包刚刚渲染出的文本断言正则,因此只是重述 `apply()`,而非检查一条独立关系。它能报出的每种失败都必须先修改本包,而这些本包的管线测试已经覆盖。仅当出现插件自身并不计算的关系时才重新引入伴生检查——例如读数将来具备可被其他包破坏的跨轮次顺序或包裹义务。
|
||||
@@ -1,6 +0,0 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-28-source-guard-staging-edit-gate.md
|
||||
2026-07-28-source-guard-staging-edit-gate.md: 8452006190166c783efafc398566ef7f4da10323
|
||||
2026-07-28-source-guard-staging-edit-gate.zh.md: 83589ce833e4aa74968b40247848685a1e030f6b
|
||||
@@ -1,76 +0,0 @@
|
||||
# Agent Note: source-guard denies direct staging-checkout edits
|
||||
|
||||
Status: implemented
|
||||
|
||||
English | [中文](2026-07-28-source-guard-staging-edit-gate.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The [`dsh-customize`](../../../../skills/dsh-customize/SKILL.md) skill governs every personal change to a dsh source checkout: implement in a task worktree branched from the staging tip, then integrate under `.agents/merge.lock`. Its central rule is negative — do not edit the personal staging checkout directly — and a negative rule delivered only as prompt text fails in exactly the case that matters. An agent that never loads the skill never sees the rule, and one that loads it early can still forget it thirty tool calls later. The failure is silent and expensive: commits land on the staging branch that the launcher runs from, outside any task branch, with no lock held and no rollback worktree.
|
||||
|
||||
Prompt guidance cannot fix this, because the guidance is what went unread. The rule needs an enforcement point.
|
||||
|
||||
## Decision
|
||||
|
||||
`@deepseek-ai/dsh-source-guard` (`packages/guard/source-guard/`) is a `tools/pre-execute` listener that returns `{kind: 'deny', reason}` for a `write` or `edit` whose target resolves inside a protected staging worktree, unless the calling session's durable log already records a successful `skill` call naming `dsh-customize`. It registers no service and contributes no prompt text or tool schema; an allowed call is indistinguishable from one made without the plugin. It is not in any shipped default composition.
|
||||
|
||||
### Git identity from files, not a path prefix and not `git`
|
||||
|
||||
Whether a path is protected is decided by reading `.git`, its `gitdir:` pointer, and `HEAD`. Three shapes resolve: a plain clone (`.git` is a directory that is its own common dir), a linked worktree (`.git` is a file pointing at `<common>/worktrees/<name>`, whose common dir is two levels up), and a detached HEAD (`HEAD` holds a raw object id and names no branch). A `gitdir:` pointer resolves whether absolute — what `git worktree add` writes — or relative, which git resolves against the worktree directory holding it.
|
||||
|
||||
Denial requires the target's worktree to match the launcher's on both identities: the same shared git directory and the same branch. Both come from resolving `protectedCheckout`, so nothing about the protected branch is configured. An earlier revision matched a `dsh-staging/*` name pattern instead; the exact-branch rule replaced it because a pattern is wrong in both directions. It denied every sibling staging worktree an old install had left behind, none of which runs a launcher, and it silently protected nothing for a maintainer whose staging branch follows no naming convention — a fatal property for a shipped default that must hold for checkouts [`scripts/install.sh`](../../../../scripts/install.sh) did not create.
|
||||
|
||||
A path-prefix rule would have been wrong, not merely imprecise. The task worktrees the skill prescribes live *inside* the protected tree at `<staging>/.worktrees/...`, so a prefix rule would deny every edit the workflow requires. Resolution walks outward from the target and stops at the first enclosing worktree, so it reports the innermost one: a nested task worktree answers with its own task branch and is allowed, while the launcher's own tree answers with the launcher's branch and is denied.
|
||||
|
||||
Two path details decide whether the gate holds at all, and both are enforcement, not polish. Repository identity is compared on symlink-resolved paths (`canonicalPath` from `dsh-sandbox`), because a session cwd under `/var/...` and a configured path under `/private/var/...` are the same macOS directory and a lexical comparison would fail open on every write. And a relative `file_path` is resolved against the calling session's workspace, exactly as `dsh-tool-fs` resolves it; judging only absolute paths would have left a relative path as an unguarded route to a protected file.
|
||||
|
||||
`protectedCheckout` names a path inside the guarded checkout, defaulting to this module's own file. That resolves the checkout the running harness was launched from — the live deployment, whatever its branch is named. A harness running from an installed copy resolves a different repository, or none, and guards nothing; the rule is meaningless outside a source checkout.
|
||||
|
||||
The shipped TUI composition loads the plugin with these defaults, so every source install is protected without configuration. It is inert for an ordinary project: a workspace in another repository, or none, never matches the launcher's identities.
|
||||
|
||||
### Satisfaction replayed from the durable log
|
||||
|
||||
The gate lifts on a `tool/call` naming the `skill` tool whose arguments parse to `{name: <requiredSkill>}`, paired by call id with a non-error `tool/result`. Both fields are already durable (`packages/core/session/src/types.ts`), so this needs no new session event and no coupling to skill-provider internals.
|
||||
|
||||
The log is the only state. In-memory satisfaction (the `WeakMap` shape [`repeat-tool-guard`](../../archived/feature/2026-07-08-repeat-tool-guard.md) uses for its chains) would be smaller, but it loses satisfaction on resume: a resumed session that already read the skill would be told to read it again, and the denial would look like a bug rather than a rule. Replay costs a scan bounded by the first hit and buys resume correctness.
|
||||
|
||||
### Fail open, deliberately
|
||||
|
||||
A path outside any worktree, a detached HEAD, a foreign repository, a malformed `gitdir:` pointer, and unreadable metadata all leave the call to the rest of the chain. The alternative — denying whenever git identity is unavailable — converts any `.git` permission problem into a harness that cannot write files at all. The guard exists to prevent one specific, recoverable mistake; it must not become a larger outage than the mistake.
|
||||
|
||||
### Narrow scope
|
||||
|
||||
`read` is never gated: inspecting staging violates nothing, and the skill explicitly permits read-only questions. `bash` is not gated either. Reliably classifying mutating shell commands is a matcher problem with no honest completion condition, so a determined model can still change staging through a shell. This is a boundary against forgetting, not a sandbox against intent.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Advisory reminder instead of denial** (`additionalContexts` on `tools/post-execute`, the `repeat-tool-guard` shape). Rejected: the write has already happened when the reminder arrives, so the violation is committed and the guidance is again just text.
|
||||
- **`{kind: 'ask'}` routed to approval.** Rejected: it prompts on every legitimate task-worktree edit in the common case, and degrades to denial in a composition without approval support, making behavior depend on unrelated plugins.
|
||||
- **Running `git rev-parse` through `ctx.subprocess`.** Rejected after measuring the alternative: two file reads answer the same question with no process spawn per gated write, no `git` on `PATH` requirement, and no subprocess dependency. Reading `.git` and `HEAD` is a stable on-disk format, not an implementation detail.
|
||||
- **Explicit `protectedRoots` config with no detection.** Rejected: it makes the common case require configuration to be correct, and a stale absolute path silently disables protection.
|
||||
- **A configurable staging-branch name pattern** (`stagingBranchPatterns`, default `dsh-staging/*`). Shipped first, then removed: it protects the wrong set in both directions — every stale sibling worktree that runs no launcher, and nothing at all for a maintainer whose branch is named otherwise. Deriving the branch from the launcher needs no configuration and cannot be misconfigured.
|
||||
- **Auto-detecting the checkout with no override.** Rejected: the detection is a default, not a law; a deployment guarding a different checkout, or running from an installed copy, needs the explicit value.
|
||||
- **Denying everything under the checkout root, `.worktrees/` included.** Rejected: it blocks the workflow the skill prescribes, so the guard would fire on every legitimate task edit.
|
||||
- **Gating `bash` with a mutating-command matcher.** Deferred, not rejected: worth revisiting if bypasses are observed in practice. A matcher that is wrong in either direction is worse than an honestly narrow gate.
|
||||
|
||||
## Consequences
|
||||
|
||||
The rule now holds without depending on the model having read it, and the denial names the path, the branch, and the skill, so the model's next action is determined rather than guessed. Enforcement sits at the operation boundary that owns the decision, so it cannot be bypassed by prompt filtering or listener order.
|
||||
|
||||
Shipping it in the TUI default means every source install is protected without configuration, and the protection follows the launcher across upgrades because the branch is derived rather than named. The cost of that reach is that the plugin loads for every user, including those whose workspace it can never match.
|
||||
|
||||
What it cost otherwise: the guard is only as complete as its tool list, and `bash` remains open. Worktree identity is cached per directory for the plugin's lifetime, so a mid-session branch switch is not observed on either side. Only the launcher's own checkout is protected, so a stale sibling stays editable. Loading the skill lifts the gate for the whole session without verifying the workflow was actually followed — the gate proves the instructions were read, not obeyed. Satisfaction is per session, so a subagent with its own session must load the skill itself.
|
||||
|
||||
## Testing
|
||||
|
||||
Unit suites drive a real agent loop against a mock adapter over real git-metadata fixtures — a staging worktree, a task worktree nested inside it, a plain clone, a foreign repository on a staging-named branch, a detached HEAD, absolute and relative `gitdir:` pointers, a symlinked route to one repository, a malformed pointer, and unreadable metadata — covering both source files to per-file 100%. A companion `invariant.ts` validates the durable denial's shape, since the refusal text is the package's only model-visible output and is actionable only when it names the path, branch, and skill.
|
||||
|
||||
The real-composition smoke boots `examples/headless-agent/tests/fixtures/guard/source-guard/cordis.yml` through the Loader and the headless app, and asserts three things about the assembled run: the tool result is an error, its text is the exact denial, and the targeted file still holds its original bytes — enforcement before dispatch, not advice after it.
|
||||
|
||||
An ACP snapshot scenario (`source-guard-staging-deny`) originally owned the assembled transcript, seeding a staging worktree in the harness's generated cwd through a new `Scenario.prepareCwd` hook — git never tracks an entry named `.git` and `.gitignore` excludes every `worktrees/` directory, so the fixture committed the two `HEAD` bodies and the hook assembled the real layout. Authoring it paid for itself immediately: it exposed both path defects above (the transcript showed `fs-policy` answering first wherever the guard had quietly declined to judge) and then caught its own first fixture, whose ignored `worktrees/` path passed locally from an untracked file. The scenario was later removed with the assembled-run evidence consolidated into the Loader-composition smoke; the `prepareCwd` hook it introduced remains part of the snapshot harness for repository-shaped fixtures.
|
||||
|
||||
## Related
|
||||
|
||||
- [The personal-staging maintenance skills Agent Note](../process/2026-07-23-personal-staging-maintenance-skills.md) — the workflow this gate enforces one rule of. That note owns the skills' content and discovery; this one owns the enforcement point and holds no authority over the workflow itself.
|
||||
- [The interception-seams Agent Note](2026-06-30-interception-seams.md) — the `tools/pre-execute` `allow`/`deny`/`ask` vocabulary this gate's denial uses.
|
||||
- [The repeat-tool-guard Agent Note](../../archived/feature/2026-07-08-repeat-tool-guard.md) — the sibling guard whose advisory shape this one deliberately does not take.
|
||||
@@ -1,76 +0,0 @@
|
||||
# Agent Note: source-guard 拒绝直接编辑 staging 检出目录
|
||||
|
||||
Status: implemented
|
||||
|
||||
[English](2026-07-28-source-guard-staging-edit-gate.md) | 中文
|
||||
|
||||
## Problem
|
||||
|
||||
[`dsh-customize`](../../../../skills/dsh-customize/SKILL.md) skill(技能)规范对 dsh 源码检出的每项个人变更:先在从 staging 分支顶端分出的任务 worktree 中实现,再于 `.agents/merge.lock` 保护下完成集成。它的核心规则是一条禁令:不得直接编辑个人 staging 检出目录;而若只通过提示词文本传达禁令,它恰好会在最要紧的场景中失效。从未加载该 skill 的 agent(智能体)根本看不到规则;即便尽早加载,也仍可能在三十次工具调用后将其忘掉。这种失败既静默又代价高昂:提交会落在启动器实际运行的 staging 分支上,不属于任何任务分支,既未持有锁,也没有用于回滚的 worktree。
|
||||
|
||||
提示词指导无法解决这个问题,因为未被阅读的正是这些指导。该规则需要一个强制执行点。
|
||||
|
||||
## Decision
|
||||
|
||||
`@deepseek-ai/dsh-source-guard`(`packages/guard/source-guard/`)是一个 `tools/pre-execute` 监听器;它会返回 `{kind: 'deny', reason}`,拒绝目标解析到受保护 staging worktree 内的 `write` 或 `edit`,除非调用会话的持久日志已经记录过一次成功的 `skill` 调用,且名称为 `dsh-customize`。它不注册服务,也不贡献提示词文本或工具 schema;获准的调用与未加载该插件时的调用没有区别。任何已交付的默认组合都不包含它。
|
||||
|
||||
### 从文件而非路径前缀或 `git` 判定 Git 身份
|
||||
|
||||
系统通过读取 `.git`、其中的 `gitdir:` 指针以及 `HEAD` 来判断路径是否受保护。它可以解析三种形态:普通克隆(`.git` 是目录,且自身就是共享目录)、链接 worktree(`.git` 是文件,指向 `<common>/worktrees/<name>`,其共享目录位于上两级)以及 HEAD 分离状态(`HEAD` 保存原始对象 id,不指向任何分支)。`gitdir:` 指针无论是绝对路径(`git worktree add` 写入的形式)还是相对路径都可以解析;Git 会以包含该指针的 worktree 目录为基准解析相对路径。
|
||||
|
||||
只有目标的 worktree 在两项身份上都与启动器的 worktree 匹配时才会拒绝:共用同一个共享 Git 目录,且分支相同。这两项身份均通过解析 `protectedCheckout` 得出,因此无需配置受保护分支的任何信息。较早版本则匹配 `dsh-staging/*` 名称模式;现已改用确切分支规则,因为模式会在两个方向上出错。它会拒绝旧安装留下的每一个同级 staging worktree,尽管其中没有任何一个运行着启动器;对于 staging 分支不遵循任何命名约定的维护者,它又会静默地完全不提供保护——而已交付的默认配置必须在 [`scripts/install.sh`](../../../../scripts/install.sh) 未创建的检出目录上也能生效,这一属性是致命的。
|
||||
|
||||
路径前缀规则不仅不精确,而且本身就是错误的。该 skill 规定的任务 worktree 位于受保护树*内部*的 `<staging>/.worktrees/...`,因此前缀规则会拒绝工作流要求的每一次编辑。解析过程从目标向外逐层查找,遇到第一个所属 worktree 时停止,因此返回最内层的 worktree:嵌套的任务 worktree 会返回自身的任务分支并获准,而启动器自身所在的树会返回启动器的分支并被拒绝。
|
||||
|
||||
有两个路径细节决定门禁究竟能否生效,二者都是强制执行要求,而非细节润色。仓库身份会按解析符号链接后的路径进行比较(使用 `dsh-sandbox` 的 `canonicalPath`),因为位于 `/var/...` 下的会话 cwd 和位于 `/private/var/...` 下的配置路径在 macOS 上是同一个目录,若按路径字符串比较,每次写入都会故障放行(fail-open)。此外,相对 `file_path` 会完全按照 `dsh-tool-fs` 的方式,相对于调用会话的工作区解析;若只判断绝对路径,相对路径就会成为绕过门禁访问受保护文件的路径。
|
||||
|
||||
`protectedCheckout` 指定受保护检出目录内的一条路径,默认值为本模块自身的文件。由此解析出运行中 harness 的启动来源检出目录——当前运行的部署,无论其分支采用什么名称。若 harness 从已安装副本运行,解析出的会是另一个仓库或没有仓库,因此不会保护任何内容;该规则在源码检出之外没有意义。
|
||||
|
||||
已交付的 TUI 组合会以这些默认值加载插件,因此每个源码安装无需配置即可受到保护。对于普通项目,它不会生效:若工作区位于其他仓库中,或不存在工作区,就绝不会匹配启动器的身份。
|
||||
|
||||
### 从持久日志回放满足状态
|
||||
|
||||
如果日志中存在一条 `tool/call`,它调用名为 `skill` 的工具,参数可解析为 `{name: <requiredSkill>}`,且按调用 id 能配对到非错误的 `tool/result`,门禁即解除。二者都已持久化(`packages/core/session/src/types.ts`),因此无需新增会话事件,也不与 skill 提供方内部实现耦合。
|
||||
|
||||
日志是唯一状态。在内存中记录满足状态(`WeakMap` 结构,[`repeat-tool-guard`](../../archived/feature/2026-07-08-repeat-tool-guard.md) 将其用于调用链)所需实现会更小,但恢复后满足状态会丢失:一个已经读取过该 skill 的恢复会话会被要求再次读取,而这次拒绝看起来会像缺陷而不是规则。回放的代价是扫描日志,但首次命中即停止,并换来恢复行为正确。
|
||||
|
||||
### 刻意采用故障放行
|
||||
|
||||
目标路径不在任何 worktree 内、HEAD 分离、属于其他仓库、`gitdir:` 指针格式错误或元数据不可读时,调用都会交给调用链的其余部分处理。反过来,只要无法判定 Git 身份就拒绝,会让任何 `.git` 权限问题都导致 harness 完全无法写文件。该 guard 旨在防止一种特定且可恢复的错误,不得造成比该错误更严重的故障。
|
||||
|
||||
### 范围收窄
|
||||
|
||||
`read` 从不受门禁限制:检查 staging 不会违反任何规则,而且该 skill 明确允许只读提问。`bash` 同样不受门禁限制。要可靠判定哪些 shell 命令会修改状态,需要构造一个无法给出可信完备标准的匹配器,因此执意修改的模型仍可通过 shell 修改 staging。这是一道防止遗忘的边界,不是阻止刻意操作的沙箱。
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **用建议性提醒代替拒绝**(使用 `additionalContexts`,挂载在 `tools/post-execute` 上,采用 `repeat-tool-guard` 的形态)。不予采纳:提醒到达时写入已经发生,违规已成事实,而指导又一次沦为纯文本。
|
||||
- **将 `{kind: 'ask'}` 交给审批。** 不予采纳:在常见场景中,它会对任务 worktree 内每次合法编辑都发起询问;在没有审批支持的组合中还会退化为拒绝,使行为取决于无关插件。
|
||||
- **运行 `git rev-parse`,并通过 `ctx.subprocess` 执行。** 对替代方案进行实测后不予采纳:读取两个文件即可回答同一问题,每次受门禁限制的写入都无需 spawn 进程,不要求 `git` 存在于 `PATH` 中,也不依赖子进程。读取 `.git` 与 `HEAD` 所依据的是稳定的磁盘格式,而非实现细节。
|
||||
- **显式配置 `protectedRoots`,不做检测。** 不予采纳:这会让常见场景的保护效果依赖配置正确性,而陈旧的绝对路径会静默禁用保护。
|
||||
- **可配置的 staging 分支名称模式**(`stagingBranchPatterns`,默认 `dsh-staging/*`)。最初随产品交付,随后删除:它从两个方向划错了保护范围——既纳入每个不运行启动器的陈旧同级 worktree,又完全不保护分支另有名称的维护者。由启动器派生分支无需配置,也不可能配置错误。
|
||||
- **自动检测检出目录,不提供覆盖项。** 不予采纳:检测只是默认行为,而非不可更改的规定;若部署要保护另一个检出目录,或自身从已安装副本运行,就需要显式值。
|
||||
- **拒绝检出根目录下的一切操作,包括 `.worktrees/`。** 不予采纳:这会阻断该 skill 规定的工作流,让 guard 在每次合法任务编辑时触发。
|
||||
- **用修改类命令匹配器把守 `bash`。** 推迟而非否决:如果实际观察到绕过行为,值得重新考虑。任一方向判断错误的匹配器,都不如如实限定范围的门禁。
|
||||
|
||||
## Consequences
|
||||
|
||||
如今,该规则无需依赖模型已经读过它也能生效;拒绝理由会列出路径、分支与 skill,让模型的下一步操作明确,无需猜测。强制执行位于拥有该决策的操作边界,因此提示词过滤或监听器顺序都无法绕过它。
|
||||
|
||||
将其纳入 TUI 默认组合意味着每个源码安装无需配置即可受到保护;由于分支是派生而非按名称指定,保护会在升级时跟随启动器。这种覆盖范围的代价是插件会为每位用户加载,包括工作区永远不可能匹配启动器身份的用户。
|
||||
|
||||
除此之外的代价是:guard 的完整程度受限于其工具列表,`bash` 仍保持开放。worktree 身份在插件生命周期内按目录缓存,因此无法观察到任一侧在会话中途切换分支。只保护启动器自身的检出目录,因此陈旧的同级检出目录仍可编辑。加载该 skill 会为整个会话解除门禁,却不会验证工作流是否确实得到遵循——门禁只能证明指令已被阅读,不能证明已被执行。满足状态按会话隔离,因此拥有独立会话的 subagent 必须自行加载该 skill。
|
||||
|
||||
## Testing
|
||||
|
||||
单元测试套件基于真实 Git 元数据 fixture(测试前置数据),使用 mock 适配器驱动真实 agent loop(智能体循环):覆盖一个 staging worktree、嵌套其中的任务 worktree、普通克隆、位于 staging 命名分支上的其他仓库、HEAD 分离状态、绝对和相对 `gitdir:` 指针、指向同一仓库的符号链接路径、格式错误的指针以及不可读元数据,使两个源码文件都达到逐文件 100% 覆盖率。配套的 `invariant.ts` 会验证持久拒绝的结构,因为拒绝文本是该包唯一面向模型的输出,且只有其中列出路径、分支和 skill 时才具有可操作性。
|
||||
|
||||
真实组合冒烟测试通过 Loader 与 headless 应用启动 `examples/headless-agent/tests/fixtures/guard/source-guard/cordis.yml`,并对组装后的运行断言三项事实:工具结果是错误、文本与拒绝理由逐字一致、目标文件仍保留原始字节。这证明系统在分发前强制执行规则,而不是事后给出建议。
|
||||
|
||||
一个 ACP(Agent Client Protocol)快照场景(`source-guard-staging-deny`)最初负责组装后的 transcript(文本记录),通过新的 `Scenario.prepareCwd` 钩子在 harness 生成的 cwd 中植入 staging worktree——Git 永远不会跟踪名为 `.git` 的条目,且 `.gitignore` 会排除所有 `worktrees/` 目录,因此 fixture 提交两个 `HEAD` 的内容,由钩子组装真实布局。编写它立刻证明了投入的价值:它暴露了上述两个路径缺陷(transcript 显示每当 guard 悄然不作判断时 `fs-policy` 都会率先响应),随后又发现了自身首版 fixture 的问题——被忽略的 `worktrees/` 路径因未跟踪文件而在本地通过。该场景后来被移除,组装运行证据合并进 Loader 组合冒烟测试;它引入的 `prepareCwd` 钩子仍留在快照 harness 中,服务于仓库形态的 fixture。
|
||||
|
||||
## Related
|
||||
|
||||
- [个人 staging 维护 skill 的 Agent Note](../process/2026-07-23-personal-staging-maintenance-skills.md):本门禁负责执行该工作流的一条规则。对方 Agent Note 负责这些 skill 的内容与发现机制;本文只负责强制执行点,对工作流本身不具有定义权。
|
||||
- [拦截 seam Agent Note](2026-06-30-interception-seams.md):本门禁拒绝时使用的 `tools/pre-execute` `allow`/`deny`/`ask` 词汇。
|
||||
- [repeat-tool-guard Agent Note](../../archived/feature/2026-07-08-repeat-tool-guard.md):同类 guard;本文刻意不采用其建议性形态。
|
||||
Reference in New Issue
Block a user