docs: trim generated prose
This commit is contained in:
@@ -33,7 +33,7 @@ The editor sets each session's `cwd` to the project it opens; both the agent's b
|
||||
|
||||
## Snapshot tests (record-once / replay-deterministic)
|
||||
|
||||
This example is the home of the harness's **snapshot tests** — they boot this server as a real subprocess, drive it with a deterministic input script, and diff its normalized output against committed golden files. The model is made deterministic by `@deepseek-ai/dsh-llm-replay`, a function/namespace plugin that installs an `llm/stream` waterfall listener and short-circuits it, serving model streams reconstructed from a recorded **session JSONL** fixture (`<scenario>/session.jsonl`) — so replay needs no API key. The fixture IS the persisted session log: its `assistant/chunk` events carry every `StreamChunk`, so grouping them by `(turn, step)` reconstructs each `stream()` call (one model call per loop step). Recording is therefore "run the real agent once and harvest the `.jsonl`"; use `pnpm run test:snapshot:record` when the model transcript itself should change, and `pnpm run test:snapshot:refresh` when the committed model transcript is still the right mock input and only the current replay output/goldens need to be rewritten. The two failure modes not expressible as logged chunks — a pure throw before any chunk, and cancel/hang — use an optional `<scenario>/replay.override.json` sidecar (a `ReplayEntry[]` that replaces the derived script). A scenario that needs the agent to operate on existing files ships an optional `<scenario>/workspace/` directory — the harness copies its contents into the temp cwd before the run (see `workspace-edit`). See [the ACP snapshot tests RFC](../../docs/rfc/implemented/testing/2026-06-19-acp-snapshot-tests.md) for the full design.
|
||||
This example hosts the ACP snapshot suite. `dsh-llm-replay` reconstructs model streams from `assistant/chunk` events in each scenario's session JSONL, so replay is keyless. Recording runs the real agent and harvests that log; refresh keeps the committed transcript as mock input and rewrites current replay outputs. `replay.override.json` covers throw and hang cases that chunks cannot express, and an optional `workspace/` seeds files. The [snapshot RFC](../../docs/rfc/implemented/testing/2026-06-19-acp-snapshot-tests.md) owns the full design.
|
||||
|
||||
## MVP limitations
|
||||
|
||||
|
||||
@@ -26,27 +26,12 @@ import {
|
||||
* WITHOUT a key, since it only needs the server to boot and answer initialize.
|
||||
*/
|
||||
|
||||
// The dsh-acp-agent bin (the demo:acp entry) and this example's cordis.yml. The
|
||||
// bin resolves its config-path arg from CWD; the subprocess runs from a temp
|
||||
// workdir, so pass the example config's ABSOLUTE path.
|
||||
// The dsh-acp-agent bin (the demo:acp entry) and this example's cordis.yml.
|
||||
const binScript = fileURLToPath(new URL('../../../packages/ui/acp-agent/src/bin.ts', import.meta.url))
|
||||
const configPath = fileURLToPath(new URL('../cordis.yml', import.meta.url))
|
||||
// Resolve tsx's loader to an ABSOLUTE path: the subprocess runs with cwd set to
|
||||
// a temp workdir (this test launches there and uses it as the session cwd; the
|
||||
// bridge no longer requires cwd === the launch dir, but a temp dir keeps the
|
||||
// test hermetic), where a bare `--import tsx` would not resolve from
|
||||
// node_modules. import.meta.resolve gives the worktree's tsx regardless of cwd.
|
||||
// Resolve tsx absolutely because the subprocess runs outside the repo.
|
||||
const tsxLoader = fileURLToPath(import.meta.resolve('tsx'))
|
||||
// Absolute path to the repo-root tsconfig. Dev/test/demo run UNBUILT: the
|
||||
// `@deepseek-ai/dsh-*` workspace imports resolve through the `paths` map in the
|
||||
// root tsconfig (tsx reads it), NOT through built `lib/` output. But tsx finds
|
||||
// that tsconfig by searching UP from the child's cwd — and the child's cwd is a
|
||||
// temp workdir OUTSIDE the repo, so the search misses and the dsh-* imports fail
|
||||
// (the child dies before writing a byte). Point tsx at the repo tsconfig
|
||||
// explicitly via TSX_TSCONFIG_PATH so resolution is cwd-independent. (Without
|
||||
// this the suite only passed by accident when a stale built `lib/` happened to
|
||||
// exist — exactly the contamination that masked the inject bug this suite now
|
||||
// guards.) The repo root is four levels up from this file (examples/acp-agent/tests).
|
||||
// Absolute path to the repo-root tsconfig.
|
||||
const repoTsconfig = fileURLToPath(new URL('../../../tsconfig.json', import.meta.url))
|
||||
|
||||
interface Spawned {
|
||||
@@ -155,9 +140,6 @@ describe('acp-agent over real stdio (no key required)', () => {
|
||||
it('emits only framed JSON-RPC on stdout', async () => {
|
||||
workdir = await mkdtemp(join(tmpdir(), 'acp-e2e-'))
|
||||
// Collect raw stdout bytes directly (bypass the SDK framing) to inspect.
|
||||
// A dummy key lets the deepseek adapter APPLY (it only checks the key is
|
||||
// present at boot, not valid — the key is used only on a real model call,
|
||||
// which this purity test never triggers). So this runs WITHOUT real creds.
|
||||
const child = spawn(process.execPath, ['--import', tsxLoader, binScript, configPath], {
|
||||
cwd: workdir,
|
||||
env: {
|
||||
@@ -196,17 +178,10 @@ describe('acp-agent over real stdio (no key required)', () => {
|
||||
}, 30_000)
|
||||
|
||||
it('session/new succeeds over real stdio (no model call)', async () => {
|
||||
// REGRESSION GUARD (this exact RPC crashed a real Zed session with
|
||||
// "cannot get property \"agents\" without inject"): `session/new` drives the
|
||||
// full bridge → `ctx.agents.create({sessionId, meta:{cwd}})` → AgentLoop →
|
||||
// registry/persistence path, ALL of which run from the JSON-RPC read loop
|
||||
// OUTSIDE the bridge plugin's injection scope. A lazy `ctx.<service>` read
|
||||
// on that path throws and the RPC fails with an Internal error — yet the
|
||||
// call never touches the model, so this reproduces WITHOUT a key. The
|
||||
// key-gated prompt test below never caught it (it needs real creds); the
|
||||
// initialize-only purity test never caught it (initialize does not reach
|
||||
// the factory). This closes that gap: boot the real subprocess and create a
|
||||
// session, asserting the RPC RESOLVES (not rejects with an inject error).
|
||||
// Regression guard (this exact RPC crashed a real Zed session with "cannot get property
|
||||
// \"agents\" without inject"): `session/new` drives the full bridge →
|
||||
// `ctx.agents.create({sessionId, meta:{cwd}})` → AgentLoop → registry/persistence path, ALL
|
||||
// of which run from the JSON-RPC read loop outside the bridge plugin's injection scope.
|
||||
workdir = await mkdtemp(join(tmpdir(), 'acp-e2e-'))
|
||||
// A dummy key lets the deepseek adapter boot (it only checks presence, not
|
||||
// validity, at apply time); no model call is made, so the key is never used.
|
||||
@@ -245,12 +220,9 @@ describe.skipIf(!process.env.DEEPSEEK_API_KEY)('acp-agent e2e: real prompt over
|
||||
const toolCalls = updates.filter(u => u.sessionUpdate === 'tool_call')
|
||||
expect(toolCalls.length).toBeGreaterThan(0)
|
||||
|
||||
// Tool-call UI quality (the tool owns its presentation): the bash tool's
|
||||
// `presentCall` sets the title to the exact command (an execute card hides
|
||||
// rawInput, so the command IS the title) — NOT the bare tool name "bash".
|
||||
// A `bash` call must therefore carry an execute kind, a non-"bash" title,
|
||||
// and a string rawInput (the command). `toolCalls` is already narrowed to
|
||||
// the `tool_call` shape by the filter above, so these fields are reachable.
|
||||
// Tool-call UI quality (the tool owns its presentation): the bash tool's `presentCall` sets
|
||||
// the title to the exact command (an execute card hides rawInput, so the command IS the
|
||||
// title) — not the bare tool name "bash".
|
||||
const bashCall = toolCalls.find(u => u.kind === 'execute')
|
||||
expect(bashCall).toBeDefined()
|
||||
if (bashCall === undefined) throw new Error('expected an execute tool_call')
|
||||
|
||||
@@ -75,31 +75,12 @@ const SCENARIOS: Scenario[] = [
|
||||
// child runs as a spawn subagent under the worker-thread engine (its session is the
|
||||
// child fixture), and the tool result carries the script's return value.
|
||||
{ name: 'workflow-run', hasModelTurn: true, recorded: true, childSessions: 1 },
|
||||
// Hook matrix — one scenario per hook point × its headline Decision outcome,
|
||||
// across BOTH bridges (Claude `hooks.json`, Codex `codex-hooks.json`, seeded in
|
||||
// workspace/). The block scenarios need no model call: a UserPromptSubmit hook
|
||||
// blocks the prompt before any step runs (keyless, authored — the derived
|
||||
// script is empty so no sidecar), yet persists a `rejected` turn carrying
|
||||
// `hook/*` events, so their logs ARE compared. Every other point fires a real
|
||||
// seam mid-turn, so its transcript is recorded WITH the hook active.
|
||||
// Hook matrix — one scenario per hook point × its headline Decision outcome, across BOTH
|
||||
// bridges (Claude `hooks.json`, Codex `codex-hooks.json`, seeded in workspace/).
|
||||
{ name: 'hook-cc-promptsubmit-block', hasModelTurn: false, comparesLog: true, recorded: false },
|
||||
{ name: 'hook-codex-promptsubmit-block', hasModelTurn: false, comparesLog: true, recorded: false },
|
||||
// The mid-turn seams fire during a real model turn, so each is recorded WITH
|
||||
// its hook active (the model's reaction to a deny/block/force-continue is part
|
||||
// of the captured transcript). The Codex bridge exercises the same seams in its
|
||||
// own snake_case dialect.
|
||||
//
|
||||
// Two hook points are deliberately NOT snapshotted, and stay on the bridges'
|
||||
// unit coverage (`bridge.spec.ts` / `coverage.spec.ts`) instead:
|
||||
// - SessionStart and SubagentStart inject context through a detached,
|
||||
// best-effort `void runPoint(...).then(agent.inject())` with no turn
|
||||
// binding, so the resulting `context/message` races the work it precedes
|
||||
// and lands at a nondeterministic log position — a recorded golden does not
|
||||
// even reproduce on its own replay.
|
||||
// - SubagentStop is observe-only with no turn and no injection, so it writes
|
||||
// NOTHING to the transcript — a golden would be byte-identical to the
|
||||
// no-hook run and could never be proven to fail.
|
||||
// See the hook-snapshot-matrix RFC for the full rationale.
|
||||
// The mid-turn seams fire during a real model turn, so each is recorded with its hook active
|
||||
// (the model's reaction to a deny/block/force-continue is part of the captured transcript).
|
||||
{ name: 'hook-cc-promptsubmit-context', hasModelTurn: true, recorded: true },
|
||||
{ name: 'hook-cc-pretool-deny', hasModelTurn: true, recorded: true },
|
||||
{ name: 'hook-cc-pretool-ask', hasModelTurn: true, recorded: true },
|
||||
@@ -114,11 +95,9 @@ const SCENARIOS: Scenario[] = [
|
||||
{ name: 'hook-codex-posttool-block', hasModelTurn: true, recorded: true },
|
||||
{ name: 'hook-codex-posttool-context', hasModelTurn: true, recorded: true },
|
||||
{ name: 'hook-codex-stop-continue', hasModelTurn: true, recorded: true },
|
||||
// Code Mode: the registry in `mode: code` — the wire tool list collapses to
|
||||
// [run_code], the tools:sdk section rides in the prompt, and the program's
|
||||
// tool calls land as tool/code-dispatch events. Each mode boots its own
|
||||
// overlay config, composes a different header by construction, and
|
||||
// therefore pins its own class.
|
||||
// Code Mode: the registry in `mode: code` — the wire tool list collapses to [run_code], the
|
||||
// tools:sdk section rides in the prompt, and the program's tool calls land as
|
||||
// tool/code-dispatch events.
|
||||
{ name: 'code-mode-turn', hasModelTurn: true, recorded: true, pinsHeader: true, headerClass: 'code', configPath: CODE_MODE_CONFIG },
|
||||
{ name: 'both-mode-turn', hasModelTurn: true, recorded: true, pinsHeader: true, headerClass: 'both', configPath: BOTH_MODE_CONFIG },
|
||||
]
|
||||
|
||||
@@ -17,21 +17,8 @@ import {
|
||||
} from '@agentclientprotocol/sdk'
|
||||
|
||||
/**
|
||||
* With-key e2e: the Claude Code hook bridge running against the REAL acp-agent
|
||||
* subprocess and the REAL model. The example `cordis.yml` loads `dsh-hooks-claude`
|
||||
* with a PROCESS-LEVEL `configPath` of `./hooks.json`, resolved once at load
|
||||
* against the ACP server's launch cwd (NOT per-session); this test sets that
|
||||
* launch cwd to the temp workspace and writes a `hooks.json` there with a
|
||||
* PreToolUse hook that BLOCKS every bash command, then asks the live model to
|
||||
* write a file — and verifies the WORLD (the file never appears on disk),
|
||||
* proving the hook actually intercepted execution rather than the agent merely
|
||||
* claiming it couldn't. (The hook itself then runs in the session cwd.)
|
||||
* Key-gated; owns and disposes its subprocess.
|
||||
*
|
||||
* A keyless companion lives in acp.e2e.ts (stdout purity + session/new); the
|
||||
* full hook-fires-end-to-end transcript is the keyless `hook-cc-promptsubmit-block`
|
||||
* snapshot scenario. This one closes the "green plumbing, broken product" gap:
|
||||
* only a real model deciding to call bash exercises the PreToolUse seam live.
|
||||
* With-key e2e: the Claude Code hook bridge running against the real acp-agent subprocess and
|
||||
* the real model.
|
||||
*/
|
||||
|
||||
const binScript = fileURLToPath(new URL('../../../packages/ui/acp-agent/src/bin.ts', import.meta.url))
|
||||
@@ -89,9 +76,7 @@ afterEach(async () => {
|
||||
describe.skipIf(!process.env.DEEPSEEK_API_KEY)('acp-agent e2e: a PreToolUse hook blocks bash (real model)', () => {
|
||||
it('denies every bash command, so the requested file is never written (verified on disk)', async () => {
|
||||
workdir = await mkdtemp(join(tmpdir(), 'acp-hooks-e2e-'))
|
||||
// A PreToolUse hook that blocks EVERY tool (exit 2, no matcher = match-all).
|
||||
// The session cwd is `workdir`, and the bridge resolves `./hooks.json` from
|
||||
// the process cwd (the launch dir = workdir), so this is the config it loads.
|
||||
// A PreToolUse hook that blocks every tool (exit 2, no matcher = match-all).
|
||||
await writeFile(join(workdir, 'hooks.json'), JSON.stringify({
|
||||
hooks: { PreToolUse: [{ hooks: [{ type: 'command', command: 'echo "bash blocked by policy" >&2; exit 2' }] }] },
|
||||
}))
|
||||
|
||||
Reference in New Issue
Block a user