Files
deepseek-harness/docs/rfc/implemented/testing/2026-06-22-subagent-snapshot-replay.md
Tianyi Cui f393043b03 Add the ACP subagent backend: out-of-process delegation (PR3)
The first OUT-OF-PROCESS subagent backend, proving the seam generalizes past the
in-process backends. @deepseek-ai/dsh-subagent-acp runs each child agent in a
spawned subprocess, driven over the Agent Client Protocol as the CLIENT — the
direction-inverted twin of the dsh-acp server bridge. Point the configured
command at the acp-agent example and the harness talks to its own process.

- Fresh process per run: start spawns, runs one ACP session (initialize →
  newSession → prompt), dispose kills the subprocess and awaits its exit.
- Minimal client stub: advertises no fs/terminal; accumulates agent_message_chunk
  text as the result output; auto-answers session/request_permission by a
  configured policy (reject default / allow). No start-time capabilities (an
  out-of-process child can't enforce the parent's depth/tool-filter); ignores
  request.parent; injects only `subagents`.
- StopReason mapping (end_turn→completed, cancelled→aborted, …); result resolves
  error/aborted on a child failure, never rejects (seam contract).
- Security: credential-shaped ambient env vars are scrubbed; the child's own key
  is forwarded only via explicit config.env. A spawn-level error (ENOENT) is
  captured and raced against the ACP drive so a bad command settles error rather
  than crashing the parent.

Testing designed at every tier: keyless integration drives a scripted mock ACP
server subprocess (cancellation incl. the pre-newSession race and a
torn-pipe-after-cancel, permission auto-answer, non-message updates, spawn
failure, HMR, export shape) at 100% coverage; a with-key e2e drives the REAL
acp-agent example process (PONG + real file write, verified on disk) — the
harness driving itself. Snapshot coverage of an ACP child is deferred as
TODO(acp-subagent-replay) (each child is its own process with its own replay).

Stayed on @agentclientprotocol/sdk 0.25.1: the proposed 0.28.x bump only
deprecates the stable ClientSideConnection/AgentSideConnection API this layer
uses (33 sites incl. the server bridge), turning no-deprecated red across code
this PR shouldn't rewrite — that fluent-API migration is its own follow-up. The
backend needs nothing 0.28.x adds.

This completes the subagent seam stack (PR1 interface → PR2 in-process → PR2.5
snapshot infra → PR3 ACP); the seam RFC moves to implemented/, amended.
2026-06-22 10:47:02 +08:00

6.8 KiB

RFC: Per-session snapshot replay for nested agents

Status: implemented

Problem

The snapshot tier (pnpm run test:snapshot) boots the real acp-agent subprocess, replays a recorded session through dsh-llm-replay, and diffs the normalized stdout transcript + re-persisted session log against committed goldens. It is the only tier that exercises the full editor-facing transcript end to end.

It was built for ONE session per process, and that assumption is wired into two places:

  • dsh-llm-replay keyed nothing. It served the Nth llm/stream call the Nth recorded entry from a single global cursor. With a parent agent AND an in-process subagent both streaming on one context, the calls interleave and the single cursor hands the child the parent's script (and vice versa).
  • The harness harvested one log. findSessionLog walked the sessions root and returned the FIRST .jsonl it found. A subagent runs as a second Session with its own log in the same cwd bucket, so the child's transcript was silently dropped.

This was the TODO(subagent-snapshots) deferral recorded in the subagent seam RFC: the in-process backends (PR2) shipped with unit + e2e coverage, but the full-transcript snapshot tier could not express a nested-agent shape until this infrastructure landed. This RFC is that stacked follow-up.

Decision

Replay is keyed per calling session, and the harness harvests every session log.

1. The calling session id rides on the model request

GenerateOptions gains an optional sessionId, stamped by the agent loop from agent.session.id at request-assembly time (where the session is already in scope). Adapters ignore it; it exists so an llm/stream listener can route a call by WHICH session issued it. It is typed Branded<'SessionId'> (from dsh-brand) rather than importing SessionId from dsh-session — that package imports Message from dsh-llm, so importing its id back would cycle. SessionId IS Branded<'SessionId'>, so a real id assigns with no cast. (A future dedicated ids package could own the brand and dissolve the note; tracked separately — it touches every id import and does not belong in this testing PR.)

2. Replay binds live sessions to recorded scripts by first-call order

A nested scenario records more than one log: the parent (session.jsonl) plus one per subagent child (session.1.jsonl, …). dsh-llm-replay loads them all, derives one script per recorded session, and orders the scripts by header createdAt (the parent is created before its children).

Live session ids are freshly random every run and never equal the recorded ones, so a live session cannot bind to a script by id equality. Instead it binds by first-call order: the first live session to make any model call claims the first ordered script (the parent — earliest createdAt, and necessarily the first to stream, because it must run a turn before it can delegate), the next new live session claims the next script, and so on. Each session then advances its own cursor independently.

This keys by WHO calls, not by global call order — so it stays correct even if subagents ever run concurrently or in the background (a global cursor would interleave them). A call carrying no sessionId (a direct unit-test stream()) is treated as one anonymous session bound to the primary script, so the single-session path is byte-for-byte the old behavior. More distinct live sessions than recorded scripts is a fail-loud error (an unrecorded subagent appeared), never a silent mis-route.

The ordering key is the session header createdAt. In the current synchronous cut this is sound because sibling children are created strictly sequentially — the subagent tool awaits one child's result and disposes it before the parent's next tool call starts the next child — so their createdAt values are strictly ordered and match first-call order exactly. A same-millisecond sibling tie is therefore unreachable; the recordedId tiebreak only keeps such a degenerate collision deterministic, it does not recover first-call order. A future cut that runs siblings concurrently/backgrounded WOULD be able to create two children in the same millisecond, and must then thread a real first-call ordinal (the order live sessions first stream) rather than leaning on createdAt — flagged with XXX(concurrent-subagents) at the sort site.

The alternative considered and rejected was a call-ordered merge of the parent and child logs into one global script (sound only because in-process subagent execution is strictly nested — the parent blocks on the child). It is simpler for today's synchronous cut but bakes in the parent-blocks-on-child invariant that a future backgrounded/concurrent subagent would break; per-session keying does not.

3. The harness harvests every log, primary-first

harvestSessionLogs collects every .jsonl across every cwd bucket under the sessions root (the JSONL backend puts a parent and its same-cwd child in the same bucket), parses each header, and orders them primary-first: the top-level session (no parentSession) leads, then each child by ascending createdAt. RunResult.sessionLogs is the plural result; the spec writes each back to its fixture on record (session.jsonl + session.<n>.jsonl) and diffs each harvested log against its fixture on replay. The normalizer already accepted plural session ids and collapses any stray UUID, so no normalizer change was needed.

4. Scenarios

Two nested scenarios were added and recorded against the real API:

  • subagent-spawn — the parent delegates one subtask via the subagent tool to a fresh spawn child (2 sessions).
  • subagent-multi — the parent delegates two subtasks, each to its own spawn child (3 sessions), stressing the per-session keying with three concurrent scripts and the createdAt ordering of two children under one parent.

Both replay keyless in the default gate.

Consequences

  • The TODO(subagent-snapshots) deferral is resolved: nested-agent transcripts are now a first-class snapshot shape.
  • GenerateOptions.sessionId is a small, honest core-seam addition useful beyond replay (telemetry, request routing).
  • The subagent tool is bound to a single provider, so both children in subagent-multi are spawn (fresh). The fork backend is loaded in the example and exercised by PR2's unit tests; a mixed spawn+fork snapshot would need a second tool instance bound to fork (pure config) and is a trivial future addition, not a gap in the keying — the keying routes by session, not by backend.
  • Out-of-process (ACP) subagents are a different replay shape entirely (each child is its own PROCESS with its own replay), tracked as TODO(acp-subagent-replay) in the PR3 plan.