Add in-process subagent backends: spawn (fresh) and fork (seeded)

The second PR of the subagent seam: the two in-process backends that run a
child agent on the same cordis context, reusing the agent factory's quiescent
AgentHandle teardown. Both register on ctx.subagents (PR1's named-provider
registry) and share one run driver.

- dsh-subagent-spawn: a FRESH child via ctx.agents.create — own session, the
  parent's model by default (overridable), zero inherited conversation. Also
  exports the shared in-process run driver (startInProcessRun): mint ids, stamp
  cwd/parentSession-lineage/depth, drive the one-shot (send → whenIdle), read
  the last assistant/message + turn/end reason, dispose to quiescence.
- dsh-subagent-fork: a child SEEDED with the parent's balanced completed-turn
  prefix (the log up to and including its last turn/end), so the child inherits
  context. The in-flight unbalanced turn is excluded — a raw seed would fail the
  invariants replay. Proven: a regression test goes red if the boundary seeds
  the open turn.
- Seam extension: CreateAgentOptions.seed, threaded through AgentLoop.createAgent
  → ctx.sessions.prepare({ seed }) (the primitive resume already used). This is
  the fork-lineage path the TODO(sub-agents) markers anticipated.
- Depth: a merge-extensible AgentOptions.subagentDepth (0 top-level, parent+1 for
  a child); the depthLimit capability refuses a spawn past request.maxDepth.

Tests: real-loop unit tests for both backends (mock MODEL only, real loop +
invariants), a multi-subagent test (one parent drives a fork AND a spawn child
then keeps working), and a with-key e2e (a real parent delegates via the
`subagent` tool to a real child that writes a file on disk — world-verified).
100% per-file coverage. The coding-agent demo wires the spawn backend + tool.

Snapshot coverage of nested agents is deferred to a stacked follow-up
(TODO(subagent-snapshots)): dsh-llm-replay is a single global positional cursor
that cannot route calls to a parent vs. a child on one context. Recorded in the
RFC's deferrals and a new AGENTS.md rule: designing a subsystem must design its
test infrastructure END TO END up front, verifying the snapshot/e2e harness can
express the new shape — a gap this plan hit.
This commit is contained in:
Tianyi Cui
2026-06-22 05:58:40 +08:00
parent 861791d2d8
commit 7aabd2a3df
28 changed files with 1296 additions and 13 deletions

View File

@@ -211,6 +211,16 @@ async function* replayEntry(entry: ReplayEntry, signal: AbortSignal | undefined)
* the snapshot harness runs one ACP session per scenario to guarantee that. The
* cursor is advanced synchronously at listener-invocation time (not lazily
* inside the generator) so call ORDER, not iteration order, fixes the mapping.
*
* TODO(subagent-snapshots): this single global cursor cannot route calls to the
* right agent when a parent and an in-process subagent both stream on one ctx.
* Snapshot coverage of nested agents needs either per-session-keyed replay (a
* `Map<sessionId, cursor>` fed by the calling agent on the `agent/request`
* waterfall, which carries the agent) or a call-ordered merge of the parent and
* child session logs (sound because subagent execution is strictly nested —
* the parent blocks on the child). Tracked as a stacked follow-up to the
* in-process subagent backends; see the subagent RFC's "Snapshot coverage of
* nested agents" deferral.
*/
export function installLlmReplay(ctx: Context, config: ReplayConfig): () => void {
const entries = loadReplayScript(config)