feat(mode): exit_plan_mode + the ACP session-mode picker + scriptable review answers
Plan mode's stage 2 (RFC 2026-07-07-plan-mode). The exit tool: one required plan argument (the durable log artifact), execute re-checks the folded mode, then conducts the review over the user-interaction seam — one single-select question (Approve / Keep planning) with free text open — so an approval appends mode/set back to default in-turn and every other outcome (keep-planning feedback verbatim, aborted, no provider) returns the corrective isError with the mode unchanged. presentCall is a generic card titled by the plan's first heading carrying the plan markdown; over ACP the review rides the ask_user elicitation flow, in the terminal the stdio prompt queue — no approval-seam dependency. The ACP bridge maps the picker 1:1 onto ctx.modes (opportunistic, a type-only peer edge): session/new + session/load advertise availableModes/currentModeId, session/set_mode validates through set() and echoes an optimistic current_mode_update (the pending mode IS the selection; the logged mode/set lands at the boundary and, matching, is not re-sent), and a session/event listener re-notifies on each logged flip that differs from the last sent — the tool-driven exit updates the picker. The feature matrix rows move from 'not modeled' to the picker-to-modes / knobs-to-config-options division, with the ACP v2 removal direction recorded as a mechanical-migration risk. The snapshot harness gains the setMode/setModeExpectError ops and a scripted elicitationAnswers FIFO (cancel on exhaustion; a stray choice string reaches the agent verbatim as a non-consenting custom answer, so a scenario bug fails safe). The suite factory's header-pin requirement now applies only to model-turn scenarios — a protocol-only suite has no header content to anchor. examples/plan-acp-agent is the live composition; its keyless modes-advertise scenario pins the wire surface (advertisement, both set_mode round-trips, unknown-id rejection). The recorded plan-mode approve/reject arc awaits a with-key recording session; its texts are pinned at the unit tier meanwhile. examples/AGENTS.md ceiling 653 → 680: the new example's required smoke row does not fit the old budget.
This commit is contained in:
@@ -35,4 +35,4 @@ A scenario booting a differently-composed tree sets its own `configPath` (an ove
|
||||
|
||||
The example also ships a `cordis.snapshot.yml` replay overlay next to its `cordis.yml` (the bin swaps them under `DSH_SNAPSHOT=replay` — [single-source replay config RFC](../../../docs/rfc/implemented/testing/2026-07-04-single-source-acp-replay-config.md)); replay fixtures are served by [`dsh-llm-replay`](../llm-replay/README.md), which this package points at via the `DSH_SNAPSHOT_*` env vars it sets on the child. Fixture roles, record/replay semantics, and scenario-table fields are documented on `Scenario` and in the [snapshot RFC](../../../docs/rfc/implemented/testing/2026-06-19-acp-snapshot-tests.md).
|
||||
|
||||
Constraints: `suite.ts` imports vitest, so the package is importable only inside a vitest run (the harness and normalizers have no such dependency but ship from the same entry). ACP-specific by design — the harness speaks the SDK's `ClientSideConnection`. Permission round-trips are scriptable: `InputScript.permissionAnswers` is a FIFO queue of option-kind selections (`allow_once`, `reject_once`, …) the client maps to the agent-issued `optionId` at answer time; an absent or exhausted queue answers `cancelled`, and a kind the request never offered rejects the run (the agent is answered `cancelled`, so a tolerant agent cannot absorb the scenario bug).
|
||||
Constraints: `suite.ts` imports vitest, so the package is importable only inside a vitest run (the harness and normalizers have no such dependency but ship from the same entry). ACP-specific by design — the harness speaks the SDK's `ClientSideConnection`. Permission round-trips are scriptable: `InputScript.permissionAnswers` is a FIFO queue of option-kind selections (`allow_once`, `reject_once`, …) the client maps to the agent-issued `optionId` at answer time; an absent or exhausted queue answers `cancelled`, and a kind the request never offered rejects the run (the agent is answered `cancelled`, so a tolerant agent cannot absorb the scenario bug). Elicitation round-trips (`ask_user_question` / the plan review) script the same way: `InputScript.elicitationAnswers` is a FIFO of `{ action, choice?, custom? }` form answers; exhaustion answers `cancel`, and a stray `choice` string reaches the agent verbatim as a non-consenting custom answer, so a scenario bug fails safe in the transcript.
|
||||
|
||||
@@ -29,6 +29,8 @@ import {
|
||||
PROTOCOL_VERSION,
|
||||
type Agent as AcpAgent,
|
||||
type Client,
|
||||
type CreateElicitationRequest,
|
||||
type CreateElicitationResponse,
|
||||
type RequestPermissionRequest,
|
||||
type RequestPermissionResponse,
|
||||
type SessionNotification,
|
||||
@@ -85,6 +87,8 @@ export type InputStep =
|
||||
| { op: 'promptExpectError'; text: string }
|
||||
| { op: 'promptAndCancel'; text: string }
|
||||
| { op: 'cancel' }
|
||||
| { op: 'setMode'; modeId: string }
|
||||
| { op: 'setModeExpectError'; modeId: string }
|
||||
|
||||
/** A scenario's `input.json`: an ordered list of input steps. */
|
||||
export interface InputScript {
|
||||
@@ -102,6 +106,16 @@ export interface InputScript {
|
||||
* agent itself just sees `cancelled`, so it cannot absorb the bug).
|
||||
*/
|
||||
permissionAnswers?: PermissionAnswer[]
|
||||
/**
|
||||
* Ordered answers for the agent's `elicitation/create` round-trips (the
|
||||
* ask_user_question / plan-review forms), consumed FIFO — the Nth request
|
||||
* gets the Nth answer. Exhaustion (or no queue) answers `cancel`, the same
|
||||
* fail-closed stub an elicitation-free scenario relies on. Unlike permission
|
||||
* kinds, the scripted strings are not validated against the offered form —
|
||||
* a stray `choice` reaches the agent verbatim, which reads it as a custom
|
||||
* (non-consenting) answer, so a scenario bug fails safe in the transcript.
|
||||
*/
|
||||
elicitationAnswers?: ElicitationAnswer[]
|
||||
}
|
||||
|
||||
/** One scripted answer to a permission request: which offered option kind to select. */
|
||||
@@ -110,6 +124,16 @@ export interface PermissionAnswer {
|
||||
kind: 'allow_once' | 'allow_always' | 'reject_once' | 'reject_always'
|
||||
}
|
||||
|
||||
/** One scripted answer to an elicitation form (accept with choice/custom content, or cancel). */
|
||||
export interface ElicitationAnswer {
|
||||
/** Accept the form with the content below, or cancel it. */
|
||||
action: 'accept' | 'cancel'
|
||||
/** The selected option label (the form's `choice` field). */
|
||||
choice?: string
|
||||
/** Free-form text (the form's `custom` field). */
|
||||
custom?: string
|
||||
}
|
||||
|
||||
/** One harvested session log plus the identifying facts off its header line. */
|
||||
export interface HarvestedLog {
|
||||
/** The recorded session id (header `id`). */
|
||||
@@ -251,6 +275,8 @@ export async function runScenario(input: InputScript, opts: RunOptions): Promise
|
||||
// Permission answers are consumed FIFO across the whole run; exhaustion
|
||||
// falls back to `cancelled` so approval-free scenarios keep the plain stub.
|
||||
const permissionQueue = [...input.permissionAnswers ?? []]
|
||||
// Elicitation answers mirror the permission queue: FIFO, cancel on exhaustion.
|
||||
const elicitationQueue = [...input.elicitationAnswers ?? []]
|
||||
// A scenario bug detected inside a client callback (a scripted permission
|
||||
// kind the agent never offered). It cannot fail the run from in there: a
|
||||
// callback throw only becomes a JSON-RPC error RESPONSE to the agent, and
|
||||
@@ -291,6 +317,17 @@ export async function runScenario(input: InputScript, opts: RunOptions): Promise
|
||||
}
|
||||
return Promise.resolve({ outcome: { outcome: 'selected', optionId: option.optionId } })
|
||||
},
|
||||
unstable_createElicitation(_params: CreateElicitationRequest): Promise<CreateElicitationResponse> {
|
||||
const answer = elicitationQueue.shift()
|
||||
if (answer === undefined || answer.action !== 'accept') return Promise.resolve({ action: 'cancel' })
|
||||
return Promise.resolve({
|
||||
action: 'accept',
|
||||
content: {
|
||||
...answer.choice !== undefined ? { choice: answer.choice } : {},
|
||||
...answer.custom !== undefined ? { custom: answer.custom } : {},
|
||||
},
|
||||
})
|
||||
},
|
||||
})
|
||||
const client = new ClientSideConnection(makeClient, stream)
|
||||
|
||||
@@ -406,6 +443,24 @@ async function runStep(
|
||||
await client.cancel({ sessionId })
|
||||
return
|
||||
}
|
||||
case 'setMode': {
|
||||
const sessionId = getSessionId()
|
||||
if (sessionId === undefined) throw new Error('snapshot-harness: setMode before newSession')
|
||||
await client.setSessionMode({ sessionId, modeId: step.modeId })
|
||||
return
|
||||
}
|
||||
case 'setModeExpectError': {
|
||||
const sessionId = getSessionId()
|
||||
if (sessionId === undefined) throw new Error('snapshot-harness: setModeExpectError before newSession')
|
||||
// The bridge rejects an unknown/uncomposed mode id with invalidParams;
|
||||
// that rejection IS the expected wire behavior — swallow it so the run
|
||||
// completes and the error frame is captured in the transcript.
|
||||
await client.setSessionMode({ sessionId, modeId: step.modeId }).then(
|
||||
() => { throw new Error('snapshot-harness: expected session/set_mode to be rejected but it succeeded') },
|
||||
() => { /* expected: the bridge rejected the mode id */ },
|
||||
)
|
||||
return
|
||||
}
|
||||
default:
|
||||
throw new Error(`snapshot-harness: unknown input op ${JSON.stringify(step)}`)
|
||||
}
|
||||
|
||||
@@ -18,6 +18,7 @@
|
||||
export {
|
||||
runScenario,
|
||||
type AgentUnderTest,
|
||||
type ElicitationAnswer,
|
||||
type HarvestedLog,
|
||||
type InputScript,
|
||||
type InputStep,
|
||||
|
||||
@@ -229,6 +229,11 @@ export function defineAcpSnapshotSuite(options: SnapshotSuiteOptions): void {
|
||||
pinningByClass.set(cls, scenario)
|
||||
}
|
||||
for (const scenario of scenarios) {
|
||||
// Only a scenario that RUNS a model turn produces request-header events
|
||||
// for the uniformity guard to compare — a protocol-only scenario's fixture
|
||||
// carries no header content, so it needs no anchor (and a suite of only
|
||||
// protocol scenarios legitimately has none).
|
||||
if (!scenario.hasModelTurn) continue
|
||||
if (!pinningByClass.has(classOf(scenario))) {
|
||||
throw new Error(`acp-snapshot: no scenario pins the request-header content of class "${classOf(scenario)}" (needed by ${scenario.name})`)
|
||||
}
|
||||
@@ -329,9 +334,9 @@ export function defineAcpSnapshotSuite(options: SnapshotSuiteOptions): void {
|
||||
// scenario's fixture) or composition became session-dependent by
|
||||
// design (give the divergent shape its own pinning scenario and
|
||||
// class).
|
||||
if (scenario.pinsHeader !== true) {
|
||||
/* v8 ignore next -- construction guarantees the pin exists; a miss would fail the one-header assertion loudly. */
|
||||
const pinningScenario = pinningByClass.get(classOf(scenario)) ?? scenario
|
||||
const classPin = pinningByClass.get(classOf(scenario))
|
||||
if (scenario.pinsHeader !== true && classPin !== undefined) {
|
||||
const pinningScenario = classPin
|
||||
const pinnedFixture = await readFile(join(snapshotsDir, pinningScenario.name, 'session.jsonl'), 'utf8')
|
||||
const pinned = normalizedHeaders(pinnedFixture, fixtureContext(pinnedFixture))
|
||||
expect(pinned.length, `the pinning fixture (${pinningScenario.name}) must carry exactly one request/header`)
|
||||
@@ -401,7 +406,7 @@ export function defineAcpSnapshotSuite(options: SnapshotSuiteOptions): void {
|
||||
}
|
||||
expect(Object.fromEntries([...pins].map(([cls, names]) => [cls, names.length]))).toEqual(
|
||||
Object.fromEntries([...pinningByClass.keys()].map(cls => [cls, 1])))
|
||||
for (const scenario of scenarios) {
|
||||
for (const scenario of scenarios.filter(s => s.hasModelTurn)) {
|
||||
expect(pinningByClass.has(classOf(scenario)), `class "${classOf(scenario)}" (scenario ${scenario.name}) has a pin`).toBe(true)
|
||||
}
|
||||
})
|
||||
|
||||
@@ -44,6 +44,10 @@ interface Behavior {
|
||||
prompt?: 'respond' | 'error' | 'hang-until-cancel'
|
||||
/** Before responding to a prompt, send a `session/request_permission` request and echo its outcome as a chunk. */
|
||||
permissionProbe?: boolean
|
||||
/** Before responding to a prompt, send an `elicitation/create` request and echo its response as a chunk. */
|
||||
elicitationProbe?: boolean
|
||||
/** How `session/set_mode` settles: an empty response (echoing the modeId as a chunk) or a JSON-RPC error. */
|
||||
setMode?: 'respond' | 'error'
|
||||
/** Echo the `DSH_SNAPSHOT_*` env the harness set as a chunk (spec-side env-plumbing assertions). */
|
||||
echoEnv?: boolean
|
||||
/** Echo the sorted cwd listing as a chunk (spec-side workspace-seeding assertions). */
|
||||
@@ -79,8 +83,8 @@ let sessionId = ''
|
||||
let sessionCwd = ''
|
||||
/** The parked prompt request id while `hang-until-cancel` waits for the cancel notification. */
|
||||
let parkedPromptId: number | string | null = null
|
||||
/** Resolvers for permission-probe responses, keyed by outbound request id. */
|
||||
const pendingPermission = new Map<number, (outcome: unknown) => void>()
|
||||
/** Resolvers for outbound probe responses (permission/elicitation), keyed by request id. */
|
||||
const pendingOutbound = new Map<number, (result: unknown) => void>()
|
||||
|
||||
function send(frame: Record<string, unknown>): void {
|
||||
process.stdout.write(`${JSON.stringify({ jsonrpc: '2.0', ...frame })}\n`)
|
||||
@@ -136,8 +140,8 @@ async function handlePrompt(id: number | string): Promise<void> {
|
||||
}
|
||||
if (behavior.permissionProbe === true) {
|
||||
const requestId = nextOutboundId++
|
||||
const outcome = await new Promise<unknown>((resolve) => {
|
||||
pendingPermission.set(requestId, resolve)
|
||||
const result = await new Promise<unknown>((resolve) => {
|
||||
pendingOutbound.set(requestId, resolve)
|
||||
send({
|
||||
id: requestId,
|
||||
method: 'session/request_permission',
|
||||
@@ -151,7 +155,24 @@ async function handlePrompt(id: number | string): Promise<void> {
|
||||
},
|
||||
})
|
||||
})
|
||||
chunk(`permission:${JSON.stringify(outcome)}`)
|
||||
chunk(`permission:${JSON.stringify((result as { outcome?: unknown } | undefined)?.outcome ?? null)}`)
|
||||
}
|
||||
if (behavior.elicitationProbe === true) {
|
||||
const requestId = nextOutboundId++
|
||||
const result = await new Promise<unknown>((resolve) => {
|
||||
pendingOutbound.set(requestId, resolve)
|
||||
send({
|
||||
id: requestId,
|
||||
method: 'elicitation/create',
|
||||
params: {
|
||||
sessionId,
|
||||
mode: 'form',
|
||||
message: 'Approve this plan and leave plan mode?',
|
||||
requestedSchema: { type: 'object', title: 'Plan review', properties: { choice: { type: 'string' }, custom: { type: 'string' } }, required: [] },
|
||||
},
|
||||
})
|
||||
})
|
||||
chunk(`elicitation:${JSON.stringify(result ?? null)}`)
|
||||
}
|
||||
switch (behavior.prompt ?? 'respond') {
|
||||
case 'respond':
|
||||
@@ -171,10 +192,10 @@ function handleFrame(frame: Record<string, unknown>): void {
|
||||
const method = frame.method as string | undefined
|
||||
const params = (frame.params ?? {}) as Record<string, unknown>
|
||||
// A response to one of OUR outbound requests (the permission probe).
|
||||
if (method === undefined && id !== undefined && typeof id === 'number' && pendingPermission.has(id)) {
|
||||
const resolve = pendingPermission.get(id) as (outcome: unknown) => void
|
||||
pendingPermission.delete(id)
|
||||
resolve((frame.result as { outcome?: unknown } | undefined)?.outcome ?? null)
|
||||
if (method === undefined && id !== undefined && typeof id === 'number' && pendingOutbound.has(id)) {
|
||||
const resolve = pendingOutbound.get(id) as (result: unknown) => void
|
||||
pendingOutbound.delete(id)
|
||||
resolve(frame.result)
|
||||
return
|
||||
}
|
||||
switch (method) {
|
||||
@@ -195,6 +216,14 @@ function handleFrame(frame: Record<string, unknown>): void {
|
||||
case 'session/prompt':
|
||||
void handlePrompt(id as number | string)
|
||||
return
|
||||
case 'session/set_mode':
|
||||
if ((behavior.setMode ?? 'respond') === 'error') {
|
||||
respondError(id as number | string, 'unknown mode')
|
||||
return
|
||||
}
|
||||
chunk(`setMode:${String(params.modeId)}`)
|
||||
respond(id as number | string, {})
|
||||
return
|
||||
case 'session/cancel':
|
||||
if (parkedPromptId !== null) {
|
||||
const parked = parkedPromptId
|
||||
|
||||
@@ -231,6 +231,69 @@ describe('runScenario', () => {
|
||||
expect(result.sessionLogs).toHaveLength(0)
|
||||
})
|
||||
|
||||
it('drives session/set_mode and swallows the expected rejection of setModeExpectError', { timeout: 20_000 }, async () => {
|
||||
const { fixtureFile } = await scenario({})
|
||||
const result = await runScenario(
|
||||
{ steps: [...boot, { op: 'setMode', modeId: 'plan' }] },
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile },
|
||||
)
|
||||
expect(result.rawStdout).toContain('setMode:plan')
|
||||
|
||||
const rejecting = await scenario({ setMode: 'error' })
|
||||
const rejected = await runScenario(
|
||||
{ steps: [...boot, { op: 'setModeExpectError', modeId: 'yolo' }] },
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile: rejecting.fixtureFile },
|
||||
)
|
||||
expect(rejected.rawStdout).toContain('unknown mode')
|
||||
})
|
||||
|
||||
it('fails the run when setModeExpectError unexpectedly succeeds, and both mode ops require a session', { timeout: 20_000 }, async () => {
|
||||
const { fixtureFile } = await scenario({})
|
||||
await expect(runScenario(
|
||||
{ steps: [...boot, { op: 'setModeExpectError', modeId: 'plan' }] },
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile },
|
||||
)).rejects.toThrow(/expected session\/set_mode to be rejected/)
|
||||
await expect(runScenario(
|
||||
{ steps: [{ op: 'initialize' }, { op: 'setMode', modeId: 'plan' }] },
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile },
|
||||
)).rejects.toThrow(/setMode before newSession/)
|
||||
await expect(runScenario(
|
||||
{ steps: [{ op: 'initialize' }, { op: 'setModeExpectError', modeId: 'plan' }] },
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile },
|
||||
)).rejects.toThrow(/setModeExpectError before newSession/)
|
||||
})
|
||||
|
||||
it('answers elicitations from the scripted queue, falling back to cancel on exhaustion', { timeout: 20_000 }, async () => {
|
||||
const { fixtureFile } = await scenario({ elicitationProbe: true })
|
||||
// Three prompts → three elicitations: an accept-with-choice, an
|
||||
// accept-with-custom (feedback), then the exhausted-queue cancel.
|
||||
const result = await runScenario(
|
||||
{
|
||||
steps: [...boot, { op: 'prompt', text: 'one' }, { op: 'prompt', text: 'two' }, { op: 'prompt', text: 'three' }],
|
||||
elicitationAnswers: [
|
||||
{ action: 'accept', choice: 'Approve' },
|
||||
{ action: 'accept', custom: 'add tests first' },
|
||||
],
|
||||
},
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile },
|
||||
)
|
||||
const first = result.rawStdout.indexOf('elicitation:{\\"action\\":\\"accept\\",\\"content\\":{\\"choice\\":\\"Approve\\"}}')
|
||||
const second = result.rawStdout.indexOf('elicitation:{\\"action\\":\\"accept\\",\\"content\\":{\\"custom\\":\\"add tests first\\"}}')
|
||||
const third = result.rawStdout.indexOf('elicitation:{\\"action\\":\\"cancel\\"}')
|
||||
expect(first).toBeGreaterThanOrEqual(0)
|
||||
expect(second).toBeGreaterThan(first)
|
||||
expect(third).toBeGreaterThan(second)
|
||||
})
|
||||
|
||||
it('a scripted elicitation cancel answers cancel', { timeout: 20_000 }, async () => {
|
||||
const { fixtureFile } = await scenario({ elicitationProbe: true })
|
||||
const result = await runScenario(
|
||||
{ steps: [...boot, { op: 'prompt', text: 'one' }], elicitationAnswers: [{ action: 'cancel' }] },
|
||||
{ agent: AGENT, mode: 'replay', fixtureFile },
|
||||
)
|
||||
expect(result.rawStdout).toContain('elicitation:{\\"action\\":\\"cancel\\"}')
|
||||
})
|
||||
|
||||
it('answers permission requests from the scripted queue by option kind, falling back to cancelled', { timeout: 20_000 }, async () => {
|
||||
const { fixtureFile } = await scenario({ permissionProbe: true })
|
||||
// Two prompts → two permission round-trips; one scripted answer, so the
|
||||
|
||||
Reference in New Issue
Block a user