Review follow-up on the residual the previous commit accepted — and the
acceptance was wrong, because the fix is clean: in Code Mode the SDK
section IS the soft surface (the wire carries only run_code), section
text resolves in assemble's base, and renderToolsSdk is an exported
pure renderer. The outermost wrapper therefore re-renders tools:sdk
from the same visibility predicate the wire filter applies (allowlist,
exit-IFF-plan, minus run_code mirroring the registry's own exclusion):
a plan-mode program is documented exactly the callable bindings — read
and the exit, never the denied write. The default mode leaves the
section untouched (absence of policy), both pinned by tests.
The soft layer's promise — the model is never encouraged toward a tool
the gate denies — now holds in Code Mode too; the only remaining
prompt-honesty residual is a prepend-after-load assemble listener,
where the gate still covers execution.
Review finding, valid: under the registry's Code Mode the assembly's
only wire tool is run_code, which the plan allowlist filtered out —
leaving the model with NO tools at all, the exit review included. The
composition exists today (the acp-agent example ships a code-mode
overlay), so plan mode bricked it outright.
run_code is a transport, not a capability: every bridged sub-call is
serialized back through ToolRegistry.execute() carrying the same agent,
so tools/pre-execute judges each capability individually — exactly like
native calls. Both layers now exempt it by name: the filter keeps it
visible (tests pin plan-mode Code Mode assembly = ['run_code']) and the
gate passes the wrapper while the same run's write sub-call still
denies with the plan-mode reason.
Documented residual, same class as the prepend-after-load one: the SDK
section renders from the registry's store, so a plan-mode program may
be offered bindings whose dispatch the gate then denies — nothing runs
that a native call could not.
Review finding with a real in-repo instance: the structured runtime's
per-spawn final-assembly wrapper (prepend, post-next) re-injects
structured_output OUTSIDE the mode filter, so a structured child in
plan mode would see a tool the gate then denies — the soft policy and
the hard gate telling different stories. The suggested fix (make the
mode filter outermost) cannot beat that instance: prepend unshifts, so
the per-spawn listener always registers later and wraps outer.
Two-part resolution instead. Semantically, structured_output enters the
shipped plan allowlist — it is a child's pure result channel, the same
ask/report class as ask_user_question and exit_plan_mode, so the
filter, the re-injection, and the gate now agree wherever a structured
child runs in plan mode. Mechanically, the filter registers with
prepend anyway: it now wraps outside every append-registered listener
regardless of load order (regression test pins a pre-registered
post-next mutator being filtered), narrowing the documented cosmetic
residual to prepend-after-load listeners only, where the gate still
covers execution. Severity note: no execution breach existed — the gate
held throughout; this closes the prompt-honesty gap.
Review hardening (the finding's ordering premise did not hold — see the
PR thread — but its failure-path kernel did): onBoundary cleared the
pending intent BEFORE appending the mode/set, so a backend rejecting
that one write lost the switch forever while the picker kept showing it
optimistically. The intent is now cleared only after the append lands;
a failed flush stays parked and the next healthy boundary converges the
log with the picker. The containment test extends to pin the re-park
and the retry.
The bridge's re-notify keeps deriving from the logged event's value —
now documented in place: the service holds ONE coalesced pending slot
(every flush reads the latest selection, so a stale flush cannot
exist), and for any other writer the logged value is the truth the
picker should track, in log order.
Live-session feedback (a real Zed elicitation round-trip): the model
presented its finished plan as a plain reply and asked the USER to
switch modes — the exact reversal the roadmap warns about — because the
shipped section's 'present it with the exit_plan_mode tool' read as a
suggestion. The section now says a finished plan is delivered by
calling exit_plan_mode, preferred over pasting it as a plain reply or
asking the user to switch modes — firmer, without imperatives.
ask_user_question enters the shipped plan allowlist (asking is
read-only-safe), and the section points a blocked decision at it. The
plan-acp-agent example composes the bash family (default mode only —
plan's allowlist keeps excluding it, so the two modes now demo a real
difference) plus tool-ask-user; both recorded scenarios re-recorded:
the pin now shows plan = [ask_user_question, exit_plan_mode, read,
todo_write] and post-exit default = the full eight-tool surface.
Review finding: the exit tool's direct mode/set append flipped the
folded mode while the loop could still execute further tool calls from
the SAME assistant response — a same-batch exit_plan_mode + write pair
would sail past tools/pre-execute under 'default' even though the
request was assembled under the plan-shaped header. That broke the
design's own invariant (a step's executions run under the mode its
assembly folded), which the pending-intent flush was built to hold for
user flips.
The tool now records the switch as a pending intent like every other
writer, flushed at this step's end (still in-turn); pending intents
carry a narrate flag so the exit's flush stays silent — the tool result
is its narration — while user flips keep the coalesced boundary notice.
The gate, folding the logged mode only, now provably covers the whole
batch: regression test pins approve-then-write-in-the-same-batch as
denied, and the widened toolset still arrives on the next step.
The recorded scenarios are re-recorded: the fixture now shows mode/set
landing after step/end, before the widened fallback header.
The proposal survives contact with the code with three amendments, per
the RFCs-are-proposals rule. (1) A mode transition logs a
request/header-delta only when expressible: adding exit_plan_mode
resorts the canonical tool list, and a pure reordering has no delta
form, so entering plan mode logs the full fallback snapshot — the
attributability claim holds either way. (2) The proposed/ skeleton
converts to the implemented grammar: Proposal → Decision, the roadmap's
staging (now history) drops to the standing Deferred list, and
Acceptance criteria + Risks fold into Consequences (what holds, then the
accepted costs, including the ACP v2 mode-removal migration). (3) The
two recorded scenarios stay pending a with-key session, recorded in
Deferred.
The docs tail completes: the cookbook's plan-mode row upgrades from
sketch to the shipped package, architecture.md gains the ctx.modes
capability row (ceiling 1640 → 1650: a new capability service's table
row does not fit the old budget), and every reference repoints to
implemented/.
Plan mode's stage 2 (RFC 2026-07-07-plan-mode). The exit tool: one
required plan argument (the durable log artifact), execute re-checks the
folded mode, then conducts the review over the user-interaction seam —
one single-select question (Approve / Keep planning) with free text open
— so an approval appends mode/set back to default in-turn and every
other outcome (keep-planning feedback verbatim, aborted, no provider)
returns the corrective isError with the mode unchanged. presentCall is a
generic card titled by the plan's first heading carrying the plan
markdown; over ACP the review rides the ask_user elicitation flow, in
the terminal the stdio prompt queue — no approval-seam dependency.
The ACP bridge maps the picker 1:1 onto ctx.modes (opportunistic, a
type-only peer edge): session/new + session/load advertise
availableModes/currentModeId, session/set_mode validates through set()
and echoes an optimistic current_mode_update (the pending mode IS the
selection; the logged mode/set lands at the boundary and, matching, is
not re-sent), and a session/event listener re-notifies on each logged
flip that differs from the last sent — the tool-driven exit updates the
picker. The feature matrix rows move from 'not modeled' to the
picker-to-modes / knobs-to-config-options division, with the ACP v2
removal direction recorded as a mechanical-migration risk.
The snapshot harness gains the setMode/setModeExpectError ops and a
scripted elicitationAnswers FIFO (cancel on exhaustion; a stray choice
string reaches the agent verbatim as a non-consenting custom answer, so
a scenario bug fails safe). The suite factory's header-pin requirement
now applies only to model-turn scenarios — a protocol-only suite has no
header content to anchor. examples/plan-acp-agent is the live
composition; its keyless modes-advertise scenario pins the wire surface
(advertisement, both set_mode round-trips, unknown-id rejection). The
recorded plan-mode approve/reject arc awaits a with-key recording
session; its texts are pinned at the unit tier meanwhile.
examples/AGENTS.md ceiling 653 → 680: the new example's required smoke
row does not fit the old budget.
Plan mode's stage 1 (RFC 2026-07-07-plan-mode): a new packages/mode/ group
with one product package owning the mode/set SessionEventMap vocabulary
(log-only, non-surface, whole-value replace), the pure foldMode, and the
ctx.modes service (list/get/set). User flips are pending intents flushed
at turn/start / step/end — turn enclosure makes an idle append illegal —
with one coalesced context/message notice when the flushed mode differs
from what the last logged request header told the model; a folded mode
the config no longer defines reads as default plus one boundary notice.
Enforcement is two covering layers: a system-prompt/assemble wrapper
filters the RETURNED assembly's tools to the mode's allowlist (and shows
exit_plan_mode IFF the folded mode is plan) beside the mode:policy
section at order 50, and a tools/pre-execute gate denies deny-by-default
against the same allowlist, judging by the logged mode only. The default
mode is the absence of policy — assemblies stay byte-identical to a
no-dsh-mode deployment.
AgentOptions.mode (declaration-merged) seeds a child's initial mode
through the same flush on agent/created; the stdio app gains /mode
(print/switch, never sent to the model) over an opportunistic
ctx.get('modes'). Config is an explicit resolve step: the built-in plan
definition (read-only allowlist; bash/subagent excluded until the
sandbox family lands) merges unless overridden, 'default' as a key
throws at load, unknown names throw at set() time.