Define the in-file RFC contract in docs/rfc/README.md § The file format: the header block (`# RFC: <title>` plus a dateless Status enum cross-checked against the lifecycle folder), the per-lifecycle body skeleton (a Problem opener everywhere; Proposal/Alternatives considered/ Acceptance criteria/Risks in proposed/; present-tense Decision/ Consequences with proposal-era headings banned in implemented/; the frozen proposal shape in rejected/), and a mandatory Alternatives considered section with a date-fenced grandfather comment for pre-format RFCs whose alternatives are not reconstructible from the record. Enforce it with a new doc-sync gate, scripts/verify-rfc-format.ts, and normalize all 112 RFCs to it: ~15 Status-line spellings collapse to the enum, 29 Context openers become Problem, the 39 legacy-format XXX debt markers are resolved and banned from reappearing, proposal-era sections in implemented RFCs are rewritten to shipped reality (including the web/fs/subagent seam RFCs' migration plans and test checklists, closing the doc-tiers deferred-work item on the web seam), every RFC gains an Alternatives considered section or the grandfather comment, and the bilingual pair is re-mirrored and re-recorded. Move the generated index tables out of README.md into a fully generated docs/rfc/INDEX.md — gen-rfc-index now writes the whole file, and verify-rfc-classification checks its freshness and rejects index-shaped rows in the curated README — which makes room for the format contract to live in the README front door instead of a separate FORMAT.md. The decision record, and the first RFC written in the new format, is docs/rfc/implemented/process/2026-07-05-uniform-rfc-format.md.
7.4 KiB
RFC: The todo_write tool — model task list as event-sourced session state
Status: implemented
Problem
The harness gives the model bash and subagent tools but no way to record a structured task list. A todo list serves two co-equal purposes: it steers the model to plan multi-step work and keep the active task unambiguous (at most one active, exactly one while work remains), and it gives the human a live progress checklist. The ACP protocol has a native plan sessionUpdate that editors (Zed) already render, but the bridge never emitted one. Every reference coding agent surveyed (claude-code, opencode, codex, oh-my-pi, pi) ships some form of this; the harness had nothing.
Decision
Add a model-facing todo_write(todos: [{ content, status }]) tool whose whole-list state lives on the event-sourced session log as a new todo/write SessionEventMap variant. Both the stdio UI and the ACP bridge render off the existing session/event — the ACP bridge maps the list to a plan sessionUpdate.
Whole-list replace, three-state status
The model sends the ENTIRE list every call; the new list replaces the old (last-write-wins on replay). This is the shape claude-code V1, opencode, and codex update_plan all use, and the shape the model is most trained on — no per-item ids, no delta protocol. status is exactly pending | in_progress | completed: the same triple as codex update_plan and, crucially, identical to the ACP PlanEntryStatus, so the bridge maps it 1:1 with no lossy translation.
State on the session log, not a service
The list is appended as a todo/write event carrying the full { todos } snapshot. The harness is event-sourced — the LLM history, tool calls, and turn structure all live on the log — so the todo list lives there too. This buys durability, replay, and session/load reconstruction for free: a reopened session re-derives the current list (the last todo/write) and the ACP bridge re-emits the plan on load, with no separate persistence backend, no in-memory service to rehydrate, and no extra wiring. An in-memory ctx.todos service would have had to reinvent all of that.
NOT a surface event
todo/write is deliberately excluded from SurfaceEventType. The surface is the projection that produces the LLM message history (deriveMessages()); a todo write produces no conversation message. So it carries no surfaceOp, never joins the surface linked list, and never reaches deriveMessages() — it is durable, replayable UI state that travels alongside the conversation without being part of it. (The dev-mode invariants still require it to sit inside an open turn, which it always does: it is appended mid-step during a tool call.)
Priority synthesized only at the ACP boundary
ACP's PlanEntry requires content + priority + status, but a TodoItem has no priority — the model never reasons about it. Rather than burden the schema with a field the model must always supply, the bridge synthesizes a constant priority: 'medium' on every entry when it builds the plan. Priority is an ACP wire requirement, not a harness concept, so it lives at exactly the boundary that needs it.
Dropped vs claude-code V1: activeForm, id, priority
claude-code V1's item is { content, status, activeForm }; later (V2) it grew ids, dependencies, and ownership — but only to support agent swarms (disk-backed, lock-guarded, per-item mutation). This tool keeps the item at the minimum: { content, status }. No activeForm (the present-continuous label) — the UI shows content; no id — whole-list replace needs no stable identity; no priority — see above. Each dropped field is one less thing the model must produce on every call.
Single owner — no swarm machinery (YAGNI)
The list belongs to the ONE agent session that called the tool (exec.agent.session); a non-agent caller is rejected. There is deliberately no shared/multi-owner scope, no capability seam (interface/impl/consumer), no scope resolver, and no delta protocol. The harness does have subagents, and a shared cross-agent list is conceivable — but building that now means designing for a form the product does not yet have. The whole-list-replace + single-owner shape is what claude-code V1, opencode, and codex all ship; if a shared list is ever needed, the on-log representation would change to per-item deltas (so concurrent writers can't clobber each other) and a scope resolver would choose the target log. That is a future RFC, not speculative scaffolding today.
Validation: the cheap middle
The schema enforces type/required/enum. Beyond that, execute rejects empty or duplicate content and more than one in_progress task. claude-code leaves single-in-progress to the prompt; oh-my-pi enforces it in code. We take the middle: enforce the cheap invariants that make a plan coherent (no blank tasks, no dupes, at most one active), but leave ordering and the discipline of keeping the list current to the model via the tool description. A rejected write returns an isError result so the model self-corrects.
Why no cordis-catalog entry / no @mode
todo/write is a member of SessionEventMap, not a first-class cordis interface Events event. The catalog generator (scripts/gen-cordis-catalog.ts) scans interface Events declarations; a SessionEventMap variant rides the existing session/event emit and produces no new catalog row. So it carries no @mode tag (which the generator requires only on interface Events members) — adding one would be meaningless.
Testing
Four tiers, designed up front:
- Unit — the session event (append/snapshot-clone/last-write-wins/not-on-surface); the tool (schema shape, arg validation via the real
ctx.tools.execute, value validation, the event append + replacement, no-agent rejection,presentCall, HMR-safety); the ACPtodosToPlanmapping; the stdio render arm. - Real-Loader path — the plugin run through
Loader.unwrapExports, asserting the namespace export shape survives (it HASinject, so a stray default would crash at load — postmortem/0001). - Full-loop integration — a scripted mock model calls
todo_writethrough the real agent loop; thetodo/writeevent lands and a second call replaces it. session/loadreplay — a persistedtodo/writere-emits theplanupdate when a fresh ACP bridge loads the session.- With-key e2e + snapshot — a real prompt induces a
todo_write; the snapshot golden gains theplannotification and the log event.
Alternatives considered
- In-memory
ctx.todosservice — would reinvent durability, replay, andsession/loadreconstruction the log gives for free. - Per-item delta protocol — only needed for a shared multi-owner list, which is out of scope; whole-list replace is simpler and matches the references.
- Tool in
core/—todo_writeis an extension tool registering onctx.tools, not part of the spine; it lives in its ownpackages/todo/group like other tool families.
Consequences
The todo list is durable, replayable session state: a persisted todo/write re-emits the editor's plan update on session/load, and the log — not plugin memory — is the single source of truth. Whole-list replace means one tool call per update with last-write-wins; there is no delta protocol to reconcile. The event stays off the surface, so a todo update never perturbs the derived model history — the model sees only its own tool call and result.