Files
deepseek-harness/docs/rfc/implemented/architecture/2026-06-20-generic-long-running-tool-runtime.md
Yichen Jiang bd59fddacd refactor(tasks): declare-then-execute — ctx.tasks.start() replaces register()
start({ kind, label, owner, run }) preflights everything that can fail
(the attachSurface fence, validation, the owner-cleanup attach) BEFORE
invoking the producer's run() starter, then commits atomically —
'work started but never got a collectable id' is now structurally
impossible instead of a producer try/catch rollback obligation (the
P1 review fix, rebuilt on #185's declare/execute split). Producers
lose their catch-wraps; the leak tests now pin the stronger property
that a failed preflight never spawns anything. TaskRegistration splits
into TaskStart (identity + run) and TaskHooks (cancel/done/readOutput);
docs, type-equiv manifest, catalogs, and both RFCs move with it.
2026-07-09 21:55:07 +08:00

28 KiB

RFC: The background task runtime (ctx.tasks) and the generic task control tools

Status: implemented

Problem

The bash capability seam supports both foreground commands and long-running background tasks. Background support was large: the abstract executor exposed start, get, ownerOf, list, readOutput, kill, and onTaskDone; the local executor tracked tasks, incremental reads, owner tokens, process cleanup, and completion listeners; the model saw three tools (bash, bash_output, bash_kill); the tool plugin injected completion notices back into the owning agent's session. The local executor fenced task access behind owner tokens because predictable global task ids are a cross-session read/kill hazard.

The tool cookbook already pointed at the real design smell: background bash is really generic long-running-tool infrastructure living inside one tool. The pressure stopped being hypothetical with background subagent tasks, which needs the same task ids, owner isolation, polling, stop, completion notices, and prompt guidance, and whose first draft answered by cloning the protocol under new names (subagent_wait, subagent_output, subagent_stop) and reshaping dsh-tool-subagent into a multi-tool plugin solely so the cloned companion tools would not collide across instances. Every future long-running capability (dev servers, watchers, remote jobs) would clone it again, and the model would learn a new collect/stop habit per capability.

The surveyed peer products converged on the opposite shape. Claude Code exposes one TaskOutput/TaskStop pair spanning seven task kinds (background shells, subagents, remote sessions, …), with its earlier per-capability BashOutput/KillShell names kept only as aliases; Kimi Code's BackgroundManager runs process, agent, and pending-question kinds behind the same two tools and a ~5-method producer interface; DeepSeek-Reasonix serves bash and delegation from one session-scoped jobs manager; OpenCode's BackgroundJob registry is kind-agnostic by construction. The lesson is that the task registry, the control tools, and the notification path are one capability, and the producers (bash, subagents) are plugins into it.

Decision

The tasks/ package group owns background-task semantics once, and bash and subagents are producers:

  • @deepseek-ai/dsh-tasks — the task registry service (ctx.tasks): branded task ids, owner-scoped authorization, status snapshots, incremental/final output reads, cancellation, wait-for-terminal, completion listeners, and the awaited owner-cleanup path.
  • @deepseek-ai/dsh-tool-tasks — the model-facing control surface: task_output, task_list, task_kill, the completion-notice injection into the owning session, and the system-prompt section that teaches the background-task habit.

Producers register running work into ctx.tasks and stay owners of their execution concerns: dsh-tool-bash's run_in_background path registers the process it started (incremental stdout, spill formatting, kill), and dsh-tool-subagent's background mode (the feature RFC) registers the child run (final output only, cancel + dispose). The bash seam carries no registry: bash_output/bash_kill no longer exist (the generic tools replaced them), and the subagent companion tools were never created. The dsh-agent-core bundle loads the pair, so every shipped deployment has the control surface.

The registry is a CONCRETE service, not an interface/implementation seam pair: there is exactly one sensible in-process implementation today, and the capability-seam convention says not to split preemptively. The pre-release stance lets a later durable/remote job system extract an interface when a second backend actually exists.

Task model

dsh-tasks owns the vocabulary (data-structure catalog). TaskId is branded, generated by the registry as <kind>-N with a per-kind counter (bash-1, subagent-1) — the kind prefix keeps ids self-describing in transcripts and preserves the pre-runtime bash-N shape. Ids are runtime-global and predictable, so every access is authorized (below).

A producer hands its work to ctx.tasks.start() in a declare-then-execute shape (the pattern the timeout-policy plugin set: the capability declares, the shared layer executes): identity first, then a run() starter the runtime invokes only once nothing can fail anymore.

interface TaskStart {
  /** Producer kind — also the id prefix ('bash', 'subagent', …). */
  kind: string
  /** One-line model-facing label (the command; the delegation description). */
  label: string
  /** The spawning agent; undefined = unowned (open access, dies with the service). */
  owner?: Agent
  /** Start the actual work; called exactly once, after preflight passed. */
  run(): TaskHooks
}

interface TaskHooks {
  /** Request termination; idempotent; must lead to `done` settling. The optional reason is `task_kill`'s logged reason, forwarded. */
  cancel(reason?: string): void
  /** Settles at QUIESCENCE — after the producer has released the task's resources. Never rejects. */
  done: Promise<TaskOutcome>
  /** OPTIONAL incremental read (stream kinds). Consecutive calls never re-deliver output; the producer owns truncation/spill formatting. Absence = final-output-only kind. */
  readOutput?(): string
}

interface TaskOutcome {
  status: 'completed' | 'killed' | 'failed'
  /** Kind-specific detail rendered into the status line ('exit code: 3', 'max-tokens'). */
  detail?: string
  /** Final output for final-only kinds; read idempotently after the task settles. */
  output?: string
}

The task status vocabulary is generic and closed: running, stopping (cancel requested, not yet settled), and the three terminal values above. Kind-specific meaning rides in detail, so the registry never learns process or agent semantics — the method presence (readOutput) is the capability, mirroring SubagentRun.sendMessage.

The registry attaches ONE continuation to done: record the terminal snapshot, then notify task-done listeners with per-listener containment (the guarantee the bash seam's notifyTaskDone used to give its own listener set). done settling at quiescence — not merely at completion — is what makes owner cleanup and service disposal awaitable without a second completion surface; this resolves the old seam's duplication of a per-task done promise AND a global onTaskDone registry by making the promise the producer contract and the listener registry the consumer surface.

Registrations are NOT effect-scoped to the registering fiber: a task belongs to its owning agent and its producing backend, not to the tool plugin whose call started it, so an HMR reload of dsh-tool-bash or dsh-tool-tasks never orphans or kills a running task (the same argument that used to keep bash ownership in the executor). The registry's own disposal cancels every live task and awaits settlement — no orphans survive fiber.dispose().

Authorization and the service surface

Cross-session isolation lives IN the runtime so every consumer gets the same rule for free: read/kill/wait/get take the caller (Agent | undefined), and a task whose owner session differs from the caller's session is rejected (!== undefined comparison — an unowned task is open, a no-agent caller cannot match an owned task). list(caller) returns only the caller-visible tasks (owned-by-caller or unowned) — a global listing would leak other sessions' labels. Owner identity is session.header.id, the canonical id every other subsystem keys on; because both sides of the comparison come from live Agents, the freestanding OwnerToken brand the bash seam used to carry became internal state rather than a seam type.

class TaskService extends Service {           // ctx.tasks
  start(spec: TaskStart): TaskId              // preflight (throws) → spec.run() starts the work → atomic commit (cannot fail)
  get(id: TaskId, caller?: Agent): TaskSnapshot            // non-consuming; throws: unknown id, foreign owner
  list(caller?: Agent): TaskSnapshot[]                     // caller-visible only
  read(id: TaskId, caller?: Agent): TaskRead               // delta (stream kinds, consuming) or final output (final kinds, idempotent) + snapshot
  kill(id: TaskId, caller?: Agent, reason?: string): 'requested' | 'already-terminal'
  wait(id: TaskId, timeoutMs: number, caller?: Agent, signal?: AbortSignal): Promise<TaskSnapshot>
  onTaskDone(listener: (snapshot: TaskSnapshot) => void): () => void   // effect-scoped, contained, never fires after dispose
  attachSurface(name: string): () => void     // the misconfiguration fence, below
}

TaskSnapshot is the read-only projection: id, kind, label, owner session, status, detail, started/finished timestamps, and the reported notice-suppression flag (below). wait resolves with the terminal snapshot, or with the still-running snapshot on timeout; aborting the wait cancels only the wait.

Misconfiguration fails loud: a deployment that loads a background-capable producer without any control surface would let the model start tasks it can never read or stop — the half-loaded failure mode the subagent RFC's first draft reshaped a whole plugin to avoid. The fence is attachSurface(): dsh-tool-tasks attaches (effect-scoped) on load, and start() throws background tasks unavailable: no control surface is attached (load @deepseek-ai/dsh-tool-tasks) when none is attached — the earliest self-contained moment, since concurrent plugin start makes a load-time check racy. The registry stays ignorant of tool names; a deployment with a custom (non-model) surface attaches its own.

The model-facing control tools

dsh-tool-tasks registers three kind-agnostic tools (ACP render intent: generic cards, kind: 'execute' for kill and 'read' for output/list, no locations):

  • task_output(task_id, wait?, timeout_ms?) — non-blocking by default: stream kinds return output produced since the previous read, final kinds return only a status line while running and the final output once terminal; every response ends with the status line ([status: running], [status: completed, exit code: 0], [status: failed, max-tokens] — generic status + producer detail). wait: true blocks until the task settles or the timeout expires (config: defaulted waitTimeoutMs, capped maxWaitTimeoutMs); a timed-out wait returns [status: running] and leaves the task alive. Polling-by-default preserves the established bash habit; wait is what a parent uses when it is genuinely blocked on a subagent's answer.
  • task_list() — the caller's tasks, one line each: <id> [<kind>] <status> — <label>; (no background tasks) when empty. Most peers make listing a human-only surface (/tasks panels) and only Gemini CLI ships a model-facing list; DSH keeps it model-facing because the harness is an SDK with no guaranteed user UI — a deployment may have no /tasks equivalent — and a caller-scoped list is one cheap registry read.
  • task_kill(task_id, reason?) — requests cancellation and returns immediately (requested cancellation of task <id>); the optional reason lands in the logged tool args and is forwarded to the producer's cancel where the underlying seam accepts one (SubagentRun.cancel(reason)). Killing an already-terminal task reports its terminal status rather than failing; a producer cancel that throws fails the call loud and leaves the task untouched (still running, notice not suppressed).

task_output's read cursor is task-scoped and CONSUMING for stream kinds: the registry keeps one cursor per task, and a read returns everything produced since the previous read, exactly like the old bash_output. v1's intended reader is the owning model — the owner fence already makes it the only model-facing one — so a non-consuming observation surface (a UI tailing a task, multiple concurrent readers) is deliberately out of scope; when one is needed, it extends the registry with a cursor/snapshot read API rather than changing task_output, because two consumers sharing the consuming cursor would silently eat each other's output.

One system-prompt section (order 106, next to tool:bash) teaches the cross-call habit the per-tool descriptions cannot: track every returned task id; you are notified in-session when a task finishes, so do not busy-poll or sleep on one — keep working on independent steps and do not duplicate a running task's work; do not produce a final answer while a relevant task still runs — call task_output (with wait when blocked) to collect it first; task_kill tasks that stopped mattering. The do-not-poll and do-not-duplicate sentences are near-verbatim convergent across Claude Code, Kimi Code, and OpenCode — they are the two failure modes every peer engineered against.

Completion notices stay durable context, not a wake-up (agent.inject() appends a logged context/message the next model request sees; it does not run the model): on onTaskDone, dsh-tool-tasks injects background task <id> (<kind>: <label>) finished [status: …]. Read its output with task_output. into the owning agent's session, with the same disposed-race containment dsh-tool-bash used to carry. Notices are deduplicated the way Claude Code and Kimi Code both learned to: a task the model explicitly killed, or whose terminal state a read/wait already returned (including a wait pending at the moment of settlement), is marked reported and its notice suppressed — never a redundant "finished" for work the model just collected or ended. The model-visible ⟺ logged invariant holds with no new session event type.

Producer opt-in and schema exposure

Whether a producer tool offers run_in_background is that producer's own defaulted config: enableRunInBackground?: boolean on dsh-tool-bash and on each dsh-tool-subagent instance (both default true — bash keeps its always-exposed behavior, and a deployment disables either per instance from cordis.yml, no code edit). A disabled producer omits the parameter from its schema entirely, so schema and capability can never disagree. ctx.tasks plays no part in schema shaping — it never rewrites or decorates a producer's tool schema (Kimi Code regex-rewrites its bash description when background is disabled; config-owns-the-schema makes that trick unnecessary) — it only provides runtime registration. The two halves compose fail-loud: the producer's config decides what the model sees, and a background call that still reaches start() without a control surface throws the load-this-package error. start() preflights every failable check (the fence, validation, the owner-cleanup attach) BEFORE invoking the producer's run() and commits atomically after — background work started without a collectable id is structurally impossible, not a producer rollback obligation.

The awaited owner-cleanup seam

A background task must not outlive its owner: the subagent case leaks live child agents/sessions otherwise, and agent/disposed is emitted synchronously inside the disposal chain without awaiting listener work, so an emit listener cannot promise quiescence (the analysis in the feature RFC). The runtime therefore needs a seam the owning agent's disposal chain actually awaits, and that seam belongs to dsh-agent, where every lifecycle consumer can reach it:

  • AgentRegistry.onCleanup(agentId, cleanup: () => Promise<void>): () => void — a per-agent cleanup registry (registrations are effects; the disposer unregisters).
  • The loop's composite disposal chain carries one link for it: after stop-and-drain and before unregister, await ctx.agents.drainCleanups(agent.id) runs every registered cleanup with per-cleanup containment (a throwing cleanup is logged and never starves later cleanups or the rest of the chain). This is a documented dsh-agent-loop change; running cleanups is part of the AgentFactory dispose contract so a replacement loop honors it too.

dsh-tasks consumes the seam: the first task registered for an owner attaches one cleanup that cancels the owner's still-live tasks, awaits each task's done (quiescence), and drops the owner's snapshots. AgentHandle.dispose() thus resolves only after the owner's background children are actually gone, and the guarantee composes transitively: a background subagent that started background tasks of its own drains them when its child agent disposes inside the parent task's settlement path (the cascade OpenCode implements with explicit parent-chain walking falls out of the seam here). This is a deliberate behavior change for bash — a background bash task used to outlive its owning agent until service disposal — adopted for uniformity: an ownerless task is the sanctioned way to outlive an agent, and a future durable-job RFC is the way to outlive the runtime.

Bash migration

dsh-bash keeps the execution contract and carries no registry. The seam is resolve, run, and start, where start(spec) returns a process handle — BashProcess: { command, status, exitCode, signal, done, readOutput(), kill() } — instead of a registry entry: get/ownerOf/list/onTaskDone, the listener machinery, BashTaskId, OwnerToken, and the spec's owner field are gone (a consumer census found get/list reached only by test harnesses and onTaskDone single-consumer — dsh-tool-bash; the hook bridges consume resolve+run only). The local executor keeps an internal table of LIVE processes solely for its own disposal quiescence (entries leave on settlement). The foreground trusted-plugin path (resolve + run with stdin/env, used by the hook bridges) is untouched and never routes through the runtime; BashExecSpec.timeoutMs stays required-but-ignored by start() (shared-spec status quo, documented in the seam JSDoc); the credential-scrub duplication between the bash and ACP spawn sites is explicitly NOT this runtime's work — the registry never touches process spawning.

dsh-tool-bash keeps the bash tool; the run_in_background path is ctx.tasks.start({ kind: 'bash', label: command, owner: exec.agent, run }) whose run() spawns through ctx.bash.start(...) and returns the hooks, where done maps the process exit to a TaskOutcome (processOutcome: completed/killed + exit-code/signal detail) and readOutput wraps the handle's incremental read with the spill/lossy formatting (renderProcessRead). The completion-notice listener left dsh-tool-bash entirely.

Subagent integration

Background subagent tasks rides this runtime; the headline consequence is that dsh-tool-subagent KEEPS its one-instance-per-provider shape — the multi-tool reshape existed only to keep cloned companion tools from colliding, and there are no companion tools to clone. Its background call is ctx.tasks.start({ kind: 'subagent', label: description, owner: parent, run }) whose run() starts the provider run and returns { cancel: run.cancel, done }, where done awaits run.result, awaits run.dispose() (quiescence), and maps the stop reason (completed → completed; aborted → killed; error/max-tokens/refusal/unknown → failed with the reason as detail) and the final text as output. No readOutput — the child session remains the detailed trace, exactly as that RFC argues.

Alternatives considered

Why not per-capability companion tools (bash_output/bash_kill + subagent_wait/subagent_output/subagent_stop)?

That is the trajectory this RFC interrupted. Each capability re-implements ids, ownership, polling, stop, notices, and guidance; the model learns N collect/stop habits and the prompt carries N near-identical tool descriptions; dsh-tool-subagent needed a structural reshape purely to de-duplicate its clones. Claude Code walked this exact path — per-capability BashOutput/KillShell first, then a generalized TaskOutput/TaskStop spanning shells, agents, and remote sessions with the old names kept as deprecated aliases — and the pre-release stance let this repo land directly on the unified shape with no alias burden.

Why not an abstract TaskRuntime seam with swappable backends?

There is one in-process implementation and no concrete second backend; the capability-seam rule is to split when the consumer and backend can actually evolve independently, not before. A durable/persistent job system is the plausible second backend, and it changes the lifecycle contract (survival across owner disposal) enough that its RFC should own the interface extraction.

Why not keep authorization in the consumers, as bash did?

The bash split put policy in dsh-tool-bash so the executor seam stayed session-free — right for a seam that may be implemented remotely. The registry is harness-local infrastructure whose entire purpose includes the isolation fence; leaving the fence to each consumer means every future surface (model tools, a UI bridge, hook bridges) re-implements it or forgets it. Centralizing it is most of the reason the runtime exists.

Why not a parallel-mode agent/cleanup event instead of the keyed registry?

An event fires for every agent at every disposal and every listener must filter; Promise.all rejection semantics need extra containment; and there is no disposer to make registrations effects. A keyed registry is targeted, contained, and disposable — and the loop already drains an ordered chain, so one more awaited link was the smaller change.

Why not blocking-by-default task_output (Claude Code's block: true)?

The established bash habit is poll-between-work, and the guidance tells the model to keep doing independent work while tasks run; defaulting to block would silently serialize the parent on its slowest child. The explicit wait: true keeps blocking a deliberate act, and the wait/read/kill trio still lands within the three-tool surface.

Why not a separate task_wait tool?

Waiting is never useful without reading the result afterwards; a separate tool doubles the calls and the schema surface for zero information. Folding it into task_output matches the only real usage pattern.

Why not ToolDefinition.timeoutMs (the timeout-policy plugin) for task_output's wait?

The timeout library gives wait() its timing internals — ctx.tasks.wait arms a deadline() and classifies wait-timeout vs caller-abort with timeoutOf scoped to TASK_WAIT_TIMEOUT — but the tool-call-level policy is deliberately NOT adopted: timeout-policy replaces a timed-out call with a structured TOOL_TIMEOUT failure, whereas a timed-out task_output(wait: true) is a SUCCESS that must still report [status: running] (the model needs the task's state either way, and the task keeps running). The wait therefore bounds its own deadline through the tool's waitTimeoutMs/maxWaitTimeoutMs config. For the same reason, no timeout policy manages a background task's LIFETIME: once the id is returned the work is off the tool-call deadline entirely — cancellation belongs to task_kill and owner cleanup.

Why not a push-sink producer contract (appendOutput/settle), as Kimi Code's manager uses?

A sink centralizes output buffering, truncation, and spill in the runtime, which is elegant when the runtime owns output storage. In this codebase those concerns already live — bounded, tested, spill-file-aware — inside dsh-bash-local, and keeping process concerns in the executor is the point of the bash seam. The pull contract (readOutput() returning a formatted delta) reuses that machinery as-is; a sink would relocate it for no v1 gain. If a durable backend later makes the runtime own output storage, the producer contract is the one seam to revisit.

Why not random task ids (Claude Code's per-kind 36^8 suffixes)?

Peers with shared registries use unguessable ids as defense-in-depth against cross-session access and predictable-path attacks. Here the owner fence is the boundary — exactly as dsh-tool-bash documented for the old predictable bash-N — and the registry hands no filesystem paths derived from the id, so sequential per-kind counters keep transcripts readable and tests deterministic. Nothing prevents switching the generator later; the id is branded and opaque to consumers.

Why not foreground→background promotion?

Claude Code, Kimi Code, and OpenCode all let a running foreground call be promoted to a background task (user action or an over-budget auto-detach). It is deliberately out of v1: promotion needs a UI/user channel the SDK does not prescribe, and it changes the foreground tools' result contracts. The registration-based design leaves the door open — a foreground execution is promotable by registering its already-running work mid-flight — and a follow-up RFC can add it without touching the model-facing control tools.

Why not new session events for task lifecycle?

Everything model-visible already lands in the log: starts and reads are tool calls/results, notices are injected context/message events. A task/* session event would duplicate facts the log carries, and live registry state (task_list) is intentionally runtime state, exactly as bash's task table was.

Testing

Unit coverage pins the registry lifecycle (register/read/kill/wait/list, owner isolation including no-agent callers, stream-vs-final read semantics, listener containment, notice suppression after an explicit kill or terminal read/wait, the attachSurface fence, start atomicity — a failed preflight mutates nothing and burns no counter; a failed producer cancel leaves the task untouched — disposal quiescence, per-kind id counters), the onCleanup drain ordering + containment (including mid-drain registration), both producers' start mapping plus the structural no-orphan guarantee (a failed preflight means the producer's run() — the spawn — was never invoked), and unchanged foreground bash/subagent behavior. Snapshot coverage pins the task tool schemas and the prompt section through the pinned-header fixture.

Consequences

One background-task contract exists instead of a per-capability clone family: a background subagent and a background bash command coexist in one session under one id namespace, one listing, one notice format, and one guidance section, and the tool cookbook points long-running tools at ctx.tasks. The cost was a wide landing change — the bash seam lost its registry surface and every test tier moved with it, model-facing tool names churned (bash_output/bash_kill deleted), and the ACP snapshot pinned header was refreshed for the new schemas — sanctioned by the pre-release stance.

Owner-scoped cleanup changed bash semantics: a task that previously outlived its agent now dies with it. Deployments that relied on fire-and-forget background commands start them unowned (a non-agent caller) or accept the new lifecycle; the uniformity was judged worth the change, and the durable-job direction remains open for real survival requirements.

wait is the first blocking tool call whose duration is model-controlled; the config cap bounds it, but a model that serializes on wait loses the parallelism the feature exists for — prompt guidance mitigates, and a future continuation-policy guard can enforce. Recording live background-flow snapshot scenarios (a polled and killed background command, a completion-notice turn) and a with-key e2e background lifecycle require a DEEPSEEK_API_KEY re-record and remain named follow-up work. The runtime deliberately defers durable/cross-restart tasks, non-consuming observation cursors, and foreground→background promotion (see Alternatives).