Files
deepseek-harness/packages/core/tools/README.md
2026-07-12 02:55:26 +08:00

27 KiB
Raw Blame History

dsh-tools

Tool registry and execution pipeline. Tool plugins register their schemas and executors; the agent loop executes each call through tools/pre-execute (the extensible allow/deny gate) → monotonic registered guards → tools/execute (an around-dispatch wrapper for timeout/retry/metrics plugins) → tools/post-execute (inspect/replace the result, attach context) → the observe-only tools/result notification. The registry also owns HOW its tools are presented to the model — its mode config selects native function calling, Code Mode, or both.

Service: ToolRegistry (ctx key: tools)

Config

tools:
  mode: native   # native (default) | code | both

native contributes the calling agent's visible end capabilities as wire function definitions. Under code, this registry's canonical contribution is the reserved run_code transport plus the generated tools:sdk prompt section (see Code Mode); an assembly listener may still deliberately add unrelated schemas. both contributes the visible native definitions, run_code, and the SDK section. In non-native modes both infrastructure pieces are protected rather than filterable capabilities: restrictions and assembly listeners cannot remove them, a scoped section cannot shadow the globally protected tools:sdk, and registering, shadowing, or explicitly filtering run_code fails loudly. Non-native modes require a loaded ctx.codeRuntime with language: 'typescript'; a missing or mismatched runtime rejects every prompt assembly with an actionable error, and a systemPrompt.toolOrder naming tools the mode no longer contributes rejects the assembly the same way.

Public API

  • ctx.tools.register(definition: ToolDefinition): () => Promise<void> | void Register a tool as a frozen snapshot. Parameters must survive lossless-JSON validation before and after cloning; scalar fields are copied, and execute/presentation callbacks are bound once to the original definition as their method receiver, so later callback-property replacement cannot change dispatch. The layer is the CALLING context's scope (dsh-scope): a plain plugin context registers globally; an agent's agent.ctx registers for that agent alone, SHADOWING a same-named global tool there (per-agent tool variants). Duplicate names within one layer throw; non-native modes also reject the reserved run_code transport name. Disposed with the calling fiber (= the agent, for scoped registrations).
  • ctx.tools.restrict(filter: ToolRestriction): () => Promise<void> | void Scoped-only (throws on a plain context): mask the GLOBAL end-capability surface for the calling agent — allow keeps only the listed tools, deny removes them; multiple restrictions intersect; scoped registrations bypass restriction as explicit grants. The reserved run_code transport remains available automatically and cannot be named explicitly. Snapshot-at-registration, loud unknown-name validation, restrict({}) rejects (the materialized-empty-config trap).
  • ctx.tools.get(name: string, scope?: ScopeKey): ToolDefinition | undefined Resolution as one scope sees it (shadowing applied; a restricted-away global reads as absent) — presenters pass the calling agent so the card matches what executed. Returned definitions are the registry's frozen snapshots.
  • ctx.tools.visible(scope?: ScopeKey): ToolDefinition[] The canonical executable view — restricted global layer the scope's own layer, plus the reserved transport in non-native modes — feeding prompt assembly, get, and execute, so presentation and dispatch resolve the same frozen definitions.
  • ctx.tools.knownNames(scope?: ScopeKey): string[] The PRE-restriction end-capability name universe restrict validates against: a typo fails loud while a restricted-away tool stays a normal absence. Presentation providers add reserved transport names separately when validating toolOrder.
  • ctx.tools.schemas(scope?: ScopeKey): ToolSchema[] Schemas of everything the scope can see (without the execute functions). The shipped tools' schemas are catalogued in docs/tool-catalog.md, generated by booting each tool plugin and harvesting this method (see the tool-schema-catalog RFC).
  • ctx.tools.guard(guard: ToolGuard): () => Promise<void> | void Register a monotonic synchronous execution guard after tools/pre-execute: returning a reason denies the call, while undefined leaves it unchanged. A plain-context guard applies globally; an agent.ctx guard applies only to that agent. Later waterfall listeners cannot turn a guard denial back into permission. Disposed with the calling fiber.
  • ctx.tools.execute(exec: ToolExecutionInput): Promise<ToolExecutionResult> Snapshot one single-use call input into a pipeline-owned execution, assign its opaque correlation token, require arguments to be losslessly JSON-serializable before and after cloning, deep-freeze the detached arguments, and protect its identity before running tools/pre-execute → guards → tools/executetools/post-execute; optional signal is the only operational field an around-dispatch wrapper may add, replace, or remove. Validate the final result as losslessly JSON-serializable and freeze the complete execution before tools/result observers run. Invalid or unstable input—including cloneable mutable exotics—and malformed or non-JSON listener/tool results normalize to isError outcomes rather than bypassing policy or failing later at the session log.

Injected services

SystemPrompt — the registry automatically feeds its tool schemas into the system-prompt assembly via ctx.systemPrompt.tools(). The approval seam is consumed opportunistically instead (ctx.get('approval'), no static inject): a deployment without it keeps the ask→deny degrade, and the registry stays active either way.

Live events

The live registry pipeline has three transformable waterfalls followed by the owner-final tools/result observation boundary; registry changes are deliberately unfiltered shared-state notifications. Exact signatures, dispatch modes, scope filtering, and failure-containment contracts live in the generated Cordis event catalog, while the complete ordering is visualized in the generated tool execution pipeline. tools/result is live and observe-only; the similarly named tool/result is the durable session event the agent loop appends afterwards.

Key types

  • ToolDefinitionToolSchema + execute(args, exec): Promise<ContentBlock[] | { content: ContentBlock[]; meta? }> (the bare array is the model-facing content; the object form additionally attaches an opaque, JSON-serializable meta presentation payload persisted on the tool/result event and handed back to presentResult), plus optional presentCall(args) / presentResult(args, result) for tool-owned UI presentation (see below). It also carries an optional cooperative timeout budget timeoutMs?: number (ms) enforced by @deepseek-ai/dsh-timeout-policy, never sent to the model. Registration stores a frozen snapshot with detached JSON parameters and once-bound callback identities.
  • ToolExecutionInput — the caller-supplied call description: { callId, name, arguments, agent?, parent?, signal? }; arguments must be losslessly JSON-serializable, and callers may pass an enclosing execution's opaque token as parent but never choose the new execution's own token.
  • ToolExecutionToken — a frozen, property-free identity value assigned by the registry. It supports equality correlation only and exposes no live outer execution state.
  • ToolExecution — the pipeline-owned call: immutable { token, callId, name, arguments, agent?, parent? } identity plus optional operational signal, which an around wrapper may add, replace, remove, and restore. A nested call's parent is a ToolExecutionToken, not an execution object.
  • ToolExecutionResult — losslessly JSON-serializable outcome: { callId, content, isError, error?, additionalContext?, meta? }. The registry validates the complete post-policy value before final observation. On failure with a HarnessError, error: { name, code } carries the structured failure class alongside the model-facing text (the loop forwards it onto the tool/result session event for retry/sandbox plugins and replay). additionalContext (a HookContext) ferries any tools/post-execute context up to the loop, which buffers it and appends it as a context/message after all tool/results in the step. meta is the tool's opaque presentation payload from a successful execute (the object return form); the loop forwards it onto the tool/result session event for result-card rendering.
  • PreToolDecision{kind:'allow'} | {kind:'deny', reason} | {kind:'ask', reason?}. The registry validates this exact union at runtime: a JavaScript/casted value with an unknown kind, a malformed reason, or extra fields fails closed as an isError; the tool body does not run and final observers still receive one result. Input rewrite (changing arguments) is deliberately NOT offered (it would desync the pre-execution audit/history/UI from what ran — its own proposed RFC); ask is serviced by ctx.approval when a deployment mounts it (allowed-once proceeds to dispatch; rejected/cancelled/unavailable deny with distinct reasons) and degrades to deny when none is mounted or the execution carries no agent.
  • PostToolDecision{kind:'accept', content?, additionalContext?} (keep the call successful, optionally replacing the model-facing content) | {kind:'block', feedback, additionalContext?} (turn it into an isError whose content is the corrective feedback). Output replacement is clean because tool/result is logged AFTER execute() returns.
  • ToolGuard(execution) => string | undefined; the returned string is a final monotonic denial reason evaluated after the reorderable pre-execute waterfall and before dispatch.
  • ToolCallView / ToolResultView — provider-neutral card-tagged render intents a tool returns from presentCall / presentResult to own how a UI renders ITS calls (see "Tool-owned UI presentation").

Extension points

  • Tool plugins call ctx.tools.register() — schemas flow into the assembly automatically.
  • tools/pre-execute is the reorderable allow/deny/ask gate (sandbox, permission, hooks): listeners receive (exec, next) and call next() to delegate to the default (allow) or return a PreToolDecision to short-circuit; a deny skips dispatch, while an ask resolves through the approval seam and dispatches only after a grant. Either non-grant path yields an isError result. ctx.tools.guard() installs scope-aware monotonic policy after that waterfall when a denial must not be overridable by listener ordering. tools/execute is the around-dispatch seam (timeout, retry, metrics): listeners receive (exec, next) and call next() to delegate to core dispatch (returning its ToolExecutionResult, optionally wrapped), or return a replacement result to short-circuit dispatch; the base next() is dispatch-with-normalization, so await next() already yields an isError result for a thrown or unknown tool. A wrapper may change only exec.signal before next()—adding a per-call deadline, replacing a caller signal, or restoring absence afterwards—because call identity is protected before policy begins. tools/post-execute is the inspect/transform seam: (exec, result, next) → a PostToolDecision that can replace content, block with feedback, or attach additionalContext. Core dispatch is the base of the tools/execute waterfall; the tool body keeps its own error boundary so a thrown tool still reaches post-execute as an isError. Finally, tools/result observes the immutable authoritative result after every transform and error boundary. All follow the typed-decision idiom shared with the agent/* seams (see dsh-agent); @deepseek-ai/dsh-timeout-policy is the reference tools/execute wrapper.
  • MCP servers: one plugin per server, discover tools, call ctx.tools.register() with the server's schemas.

Typed tool parameter schemas

First-party plugin authors can use the defineTool() helper (exported from this package) for typed tool parameter schemas:

import { readFile } from 'node:fs/promises'
import type { Context } from 'cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'

declare const ctx: Context

ctx.tools.register(defineTool({
  name: 'read_file',
  description: 'Read a file from disk.',
  parameters: {
    path: { type: 'string', required: true, description: 'Absolute file path' },
    offset: { type: 'number' },
    limit: { type: 'number' },
  },
  async execute(args, exec) {
    // args is typed: { path: string; offset?: number; limit?: number }
    const text = await readFile(args.path, 'utf8')
    return [{ type: 'text', text }]
  },
}))

The helper converts the author-facing SchemaSpec (with required: true as a per-property boolean) to standard JSON Schema for the wire format. Raw JSON-Schema tool definitions (from MCP servers) are still accepted by the registry directly.

A defineTool tool also validates the model-generated arguments against its SchemaSpec before execute runs (validateArgs). The model's JSON is untrusted — InferArgs<S> is a compile-time claim, not a runtime guarantee — so on a mismatch (missing required key, wrong primitive, bad enum member, nested violation) the tool throws a ToolArgsError (code: 'INVALID_ARGS'); the registry turns it into an isError result whose text lists the violations, which the model sees and self-corrects from. Validation mirrors the JSON Schema conversion exactly: extra keys are allowed, default is not applied, and an object/array prop without properties/items only type-checks. Raw-registered tools (MCP) are not validated by the harness — they validate their own input.

See defineTool, validateArgs, ToolArgsError, SchemaSpec, InferArgs, and schemaSpecToJsonSchema in the public API for details.

defineTool also validates an optional timeoutMs at definition time when present: it must be a positive finite number, or the helper throws — the budget is attached to the produced ToolDefinition (for @deepseek-ai/dsh-timeout-policy) and never reaches the model.

Structured-output schema subset

A separate vocabulary for callers that DEMAND a machine-readable value from an agent — the subagent seam's SubagentStartRequest.outputSchema (and, by extension, a workflow's agent({ schema })). Unlike SchemaSpec (the author-facing DSL for tool parameters), a StructuredOutputSchema is an object-rooted raw JSON Schema subset as data: it travels verbatim to the model as a forced tool's parameters, and the produced value is validated against it.

The subset is deliberately narrow and REJECTS LOUD outside it — accepting a keyword the validator doesn't enforce would validate less than the schema promises (accepted-then-ignored). Supported: single-string type (object/array/string/number/integer/boolean/null; type arrays rejected), properties/required/additionalProperties (boolean; every required key must be declared), items, scalar-only enum/const; annotations (description/title/default/examples) are ignored but must still be JSON data. assertSupportedOutputSchema(schema) throws OutputSchemaError (code: 'UNSUPPORTED_SCHEMA', listing every violation) for anything else; validateStructuredValue(schema, value) returns path-qualified violations (empty = valid, total — never throws).

Tool-owned UI presentation

A tool owns how ITS calls render in a UI (an editor's tool-call card, a CLI log line) — a UI plugin must NOT special-case tool names. A ToolDefinition may declare two optional, pure, display-only methods that return a card-tagged render intent (a discriminated union — a tool declares its card kind once and a UI bridge switches on card):

  • presentCall(args): ToolCallView | undefined — the PENDING state, one of:
    • { card: 'generic', title, kind?, rawInput?, content?, locations? } — the default card: a human-readable title, an optional kind (read/edit/execute/… for icon/treatment, default other), an optional rawInput (the salient input to show in a detail view — e.g. a background task id, NOT the whole args object), optional content (extra UI content blocks), and optional locations ({ path, line? }[] — files this call reads/modifies, so a capable UI can follow along; the ACP bridge forwards them as tool_call.locations).
    • { card: 'terminal', title, description?, cwd? } — a shell command: a capable UI renders a terminal card (the title is the command, description renders above it, cwd heads it); an incapable UI falls back to a generic execute card.
    • { card: 'diff', title, diffs, locations? } — a file create/modify: a capable UI renders an inline diff card from diffs ({ path, oldText, newText }[]; oldText: null for a new file). Used by write/edit.
  • presentResult(args, result): ToolResultView | undefined — the COMPLETED state, given the same args and the { content, isError, meta? } result, one of:
    • { card: 'generic', title?, content? } — an optional replacement title and reformatted content.
    • { card: 'terminal', title?, output?, exitCode?, signal? } — a terminal run's captured output and exit status. A capable UI shows an exit-status pill; an incapable UI gets a fenced ```console fallback the BRIDGE derives from output (the tool does not encode the fences).
    • { card: 'diff', title?, diffs } — a completed file mutation as an inline diff. diffs is FileDiff[] — typically the applied hunks with surrounding context computed from the before/after content, or a whole-file diff (oldText: null) when there is no before-image (a file create). Used by write/edit; a tool_call_update.content replaces the call's content, so a mutation tool returns this even when it duplicates the call-time snippet (else the result text would clobber the pending diff).

Returning undefined (or omitting a method) tells a UI to fall back to a generic presentation (title = tool name, raw args as input, raw result content). Both methods must be pure and side-effect-free: a UI may call them during live streaming AND during a session-log replay, so they depend only on their arguments. result.meta is the tool's own optional presentation payload (opaque unknown, JSON-serializable), attached by execute (see below) and persisted on the tool/result event, so a presentResult reading it stays replay-deterministic (the same meta is read back from the log). With defineTool, args is the typed InferArgs<S> shape; the helper soft-validates before calling (a malformed/older logged arg shape yields undefined rather than throwing, since display must never crash a replay). The views are provider-neutral — the ACP bridge (dsh-acp) maps each card to ACP tool_call/tool_call_update wire fields (a diff card to a { type: 'diff' } content block, a terminal card to the _meta terminal convention), and relativizes a file card's title against the session cwd. See the render-intent-union RFC (docs/rfc/implemented/architecture/2026-07-02-tool-render-intent-union.md) and the applied-hunk-diffs RFC (docs/rfc/implemented/architecture/2026-07-02-result-time-applied-hunk-diffs.md); dsh-tool-bash (terminal) and dsh-tool-fs (diff/generic) are the reference implementations.

import { defineTool } from '@deepseek-ai/dsh-tools'

const bash = defineTool({
  name: 'bash',
  description: 'Run a shell command.',
  parameters: {
    command: { type: 'string', required: true, description: 'The command to run.' },
    description: { type: 'string', required: true, description: 'One-line summary shown in the UI.' },
  },
  async execute(args) {
    return [{ type: 'text', text: `ran: ${args.command}` }]
  },
  // A terminal card: the command is the title, the description renders above it.
  presentCall: args => ({ card: 'terminal', title: args.command, description: args.description }),
  // A terminal result: the raw output + exit; the bridge derives the fenced fallback.
  presentResult: (_args, result) => {
    const block = result.content.length === 1 ? result.content[0] : undefined
    if (block === undefined || block.type !== 'text') return undefined
    return { card: 'terminal', output: block.text }
  },
})

Code Mode

Under mode: code (or both) the registry turns the tool surface into a programming API, per the Code Mode RFC: the model writes a TypeScript program (the body of an async function) and passes it to the reserved wire transport run_code; the program runs in ctx.codeRuntime (the code-execution seam — the shipped backend is a worker thread) with one async binding per visible end-capability tool (await tools.bash({...})), and ONLY what it prints or returns re-enters the model's context. Scope restrictions change those SDK bindings but cannot remove or replace the transport itself.

  • The SDK section (tools:sdk, order 150): a lazy prompt section regenerating, at each assembly, a declare const tools: {...} TypeScript declaration of the calling scope's visible end capabilities (exotic names via quoted keys), plus fixed usage instructions. The registry protects this section and the run_code wire schema after the assembly waterfall, so Code Mode cannot silently lose either half of its transport. Deterministic — lexicographic tool order, byte-identical text for an unchanged tool set (prefix-cache-friendly). The codegen (jsonSchemaToTs, exported) is total: constructs outside the defineTool subset degrade to unknown, never throw.
  • The dispatch bridge (run_code's execute): every binding call is JSON-normalized before dispatch (a value that does not survive — BigInt, circulars — rejects that one call, so the dispatched form and logged form are the same JSON value by construction), serialized through a per-run queue (even Promise.all executes underlying calls one at a time in submission order), given the outer execution's opaque token as parent, and run through the complete pre-execute → guards → execute → post-execute → result pipeline. A denial reaches the program as a binding rejection, and each sub-call is logged as a tool/code-dispatch session event with deterministic id <parent>:code:<n>; deriveMessages() does not surface that event. Token correlation lets commit-style observers defer an inner success until the final run_code result without exposing the live outer execution; ordinary tool side effects are not rolled back. A sub-call's additionalContext is deliberately dropped because inserting it inside a running parent call would break tool-call/result adjacency.
  • Settlement discipline: the bridge owns a run-scoped abort that follows the outer signal in and fires when the run settles for any reason, so a budget expiry aborts an in-flight sub-tool instead of orphaning it; the bridge then drains its queue BEFORE returning, so every tool/code-dispatch lands inside the open turn. A failed run throws CodeRunFailedError (code: 'CODE_RUN_FAILED', message = the failure kind + captured logs), which the pipeline converts to a structured isError the model self-corrects from.

The wire collapse is the registry's own contribution (systemPrompt.tools() is mode-aware), so the logged request/header records it for free. With no deliberate schema-adding assembly listener, code assembles exactly [run_code], pinned by tests and the snapshot goldens; protection guarantees that run_code and tools:sdk remain present, not that unrelated listener additions are erased. Try it: pnpm run demo:code-mode (the coding-agent example's Code Mode overlay); pnpm run demo:code-mode acp serves the same mode over ACP instead of the REPL.

Model Experience

Context surface What the model sees Token effect
Tool schemas and Code Mode SDK In normal mode the model sees each visible definition's name, description, and JSON schema. Code Mode instead protects one run_code wire schema and adds a generated TypeScript tools SDK section; both exposes both forms. Agent-scoped restrictions and shadows change that agent's set. Fixed per-request cost proportional to the visible definitions. Code Mode trades end-tool schemas for generated SDK text plus one transport schema rather than promising a universal reduction.
Tool-call history and results The loop retains model-emitted arguments and the registry's final normalized content or structured error. Post-execute listeners may append source-attributed context after the result. Code Mode exposes only the outer program's printed or returned value; inner dispatch events stay log-only. Arguments, results, and additional context are data-dependent and resent until compaction. Restrictions that hide tools also remove their schemas before the model can call them.

Known Limitations and Deferred Work

  • Native tool calls execute sequentiallyToolDefinition carries no concurrency-safety metadata; adding it (and parallel execution in the loop) waits on the deferred tool-shapes review (TODO(review)).
  • tools/pre-execute deliberately cannot rewrite exec.arguments — logged and rendered args would desync from what ran; the rewrite design is a proposed RFC.
  • defineTool's schema DSL is a deliberate subset — string/number/boolean/object/array with string-only enum; validateArgs tolerates extra keys and never applies default (XXX(unused-default) flags removing that field); raw-registered JSON-Schema tools validate their own input.
  • timeoutMs on a definition is declarative only — the registry never enforces deadlines; enforcement requires the @deepseek-ai/dsh-timeout-policy wrapper.
  • Code Mode is TypeScript-only and the presentation mode is service-widemode: code/both rejects prompt assembly unless ctx.codeRuntime.language === 'typescript'; scoped restrictions/shadows still choose each agent's visible bindings, but one tool cannot be native-only while another is code-only.
  • Code Mode bindings return text only — non-text content blocks in a sub-call result collapse to [<type> content] placeholders.
  • run_code state is fresh per run — a persistent REPL-style kernel is rejected for the MVP (cross-call state would be invisible to the log); see the Code Mode RFC.