Add examples/coding-agent and the docs cookbook
The first real agent wiring: DeepSeek V4 + the bash tool suite + stdio chat + JSONL persistence, runnable via yarn demo:coding (reads the gitignored repo-root .env through process.loadEnvFile). - examples/coding-agent: cordis.yml wiring both real plugin families (llm-deepseek with !!js env secrets; bash-local + tool-bash), a bash-only coding system prompt, a max-steps-guard plugin (bounds runaway turns via the agent/turn-continuation waterfall — abort() from step-end is a no-op by then), and a stdio UI with dimmed reasoning and exit-on-idle for piped stdin. - e2e (yarn test:e2e, key-gated): full-loop.e2e.ts runs a real model against the real bash tool; coding-task.e2e.ts is the swebench-style smoke — the model fixes a buggy add.js in a temp dir and the test re-runs node add.test.js itself rather than trusting the agent. - docs/cookbook: adding-a-package (the verified checklist), adding-a-tool (execute() contract, background pattern, seams), adding-an-llm-adapter (protocol obligations, mock-server testing, e2e policy). AGENTS.md layout/commands/secrets sections updated; architecture.md points at both examples and the cookbook. - vitest.e2e.config.ts: serialize test files + retry twice — parallel e2e files trip the shared internal key's concurrency quota. - fix: the !js YAML tag spelling in docs/JSDoc is actually !!js (js-yaml resolves custom tags under tag:yaml.org,2002:js).
This commit is contained in:
77
docs/cookbook/adding-a-tool.md
Normal file
77
docs/cookbook/adding-a-tool.md
Normal file
@@ -0,0 +1,77 @@
|
||||
# Cookbook: adding a tool
|
||||
|
||||
How to give the model a new capability. Reference implementations:
|
||||
`examples/echo-agent/src/echo-tool.ts` (minimal) and
|
||||
`packages/tool-bash` (production-grade, three-package seam).
|
||||
|
||||
## The minimal shape
|
||||
|
||||
```ts
|
||||
import type { Context } from 'cordis'
|
||||
import { defineTool } from '@deepseek-ai/dsh-tools'
|
||||
|
||||
export const name = 'my-tool'
|
||||
export const inject = ['tools']
|
||||
|
||||
export function apply(ctx: Context) {
|
||||
ctx.tools.register(defineTool({
|
||||
name: 'read_file',
|
||||
description: 'Read a file from disk.', // what the model sees
|
||||
parameters: {
|
||||
path: { type: 'string', required: true, description: 'Absolute path' },
|
||||
limit: { type: 'number' }, // optional by default
|
||||
},
|
||||
async execute(args, exec) {
|
||||
// args is TYPED from the schema: { path: string; limit?: number }
|
||||
// exec carries { callId, name, arguments, agent?, signal? }
|
||||
return [{ type: 'text', text: await readFile(args.path, 'utf8') }]
|
||||
},
|
||||
}))
|
||||
}
|
||||
```
|
||||
|
||||
Registration is effect-based: disposing the plugin fiber unregisters the
|
||||
tool (write the HMR test). Schemas flow into the system-prompt assembly
|
||||
automatically.
|
||||
|
||||
## Rules of the execute() contract
|
||||
|
||||
- **Validate args at runtime.** `defineTool`'s `InferArgs` typing is
|
||||
compile-time only; at runtime `arguments` is whatever JSON the model
|
||||
emitted. Check every field; throw a descriptive Error for bad input.
|
||||
- **Throwing means isError.** The registry catches anything `execute()`
|
||||
throws and returns `{isError: true}` to the model. Use that for
|
||||
infrastructure failures (bad input, spawn errors, aborts) — but REPORT
|
||||
domain failures in the result text instead (e.g. tool-bash returns
|
||||
`[exit code: 9]` with `isError: false`: the model decides what a failing
|
||||
command means).
|
||||
- **Honor `exec.signal`.** Cancel in-flight work when it fires.
|
||||
- **Use `exec.agent` for async notifications.** `agent.inject(content,
|
||||
{source: {kind: 'plugin', plugin: '<name>'}})` appends durable context the
|
||||
NEXT model request sees — it is not a wake-up (an idle agent stays idle).
|
||||
Guard against disposed agents (try/catch).
|
||||
|
||||
## Long-running work
|
||||
|
||||
Follow tool-bash's background pattern: a `run_in_background` flag returns a
|
||||
task id immediately; companion tools poll incrementally and kill; completion
|
||||
notices arrive via `agent.inject()`. Bound buffers and spill full output to
|
||||
disk so nothing is silently lost.
|
||||
|
||||
> TODO: each tool reimplements this background pattern by hand today. At some
|
||||
> point we need a generic long-running-tool layer that handles task ids,
|
||||
> incremental polling, kill, and completion notices uniformly.
|
||||
|
||||
## Permissions / sandboxing
|
||||
|
||||
Prefer not to build policy into the tool. The seam is the `tools/execute` waterfall
|
||||
(veto or wrap — see the permission-gate example in docs/architecture.md), or
|
||||
a sandboxing implementation behind the tool's executor seam.
|
||||
|
||||
## Tests every tool needs
|
||||
|
||||
Arg-validation rejections, result shaping for every outcome, the HMR
|
||||
disposal test, and — for tools with side effects — an integration spec that
|
||||
drives the tool through the agent loop with a scripted `MockAdapter`
|
||||
(`packages/agent-loop/tests/mock-adapter.ts`), asserting the `tool/call` /
|
||||
`tool/result` session events.
|
||||
Reference in New Issue
Block a user