Add examples/coding-agent and the docs cookbook

The first real agent wiring: DeepSeek V4 + the bash tool suite + stdio
chat + JSONL persistence, runnable via yarn demo:coding (reads the
gitignored repo-root .env through process.loadEnvFile).

- examples/coding-agent: cordis.yml wiring both real plugin families
  (llm-deepseek with !!js env secrets; bash-local + tool-bash), a
  bash-only coding system prompt, a max-steps-guard plugin (bounds
  runaway turns via the agent/turn-continuation waterfall — abort()
  from step-end is a no-op by then), and a stdio UI with dimmed
  reasoning and exit-on-idle for piped stdin.
- e2e (yarn test:e2e, key-gated): full-loop.e2e.ts runs a real model
  against the real bash tool; coding-task.e2e.ts is the swebench-style
  smoke — the model fixes a buggy add.js in a temp dir and the test
  re-runs node add.test.js itself rather than trusting the agent.
- docs/cookbook: adding-a-package (the verified checklist),
  adding-a-tool (execute() contract, background pattern, seams),
  adding-an-llm-adapter (protocol obligations, mock-server testing,
  e2e policy). AGENTS.md layout/commands/secrets sections updated;
  architecture.md points at both examples and the cookbook.
- vitest.e2e.config.ts: serialize test files + retry twice — parallel
  e2e files trip the shared internal key's concurrency quota.
- fix: the !js YAML tag spelling in docs/JSDoc is actually !!js
  (js-yaml resolves custom tags under tag:yaml.org,2002:js).
This commit is contained in:
Tianyi Cui
2026-06-13 00:47:32 +08:00
parent ab19fed77c
commit e98c1c5d42
17 changed files with 810 additions and 11 deletions

View File

@@ -371,14 +371,22 @@ export function apply(ctx: Context) {
}
```
A complete runnable wiring lives in [`examples/echo-agent`](../examples/echo-agent)
(mock model + echo tool + stdio UI + JSONL persistence, loaded from
`cordis.yml` with HMR).
Two complete runnable wirings exist: [`examples/echo-agent`](../examples/echo-agent)
(mock model + echo tool — the all-mock skeleton check) and
[`examples/coding-agent`](../examples/coding-agent) (DeepSeek V4 + the bash
tool suite — the real thing; `yarn demo:coding`). Both load from `cordis.yml`
with HMR.
Step-by-step guides live in [`docs/cookbook`](./cookbook): adding a package,
adding a tool, adding an LLM adapter.
## Deferred work (TODO)
Tracked here deliberately — each is designed-for but not implemented:
- **Restructure this document** — it has grown long; split it into focused
sections (or per-area files) so readers can navigate it without scrolling
the whole thing.
- **Sub-agent spawn/fork semantics** (seam: `AgentLoop.create()`); inter-agent
channels beyond `send`/`steer`/events.
- **Persistence backends** (JSONL session dirs, sqlite) on the

View File

@@ -0,0 +1,55 @@
# Cookbook: adding a workspace package
The file-by-file checklist for a new `@deepseek-ai/dsh-<name>` package.
(Verified by the bash and adapter packages; if it drifts, fix it here.)
## 1. Create the package
```
packages/<name>/
package.json # copy from packages/tools, adjust name/description/deps
tsconfig.json # extends ../../tsconfig.base.json, rootDir src, outDir lib,
# references: vendor/cosmokit, vendor/cordis (+ vendor/schemastery
# if you use Config, + ../<dep> for each dsh dependency)
src/index.ts # service default export or plugin (name/inject/apply/Config)
tests/<x>.spec.ts
README.md # service API, events, extension points, design notes
```
package.json invariants (enforced by `yarn constraints` / yarn.config.cjs):
`private: true`, `version: 0.0.1`, `type: module`, `cordis` in BOTH
peerDependencies and devDependencies (same range). Mirror every dsh peer
dependency in devDependencies. `schemastery` goes in `dependencies` (it is a
runtime validator), matching agent-loop.
## 2. Register it in the root configs
| File | Change |
|---|---|
| `tsconfig.base.json` | add `"@deepseek-ai/dsh-<name>": ["./packages/<name>/src"]` to `paths` |
| `tsconfig.typecheck.json` | same entry (this file overrides the map wholesale) |
| `tsconfig.build.json` | add `{ "path": "./packages/<name>" }` to `references` |
| `scripts/publint-all.ts` | add `'packages/<name>'` to the array |
| `knip.json` | only if the package has non-`*.spec.ts` entries (e.g. `*.e2e.ts` → add a per-workspace override like `packages/llm-deepseek`) |
Covered automatically by globs — no edits needed: root `package.json`
workspaces, `tsdown.config.ts`, `vitest.config.ts`, `eslint.config.mjs`.
## 3. Decide the package topology
For a swappable capability, split interface / implementation / consumer into
separate packages (see docs/architecture.md § "Capability seams" — the bash
trio is the template). A single-purpose plugin stays one package.
## 4. Verify
```sh
yarn install # registers the workspace
yarn constraints && yarn typecheck && yarn lint
yarn test:coverage # 100% per-file over src (types.ts exempt)
yarn build && yarn knip && yarn publint
```
Test expectations: every registry/registration needs an HMR-safety test
(register from a child fiber, dispose it, assert cleanup). Excessive tests
are welcome — see AGENTS.md.

View File

@@ -0,0 +1,77 @@
# Cookbook: adding a tool
How to give the model a new capability. Reference implementations:
`examples/echo-agent/src/echo-tool.ts` (minimal) and
`packages/tool-bash` (production-grade, three-package seam).
## The minimal shape
```ts
import type { Context } from 'cordis'
import { defineTool } from '@deepseek-ai/dsh-tools'
export const name = 'my-tool'
export const inject = ['tools']
export function apply(ctx: Context) {
ctx.tools.register(defineTool({
name: 'read_file',
description: 'Read a file from disk.', // what the model sees
parameters: {
path: { type: 'string', required: true, description: 'Absolute path' },
limit: { type: 'number' }, // optional by default
},
async execute(args, exec) {
// args is TYPED from the schema: { path: string; limit?: number }
// exec carries { callId, name, arguments, agent?, signal? }
return [{ type: 'text', text: await readFile(args.path, 'utf8') }]
},
}))
}
```
Registration is effect-based: disposing the plugin fiber unregisters the
tool (write the HMR test). Schemas flow into the system-prompt assembly
automatically.
## Rules of the execute() contract
- **Validate args at runtime.** `defineTool`'s `InferArgs` typing is
compile-time only; at runtime `arguments` is whatever JSON the model
emitted. Check every field; throw a descriptive Error for bad input.
- **Throwing means isError.** The registry catches anything `execute()`
throws and returns `{isError: true}` to the model. Use that for
infrastructure failures (bad input, spawn errors, aborts) — but REPORT
domain failures in the result text instead (e.g. tool-bash returns
`[exit code: 9]` with `isError: false`: the model decides what a failing
command means).
- **Honor `exec.signal`.** Cancel in-flight work when it fires.
- **Use `exec.agent` for async notifications.** `agent.inject(content,
{source: {kind: 'plugin', plugin: '<name>'}})` appends durable context the
NEXT model request sees — it is not a wake-up (an idle agent stays idle).
Guard against disposed agents (try/catch).
## Long-running work
Follow tool-bash's background pattern: a `run_in_background` flag returns a
task id immediately; companion tools poll incrementally and kill; completion
notices arrive via `agent.inject()`. Bound buffers and spill full output to
disk so nothing is silently lost.
> TODO: each tool reimplements this background pattern by hand today. At some
> point we need a generic long-running-tool layer that handles task ids,
> incremental polling, kill, and completion notices uniformly.
## Permissions / sandboxing
Prefer not to build policy into the tool. The seam is the `tools/execute` waterfall
(veto or wrap — see the permission-gate example in docs/architecture.md), or
a sandboxing implementation behind the tool's executor seam.
## Tests every tool needs
Arg-validation rejections, result shaping for every outcome, the HMR
disposal test, and — for tools with side effects — an integration spec that
drives the tool through the agent loop with a scripted `MockAdapter`
(`packages/agent-loop/tests/mock-adapter.ts`), asserting the `tool/call` /
`tool/result` session events.

View File

@@ -0,0 +1,75 @@
# Cookbook: adding an LLM adapter
How to connect a new model provider. Reference implementations:
`packages/llm-deepseek` (hand-rolled HTTP/SSE) and `packages/llm-pi-ai`
(wrapping an LLM library). Read the `StreamChunk` doc in
`packages/llm/src/types.ts` first — it records the protocol conventions both
adapters were verified against.
## The shape
```ts
class MyAdapter extends LlmAdapter {
async * stream(options: GenerateOptions): AsyncIterable<StreamChunk> { }
}
export const name = 'llm-myprovider'
export const inject = ['llm']
export const Config: z<Config> = z.object({ apiKey: z.string(), })
export function apply(ctx: Context, config: Config) {
ctx.llm.registerAdapter(['model-a', 'model-b'], new MyAdapter())
}
```
Registration is effect-based (HMR-safe); one adapter per model name —
duplicates throw. Secrets are cordis-native: schemastery Config with env
fallbacks, fed from cordis.yml via `!!js process.env.MY_KEY`. Never read
ad-hoc key files in code.
## Protocol obligations (the contract two implementations verified)
- Emit `usage` BEFORE `finish`; emit NOTHING after `finish`. The robust way:
buffer finish/usage until the provider's end-of-stream marker, then flush
(handles providers that send trailing usage-only chunks).
- Tool-call `arguments` are RAW JSON strings end-to-end; stream fragments as
`argumentsDelta`. If your provider hands back parsed objects, re-stringify
at `block-end`.
- Allocate block `index`es in first-seen stream order; reuse the index for
every delta of the same block.
- Errors have exactly two sanctioned paths: THROW from `stream()` (transport
and protocol failures — use `LlmError` with a stable code), or end the
stream with `finish {kind: 'error' | 'aborted'}` (provider in-band
failures). Consumers handle both; pick per failure class and document it.
- Honor `options.signal` (pass it to fetch / your SDK).
- `prefill` and other unsupported `GenerateOptions` fields: throw
`LlmError(..., 'UNSUPPORTED')` rather than silently dropping.
Provider-specific request knobs (thinking modes, effort levels) belong in
the ADAPTER's Config, not in `GenerateOptions` — the core vocabulary stays
provider-neutral.
## Structure that worked
Split the adapter into testable stages (llm-deepseek's layout): wire types
(`types.ts`, coverage-exempt) → request serializer → SSE/transport parser →
chunk-translation state machine → a thin adapter class wiring them. Each
stage gets its own unit suite.
## Testing
- **Unit: mock the provider, not the harness.** A scripted `node:http`
server speaking the provider's wire format covers happy paths, every error
status, malformed payloads, premature closes, and aborts — no network, and
it drives the 100% per-file coverage gate. Works for SDK-backed adapters
too (point the SDK's baseURL at the mock).
- **Hostile framing tests.** Split stream payloads at arbitrary byte
positions (including mid-UTF-8) — real networks do.
- **E2E: `tests/*.e2e.ts`** under `yarn test:e2e`, gated with
`describe.skipIf(!process.env.MY_KEY)` so CI (no secrets) stays green.
Cover each model × each provider mode you map (thinking on/off, effort
levels), a tool-call round trip INCLUDING the follow-up turn with results
in history, and loose assertions only (substring/structure, bounded
maxTokens — real models are nondeterministic).
- Register the e2e file pattern in `knip.json` (per-workspace `entry`
override) or knip flags it unused.