Merge origin/master (generic task runtime #219, single-exe closure, package renames)

Semantic resolutions beyond line merges:
- The bash seam keeps resolveMode + the bash/resolve-mode waterfall on
  master's task-free BashExecutor (run/start/resolve only; BashProcess
  handles); tool-bash consults it at its stamping site and escalation
  baseline on master's render/background split, with a waterfall test on
  the recording executor.
- dsh-mode's BASH_FAMILY narrows to ['bash']: bash_output/bash_kill are
  replaced by the kind-generic task_output/task_kill, which span every
  task kind and only observe or stop work, so the access cap withholds
  only the starter it can reason about.
- The plan-mode snapshot suite adopts master's pin grammar (tool-schema
  sidecars; the expectedHeaderSnapshots extension is gone — the exit
  transition deltas, and entering-before-turn-1 needs no second
  snapshot); modes-advertise joins the plan header class (no-model, so
  membership is vacuous). Fixtures re-recorded on the acp-demo bin;
  the replay overlay gains the passthrough sandbox runner.
- examples/plan-acp-agent rewires to @deepseek-ai/dsh-acp-demo and drops
  its tool-bash entry (the spine bundle now composes it); dsh-stdio (the
  renamed stdio-chat home) keeps its /mode command and gains the dsh-mode
  peer edge; the acp bridge keeps the modes surface beside master's
  permission presets.
- mode README gains the Model Experience / Known Limitations sections the
  new README gates require; AGENTS.md ceiling 1370 → 1440 for the kept
  mode/ layout line and Agent efficiency section.
This commit is contained in:
kingwl
2026-07-15 23:17:18 +08:00
935 changed files with 27831 additions and 21211 deletions

View File

@@ -41,4 +41,31 @@ The model-facing exit tool. Its single required argument is the plan text — a
Definitions are validated at load (`resolveConfig`): the built-in `plan` (the shipped guidance section plus `access: read-only`) merges unless overridden, `default` is rejected as a key, an `access` outside the `SANDBOX_MODES` ladder throws, and any other key — a `tools` list included — fails loud. An unknown mode name fails loudly at `set()` time.
## Model Experience
### System prompt
**What the model sees**: In the default mode, nothing — assemblies are byte-identical to a deployment without this plugin (the always-registered `exit_plan_mode` tool is filtered from the wire and the Code Mode SDK). In a non-default mode, the mode's `section` renders as the `mode:policy` section (order 50) and, in plan mode, the `exit_plan_mode` tool joins the toolset; under an unhonorable `access` cap the `bash` tool disappears. A mode flip mid-session appends one coalesced `context/message` notice when the last header disagrees.
**Token effect**: Zero in the default mode. In plan mode, the section text plus one tool schema per request; every mode transition changes the logged header and therefore resets the provider prefix cache.
#### Plan-mode policy section
```markdown
You are in plan mode: a read-only planning state. Explore, analyze, and design; do not modify anything yet — edits and other side effects belong in the plan and run after its approval, not in this mode. Where a bash tool is present it runs under a read-only sandbox: commands that only read work normally, while a command that writes is denied by the sandbox — that denial marks the edge of plan mode rather than a bug, and sandbox escalation is not offered here; put the step in the plan for after approval instead. When a decision or a missing detail blocks the plan, ask the user through the ask_user_question tool where it is available. A finished plan is delivered by calling exit_plan_mode — that call is what puts it in front of the user for review, so prefer it over pasting the plan as a plain reply or asking the user to switch modes themselves. If exit_plan_mode is unavailable or its review fails, ask the user to switch the session out of plan mode instead of pressing on.
```
### Exit review
**What the model sees**: The `exit_plan_mode` call carries the full plan markdown as its argument (retained in context as ordinary tool args); its result is one short confirmation or the reviewer's feedback verbatim.
**Token effect**: The plan text is paid once as tool args and stays in the conversation; a keep-planning round adds one feedback-sized result per revision.
## Known Limitations and Deferred Work
- **Non-shell restraint is guidance-only** — until effects self-declaration lands on tool definitions, a non-default mode restrains tools other than `bash` by its section text alone; the [plan-mode RFC](../../../docs/rfc/implemented/feature/2026-07-07-plan-mode.md) archives the removed interim allowlist and its restart trigger.
- **The cap guard names `bash` rather than deriving it** — the same effects item generalizes it.
- **A pending flip set while idle dies with the process** — the UI re-applies; the idle-record primitive is the escape hatch if this bites.
- **Subagent mode inheritance is deferred** — a fork child inherits via the seeded prefix; a spawn child starts default unless its creator seeds `AgentOptions.mode`.
RFC: [plan mode](../../../docs/rfc/implemented/feature/2026-07-07-plan-mode.md).

View File

@@ -146,12 +146,12 @@ const PLAN_SECTION
+ 'of pressing on.'
/**
* The three bash tools an `access` cap conditions on a confining executor:
* `bash` runs commands under the capped sandbox; `bash_output`/`bash_kill`
* only observe and stop tasks that ran under it. Both policy layers withhold
* the trio when no mounted executor can honor the cap.
* The bash tools an `access` cap conditions on a confining executor: today
* exactly the starter, `bash` — the generic task controls (`task_output`/
* `task_kill`, from dsh-tool-tasks) span every task kind and only observe or
* stop work, so they stay visible regardless of the cap.
*/
const BASH_FAMILY = ['bash', 'bash_output', 'bash_kill']
const BASH_FAMILY = ['bash']
/** The review question's approve option label — the answer item is matched by it. */
const APPROVE_LABEL = 'Approve'
@@ -357,7 +357,7 @@ export class ModesService extends Service {
* `run_code` itself, mirroring the registry's own exclusion. A no-op when
* the section is absent (native mode).
*/
function rerenderSdk(result: { sections: { name: string; order: number; text: string }[] }, include: (name: string) => boolean): void {
function rerenderSdk(result: { sections: { name: string; text: string }[] }, include: (name: string) => boolean): void {
const sdkIndex = result.sections.findIndex(section => section.name === 'tools:sdk')
if (sdkIndex < 0) return
const sdkText = renderToolsSdk(ctx.tools.schemas().filter(schema =>

View File

@@ -8,7 +8,7 @@ import { AgentId, agentEvents, type Agent } from '@deepseek-ai/dsh-agent'
import UserInteractionService, { type AskUserQuestionRequest } from '@deepseek-ai/dsh-user-interaction'
import { CodeRuntime, type CodeRunRequest, type CodeRunResult } from '@deepseek-ai/dsh-code-runtime'
import { BashExecutor, setSandboxMode } from '@deepseek-ai/dsh-bash'
import type { BashExecRequest, BashExecSpec, BashRunResult, BashTask, BashTaskRead, OwnerToken } from '@deepseek-ai/dsh-bash'
import type { BashExecRequest, BashExecSpec, BashProcess, BashRunResult } from '@deepseek-ai/dsh-bash'
import ModesService, { DEFAULT_MODE, EXIT_PLAN_MODE, PLAN_MODE, foldMode, resolveConfig } from '../src/index.ts'
import type { ModeConfig } from '../src/index.ts'
@@ -710,7 +710,7 @@ describe('exit_plan_mode', () => {
/**
* A minimal confining executor for the access-cap tests: only `sandboxMode`
* (the capability fact both policy layers and `resolveMode` read) matters;
* the task API is never exercised here.
* the process API is never exercised here.
*/
class FakeSandboxExecutor extends BashExecutor {
constructor(ctx: Context, private readonly config: { mode?: 'read-only' | 'workspace-write' | 'danger-full-access' } = {}) {
@@ -722,7 +722,7 @@ class FakeSandboxExecutor extends BashExecutor {
}
resolve(request: BashExecRequest): BashExecSpec {
return { command: request.command, workdir: '/w', timeoutMs: 1000, owner: request.owner, sandboxMode: request.sandboxMode }
return { command: request.command, workdir: '/w', timeoutMs: 1000, sandboxMode: request.sandboxMode }
}
run(_spec: BashExecSpec): Promise<BashRunResult> {
@@ -737,12 +737,7 @@ class FakeSandboxExecutor extends BashExecutor {
})
}
start(_spec: BashExecSpec): BashTask { throw new Error('unused in access-cap tests') }
get(): BashTask | undefined { return undefined }
ownerOf(): OwnerToken | undefined { return undefined }
list(): BashTask[] { return [] }
readOutput(): BashTaskRead { throw new Error('unused in access-cap tests') }
kill(): boolean { return false }
start(_spec: BashExecSpec): BashProcess { throw new Error('unused in access-cap tests') }
}
describe('the access cap (bash/resolve-mode clamp)', () => {
@@ -804,26 +799,42 @@ describe('the access cap (bash/resolve-mode clamp)', () => {
})
})
describe('the bash trio under an access cap', () => {
const TRIO = ['bash', 'bash_output', 'bash_kill']
it('exposes and admits the trio in plan mode under a confining executor', async () => {
describe('the bash tool under an access cap', () => {
it('exposes and admits bash — and leaves the generic task controls alone — under a confining executor', async () => {
const ctx = await setup()
await ctx.plugin(FakeSandboxExecutor, { mode: 'workspace-write' })
registerNamedTools(ctx, ['read', 'write', ...TRIO])
registerNamedTools(ctx, ['read', 'write', 'bash', 'task_output', 'task_kill'])
const agent = agentWithSession()
agent.session.append('mode/set', { mode: PLAN_MODE })
const assembly = await ctx.systemPrompt.assemble({ agent })
expect(assembly.tools.map(tool => tool.name).sort()).toEqual(['bash', 'bash_kill', 'bash_output', EXIT_PLAN_MODE, 'read', 'write'])
for (const name of TRIO) {
expect(assembly.tools.map(tool => tool.name).sort()).toEqual(['bash', EXIT_PLAN_MODE, 'read', 'task_kill', 'task_output', 'write'])
for (const name of ['bash', 'task_output', 'task_kill']) {
const result = await execute(ctx, name, agent)
expect(result.isError).toBe(false)
}
})
it('hides and denies the trio in plan mode without any executor', async () => {
it('hides and denies bash in plan mode without any executor; the kind-generic task controls stay', async () => {
const ctx = await setup()
registerNamedTools(ctx, ['read', ...TRIO])
registerNamedTools(ctx, ['read', 'bash', 'task_output', 'task_kill'])
const agent = agentWithSession()
agent.session.append('mode/set', { mode: PLAN_MODE })
const assembly = await ctx.systemPrompt.assemble({ agent })
// task_output/task_kill span every task kind (subagents included), so the
// cap withholds only the starter it can reason about.
expect(assembly.tools.map(tool => tool.name).sort()).toEqual([EXIT_PLAN_MODE, 'read', 'task_kill', 'task_output'])
const denied = await execute(ctx, 'bash', agent)
expect(denied.isError).toBe(true)
expect(denied.content).toEqual([{
type: 'text',
text: 'Error: tool "bash" is not available in plan mode: its read-only sandbox cap needs a sandboxing bash executor, and none is mounted',
}])
})
it('hides and denies bash in plan mode under a never-confining executor', async () => {
const ctx = await setup()
await ctx.plugin(FakeSandboxExecutor, {})
registerNamedTools(ctx, ['read', 'bash'])
const agent = agentWithSession()
agent.session.append('mode/set', { mode: PLAN_MODE })
const assembly = await ctx.systemPrompt.assemble({ agent })
@@ -836,26 +847,10 @@ describe('the bash trio under an access cap', () => {
}])
})
it('hides and denies the trio in plan mode under a never-confining executor', async () => {
const ctx = await setup()
await ctx.plugin(FakeSandboxExecutor, {})
registerNamedTools(ctx, ['read', ...TRIO])
const agent = agentWithSession()
agent.session.append('mode/set', { mode: PLAN_MODE })
const assembly = await ctx.systemPrompt.assemble({ agent })
expect(assembly.tools.map(tool => tool.name).sort()).toEqual([EXIT_PLAN_MODE, 'read'])
const denied = await execute(ctx, 'bash_output', agent)
expect(denied.isError).toBe(true)
expect(denied.content).toEqual([{
type: 'text',
text: 'Error: tool "bash_output" is not available in plan mode: its read-only sandbox cap needs a sandboxing bash executor, and none is mounted',
}])
})
it('denies a bash call carrying sandbox_permissions under the cap (no widening mid-mode)', async () => {
const ctx = await setup()
await ctx.plugin(FakeSandboxExecutor, { mode: 'workspace-write' })
registerNamedTools(ctx, TRIO)
registerNamedTools(ctx, ['bash'])
const agent = agentWithSession()
agent.session.append('mode/set', { mode: PLAN_MODE })
const denied = await ctx.tools.execute({