Merge origin/master (generic task runtime #219, single-exe closure, package renames)
Semantic resolutions beyond line merges: - The bash seam keeps resolveMode + the bash/resolve-mode waterfall on master's task-free BashExecutor (run/start/resolve only; BashProcess handles); tool-bash consults it at its stamping site and escalation baseline on master's render/background split, with a waterfall test on the recording executor. - dsh-mode's BASH_FAMILY narrows to ['bash']: bash_output/bash_kill are replaced by the kind-generic task_output/task_kill, which span every task kind and only observe or stop work, so the access cap withholds only the starter it can reason about. - The plan-mode snapshot suite adopts master's pin grammar (tool-schema sidecars; the expectedHeaderSnapshots extension is gone — the exit transition deltas, and entering-before-turn-1 needs no second snapshot); modes-advertise joins the plan header class (no-model, so membership is vacuous). Fixtures re-recorded on the acp-demo bin; the replay overlay gains the passthrough sandbox runner. - examples/plan-acp-agent rewires to @deepseek-ai/dsh-acp-demo and drops its tool-bash entry (the spine bundle now composes it); dsh-stdio (the renamed stdio-chat home) keeps its /mode command and gains the dsh-mode peer edge; the acp bridge keeps the modes surface beside master's permission presets. - mode README gains the Model Experience / Known Limitations sections the new README gates require; AGENTS.md ceiling 1370 → 1440 for the kept mode/ layout line and Agent efficiency section.
This commit is contained in:
@@ -41,4 +41,31 @@ The model-facing exit tool. Its single required argument is the plan text — a
|
||||
|
||||
Definitions are validated at load (`resolveConfig`): the built-in `plan` (the shipped guidance section plus `access: read-only`) merges unless overridden, `default` is rejected as a key, an `access` outside the `SANDBOX_MODES` ladder throws, and any other key — a `tools` list included — fails loud. An unknown mode name fails loudly at `set()` time.
|
||||
|
||||
## Model Experience
|
||||
|
||||
### System prompt
|
||||
|
||||
**What the model sees**: In the default mode, nothing — assemblies are byte-identical to a deployment without this plugin (the always-registered `exit_plan_mode` tool is filtered from the wire and the Code Mode SDK). In a non-default mode, the mode's `section` renders as the `mode:policy` section (order 50) and, in plan mode, the `exit_plan_mode` tool joins the toolset; under an unhonorable `access` cap the `bash` tool disappears. A mode flip mid-session appends one coalesced `context/message` notice when the last header disagrees.
|
||||
|
||||
**Token effect**: Zero in the default mode. In plan mode, the section text plus one tool schema per request; every mode transition changes the logged header and therefore resets the provider prefix cache.
|
||||
|
||||
#### Plan-mode policy section
|
||||
|
||||
```markdown
|
||||
You are in plan mode: a read-only planning state. Explore, analyze, and design; do not modify anything yet — edits and other side effects belong in the plan and run after its approval, not in this mode. Where a bash tool is present it runs under a read-only sandbox: commands that only read work normally, while a command that writes is denied by the sandbox — that denial marks the edge of plan mode rather than a bug, and sandbox escalation is not offered here; put the step in the plan for after approval instead. When a decision or a missing detail blocks the plan, ask the user through the ask_user_question tool where it is available. A finished plan is delivered by calling exit_plan_mode — that call is what puts it in front of the user for review, so prefer it over pasting the plan as a plain reply or asking the user to switch modes themselves. If exit_plan_mode is unavailable or its review fails, ask the user to switch the session out of plan mode instead of pressing on.
|
||||
```
|
||||
|
||||
### Exit review
|
||||
|
||||
**What the model sees**: The `exit_plan_mode` call carries the full plan markdown as its argument (retained in context as ordinary tool args); its result is one short confirmation or the reviewer's feedback verbatim.
|
||||
|
||||
**Token effect**: The plan text is paid once as tool args and stays in the conversation; a keep-planning round adds one feedback-sized result per revision.
|
||||
|
||||
## Known Limitations and Deferred Work
|
||||
|
||||
- **Non-shell restraint is guidance-only** — until effects self-declaration lands on tool definitions, a non-default mode restrains tools other than `bash` by its section text alone; the [plan-mode RFC](../../../docs/rfc/implemented/feature/2026-07-07-plan-mode.md) archives the removed interim allowlist and its restart trigger.
|
||||
- **The cap guard names `bash` rather than deriving it** — the same effects item generalizes it.
|
||||
- **A pending flip set while idle dies with the process** — the UI re-applies; the idle-record primitive is the escape hatch if this bites.
|
||||
- **Subagent mode inheritance is deferred** — a fork child inherits via the seeded prefix; a spawn child starts default unless its creator seeds `AgentOptions.mode`.
|
||||
|
||||
RFC: [plan mode](../../../docs/rfc/implemented/feature/2026-07-07-plan-mode.md).
|
||||
|
||||
@@ -146,12 +146,12 @@ const PLAN_SECTION
|
||||
+ 'of pressing on.'
|
||||
|
||||
/**
|
||||
* The three bash tools an `access` cap conditions on a confining executor:
|
||||
* `bash` runs commands under the capped sandbox; `bash_output`/`bash_kill`
|
||||
* only observe and stop tasks that ran under it. Both policy layers withhold
|
||||
* the trio when no mounted executor can honor the cap.
|
||||
* The bash tools an `access` cap conditions on a confining executor: today
|
||||
* exactly the starter, `bash` — the generic task controls (`task_output`/
|
||||
* `task_kill`, from dsh-tool-tasks) span every task kind and only observe or
|
||||
* stop work, so they stay visible regardless of the cap.
|
||||
*/
|
||||
const BASH_FAMILY = ['bash', 'bash_output', 'bash_kill']
|
||||
const BASH_FAMILY = ['bash']
|
||||
|
||||
/** The review question's approve option label — the answer item is matched by it. */
|
||||
const APPROVE_LABEL = 'Approve'
|
||||
@@ -357,7 +357,7 @@ export class ModesService extends Service {
|
||||
* `run_code` itself, mirroring the registry's own exclusion. A no-op when
|
||||
* the section is absent (native mode).
|
||||
*/
|
||||
function rerenderSdk(result: { sections: { name: string; order: number; text: string }[] }, include: (name: string) => boolean): void {
|
||||
function rerenderSdk(result: { sections: { name: string; text: string }[] }, include: (name: string) => boolean): void {
|
||||
const sdkIndex = result.sections.findIndex(section => section.name === 'tools:sdk')
|
||||
if (sdkIndex < 0) return
|
||||
const sdkText = renderToolsSdk(ctx.tools.schemas().filter(schema =>
|
||||
|
||||
@@ -8,7 +8,7 @@ import { AgentId, agentEvents, type Agent } from '@deepseek-ai/dsh-agent'
|
||||
import UserInteractionService, { type AskUserQuestionRequest } from '@deepseek-ai/dsh-user-interaction'
|
||||
import { CodeRuntime, type CodeRunRequest, type CodeRunResult } from '@deepseek-ai/dsh-code-runtime'
|
||||
import { BashExecutor, setSandboxMode } from '@deepseek-ai/dsh-bash'
|
||||
import type { BashExecRequest, BashExecSpec, BashRunResult, BashTask, BashTaskRead, OwnerToken } from '@deepseek-ai/dsh-bash'
|
||||
import type { BashExecRequest, BashExecSpec, BashProcess, BashRunResult } from '@deepseek-ai/dsh-bash'
|
||||
import ModesService, { DEFAULT_MODE, EXIT_PLAN_MODE, PLAN_MODE, foldMode, resolveConfig } from '../src/index.ts'
|
||||
import type { ModeConfig } from '../src/index.ts'
|
||||
|
||||
@@ -710,7 +710,7 @@ describe('exit_plan_mode', () => {
|
||||
/**
|
||||
* A minimal confining executor for the access-cap tests: only `sandboxMode`
|
||||
* (the capability fact both policy layers and `resolveMode` read) matters;
|
||||
* the task API is never exercised here.
|
||||
* the process API is never exercised here.
|
||||
*/
|
||||
class FakeSandboxExecutor extends BashExecutor {
|
||||
constructor(ctx: Context, private readonly config: { mode?: 'read-only' | 'workspace-write' | 'danger-full-access' } = {}) {
|
||||
@@ -722,7 +722,7 @@ class FakeSandboxExecutor extends BashExecutor {
|
||||
}
|
||||
|
||||
resolve(request: BashExecRequest): BashExecSpec {
|
||||
return { command: request.command, workdir: '/w', timeoutMs: 1000, owner: request.owner, sandboxMode: request.sandboxMode }
|
||||
return { command: request.command, workdir: '/w', timeoutMs: 1000, sandboxMode: request.sandboxMode }
|
||||
}
|
||||
|
||||
run(_spec: BashExecSpec): Promise<BashRunResult> {
|
||||
@@ -737,12 +737,7 @@ class FakeSandboxExecutor extends BashExecutor {
|
||||
})
|
||||
}
|
||||
|
||||
start(_spec: BashExecSpec): BashTask { throw new Error('unused in access-cap tests') }
|
||||
get(): BashTask | undefined { return undefined }
|
||||
ownerOf(): OwnerToken | undefined { return undefined }
|
||||
list(): BashTask[] { return [] }
|
||||
readOutput(): BashTaskRead { throw new Error('unused in access-cap tests') }
|
||||
kill(): boolean { return false }
|
||||
start(_spec: BashExecSpec): BashProcess { throw new Error('unused in access-cap tests') }
|
||||
}
|
||||
|
||||
describe('the access cap (bash/resolve-mode clamp)', () => {
|
||||
@@ -804,26 +799,42 @@ describe('the access cap (bash/resolve-mode clamp)', () => {
|
||||
})
|
||||
})
|
||||
|
||||
describe('the bash trio under an access cap', () => {
|
||||
const TRIO = ['bash', 'bash_output', 'bash_kill']
|
||||
|
||||
it('exposes and admits the trio in plan mode under a confining executor', async () => {
|
||||
describe('the bash tool under an access cap', () => {
|
||||
it('exposes and admits bash — and leaves the generic task controls alone — under a confining executor', async () => {
|
||||
const ctx = await setup()
|
||||
await ctx.plugin(FakeSandboxExecutor, { mode: 'workspace-write' })
|
||||
registerNamedTools(ctx, ['read', 'write', ...TRIO])
|
||||
registerNamedTools(ctx, ['read', 'write', 'bash', 'task_output', 'task_kill'])
|
||||
const agent = agentWithSession()
|
||||
agent.session.append('mode/set', { mode: PLAN_MODE })
|
||||
const assembly = await ctx.systemPrompt.assemble({ agent })
|
||||
expect(assembly.tools.map(tool => tool.name).sort()).toEqual(['bash', 'bash_kill', 'bash_output', EXIT_PLAN_MODE, 'read', 'write'])
|
||||
for (const name of TRIO) {
|
||||
expect(assembly.tools.map(tool => tool.name).sort()).toEqual(['bash', EXIT_PLAN_MODE, 'read', 'task_kill', 'task_output', 'write'])
|
||||
for (const name of ['bash', 'task_output', 'task_kill']) {
|
||||
const result = await execute(ctx, name, agent)
|
||||
expect(result.isError).toBe(false)
|
||||
}
|
||||
})
|
||||
|
||||
it('hides and denies the trio in plan mode without any executor', async () => {
|
||||
it('hides and denies bash in plan mode without any executor; the kind-generic task controls stay', async () => {
|
||||
const ctx = await setup()
|
||||
registerNamedTools(ctx, ['read', ...TRIO])
|
||||
registerNamedTools(ctx, ['read', 'bash', 'task_output', 'task_kill'])
|
||||
const agent = agentWithSession()
|
||||
agent.session.append('mode/set', { mode: PLAN_MODE })
|
||||
const assembly = await ctx.systemPrompt.assemble({ agent })
|
||||
// task_output/task_kill span every task kind (subagents included), so the
|
||||
// cap withholds only the starter it can reason about.
|
||||
expect(assembly.tools.map(tool => tool.name).sort()).toEqual([EXIT_PLAN_MODE, 'read', 'task_kill', 'task_output'])
|
||||
const denied = await execute(ctx, 'bash', agent)
|
||||
expect(denied.isError).toBe(true)
|
||||
expect(denied.content).toEqual([{
|
||||
type: 'text',
|
||||
text: 'Error: tool "bash" is not available in plan mode: its read-only sandbox cap needs a sandboxing bash executor, and none is mounted',
|
||||
}])
|
||||
})
|
||||
|
||||
it('hides and denies bash in plan mode under a never-confining executor', async () => {
|
||||
const ctx = await setup()
|
||||
await ctx.plugin(FakeSandboxExecutor, {})
|
||||
registerNamedTools(ctx, ['read', 'bash'])
|
||||
const agent = agentWithSession()
|
||||
agent.session.append('mode/set', { mode: PLAN_MODE })
|
||||
const assembly = await ctx.systemPrompt.assemble({ agent })
|
||||
@@ -836,26 +847,10 @@ describe('the bash trio under an access cap', () => {
|
||||
}])
|
||||
})
|
||||
|
||||
it('hides and denies the trio in plan mode under a never-confining executor', async () => {
|
||||
const ctx = await setup()
|
||||
await ctx.plugin(FakeSandboxExecutor, {})
|
||||
registerNamedTools(ctx, ['read', ...TRIO])
|
||||
const agent = agentWithSession()
|
||||
agent.session.append('mode/set', { mode: PLAN_MODE })
|
||||
const assembly = await ctx.systemPrompt.assemble({ agent })
|
||||
expect(assembly.tools.map(tool => tool.name).sort()).toEqual([EXIT_PLAN_MODE, 'read'])
|
||||
const denied = await execute(ctx, 'bash_output', agent)
|
||||
expect(denied.isError).toBe(true)
|
||||
expect(denied.content).toEqual([{
|
||||
type: 'text',
|
||||
text: 'Error: tool "bash_output" is not available in plan mode: its read-only sandbox cap needs a sandboxing bash executor, and none is mounted',
|
||||
}])
|
||||
})
|
||||
|
||||
it('denies a bash call carrying sandbox_permissions under the cap (no widening mid-mode)', async () => {
|
||||
const ctx = await setup()
|
||||
await ctx.plugin(FakeSandboxExecutor, { mode: 'workspace-write' })
|
||||
registerNamedTools(ctx, TRIO)
|
||||
registerNamedTools(ctx, ['bash'])
|
||||
const agent = agentWithSession()
|
||||
agent.session.append('mode/set', { mode: PLAN_MODE })
|
||||
const denied = await ctx.tools.execute({
|
||||
|
||||
Reference in New Issue
Block a user