Merge refreshed docs/i18n-batch-cds-postmortem into docs/i18n-batch-rfc
# Conflicts: # .agents/notes/implemented/architecture/2026-06-11-content-block-vocabulary.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-content-block-vocabulary.zh.md # .agents/notes/implemented/architecture/2026-06-11-custom-schema-dsl.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-custom-schema-dsl.zh.md # .agents/notes/implemented/architecture/2026-06-11-dev-invariants-over-deep-readonly.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-dev-invariants-over-deep-readonly.zh.md # .agents/notes/implemented/architecture/2026-06-11-event-sourced-sessions.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-event-sourced-sessions.zh.md # .agents/notes/implemented/architecture/2026-06-11-microkernel-event-taxonomy.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-microkernel-event-taxonomy.zh.md # .agents/notes/implemented/architecture/2026-06-11-runtime-arg-validation.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-runtime-arg-validation.zh.md # .agents/notes/implemented/architecture/2026-06-11-structured-error-taxonomy.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-structured-error-taxonomy.zh.md # .agents/notes/implemented/architecture/2026-06-11-tool-schemas-in-prompt-assembly.i18n.yaml # .agents/notes/implemented/architecture/2026-06-11-tool-schemas-in-prompt-assembly.zh.md # .agents/notes/implemented/architecture/2026-06-13-capability-seams.i18n.yaml # .agents/notes/implemented/architecture/2026-06-13-capability-seams.zh.md # .agents/notes/implemented/architecture/2026-06-13-twin-llm-adapters.i18n.yaml # .agents/notes/implemented/architecture/2026-06-13-twin-llm-adapters.zh.md # .agents/notes/implemented/architecture/2026-06-14-session-persistence.i18n.yaml # .agents/notes/implemented/architecture/2026-06-14-session-persistence.zh.md # .agents/notes/implemented/architecture/2026-06-15-turn-enclosure-invariant.i18n.yaml # .agents/notes/implemented/architecture/2026-06-15-turn-enclosure-invariant.zh.md # .agents/notes/implemented/architecture/2026-06-17-filesystem-capability-seam.i18n.yaml # .agents/notes/implemented/architecture/2026-06-17-filesystem-capability-seam.zh.md # .agents/notes/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.i18n.yaml # .agents/notes/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.zh.md # .agents/notes/implemented/architecture/2026-06-18-session-surface.i18n.yaml # .agents/notes/implemented/architecture/2026-06-18-session-surface.zh.md # .agents/notes/implemented/architecture/2026-06-18-shared-persistence-write-coordinator.i18n.yaml # .agents/notes/implemented/architecture/2026-06-18-shared-persistence-write-coordinator.zh.md # .agents/notes/implemented/architecture/2026-06-20-branded-ids.i18n.yaml # .agents/notes/implemented/architecture/2026-06-20-branded-ids.zh.md # .agents/notes/implemented/architecture/2026-06-20-extract-example-app-packages.i18n.yaml # .agents/notes/implemented/architecture/2026-06-20-extract-example-app-packages.zh.md # .agents/notes/implemented/architecture/2026-06-20-package-hierarchy.i18n.yaml # .agents/notes/implemented/architecture/2026-06-20-package-hierarchy.md # .agents/notes/implemented/architecture/2026-06-20-package-hierarchy.zh.md # .agents/notes/implemented/architecture/2026-06-21-mandatory-app-attribution-headers.i18n.yaml # .agents/notes/implemented/architecture/2026-06-21-mandatory-app-attribution-headers.zh.md # .agents/notes/implemented/architecture/2026-06-24-web-capability-seam.i18n.yaml # .agents/notes/implemented/architecture/2026-06-24-web-capability-seam.zh.md # .agents/notes/implemented/architecture/2026-06-26-file-context-as-event-gate.i18n.yaml # .agents/notes/implemented/architecture/2026-06-26-file-context-as-event-gate.zh.md # .agents/notes/implemented/architecture/2026-06-30-bash-stdin-env-trusted-plugin-surface.i18n.yaml # .agents/notes/implemented/architecture/2026-06-30-bash-stdin-env-trusted-plugin-surface.zh.md # .agents/notes/implemented/architecture/2026-06-30-event-domain-semantics.i18n.yaml # .agents/notes/implemented/architecture/2026-06-30-event-domain-semantics.zh.md # .agents/notes/implemented/architecture/2026-07-02-fs-per-session-cwd.i18n.yaml # .agents/notes/implemented/architecture/2026-07-02-fs-per-session-cwd.zh.md # .agents/notes/implemented/architecture/2026-07-02-result-time-applied-hunk-diffs.i18n.yaml # .agents/notes/implemented/architecture/2026-07-02-result-time-applied-hunk-diffs.zh.md # .agents/notes/implemented/architecture/2026-07-02-tool-render-intent-union.i18n.yaml # .agents/notes/implemented/architecture/2026-07-02-tool-render-intent-union.zh.md # .agents/notes/implemented/architecture/2026-07-03-filesystem-directory-listing-seam.i18n.yaml # .agents/notes/implemented/architecture/2026-07-03-filesystem-directory-listing-seam.zh.md # .agents/notes/implemented/architecture/2026-07-05-prompt-variables-and-tool-guidance-ownership.i18n.yaml # .agents/notes/implemented/architecture/2026-07-05-prompt-variables-and-tool-guidance-ownership.zh.md # .agents/notes/implemented/architecture/2026-07-05-reconstructable-requests.i18n.yaml # .agents/notes/implemented/architecture/2026-07-05-reconstructable-requests.zh.md # .agents/notes/implemented/architecture/2026-07-05-subagent-provider-lifecycle-events.i18n.yaml # .agents/notes/implemented/architecture/2026-07-05-subagent-provider-lifecycle-events.zh.md # .agents/notes/implemented/architecture/2026-07-06-timeout-deadline-library.i18n.yaml # .agents/notes/implemented/architecture/2026-07-06-timeout-deadline-library.zh.md # .agents/notes/implemented/architecture/2026-07-07-tool-call-timeout-policy.i18n.yaml # .agents/notes/implemented/architecture/2026-07-07-tool-call-timeout-policy.zh.md # .agents/notes/implemented/architecture/2026-07-08-agent-scope-contexts.i18n.yaml # .agents/notes/implemented/architecture/2026-07-08-agent-scope-contexts.zh.md # .agents/notes/implemented/architecture/2026-07-12-agent-scope-runtime-design.i18n.yaml # .agents/notes/implemented/architecture/2026-07-12-agent-scope-runtime-design.zh.md # .agents/notes/implemented/feature/2026-06-14-acp-agent-client-protocol.i18n.yaml # .agents/notes/implemented/feature/2026-06-14-acp-agent-client-protocol.zh.md # .agents/notes/implemented/feature/2026-06-14-acp-multi-session.i18n.yaml # .agents/notes/implemented/feature/2026-06-14-acp-multi-session.zh.md # .agents/notes/implemented/feature/2026-06-15-code-mode.i18n.yaml # .agents/notes/implemented/feature/2026-06-15-code-mode.zh.md # .agents/notes/implemented/feature/2026-06-17-filesystem-tool-schemas.i18n.yaml # .agents/notes/implemented/feature/2026-06-17-filesystem-tool-schemas.zh.md # .agents/notes/implemented/feature/2026-06-18-acp-terminal-and-tool-rendering.i18n.yaml # .agents/notes/implemented/feature/2026-06-18-acp-terminal-and-tool-rendering.zh.md # .agents/notes/implemented/feature/2026-06-18-compaction-capability-seam.i18n.yaml # .agents/notes/implemented/feature/2026-06-18-compaction-capability-seam.zh.md # .agents/notes/implemented/feature/2026-06-21-subagent-capability-seam.i18n.yaml # .agents/notes/implemented/feature/2026-06-21-subagent-capability-seam.md # .agents/notes/implemented/feature/2026-06-21-subagent-capability-seam.zh.md # .agents/notes/implemented/feature/2026-06-22-acp-subagent-backend.i18n.yaml # .agents/notes/implemented/feature/2026-06-22-acp-subagent-backend.zh.md # .agents/notes/implemented/feature/2026-06-25-ask-user-question.i18n.yaml # .agents/notes/implemented/feature/2026-06-25-ask-user-question.zh.md # .agents/notes/implemented/feature/2026-06-29-todo-write-tool.i18n.yaml # .agents/notes/implemented/feature/2026-06-29-todo-write-tool.zh.md # .agents/notes/implemented/feature/2026-06-30-hook-bridges.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-hook-bridges.zh.md # .agents/notes/implemented/feature/2026-06-30-hook-protocol-lib.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-hook-protocol-lib.zh.md # .agents/notes/implemented/feature/2026-06-30-interception-seams.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-interception-seams.zh.md # .agents/notes/implemented/feature/2026-06-30-session-store-fork-api.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-session-store-fork-api.zh.md # .agents/notes/implemented/feature/2026-06-30-subagent-observe-enrich.i18n.yaml # .agents/notes/implemented/feature/2026-06-30-subagent-observe-enrich.zh.md # .agents/notes/implemented/feature/2026-07-05-dynamic-workflows.i18n.yaml # .agents/notes/implemented/feature/2026-07-05-dynamic-workflows.zh.md # .agents/notes/implemented/feature/2026-07-05-skill-system.i18n.yaml # .agents/notes/implemented/feature/2026-07-05-skill-system.zh.md # .agents/notes/implemented/feature/2026-07-06-approval-seam.i18n.yaml # .agents/notes/implemented/feature/2026-07-06-approval-seam.zh.md # .agents/notes/implemented/feature/2026-07-06-explicit-tool-order.i18n.yaml # .agents/notes/implemented/feature/2026-07-06-explicit-tool-order.zh.md # .agents/notes/implemented/feature/2026-07-06-sandbox.i18n.yaml # .agents/notes/implemented/feature/2026-07-06-sandbox.zh.md # .agents/notes/implemented/feature/2026-07-07-mcp-client-plugin.i18n.yaml # .agents/notes/implemented/feature/2026-07-07-mcp-client-plugin.zh.md # .agents/notes/implemented/feature/2026-07-07-session-prefix.i18n.yaml # .agents/notes/implemented/feature/2026-07-07-session-prefix.zh.md # .agents/notes/implemented/feature/2026-07-08-repeat-tool-guard.i18n.yaml # .agents/notes/implemented/feature/2026-07-08-repeat-tool-guard.zh.md # .agents/notes/implemented/feature/2026-07-08-self-referential-cordis-toolset.i18n.yaml # .agents/notes/implemented/feature/2026-07-08-self-referential-cordis-toolset.zh.md # .agents/notes/implemented/feature/2026-07-10-session-query-service.i18n.yaml # .agents/notes/implemented/feature/2026-07-10-session-query-service.zh.md # .agents/notes/implemented/feature/2026-07-12-subagent-persona-tool-filter-and-depth.i18n.yaml # .agents/notes/implemented/feature/2026-07-12-subagent-persona-tool-filter-and-depth.zh.md # .agents/notes/implemented/process/2026-06-11-doc-sync-enforcement.i18n.yaml # .agents/notes/implemented/process/2026-06-11-doc-sync-enforcement.zh.md # .agents/notes/implemented/process/2026-06-11-quality-gates.i18n.yaml # .agents/notes/implemented/process/2026-06-11-quality-gates.md # .agents/notes/implemented/process/2026-06-11-quality-gates.zh.md # .agents/notes/implemented/process/2026-06-11-tsdown-over-dumble.i18n.yaml # .agents/notes/implemented/process/2026-06-11-tsdown-over-dumble.zh.md # .agents/notes/implemented/process/2026-06-11-vendor-cordis-as-source.i18n.yaml # .agents/notes/implemented/process/2026-06-11-vendor-cordis-as-source.zh.md # .agents/notes/implemented/process/2026-06-16-pnpm-over-yarn.i18n.yaml # .agents/notes/implemented/process/2026-06-16-pnpm-over-yarn.zh.md # .agents/notes/implemented/process/2026-06-17-ts-build-config.i18n.yaml # .agents/notes/implemented/process/2026-06-17-ts-build-config.zh.md # .agents/notes/implemented/process/2026-06-18-markdown-cross-link-lint.i18n.yaml # .agents/notes/implemented/process/2026-06-18-markdown-cross-link-lint.zh.md # .agents/notes/implemented/process/2026-06-20-core-data-structures-catalog.i18n.yaml # .agents/notes/implemented/process/2026-06-20-core-data-structures-catalog.zh.md # .agents/notes/implemented/process/2026-06-20-generated-cordis-catalog.i18n.yaml # .agents/notes/implemented/process/2026-06-20-generated-cordis-catalog.zh.md # .agents/notes/implemented/process/2026-06-20-rfc-classification.i18n.yaml # .agents/notes/implemented/process/2026-06-20-rfc-classification.zh.md # .agents/notes/implemented/process/2026-07-02-tool-schema-catalog.i18n.yaml # .agents/notes/implemented/process/2026-07-02-tool-schema-catalog.zh.md # .agents/notes/implemented/process/2026-07-03-documentation-graph-atlas.i18n.yaml # .agents/notes/implemented/process/2026-07-03-documentation-graph-atlas.zh.md # .agents/notes/implemented/process/2026-07-04-cordis-jsdoc-completeness-gate.i18n.yaml # .agents/notes/implemented/process/2026-07-04-cordis-jsdoc-completeness-gate.zh.md # .agents/notes/implemented/process/2026-07-04-doc-tiers-and-budgets.i18n.yaml # .agents/notes/implemented/process/2026-07-04-doc-tiers-and-budgets.zh.md # .agents/notes/implemented/process/2026-07-04-generate-rfc-index-tables.i18n.yaml # .agents/notes/implemented/process/2026-07-04-generate-rfc-index-tables.zh.md # .agents/notes/implemented/process/2026-07-04-persistence-log-catalog.i18n.yaml # .agents/notes/implemented/process/2026-07-04-persistence-log-catalog.zh.md # .agents/notes/implemented/process/2026-07-05-uniform-rfc-format.i18n.yaml # .agents/notes/implemented/process/2026-07-05-uniform-rfc-format.zh.md # .agents/notes/implemented/process/2026-07-06-export-surface-jsdoc-gate.i18n.yaml # .agents/notes/implemented/process/2026-07-06-export-surface-jsdoc-gate.zh.md # .agents/notes/implemented/process/2026-07-06-generated-config-catalog.i18n.yaml # .agents/notes/implemented/process/2026-07-06-generated-config-catalog.zh.md # .agents/notes/implemented/process/2026-07-06-node-engine-floor.i18n.yaml # .agents/notes/implemented/process/2026-07-06-node-engine-floor.zh.md # .agents/notes/implemented/process/2026-07-06-parallel-github-ci-gates.i18n.yaml # .agents/notes/implemented/process/2026-07-06-parallel-github-ci-gates.zh.md # .agents/notes/implemented/process/2026-07-06-parallel-pre-push-gates.i18n.yaml # .agents/notes/implemented/process/2026-07-06-parallel-pre-push-gates.zh.md # .agents/notes/implemented/process/2026-07-10-readme-known-limitations-gate.i18n.yaml # .agents/notes/implemented/process/2026-07-10-readme-known-limitations-gate.zh.md # .agents/notes/implemented/process/2026-07-12-package-model-experience-contract.i18n.yaml # .agents/notes/implemented/process/2026-07-12-package-model-experience-contract.zh.md # .agents/notes/implemented/simplification/2026-06-19-drop-mutable-session-summary.i18n.yaml # .agents/notes/implemented/simplification/2026-06-19-drop-mutable-session-summary.zh.md # .agents/notes/implemented/simplification/2026-06-20-collapse-trace-only-session-events.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-collapse-trace-only-session-events.zh.md # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-adapter-change-event.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-adapter-change-event.zh.md # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-assembled-surfaces.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-drop-unconsumed-llm-assembled-surfaces.zh.md # .agents/notes/implemented/simplification/2026-06-20-prune-dead-seam-methods.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-prune-dead-seam-methods.md # .agents/notes/implemented/simplification/2026-06-20-prune-dead-seam-methods.zh.md # .agents/notes/implemented/simplification/2026-06-20-public-agent-stop-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-public-agent-stop-surface.zh.md # .agents/notes/implemented/simplification/2026-06-20-remove-agent-boundary-mirror-events.i18n.yaml # .agents/notes/implemented/simplification/2026-06-20-remove-agent-boundary-mirror-events.zh.md # .agents/notes/implemented/simplification/2026-06-26-fsspec-style-fs-seam.i18n.yaml # .agents/notes/implemented/simplification/2026-06-26-fsspec-style-fs-seam.zh.md # .agents/notes/implemented/simplification/2026-07-02-remove-stream-chunk-mirror.i18n.yaml # .agents/notes/implemented/simplification/2026-07-02-remove-stream-chunk-mirror.zh.md # .agents/notes/implemented/simplification/2026-07-04-drop-image-content-block.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-drop-image-content-block.zh.md # .agents/notes/implemented/simplification/2026-07-04-drop-inert-request-knobs.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-drop-inert-request-knobs.zh.md # .agents/notes/implemented/simplification/2026-07-04-drop-unconsumed-web-observation-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-drop-unconsumed-web-observation-surface.zh.md # .agents/notes/implemented/simplification/2026-07-04-fold-stdio-ui-helper.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-fold-stdio-ui-helper.md # .agents/notes/implemented/simplification/2026-07-04-fold-stdio-ui-helper.zh.md # .agents/notes/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.zh.md # .agents/notes/implemented/simplification/2026-07-04-prune-write-only-fs-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-prune-write-only-fs-surface.zh.md # .agents/notes/implemented/simplification/2026-07-04-remove-agent-steering-mirror.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-remove-agent-steering-mirror.zh.md # .agents/notes/implemented/simplification/2026-07-04-share-app-bin-boot-glue.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-share-app-bin-boot-glue.zh.md # .agents/notes/implemented/simplification/2026-07-04-tighten-hook-protocol-contract.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-tighten-hook-protocol-contract.zh.md # .agents/notes/implemented/simplification/2026-07-04-trim-acp-bridge-unreachable-surface.i18n.yaml # .agents/notes/implemented/simplification/2026-07-04-trim-acp-bridge-unreachable-surface.zh.md # .agents/notes/implemented/simplification/2026-07-12-drop-unconsumed-skill-provider-events.i18n.yaml # .agents/notes/implemented/simplification/2026-07-12-drop-unconsumed-skill-provider-events.zh.md # .agents/notes/implemented/simplification/2026-07-12-prune-unused-web-seam-fields.i18n.yaml # .agents/notes/implemented/simplification/2026-07-12-prune-unused-web-seam-fields.zh.md # .agents/notes/implemented/testing/2026-06-11-property-based-testing.i18n.yaml # .agents/notes/implemented/testing/2026-06-11-property-based-testing.zh.md # .agents/notes/implemented/testing/2026-06-19-acp-snapshot-tests.i18n.yaml # .agents/notes/implemented/testing/2026-06-19-acp-snapshot-tests.zh.md # .agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.i18n.yaml # .agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.zh.md # .agents/notes/implemented/testing/2026-06-20-remove-redundant-snapshot-log-goldens.i18n.yaml # .agents/notes/implemented/testing/2026-06-20-remove-redundant-snapshot-log-goldens.zh.md # .agents/notes/implemented/testing/2026-06-22-fork-child-replay-seed-boundary.i18n.yaml # .agents/notes/implemented/testing/2026-06-22-fork-child-replay-seed-boundary.zh.md # .agents/notes/implemented/testing/2026-06-22-fork-snapshot-scenarios.i18n.yaml # .agents/notes/implemented/testing/2026-06-22-fork-snapshot-scenarios.zh.md # .agents/notes/implemented/testing/2026-06-22-subagent-snapshot-replay.i18n.yaml # .agents/notes/implemented/testing/2026-06-22-subagent-snapshot-replay.zh.md # .agents/notes/implemented/testing/2026-07-04-hook-snapshot-matrix.i18n.yaml # .agents/notes/implemented/testing/2026-07-04-hook-snapshot-matrix.zh.md # .agents/notes/implemented/testing/2026-07-04-single-source-acp-replay-config.i18n.yaml # .agents/notes/implemented/testing/2026-07-04-single-source-acp-replay-config.zh.md # .agents/notes/implemented/testing/2026-07-06-pin-request-header-content-in-one-scenario.i18n.yaml # .agents/notes/implemented/testing/2026-07-06-pin-request-header-content-in-one-scenario.zh.md # .agents/notes/implemented/testing/2026-07-08-shared-acp-snapshot-package.i18n.yaml # .agents/notes/implemented/testing/2026-07-08-shared-acp-snapshot-package.zh.md # .agents/notes/proposed/architecture/2026-06-16-typed-event-schemas.i18n.yaml # .agents/notes/proposed/architecture/2026-06-16-typed-event-schemas.zh.md # .agents/notes/proposed/architecture/2026-06-20-generic-long-running-tool-runtime.i18n.yaml # .agents/notes/proposed/architecture/2026-06-20-generic-long-running-tool-runtime.zh.md # .agents/notes/proposed/feature/2026-06-30-pre-tool-input-rewrite.i18n.yaml # .agents/notes/proposed/feature/2026-06-30-pre-tool-input-rewrite.zh.md # .agents/notes/proposed/feature/2026-07-07-claude-code-and-codex-subagent-backends.i18n.yaml # .agents/notes/proposed/feature/2026-07-07-claude-code-and-codex-subagent-backends.zh.md # .agents/notes/proposed/feature/2026-07-08-interactive-side-sessions.i18n.yaml # .agents/notes/proposed/feature/2026-07-08-interactive-side-sessions.zh.md # .agents/notes/proposed/feature/2026-07-10-sqlite-session-query-provider.i18n.yaml # .agents/notes/proposed/feature/2026-07-10-sqlite-session-query-provider.zh.md # .agents/notes/proposed/feature/2026-07-13-stream-workflow-progress-through-tool-calls.i18n.yaml # .agents/notes/proposed/feature/2026-07-13-stream-workflow-progress-through-tool-calls.zh.md # .agents/notes/proposed/process/2026-06-11-api-extractor-reports.i18n.yaml # .agents/notes/proposed/process/2026-06-11-api-extractor-reports.md # .agents/notes/proposed/process/2026-06-11-api-extractor-reports.zh.md # .agents/notes/proposed/process/2026-06-11-architectural-conformance.i18n.yaml # .agents/notes/proposed/process/2026-06-11-architectural-conformance.zh.md # .agents/notes/proposed/process/2026-06-11-supply-chain-and-vendor-drift.i18n.yaml # .agents/notes/proposed/process/2026-06-11-supply-chain-and-vendor-drift.zh.md # .agents/notes/proposed/process/2026-06-20-discover-package-inventory.i18n.yaml # .agents/notes/proposed/process/2026-06-20-discover-package-inventory.zh.md # .agents/notes/proposed/simplification/2026-06-20-unify-agent-and-session-id.i18n.yaml # .agents/notes/proposed/simplification/2026-06-20-unify-agent-and-session-id.zh.md # .agents/notes/proposed/simplification/2026-07-04-prune-dead-core-spine-surface.i18n.yaml # .agents/notes/proposed/simplification/2026-07-04-prune-dead-core-spine-surface.zh.md # .agents/notes/proposed/simplification/2026-07-12-simplify-session-log-representation.i18n.yaml # .agents/notes/proposed/simplification/2026-07-12-simplify-session-log-representation.zh.md # .agents/notes/proposed/testing/2026-06-11-deterministic-and-stress-testing.i18n.yaml # .agents/notes/proposed/testing/2026-06-11-deterministic-and-stress-testing.zh.md # .agents/notes/proposed/testing/2026-06-11-mutation-testing.i18n.yaml # .agents/notes/proposed/testing/2026-06-11-mutation-testing.zh.md # .agents/notes/rejected/architecture/2026-06-11-immutable-public-surfaces.i18n.yaml # .agents/notes/rejected/architecture/2026-06-11-immutable-public-surfaces.zh.md # .agents/notes/rejected/architecture/2026-06-20-providerless-example-base.i18n.yaml # .agents/notes/rejected/architecture/2026-06-20-providerless-example-base.zh.md # .agents/notes/rejected/simplification/2026-06-20-assembled-assistant-messages-only.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-assembled-assistant-messages-only.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-acp-session-load.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-acp-session-load.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-acp-terminal-meta.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-acp-terminal-meta.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-bash-output-spill-files.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-bash-output-spill-files.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-durable-step-boundaries.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-durable-step-boundaries.zh.md # .agents/notes/rejected/simplification/2026-06-20-drop-unused-session-lineage.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-drop-unused-session-lineage.zh.md # .agents/notes/rejected/simplification/2026-06-20-fold-session-persistence-interface.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-fold-session-persistence-interface.zh.md # .agents/notes/rejected/simplification/2026-06-20-generic-tool-rendering.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-generic-tool-rendering.zh.md # .agents/notes/rejected/simplification/2026-06-20-retire-mid-turn-steering.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-retire-mid-turn-steering.zh.md # .agents/notes/rejected/simplification/2026-06-20-single-session-acp-bridge.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-single-session-acp-bridge.zh.md # .agents/notes/rejected/simplification/2026-06-20-truncate-interrupted-turns.i18n.yaml # .agents/notes/rejected/simplification/2026-06-20-truncate-interrupted-turns.zh.md # .agents/notes/rejected/simplification/2026-07-04-prune-unimplemented-subagent-vocabulary.i18n.yaml # .agents/notes/rejected/simplification/2026-07-04-prune-unimplemented-subagent-vocabulary.zh.md # .agents/notes/rejected/simplification/2026-07-12-collapse-workflow-to-foreground-core.i18n.yaml # .agents/notes/rejected/simplification/2026-07-12-collapse-workflow-to-foreground-core.zh.md # .agents/notes/rejected/simplification/2026-07-12-prune-unused-skill-registry-surface.i18n.yaml # .agents/notes/rejected/simplification/2026-07-12-prune-unused-skill-registry-surface.zh.md # docs/rfc/implemented/architecture/2026-06-18-agent-lifecycle-and-ownership-seams.md # docs/rfc/implemented/architecture/2026-06-18-session-surface.md # docs/rfc/implemented/architecture/2026-06-20-branded-ids.md # docs/rfc/implemented/architecture/2026-07-02-fs-per-session-cwd.md # docs/rfc/implemented/feature/2026-06-18-compaction-capability-seam.md # docs/rfc/implemented/feature/2026-07-07-session-prefix.md # docs/rfc/implemented/process/2026-06-20-rfc-classification.md # docs/rfc/implemented/process/2026-07-04-generate-rfc-index-tables.md # docs/rfc/implemented/process/2026-07-05-uniform-rfc-format.md # docs/rfc/implemented/process/2026-07-06-parallel-github-ci-gates.md # docs/rfc/implemented/process/2026-07-06-parallel-pre-push-gates.md # docs/rfc/implemented/process/2026-07-12-package-model-experience-contract.md # docs/rfc/implemented/simplification/2026-06-20-remove-agent-boundary-mirror-events.md # docs/rfc/implemented/simplification/2026-07-04-prune-producerless-vocabulary-variants.md # docs/rfc/implemented/testing/2026-06-20-remove-redundant-snapshot-log-goldens.md # docs/rfc/implemented/testing/2026-07-08-shared-acp-snapshot-package.md # docs/rfc/proposed/architecture/2026-06-20-generic-long-running-tool-runtime.md # docs/rfc/proposed/simplification/2026-06-20-unify-agent-and-session-id.md # docs/rfc/proposed/simplification/2026-07-12-simplify-session-log-representation.md # docs/rfc/rejected/simplification/2026-07-04-prune-unimplemented-subagent-vocabulary.md # scripts/translation-pairing.manifest.json
This commit is contained in:
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-06-30-pre-tool-input-rewrite.md: 85ece78f3bf188b3b702b1af539256747c1cfab2
|
||||
2026-06-30-pre-tool-input-rewrite.zh.md: 6a5c2b52627d21476b96448dc120155eab7f2223
|
||||
@@ -0,0 +1,53 @@
|
||||
# Agent Note: Pre-tool input rewrite — a consistent design
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-06-30-pre-tool-input-rewrite.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The [interception-seams Agent Note](../../implemented/feature/2026-06-30-interception-seams.md) defines `tools/pre-execute` as an allow/deny/ask gate over an execution whose identity is already protected and whose arguments are deeply frozen. Claude Code's `PreToolUse` hook also offers `updatedInput`, so a faithful bridge needs an explicit rewrite mechanism. A rewrite cannot be a mutation escape hatch on the existing execution object: it must keep the durable history, audit record, presentation, and executed value consistent.
|
||||
|
||||
## The problem: three readers of pre-execution arguments
|
||||
|
||||
In the loop, a tool call's arguments are committed to the log and read by live consumers BEFORE the tool executes:
|
||||
|
||||
1. **`assistant/message`** is appended before tool dispatch — it is the model-history source `deriveMessages()` replays, so it carries the tool-call arguments the model itself emitted.
|
||||
2. **`tool/call`** is the durable AUDIT record, appended before `ctx.tools.execute()`.
|
||||
3. **Live presentation reads `tool/call.arguments`**: the ACP bridge remembers them and passes them to `presentResult`; `dsh-tool-bash` derives the card title, the rawInput, the cwd, and the terminal-vs-background treatment from them.
|
||||
|
||||
An execution-only rewrite would make the UI show one command while another ran and render the result against the wrong arguments. The registry prevents that failure mode today: it structured-clones and deep-freezes `arguments`, makes the execution identity properties non-writable, and exposes no test shim or listener path that can replace them. The rewrite design must preserve that protected-identity boundary rather than weaken it.
|
||||
|
||||
## Proposal
|
||||
|
||||
A rewrite is a pre-identity consistency transaction. When a hook supplies `updatedInput`, the effective value must be chosen before the registry constructs its immutable `ToolExecution`, and it must be reflected in all three readers atomically:
|
||||
|
||||
- The `tool/call` audit event records the REWRITTEN arguments (with the original retained in a sidecar field for the audit trail — a hook changed the call, and both the original and the effective arguments are facts worth keeping).
|
||||
- The `assistant/message` in derived history must agree with what executed — options to evaluate: rewrite the assistant message's tool-call block in place (changes what the model "sees it said"), or record a separate correction the next request carries. The CC model is that the model sees the rewrite took effect.
|
||||
- Presentation (`presentCall`/`presentResult`) reads the rewritten arguments, so the UI shows what actually ran.
|
||||
|
||||
Extending `PreToolDecision` at its current firing point is insufficient: both durable records already exist by then, and the execution identity is protected. The implementation must either move the relevant decision before the log commit or add a dedicated earlier rewrite decision over the pending model call. After the loop commits the effective arguments to history and audit, it constructs the ordinary immutable execution and runs the existing allow/deny/ask and tool pipeline unchanged.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
### Why not mutate the execution object?
|
||||
|
||||
Allowing a pre-execute listener to assign `exec.arguments` would provide only an execution rewrite, leaving model history, audit, and presentation unchanged. Keeping the identity protected makes such partial behavior unrepresentable. Until the consistency transaction exists, a CC/Codex bridge logs and warns about `updatedInput` rather than claiming it was honored; `TODO(pre-tool-input-rewrite)` at the loop dispatch site anchors the missing earlier phase.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- A requested rewrite is resolved before `ToolExecution` identity is created and reflected in all three readers atomically: the `tool/call` audit records the rewritten arguments (the original retained in a sidecar field), derived history agrees with what executed, and presentation renders the rewritten arguments.
|
||||
- The effective `ToolExecution.arguments` remains deeply frozen and non-writable throughout pre-policy, guards, dispatch, post-policy, and final observation; no mutation shim is introduced.
|
||||
- The CC/Codex bridges honor `updatedInput` instead of logging the faithful-but-degraded warning.
|
||||
|
||||
## Risks
|
||||
|
||||
- Rewriting the `assistant/message` tool-call block changes what the model "sees it said"; whether any provider rejects that on replay is the open question that must be settled empirically before the decision shape freezes.
|
||||
- An earlier rewrite phase changes the ordering relationship among `assistant/message`, `tool/call`, hook audit events, and execution; the design must pin that ordering without weakening turn enclosure or call/result adjacency.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Does rewriting the `assistant/message` tool-call block corrupt any provider's expectation on replay, or is a separate correction safer?
|
||||
- Should the original arguments be preserved on the `tool/call` event (audit) and, if so, under what field?
|
||||
- Does the rewrite decision move before the log commit or become a dedicated earlier seam, and how do existing pre-tool allow/deny hooks avoid running twice?
|
||||
- How does this interact with a future permission `ask` flow (a user approving a rewritten call)?
|
||||
@@ -0,0 +1,53 @@
|
||||
# RFC: 工具执行前输入重写——一致性设计
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-06-30-pre-tool-input-rewrite.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
[拦截 seam RFC](../../implemented/feature/2026-06-30-interception-seams.md) 将 `tools/pre-execute` 定义为一道针对执行的允许/拒绝/询问门禁,此时执行的身份标识已受保护、参数已被深度冻结。Claude Code 的 `PreToolUse` 钩子还提供了 `updatedInput`,因此忠实的桥接需要一个显式的重写机制。重写不能是对现有执行对象的可变逃逸口:它必须保持持久化历史、审计记录、展示层与实际执行值之间的一致性。
|
||||
|
||||
## 问题本质:执行前参数的三个读取方
|
||||
|
||||
在 agent loop(智能体循环)中,工具调用的参数在工具执行之前就已提交到日志并被实时消费方读取:
|
||||
|
||||
1. **`assistant/message`** 在工具分发之前追加——它是 `deriveMessages()` 回放时的模型历史来源,因此携带模型自身输出的工具调用参数。
|
||||
2. **`tool/call`** 是持久化的审计记录,在 `ctx.tools.execute()` 之前追加。
|
||||
3. **展示层实时读取 `tool/call.arguments`**:ACP(Agent Client Protocol)桥接记住这些参数并传给 `presentResult`;`dsh-tool-bash` 从中派生卡片标题、rawInput、cwd 以及终端/后台处理方式。
|
||||
|
||||
如果只做执行层面的重写,UI 会显示一条命令而实际运行的是另一条,并且结果会对着错误的参数渲染。注册表目前通过以下方式防止这种失败模式:对 `arguments` 做 structured-clone 并深度冻结,将执行身份属性设为不可写,且不暴露任何可替换它们的测试 shim 或监听路径。重写设计必须维护这一受保护的身份边界,而非削弱它。
|
||||
|
||||
## 提案
|
||||
|
||||
重写是一个「身份标识创建前的一致性事务」。当钩子提供 `updatedInput` 时,有效值必须在注册表构造其不可变的 `ToolExecution` 之前确定,并且必须原子地反映到全部三个读取方:
|
||||
|
||||
- `tool/call` 审计事件记录重写后的参数(原始参数保留在一个伴随字段中,作为审计线索——钩子修改了调用,原始参数与生效参数都是值得保留的事实)。
|
||||
- 派生历史中的 `assistant/message` 必须与实际执行一致。待评估的选项:就地重写 assistant 消息中的工具调用块(改变模型「看到自己说了什么」),或记录一条单独的修正让下一次请求携带。Claude Code 的模型是让模型看到重写已生效。
|
||||
- 展示层(`presentCall`/`presentResult`)读取重写后的参数,使 UI 显示实际运行的内容。
|
||||
|
||||
在 `PreToolDecision` 当前的触发点上做扩展是不够的:此时两条持久化记录已经存在,执行身份已受保护。实现必须将相关决策移到日志提交之前,或者增加一个专门的、更早的重写决策点来处理待定的模型调用。agent loop 将生效参数提交到历史和审计之后,再构造普通的不可变执行对象,并照常运行现有的允许/拒绝/询问和工具流水线。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
### 为什么不直接修改执行对象?
|
||||
|
||||
允许 pre-execute 监听器赋值 `exec.arguments` 只能提供执行层面的重写,模型历史、审计和展示层不会随之改变。保持身份标识受保护使得这种局部行为不可表达。在一致性事务实现之前,CC/Codex 桥接对 `updatedInput` 记录日志并发出警告,而非声称已兑现;循环分发点的 `TODO(pre-tool-input-rewrite)` 标记了缺失的更早阶段。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- 请求的重写在 `ToolExecution` 身份标识创建之前解决,并原子地反映到全部三个读取方:`tool/call` 审计记录重写后的参数(原始参数保留在伴随字段中)、派生历史与实际执行一致、展示层渲染重写后的参数。
|
||||
- 生效的 `ToolExecution.arguments` 在 pre-policy、守卫、分发、post-policy 和最终观测全程保持深度冻结且不可写;不引入任何可变 shim。
|
||||
- CC/Codex 桥接兑现 `updatedInput`,不再记录忠实但降级的警告。
|
||||
|
||||
## 风险
|
||||
|
||||
- 重写 `assistant/message` 中的工具调用块会改变模型「看到自己说了什么」;是否有提供方在回放时拒绝这种改动,是一个需要通过实验确定的开放问题,必须在决策形状冻结之前解决。
|
||||
- 更早的重写阶段改变了 `assistant/message`、`tool/call`、钩子审计事件与执行之间的顺序关系;设计必须固定这一顺序,同时不削弱轮次封闭性或调用/结果邻接性。
|
||||
|
||||
## 开放问题
|
||||
|
||||
- 重写 `assistant/message` 中的工具调用块是否会破坏某些提供方在回放时的预期?还是单独的修正更安全?
|
||||
- 原始参数是否应保留在 `tool/call` 事件(审计)上?如果是,放在什么字段?
|
||||
- 重写决策是移到日志提交之前,还是成为一个专门的更早 seam?现有的 pre-tool 允许/拒绝钩子如何避免运行两次?
|
||||
- 这与未来的权限 `ask` 流程(用户批准一个被重写的调用)如何交互?
|
||||
@@ -0,0 +1,108 @@
|
||||
# Agent Note: Recallable compaction — index checkpoints, a state checkpoint, and in-session history recall
|
||||
|
||||
Status: proposed
|
||||
|
||||
## Problem
|
||||
|
||||
Compaction is a one-way door. The summary the model sees carries no reference to what it shadows — the `shadowedRange` provenance lives only on the log-only `compact/summary` event — and no tool lets the model read a shadowed span back. Whatever the summarizer drops is gone from the model's reachable world, even though the append-only log holds every byte. Repeated compaction compounds this: the head checkpoint is rewritten every pass, so the request prefix takes a full prompt-cache miss each time, and earlier summaries are re-summarized generation after generation.
|
||||
|
||||
The root cause is one artifact playing two conflicting roles. An **index** wants to be frozen, chronological, and cheap; the model's **working memory** wants a global view, re-prioritization, and mutability. A single summary can be neither well.
|
||||
|
||||
No mainstream coding harness gives the model in-loop recall, and none of the surveyed implementations makes compaction prefix-cache-aware. An event-sourced session — originals durable, seq-addressable, replay-exact — is the natural substrate for both.
|
||||
|
||||
## Proposal
|
||||
|
||||
Split the checkpoint into two classes and make shadowed history reachable.
|
||||
|
||||
### Frozen index checkpoints
|
||||
|
||||
Newly stale history splits into chunks by deterministic policy: accumulate toward `chunkTokens`, snap edges with `toolPairingBalancedBefore` / `toolPairingBalancedAfter`, prefer turn boundaries, and place the final boundary as close to the retain boundary as balance allows, so the trailing slice shrinks to roughly one turn. Each chunk is compacted by one `compactRegion` call into an **index stub** (`stubTokens`, ~100–200 tokens):
|
||||
|
||||
- two or three lines of what happened;
|
||||
- a keyword line of low-frequency literal anchors — exact error strings, values, config keys — grouped by kind;
|
||||
- a code-composed footer: `[checkpoint c<summarySeq>: shadows conversation span #<start>–#<end>; originals retrievable via history_read]`. Pointers are assembled from provenance, never model-authored.
|
||||
|
||||
A committed stub is never rewritten and never re-enters a later compaction region. A stub call's input is layered: the fixed preamble and the byte-identical pass-start state checkpoint (the shared prefix across all calls in the phase), then the keyword lines of all previously committed stubs — so a new entry indexes what is distinctive to its chunk instead of repeating the directory — the one or two most recent committed stubs for chronological continuity, and the slice itself. Sibling stubs from the same pass are not inputs (the concurrent phase forbids it; turn-aligned boundaries carry local continuity instead), and the state checkpoint is background only, never material to summarize into the stub. A slice consisting of recalled content is stubbed by code alone — a pointer line, no LLM call. A failed stub call degrades the same way: its slice gets a code-only pointer stub and the pass continues, making the state rewrite the only hard LLM dependency in a pass.
|
||||
|
||||
### The state checkpoint
|
||||
|
||||
One mutable working-memory document (at most one; zero before the first pass), positioned after all stubs and before the retained tail. Each pass rewrites it from the previous state plus this pass's staled content — O(previous + new), under the merge-don't-restate rule already in the summarization prompt — covering decisions, current state, constraints, and next steps. It carries its own footer and a size cap at the scale of today's summary.
|
||||
|
||||
An inflation guard bounds the whole pass: if the post-compaction size is not strictly below the pre-compaction size, nothing commits and the turn proceeds; the attempt defers until more stale history accumulates. The guard compares one metric on both sides — provider-reported usage from the request path, falling back to the character estimator on both sides.
|
||||
|
||||
### Pass execution
|
||||
|
||||
- Chunk slices are surface position ranges. A pass runs two phases: all summarize calls execute concurrently, buffered off-surface; then regions commit strictly left to right — chunks first, trailing slice last — so the state checkpoint lands after every stub through contiguous single-node replaces. Wall-clock stays near one summarize call.
|
||||
- The superseded state checkpoint folds into the next pass's first chunk as ordinary history: no tombstone, no new primitive. Its stub omits it, `history_read` renders it labeled `[prior state checkpoint]`, and its footer travels with the rendered text, keeping every trailing slice reachable through the two-hop chain.
|
||||
- Range selection is frozen-aware: the compactable span begins after the last committed index checkpoint, at the surface head only when none exists. A legacy session's existing head checkpoint is adopted as state-class — its text the merge base, its node folded like any superseded state.
|
||||
- A crash in the summarize phase commits nothing; a crash mid-commit leaves a left-to-right prefix committed, and the resumed pass reads its merge base from the log's latest state-class `compact/summary` event and commits the remaining regions unconditionally — restoring `[stubs…][state][tail]` outranks shrinking.
|
||||
|
||||
### The recall tools
|
||||
|
||||
A new package `@deepseek-ai/dsh-tool-recall` (consumer-only, over the `dsh-session` and `dsh-compact` vocabularies) registers two model-facing tools:
|
||||
|
||||
- `history_read(checkpoint, offset?)` — renders the shadowed span of any checkpoint in the log, including superseded ones, as `User:`/`Assistant:`/`Tool result:` transcript, paginated by a configured budget with a continuation cursor.
|
||||
- `history_search(query, checkpoint?, limit?)` — case-insensitive literal scan over every shadowed span; returns snippets with checkpoint ids and coverage metadata (`scanned`/`matched`/`truncated`). The zero-match hint notes the scan is literal and points at direct `history_read` of a plausible checkpoint.
|
||||
|
||||
Both read `exec.agent.session.events` (the tool-todo access pattern; non-agent callers rejected), render only surface-type message events, and return ordinary `tool/result`s — recalled bytes land at the context tail, logged, so reconstructability holds with no special casing. There is no new storage and no sidecar index: the session log is the archive, `compact/summary` provenance is the index metadata, and the tools are a read path over both. The tool schemas and the package's one system-prompt section are static strings; checkpoint ids reach the model only through footers. The transcript renderer moves from `compact-basic` into `dsh-session`, shared by summarizer and tools.
|
||||
|
||||
### Cache and cost
|
||||
|
||||
The request prefix after a pass is `[system][stubs…][state][tail]`. Frozen stubs are byte-stable across passes, so the miss begins at the token replacing the previous state checkpoint and stays O(new chunks + state + tail) — against position zero today. Recall output lands at the tail, leaving the prefix untouched. Per-pass summarize input is roughly twice today's plus an m·S background term, bounded by a `chunkTokens` floor (a small multiple of the state cap) and a validated `stubTokens`/`chunkTokens` ratio ceiling; a shared-prefix input layout (preamble, then the byte-identical pass-start state, slice content in the tail) lets sibling calls earn cached-rate rereads.
|
||||
|
||||
### Packaging
|
||||
|
||||
The design ships as a new backend `dsh-compact-recallable` on the existing `ctx.compact` seam, enabled by default in the shipped example configs; `compact-basic` remains as the reference implementation and the seam's design twin, in the pattern of the paired LLM adapters. The seam JSDoc's "at most one auto-generated checkpoint, always at the head" clause is relaxed to name both backend behaviors.
|
||||
|
||||
### Relation to in-flight work
|
||||
|
||||
- **Tool-result pruning** (the in-flight pruning service): its replacement nodes carry `sourceEventSeqs`; the same registry fold lists pruned results as recallable. Follow-up scope; neither blocks the other.
|
||||
- **Provider-usage token accounting** (the in-flight move of compaction pressure onto provider-reported usage): supplies the guard's accounting; the implementation stacks after it.
|
||||
- **"Query sessions" backlog item**: the cross-session generalization; this Agent Note scopes to the live session with tool names and rendering chosen so that work extends rather than collides.
|
||||
- **Training**: when to recall is a learned behavior. The deterministic footers and keyword anchors give training a stable target, and recall usage is fully visible in the session log for trajectory export; benchmark and RL design proceed with the post-training side.
|
||||
|
||||
### Follow-ups
|
||||
|
||||
Specified during review, deferred until observation calls for them:
|
||||
|
||||
- Guard degradation ladder (code-only rollup of the oldest stub prefix, footers preserved, rolled-up ids remain recall targets; then one summary after the frozen boundary) — on observed guard livelock or stub-region pressure.
|
||||
- Echo detection on stub outputs (sentence-scale n-grams, short literals exempt, retry then strip) — on observed division-of-labor leakage.
|
||||
- Periodic state refresh from chunk originals — on observed drift in the handoff probe.
|
||||
- `stateFallbackThreshold` (full-detail state prompt below a stub count) — on short-session regression.
|
||||
- Lazy registration of the recall tools — on measured context tax in never-compacting sessions.
|
||||
- Amortized stub drafting at pre-step: as soon as stale-but-uncompacted content accumulates past `chunkTokens`, draft that chunk's stub at the next pre-step (a log-only draft event, written while the chunk's surrounding context is still live) and let the compaction pass commit drafts instead of summarizing in bulk — the deterministic, replay-exact equivalent of background compaction (the Claude Code session-memory pattern; OpenClaw demonstrates the synchronous semantics are identical). Trigger: observed pass latency, or stub-quality gains from drafting near-live proving out.
|
||||
- Split summarizer models; model-chosen chunk boundaries; cross-session recall; semantic search fallback — each behind its own evidence.
|
||||
- Richer `history_search` query forms — regex, and structured queries over logged JSON tool results (sql/jq-style, or agent-authored queries against an indexed store) — on demand from observed search misses; literal matching ships first because the recall path stays a pure function of the log.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Staged delivery** (ship recall tools alone over today's backend; gate the checkpoint split on observed recall usage) — rejected: untrained models under-use any new tool, so the gate would measure training absence rather than design value, while the training side needs the complete mechanism to build environments against; the pre-release window is when persisted-format changes are cheapest; and the cache economics are first-party knowledge, not a hypothesis awaiting telemetry. The implementation still lands as stacked PRs with the recall tools first — construction order, not a decision gate.
|
||||
- **All-frozen full-size summaries, no state checkpoint** — rejected: unbounded permanent-prefix growth, self-accelerating toward thrashing, with nothing left to re-prioritize.
|
||||
- **Pure stubs, no state checkpoint** — rejected: presumes the model knows what it is missing; fails on unknown unknowns.
|
||||
- **LLM aging/consolidation of frozen chunks** — rejected as a routine mechanism: summary-of-summary loss and frozen-prefix churn; the code-only rollup is its surviving form, deferred.
|
||||
- **Full prefix as chunk-summarizer input** — rejected: O(N²); the state document gives the same background at O(state).
|
||||
- **One summarize call emitting all outputs** — rejected: the summarize path has no structured-output enforcement; parsing one free-text response apart is the fragile seam the fail-closed design avoids.
|
||||
- **Model-chosen chunk boundaries** — deferred: parse-and-validate cost against unproven value; chunk policy sits behind config.
|
||||
- **Model-authored pointers** — rejected: pointers must be exact; deterministic assembly is.
|
||||
- **FTS/vector index sidecar** — rejected in-session: the live log is in memory and bounded, a literal scan under budget suffices; an index earns its keep at cross-session scope.
|
||||
- **Semantic search fallback / secondary-model extraction in the recall path** — rejected: an LLM or embedding call there breaks keyless replay determinism; recall stays a pure function of the log.
|
||||
- **Raw events instead of rendered transcript** — rejected: leaks log-only vocabulary and chunk noise; the model reads what a model once saw.
|
||||
- **Doing nothing (resume/fork as recovery)** — rejected: it makes recovery a human act.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- Auto-compaction over a long session yields `[stubs…][state][tail]` after every completed pass; prior stubs stay byte-identical across passes; committed stubs never fall inside a later region; the superseded state checkpoint folds without a tombstone, renders labeled, and stays reachable and searchable through the two-hop chain.
|
||||
- Every checkpoint's surface text ends with the deterministic footer; footers round-trip through replay byte-identically; the state checkpoint's provenance records its wider input range.
|
||||
- Nothing commits before all summaries exist and the guard passes on like-for-like accounting; a guard failure commits nothing and does not fail the turn; a mid-commit kill resumed at the next pre-step completes the pass with the state region committed unconditionally, merge base read from the log; a legacy head checkpoint is adopted as state-class.
|
||||
- `history_read` renders any logged checkpoint's span under budget with a working cursor; `history_search` covers every shadowed span with checkpoint-id snippets and coverage metadata, asserted in particular by finding content that exists only in a span shadowed by a superseded state checkpoint — the regression pin for trailing-slice reachability; both reject non-agent callers and never-existing ids or orphaned `compact/start` with typed errors; recalled content appears as ordinary `tool/result`s; request-reconstruction invariants pass over sessions with compaction plus recall; one keyless snapshot scenario covers compact-then-recall end to end; tool schemas and the prompt section are byte-identical across passes.
|
||||
- On the long-horizon bench suite: task success does not regress against `compact-basic` at equal budgets; a handoff-fidelity probe (restate K known decisions and constraints after a pass) scores no worse; recall usage frequency and hit usefulness are reported per run via the dsh bench report pipeline, alongside the stub-directory attention measurement and cache-hit telemetry.
|
||||
- Seam JSDoc, the compaction capability-seam Agent Note, `architecture.md`, and the generated tool, config, persistence, and module-graph catalogs update in the same change; all budgets live in config; new source directories hold per-file 100% coverage with HMR disposal tests.
|
||||
|
||||
## Risks
|
||||
|
||||
- **Recall is a learned behavior**: untrained models will under-use it, and the bench report exists to track the gap while training closes it. Until then the state checkpoint keeps the floor at today's summary quality.
|
||||
- **Unknown unknowns remain**: a detail absent from summaries and keywords draws no recall. Recall converts "unreachable even when suspected" into "reachable when suspected".
|
||||
- **The stub directory occupies attention**: dozens of stable index cards per request may dilute focus; the bench measurement in the acceptance criteria tracks it against `compact-basic`.
|
||||
- **Cost**: per-pass summarize input is roughly twice today's; short sessions sit near today's cost and quality, and the design pays off with session length.
|
||||
- **State drift and division-of-labor leakage** are observable through the handoff probe and stub review; their counters are specified follow-ups.
|
||||
- **Two backends** are a maintenance surface; the seam contract and the shared recall consumer bound it, and the bench comparison decides the default over time.
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-07-claude-code-and-codex-subagent-backends.md: 5585ea30a5ba1b4200f096069a28ccf1c3cef727
|
||||
2026-07-07-claude-code-and-codex-subagent-backends.zh.md: 2a6dd2cdea34cb777df5455f0bc12dc7073ff2a2
|
||||
@@ -0,0 +1,89 @@
|
||||
# Agent Note: Claude Code and Codex subagent backends (out-of-process delegation to external coding agents)
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-07-claude-code-and-codex-subagent-backends.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The subagent seam ([the seam Agent Note](../../implemented/feature/2026-06-21-subagent-capability-seam.md)) hosts multiple named providers on `ctx.subagents`, and the ACP backend ([the ACP backend Agent Note](../../implemented/feature/2026-06-22-acp-subagent-backend.md)) proved the seam generalizes across a process boundary; its Future-providers section explicitly named the Codex app-server and the Claude Code Agent SDK as mechanically similar siblings. Those two are the engines actually worth delegating to today: a harness turn should be able to hand a self-contained task to a real Claude Code or a real Codex — a separate product with its own model, tools, and sandbox — and get back one final answer, without the parent deployment leaking its secrets into the child or the child's behavior silently depending on whatever `~/.claude` / `~/.codex` state exists on the host machine.
|
||||
|
||||
## Proposal
|
||||
|
||||
Two sibling provider packages, structural variants of the ACP backend, plus one extraction:
|
||||
|
||||
- `@deepseek-ai/dsh-subagent-claude-code` — drives a Claude Code child through `@anthropic-ai/claude-agent-sdk`'s `query()` (the SDK runs in the parent process and spawns its bundled `claude` CLI as the subprocess). Provider name `claude-code`: the child is the Claude Code *product*, not an Anthropic model adapter — "claude" stays reserved for a future `dsh-llm` adapter.
|
||||
- `@deepseek-ai/dsh-subagent-codex` — spawns `codex app-server` and drives one thread/turn over its JSON-RPC-over-stdio protocol with a hand-rolled newline-JSON client (~200–300 lines) in the package.
|
||||
- `@deepseek-ai/dsh-subagent-process` — a pure library (the `subagent-inprocess` precedent) extracting what `dsh-subagent-acp` already carries and both new backends need: the credential env scrub (`buildChildEnv`), the EOF → SIGTERM → SIGKILL dispose ladder, and new isolated-config-dir helpers (`mkdtemp` create, best-effort remove). The ACP backend migrates onto it; `bash-local`'s sibling copy is left alone to bound the change.
|
||||
|
||||
Both providers copy the ACP backend's seam posture verbatim: fresh child per `start`, exactly one prompt round-trip, capabilities all `false`, `inheritsParentContext: false`, `request.parent`/`request.agentOptions` ignored, `id = SessionId(randomUUID())`, `result` never rejects — child-level failure flattens to a stop reason and the original error goes to `ctx.logger` via an `onError` spec callback. Model exposure is zero new code: `dsh-tool-subagent` is loaded once per provider with a distinct `toolName` (`subagent_claude_code`, `subagent_codex`). No new session events are needed — the only model-visible artifact is the tool result, so reconstructability holds exactly as it did for ACP. To be explicit about the boundary: the session log reconstructs the model-visible transcript, not workspace mutation history — a child granted write access mutates files as an ambient side effect outside the log, exactly as the bash tools and the ACP backend already do; replay reproduces requests, not the disk.
|
||||
|
||||
## Verified interface facts (pinned versions)
|
||||
|
||||
Both integration surfaces were verified against pinned implementations before this proposal — types and bundled source read, keyless spikes run — not from vendor docs alone. The pins are the verification baseline, not a runtime contract: the backends perform no runtime version probe (no `codex --version` gate, no SDK version sniffing). Compatibility is enforced at development time — every dependency bump re-runs the keyless suites against the real load path — and at runtime by failing loudly: a protocol-level surprise settles `error` via `onError`, never a silent misbehavior.
|
||||
|
||||
**`@anthropic-ai/claude-agent-sdk` 0.3.202.** `options.env` REPLACES the child environment (no merge with `process.env`), which is exactly what the scrub needs. `settingSources` defaults to loading ALL filesystem settings — isolation requires explicitly passing `[]`. Result subtypes are `success` | `error_during_execution` | `error_max_turns` | `error_max_budget_usd` | `error_max_structured_output_retries`. On abort the SDK escalates the CLI child itself: stdin EOF immediately, SIGTERM ~2s later if the child ignores it (observed; no leftover processes) — no bespoke kill fallback needed. `outputFormat: {type: 'json_schema'}` and an `agents` option exist, giving future landing points for the seam's `outputSchema` capability and named subagent types; both are out of scope here.
|
||||
|
||||
**codex CLI 0.142.5, `codex app-server` (v2 vocabulary).** LF-delimited JSON, JSON-RPC 2.0 shapes with the `"jsonrpc"` header omitted.
|
||||
|
||||
- Lifecycle: `initialize{clientInfo}` + `initialized` → `thread/start` (accepts `cwd`, `model`, `sandbox`, `approvalPolicy`, `ephemeral`; succeeds unauthenticated) → `turn/start{threadId, input:[{type:'text',text}]}` returns an `inProgress` turn immediately; the terminal signal is the `turn/completed` notification carrying `Turn{status: completed|interrupted|failed|inProgress, error}`.
|
||||
- Approvals are server-initiated requests — `item/commandExecution/requestApproval`, `item/fileChange/requestApproval`, `item/permissions/requestApproval`, `item/tool/requestUserInput`, `mcpServer/elicitation/request` — answered with `accept`/`decline`-family decisions.
|
||||
- Auth: `account/login/start{type:'apiKey', apiKey}` is a first-class RPC and `account/read` reports `requiresOpenaiAuth` — and an unauthenticated `turn/start` does NOT fail fast (it hangs in retry), so the backend MUST pre-check auth and settle `error` loudly instead of waiting on the turn.
|
||||
- Isolation: `CODEX_HOME` redirection is honored (the `initialize` response echoes it, so tests can assert isolation), and `ephemeral: true` threads leave no session files at all.
|
||||
|
||||
## Isolation and credentials
|
||||
|
||||
Deployments authenticate with API keys only, and the child must not see the host user's Claude Code / Codex configuration: behavior has to be a function of `cordis.yml` alone. Each run gets a fresh `mkdtemp` config dir — `CLAUDE_CONFIG_DIR` for Claude Code (paired with an explicit `settingSources: []`), `CODEX_HOME` for Codex — removed best-effort on dispose; a config field can pin a persistent dir instead. The child env reuses the ACP backend's `buildChildEnv` semantics verbatim via the extraction: the ambient env is forwarded MINUS credential-shaped vars (`/KEY|SECRET|TOKEN/i`), with `config.env` layered on top — so `PATH`, `HOME`, `TMPDIR`, locale, and proxy vars survive and the CLIs run normally, while only credential-shaped ambient vars are scrubbed (`ANTHROPIC_API_KEY` enters explicitly through `config.env` for Claude Code), and the Codex key travels via the `account/login/start` RPC into the isolated `CODEX_HOME` rather than a hand-written `auth.json`.
|
||||
|
||||
## Permission and approval policy
|
||||
|
||||
Instead of collapsing to ACP's single `permission: allow|reject` knob, each backend exposes its engine's native vocabulary as config, with conservative defaults: Claude Code gets `permissionMode` (default `default`) plus `permission: allow|reject` (default `reject`) as the `canUseTool` auto-answer for whatever falls through; Codex gets `sandboxMode` (default `read-only`) and `approvalPolicy` (default `never`) plus the same `permission` fallback for approval requests that still arrive. Defaults are deliberately do-no-harm (the out-of-box child cannot write files); examples demonstrate opening up (`acceptEdits` / `workspace-write`). The mechanical rule: EVERY server-initiated request is settled programmatically and promptly — the enumerated approval/user-input/elicitation requests by the configured policy, an unknown request method with a JSON-RPC method-not-found error response (never left pending), unknown notifications consumed — so no child request can wedge a turn waiting on an answer that will never come. Prompts never reach a human in this cut, matching ACP.
|
||||
|
||||
## StopReason mapping
|
||||
|
||||
Claude Code: `success` → `completed`; `error_max_turns`, `error_during_execution`, `error_max_budget_usd`, `error_max_structured_output_retries` → `error` (aligning with the ACP call on `max_turn_requests`: an unfinished task is not success); generator abort → `aborted`; anything unknown → `error`. Codex: `Turn.status` `completed` → `completed`; `interrupted` → `aborted`; `failed` with `codexErrorInfo: 'contextWindowExceeded'` → `max-tokens`, any other `failed` → `error`; transport/spawn/auth-precheck failure → `error` (or `aborted` if cancel was requested). In both, `cancel()` is the ACP shape: flag + abort/interrupt + a cancel-settled race arm so an uncooperative child cannot stall the result.
|
||||
|
||||
Liveness posture, stated explicitly: teardown timing is config, turn duration is not. Both backends take the dispose ladder's grace periods as defaulted validated config fields (the ACP backend's `disposeEofGraceMs`/`disposeGraceMs` shape, carried by the extraction), but there is deliberately NO turn-duration or startup timeout — matching ACP, liveness during a turn belongs to the caller via `cancel()`/the abort signal, a subagent turn is legitimately minutes long, and the Codex auth precheck removes the one verified guaranteed-hang; a deployment wanting a wall-clock bound cancels from the parent.
|
||||
|
||||
## Testing
|
||||
|
||||
Named at every tier per the root AGENTS.md rule, and de-risked up front:
|
||||
|
||||
- **Keyless unit/integration**, mirroring the ACP spec list per backend (round-trip and output accumulation, every stop mapping, both cancel paths, already-aborted, permission auto-answer under both policies, unknown-message tolerance, bad-command spawn failure, HMR provider cleanup, export shape, isolation assertions on child env and temp-dir removal; Codex adds the auth-precheck failure path). Claude Code's harness is a scripted fake `claude` executable behind `pathToClaudeCodeExecutable` driven by the REAL SDK — a spike already passed end-to-end keyless in 24ms (the fake CLI answers one `control_request/initialize` and speaks plain stream-json, ~40 lines). Codex's harness is a scripted mock app-server subprocess speaking the verified wire protocol, the `mock-acp-server.ts` shape.
|
||||
- **With-key e2e** per backend: the real engine does real file work verified on disk, under a pinned opened-up config so acceptance and the do-no-harm defaults don't collide — `permissionMode: 'acceptEdits'` for Claude Code, `sandboxMode: 'workspace-write'` + `approvalPolicy: 'never'` for Codex; self-skips report exactly what is missing (binary vs key). CI has no secrets, so these run locally per the with-key policy.
|
||||
- **Snapshot**: deferred as `TODO(claude-code-subagent-replay)` / `TODO(codex-subagent-replay)` — the same distinct replay shape the ACP backend deferred ([the per-session replay Agent Note](../../implemented/testing/2026-06-22-subagent-snapshot-replay.md)); the keyless suites carry deterministic coverage meanwhile.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
### Why not the official `@openai/codex-sdk` instead of a hand-rolled client?
|
||||
|
||||
The dispose ladder and env scrub require owning the child process (spawn args, env, signals, exit await); the SDK hides the process. The wire format is trivial to frame (LF JSON), the shapes are generatable per pinned version (`codex app-server generate-json-schema`), and the repo precedent (`hook-protocol`) is to own thin protocol cores rather than wrap someone's runtime. The SDK would save protocol-evolution maintenance but costs the exact control this backend exists to have.
|
||||
|
||||
### Why not a model-visible `subagent_type` parameter (one Task-style tool)?
|
||||
|
||||
Claude Code's own Task tool puts the subagent type in the model-facing schema, selecting a prompt-plus-toolset persona. Here the choice is between EXECUTION ENGINES, and only the deployer knows which engines have credentials configured — so selection stays deployment config, preserving `dsh-tool-subagent`'s documented one-provider-per-tool contract. A persona-style type selector would be a separate Agent Note against the tool, not the backends.
|
||||
|
||||
### Why not login-state credentials and the user's own config?
|
||||
|
||||
Inheriting `~/.claude` / `~/.codex` (subscription login, user settings, skills, MCP servers) would make child behavior depend on host-machine state and punch an implicit exception through the "credentials enter explicitly via `config.env`, never ambiently" rule the ACP backend and bash executor established. API-key-only plus forced config-dir isolation keeps runs reproducible; deployments wanting shared state can point the config-dir field at a persistent directory deliberately.
|
||||
|
||||
### Why not a driver-injection seam for the Claude Code keyless tests?
|
||||
|
||||
Injecting a fake `query()` would mock our own boundary and leave the real SDK load path untested (the real-over-mock policy in docs/testing.md). The risk that justified considering it — the SDK↔CLI stream-json control protocol being internal — was retired by the spike: the fake-CLI harness works against the real pinned SDK today. If an SDK upgrade breaks the mock, the keyless suite fails the upgrade PR, which is the gate working.
|
||||
|
||||
### Why not ACP adapters (e.g. `claude-code-acp`) reusing the existing backend?
|
||||
|
||||
Community shims wrap both engines in ACP, which would make them "just config" on `dsh-subagent-acp`. But that inserts an unofficial third-party layer between the harness and the engine, erases the native control surfaces this Agent Note exposes (permissionMode, sandboxMode/approvalPolicy, config-dir isolation, apiKey RPC), and trades first-party protocol stability for a shim's release cadence. First-party surfaces — the Agent SDK and the app-server — are the supported integration points.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
On a machine with both engines and keys configured: a REPL-driven model completes one real file task through `subagent_claude_code` and one through `subagent_codex`, the tool result being the child's final answer, with only `tool/call` + `tool/result` in the parent session log. Keyless suites pass at 100% per-file coverage in a credential-less environment, asserting isolation (scrubbed child env, no temp config dirs left after dispose) and that child behavior is unchanged by the presence or absence of `~/.claude` / `~/.codex`. Cancelling a parent turn quiesces both backends in bounded time with no leftover child processes. E2e suites self-skip cleanly, naming the missing prerequisite.
|
||||
|
||||
## Risks
|
||||
|
||||
- `codex app-server` is CLI-flagged experimental and its v1/v2 vocabularies coexist; the client pins 0.142.5, implements v2 only, and consumes unknown methods/notifications without crashing, but a future codex bump can still force rework (regenerate schemas and re-run the keyless suite on every bump — the development-time enforcement behind the no-runtime-version-probe stance above).
|
||||
- The Claude Code fake-CLI mock rides an internal protocol: any SDK upgrade must go through the keyless suite, and a breaking control-protocol change means reworking the mock (fallback: the driver-injection seam rejected above becomes the escape hatch).
|
||||
- The SDK's optionalDependencies weigh ~280MB per platform — accepted, and confined to the one backend package.
|
||||
- The SDK's SIGKILL branch beyond EOF→SIGTERM was not observed and is trusted; e2e keeps a no-leftover-process assertion.
|
||||
- Codex is a deployment prerequisite (no npm-bundled binary); a missing or incompatible binary surfaces as a loud spawn/protocol `error`, not a version probe.
|
||||
- Every run pays a fresh child process and only the final answer surfaces — thoughts, tool cards, and usage are consumed and dropped; pooling, intermediate-progress surfacing, `sendMessage`/`resume`, `outputSchema` via the SDK's `outputFormat`, and named subagent types via the SDK's `agents` option are all deliberate deferrals.
|
||||
@@ -0,0 +1,89 @@
|
||||
# RFC: Claude Code 与 Codex subagent 后端(向外部编码 agent 的进程外委派)
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-07-claude-code-and-codex-subagent-backends.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
为 Claude Code 和 Codex 添加隔离的 subagent 提供方。既有的[命名提供方 seam](../../implemented/feature/2026-06-21-subagent-capability-seam.md) 和 [ACP 后端](../../implemented/feature/2026-06-22-acp-subagent-backend.md)已确立了进程边界的形状。harness 的一个轮次应能将一个自包含任务委派给上述任一产品,并接收其最终答案,同时不暴露父进程的密钥,也不继承来自 `~/.claude` 或 `~/.codex` 的宿主配置。
|
||||
|
||||
## 提案
|
||||
|
||||
两个兄弟提供方包(ACP 后端的结构变体),加一次提取:
|
||||
|
||||
- `@deepseek-ai/dsh-subagent-claude-code`:通过 `@anthropic-ai/claude-agent-sdk` 的 `query()` 驱动一个 Claude Code 子进程(SDK 在父进程中运行,并将其内置的 `claude` CLI 作为子进程 spawn)。提供方名称为 `claude-code`:子进程是 Claude Code 这个*产品*,而非 Anthropic 模型适配器——"claude" 保留给未来的 `dsh-llm` 适配器。
|
||||
- `@deepseek-ai/dsh-subagent-codex`:spawn `codex app-server`,通过其 JSON-RPC-over-stdio 协议驱动一个 thread/turn,使用包内一个手写的换行 JSON 客户端(约 200–300 行)。
|
||||
- `@deepseek-ai/dsh-subagent-process`:纯库(沿用 `subagent-inprocess` 的先例),提取 `dsh-subagent-acp` 已有且两个新后端都需要的内容:凭证环境清洗(`SENSITIVE_ENV_PATTERN`/`buildChildEnv`)、EOF → SIGTERM → SIGKILL 的 dispose 阶梯,以及新的隔离配置目录辅助函数(`mkdtemp` 创建、尽力删除)。ACP 后端迁移到该库上;`bash-local` 的兄弟副本保持不动以限制变更范围。
|
||||
|
||||
两个提供方遵循 ACP 后端契约:每次 `start` 创建一个全新子进程、一次 prompt 往返、不继承父上下文也不声明可选能力、忽略 `request.parent` 和 `request.agentOptions`、使用随机的品牌化 agent id。`result` 从不 reject;子进程失败映射为 stop reason,原始错误送入 logger。每个提供方以不同的工具名挂载 `dsh-tool-subagent`。工具结果是唯一新增的模型可见产物,因此无需新的会话事件;工作区变更仍是 transcript(文本记录)回放之外的环境副作用。
|
||||
|
||||
## 已验证的接口事实(固定版本)
|
||||
|
||||
两个集成面在本提案之前均已针对固定版本进行了验证——阅读类型与打包源码、运行无需密钥的 spike——而非仅依赖厂商文档。固定版本是验证基线,不是运行时契约:后端不执行运行时版本探测(无 `codex --version` 门禁、无 SDK 版本嗅探)。兼容性在开发时强制执行——每次依赖升级都会针对真实加载路径重跑无密钥套件——在运行时则通过大声失败来保障:协议层面的意外通过 `onError` 结算为 `error`,绝不静默异常。
|
||||
|
||||
**`@anthropic-ai/claude-agent-sdk` 0.3.202。** `options.env` 会替换子进程环境(不与 `process.env` 合并),恰好满足清洗需求。`settingSources` 默认加载所有文件系统设置——隔离要求显式传入 `[]`。结果子类型为 `success` | `error_during_execution` | `error_max_turns` | `error_max_budget_usd` | `error_max_structured_output_retries`。中止时 SDK 自行升级 CLI 子进程:立即关闭 stdin,约 2 秒后若子进程未退出则发送 SIGTERM(已观察到;无残留进程)——无需自定义 kill 回退。`outputFormat: {type: 'json_schema'}` 和 `agents` 选项已存在,为 seam 的 `outputSchema` 能力和命名 subagent 类型提供了未来着陆点;两者均不在本 RFC 范围内。
|
||||
|
||||
**codex CLI 0.142.5,`codex app-server`(v2 词汇)。** LF 分隔的 JSON,JSON-RPC 2.0 形状但省略 `"jsonrpc"` 头。
|
||||
|
||||
- 生命周期:`initialize{clientInfo}` + `initialized` → `thread/start`(接受 `cwd`、`model`、`sandbox`、`approvalPolicy`、`ephemeral`;未认证即可成功)→ `turn/start{threadId, input:[{type:'text',text}]}` 立即返回一个 `inProgress` 的 turn;终止信号是携带 `Turn{status: completed|interrupted|failed|inProgress, error}` 的 `turn/completed` 通知。
|
||||
- 审批是服务端发起的请求——`item/commandExecution/requestApproval`、`item/fileChange/requestApproval`、`item/permissions/requestApproval`、`item/tool/requestUserInput`、`mcpServer/elicitation/request`——以 `accept`/`decline` 系列决策应答。
|
||||
- 认证:`account/login/start{type:'apiKey', apiKey}` 是一等 RPC,`account/read` 报告 `requiresOpenaiAuth`——且未认证的 `turn/start` 不会快速失败(它会挂在重试中),因此后端必须预检认证状态,并在失败时大声结算为 `error`,而非等待 turn。
|
||||
- 隔离:`CODEX_HOME` 重定向被尊重(`initialize` 响应会回显它,测试可据此断言隔离),`ephemeral: true` 的 thread 不留任何会话文件。
|
||||
|
||||
## 隔离与凭证
|
||||
|
||||
认证方式仅限 API key。每次运行使用一个全新的配置目录(Claude Code 用 `CLAUDE_CONFIG_DIR` 配合 `settingSources: []`,Codex 用 `CODEX_HOME`),dispose 时尽力删除;配置也可以选择一个持久目录。共享的子进程环境辅助函数转发 `PATH`、`HOME`、`TMPDIR`、locale 和代理设置等普通值,移除凭证形态的名称,并叠加显式的 `config.env`。Claude Code 通过该叠加接收 API key,而 Codex 通过 `account/login/start` 接收,而非手写认证文件。
|
||||
|
||||
## 权限与审批策略
|
||||
|
||||
每个后端暴露其引擎原生的策略词汇。Claude Code 默认 `permissionMode: default` 配合 `permission: reject`;Codex 默认 `sandboxMode: read-only`、`approvalPolicy: never`,以及相同的拒绝回退。示例可选择启用 `acceptEdits` 或 `workspace-write`。已知的审批、用户输入和 elicitation 请求接收配置的应答;未知方法接收 method-not-found,未知通知被消费。没有 prompt 到达人类,子进程也不会因等待不可用的输入而无限挂起。
|
||||
|
||||
## StopReason 映射
|
||||
|
||||
Claude Code:`success` → `completed`;`error_max_turns`、`error_during_execution`、`error_max_budget_usd`、`error_max_structured_output_retries` → `error`(与 ACP 对 `max_turn_requests` 的处理对齐:未完成的任务不是成功);生成器中止 → `aborted`;未知值 → `error`。Codex:`Turn.status` 为 `completed` → `completed`;`interrupted` → `aborted`;`failed` 且 `codexErrorInfo: 'contextWindowExceeded'` → `max-tokens`,其他 `failed` → `error`;传输/spawn/认证预检失败 → `error`(若已请求取消则为 `aborted`)。两者中,`cancel()` 采用 ACP 形状:标志位 + abort/interrupt + 一个 cancel-settled 竞争分支,使不合作的子进程无法阻塞结果。
|
||||
|
||||
活性姿态,明确声明:teardown 时序是配置项,turn 时长不是。两个后端将 dispose 阶梯的宽限期作为带默认值的已验证配置字段(ACP 后端的 `disposeEofGraceMs`/`disposeGraceMs` 形状,由提取库承载),但刻意不设 turn 时长或启动超时——与 ACP 一致:turn 期间的活性由调用方通过 `cancel()`/abort signal 掌控,subagent turn 合理地可达数分钟,而 Codex 认证预检消除了唯一已验证的必然挂起场景;需要墙钟上限的部署从父侧取消即可。
|
||||
|
||||
## 测试
|
||||
|
||||
每个适用层级都要求覆盖:
|
||||
|
||||
- **无密钥单元/集成测试:** 通过真实 SDK 驱动一个假 Claude CLI,通过真实协议客户端驱动一个脚本化的 Codex app-server。在逐文件 100% 覆盖率下,验证往返、每种 stop 映射、两条取消路径及预中止、权限策略、未知消息、spawn 失败、reload 清理、导出形状、清洗后的环境、临时目录删除,以及 Codex 认证预检失败。
|
||||
- **有密钥 e2e 测试:** 每个真实引擎在 `acceptEdits` 或 `workspace-write` 下执行文件操作;跳过时命名缺失的二进制或密钥,并断言无残留子进程。
|
||||
- **快照测试:** 标记为 `TODO(claude-code-subagent-replay)` 和 `TODO(codex-subagent-replay)` 推迟,等待 [subagent 回放 RFC](../../implemented/testing/2026-06-22-subagent-snapshot-replay.md) 描述的进程特定回放形状。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
### 为什么不用官方 `@openai/codex-sdk` 而手写客户端?
|
||||
|
||||
dispose 阶梯和环境清洗要求拥有子进程(spawn 参数、env、信号、exit 等待);SDK 隐藏了进程。协议格式极其简单(LF JSON),形状可按固定版本生成(`codex app-server generate-json-schema`),仓库先例(`hook-protocol`)是拥有薄协议核心而非包装他人的运行时。SDK 能节省协议演进的维护成本,但代价是失去本后端存在的意义所在的精确控制。
|
||||
|
||||
### 为什么不用模型可见的 `subagent_type` 参数(单一 Task 风格工具)?
|
||||
|
||||
Claude Code 自身的 Task 工具将 subagent 类型放在模型可见的 schema 中,选择一个 prompt + 工具集人格。这里的选择是在执行引擎之间做出的,而只有部署者知道哪些引擎配置了凭证——因此选择留在部署配置层,保持 `dsh-tool-subagent` 文档中的「一个提供方对应一个工具」契约。人格风格的类型选择器应是针对工具的另一个 RFC,而非针对后端。
|
||||
|
||||
### 为什么不用登录态凭证和用户自身的配置?
|
||||
|
||||
继承 `~/.claude` / `~/.codex`(订阅登录、用户设置、skill、MCP 服务器)会使子进程行为依赖宿主机状态,并在 ACP 后端和 bash 执行器确立的「凭证通过 `config.env` 显式进入,绝不隐式继承」规则上打开一个隐式例外。仅 API key 加强制配置目录隔离使运行可复现;需要共享状态的部署可以有意将配置目录字段指向一个持久目录。
|
||||
|
||||
### 为什么不为 Claude Code 无密钥测试注入驱动层 seam?
|
||||
|
||||
注入假的 `query()` 会 mock 我们自己的边界,使真实 SDK 加载路径未被测试(docs/testing.md 中的 real-over-mock 策略)。曾考虑此方案的风险——SDK↔CLI 的 stream-json 控制协议是内部实现——已被 spike 消除:假 CLI harness 今天能对真实固定版本的 SDK 正常工作。如果 SDK 升级破坏了 mock,无密钥套件会让升级 PR 失败,这正是门禁在发挥作用。
|
||||
|
||||
### 为什么不用 ACP 适配器(如 `claude-code-acp`)复用既有后端?
|
||||
|
||||
社区 shim 将两个引擎包装为 ACP,这会使它们在 `dsh-subagent-acp` 上变成「仅配置」。但这在 harness 与引擎之间插入了一个非官方的第三方层,抹去了本 RFC 暴露的原生控制面(permissionMode、sandboxMode/approvalPolicy、配置目录隔离、apiKey RPC),并以 shim 的发布节奏替换了第一方协议的稳定性。第一方接口——Agent SDK 和 app-server——才是受支持的集成点。
|
||||
|
||||
## 验收标准
|
||||
|
||||
在两个引擎和密钥均已配置的机器上:一个 REPL 驱动的模型通过 `subagent_claude_code` 完成一个真实文件任务,通过 `subagent_codex` 完成另一个,工具结果为子进程的最终答案,父会话日志中仅有 `tool/call` + `tool/result`。无密钥套件在无凭证环境下以逐文件 100% 覆盖率通过,断言隔离(清洗后的子进程环境、dispose 后无残留临时配置目录),并断言 `~/.claude` / `~/.codex` 的存在与否不影响子进程行为。取消父轮次后,两个后端在有界时间内静默,无残留子进程。e2e 套件干净地自跳过,命名缺失的前置条件。
|
||||
|
||||
## 风险
|
||||
|
||||
- `codex app-server` 被 CLI 标记为实验性,其 v1/v2 词汇共存;客户端固定 0.142.5、仅实现 v2、对未知方法/通知消费而不崩溃,但未来 codex 升级仍可能迫使返工(每次升级重新生成 schema 并重跑无密钥套件——这是上述「不做运行时版本探测」立场背后的开发时强制执行)。
|
||||
- Claude Code 假 CLI mock 依赖一个内部协议:任何 SDK 升级都必须通过无密钥套件,控制协议的破坏性变更意味着返工 mock(回退方案:上面否决的驱动注入 seam 成为逃生舱口)。
|
||||
- SDK 的 optionalDependencies 每平台约 280MB——已接受,限制在单个后端包内。
|
||||
- SDK 的 SIGKILL 分支(EOF→SIGTERM 之后)未被观察到,信任其实现;e2e 保留无残留进程断言。
|
||||
- Codex 是部署前置条件(无 npm 内置二进制);缺失或不兼容的二进制以大声的 spawn/协议 `error` 呈现,而非版本探测。
|
||||
- 每次运行付出一个全新子进程的代价,且仅最终答案浮出——思考、工具卡片和用量被消费后丢弃;连接池、中间进度浮出、`sendMessage`/`resume`、通过 SDK 的 `outputFormat` 实现 `outputSchema`、以及通过 SDK 的 `agents` 选项实现命名 subagent 类型,均为刻意推迟。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-08-interactive-side-sessions.md: ac29f80b31492ce79512cc4d08e33480e0ac6258
|
||||
2026-07-08-interactive-side-sessions.zh.md: 5d0ef101dc23febefec881b12fcbb5ba4dcf8be9
|
||||
@@ -0,0 +1,41 @@
|
||||
# Agent Note: Interactive side sessions and merge-back
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-08-interactive-side-sessions.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
A user may want to explore a question from a live session without changing its main context. Existing primitives do not expose that product shape: [session-store fork](../../implemented/feature/2026-06-30-session-store-fork-api.md) creates an unattached session, while [fork subagents](../../implemented/feature/2026-06-21-subagent-capability-seam.md) are model-driven tasks whose transcript collapses into one tool result. Neither gives the user a separate conversation, and neither records a conclusion back into the parent with provenance.
|
||||
|
||||
## Proposal
|
||||
|
||||
A **side session** is an ordinary live session forked at the source's last completed turn, attached to its own agent, framed as a read-only advisor, and able to **merge back** one condensed note.
|
||||
|
||||
- **Fork and attach:** create the child with the parent's balanced completed-turn prefix and stamp `parentSession` and `seedLength` in its metadata. This composes `ctx.agents.create({ seed, meta })`; it adds no core service or session-store method.
|
||||
- **Advisor framing:** inject one plugin-sourced `context/message` after creation that tells the child to explain without mutating or continuing the task. Keeping the system prompt byte-identical preserves the provider prefix cache over inherited history.
|
||||
- **Merge-back:** ask the child for a length-capped handback, then inject one plugin-sourced `context/message` into the parent. The next parent request sees it at its logged position, preserving replay and [request reconstructability](../../implemented/architecture/2026-07-05-reconstructable-requests.md) without a new session event.
|
||||
- **Presentation:** invocation, session switching, and handback rendering belong to the first client-owned surface. This Agent Note specifies only the surface-independent mechanics.
|
||||
|
||||
Rewind productization, session-tree views, a model-facing side-session tool, and `forkName`/`mergedInto` metadata are out of scope. A live-adapter spike validated source-log isolation, inherited context, a multi-turn child exchange, and merge-back visibility in the parent's next turn.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Use the subagent seam:** rejected because side sessions are user-driven, client-visible, and may outlive a parent turn; subagents are model-driven runs returning one tool result.
|
||||
- **Change the child system prompt:** rejected by default because any byte change invalidates the prefix cache from token zero. Deployments may still prefer that stronger separation.
|
||||
- **Add `sidechat/*` events:** deferred because a sourced `context/message` already provides durability, provenance, and replay. A dedicated event is justified only by a surface that needs distinct rendering.
|
||||
- **Bind a protocol surface now:** rejected because current UIs are client-owned. Live presentation must eventually derive from the durable message so replay renders the same record.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- Forking leaves the source untouched and creates a child with the balanced completed-turn prefix, `parentSession`, `seedLength`, and a byte-identical system prompt.
|
||||
- Advisor framing adds exactly one plugin-sourced `context/message` at the head of the child's appended history, rather than changing its system prompt.
|
||||
- Merge-back adds exactly one length-capped `context/message` with source `plugin: sidechat`; the next parent request and replay see it at the same position.
|
||||
- Parent and child run concurrently without log or stream cross-talk.
|
||||
- Unit tests cover fork/attach and merge-back; snapshot coverage lands with the first bound surface.
|
||||
|
||||
## Risks
|
||||
|
||||
- Read-only behavior is advisory until a `tools/pre-execute` deny gate enforces it; [the interception seam](../../implemented/feature/2026-06-30-interception-seams.md) can add that gate without changing these mechanics.
|
||||
- A compacted source forks its compacted view, so a bound surface should disclose that the child inherits summaries rather than replaced turns.
|
||||
- Repeated handbacks consume parent context. The per-merge length cap bounds each note; later consolidation belongs to compaction.
|
||||
@@ -0,0 +1,41 @@
|
||||
# RFC: 交互式侧会话与合并回写
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-08-interactive-side-sessions.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
用户可能希望在不改变当前会话主上下文的前提下,探索一个来自活跃会话的问题。现有原语无法提供这种产品形态:[session-store fork](../../implemented/feature/2026-06-30-session-store-fork-api.md) 创建的是一个无关联的会话,而 [fork subagent](../../implemented/feature/2026-06-21-subagent-capability-seam.md) 是模型驱动的任务,其 transcript(文本记录)会折叠为一条工具结果。两者都不能给用户一个独立的对话,也都不能将结论带着出处信息记录回父会话。
|
||||
|
||||
## 提案
|
||||
|
||||
**侧会话(side session)** 是一个普通的活跃会话,从源会话的最后一个已完成轮次 fork 而来,绑定到自己的 agent,定位为只读顾问,并能**合并回写**一条精简笔记。
|
||||
|
||||
- **Fork 并绑定:** 以父会话的平衡已完成轮次前缀创建子会话,并在其元数据中标记 `parentSession` 与 `seedLength`。这组合了 `ctx.agents.create({ seed, meta })`;不新增核心服务或 session-store 方法。
|
||||
- **顾问定位:** 创建后注入一条插件来源的 `context/message`,告知子会话只做解释,不执行变更或继续任务。保持系统提示词逐字节一致,可在继承的历史上保留提供方的前缀缓存。
|
||||
- **合并回写:** 向子会话请求一条有长度上限的 handback,然后向父会话注入一条插件来源的 `context/message`。父会话的下一次请求在其日志位置看到该消息,保持回放与[请求可重建性](../../implemented/architecture/2026-07-05-reconstructable-requests.md),无需新增会话事件。
|
||||
- **呈现:** 调用方式、会话切换与 handback 渲染属于首个客户端拥有的界面。本 RFC 仅规定与界面无关的机制。
|
||||
|
||||
回退产品化、会话树视图、面向模型的侧会话工具,以及 `forkName`/`mergedInto` 元数据均不在本 RFC 范围内。一次 live-adapter spike 已验证了源日志隔离、继承上下文、多轮子会话交互,以及合并回写在父会话下一轮次中的可见性。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **使用 subagent seam:** 否决。侧会话是用户驱动的、客户端可见的,且可能存活超过父会话的一个轮次;subagent 是模型驱动的运行,返回一条工具结果。
|
||||
- **修改子会话的系统提示词:** 默认否决,因为任何字节变化都会从第零个 token 起使前缀缓存失效。部署方仍可选择这种更强的隔离方式。
|
||||
- **新增 `sidechat/*` 事件:** 延后。插件来源的 `context/message` 已提供持久性、出处与回放能力;只有当某个界面需要差异化渲染时,专用事件才有正当理由。
|
||||
- **现在就绑定一个协议界面:** 否决。当前 UI 由客户端拥有。实时呈现最终必须从持久消息派生,以使回放渲染出相同的记录。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- Fork 不改变源会话,创建的子会话具有平衡的已完成轮次前缀、`parentSession`、`seedLength`,以及逐字节一致的系统提示词。
|
||||
- 顾问定位在子会话追加历史的头部恰好添加一条插件来源的 `context/message`,而非修改其系统提示词。
|
||||
- 合并回写恰好添加一条有长度上限的 `context/message`,来源为 `plugin: sidechat`;父会话的下一次请求与回放在相同位置看到它。
|
||||
- 父会话与子会话并发运行,日志和流之间无串扰。
|
||||
- 单元测试覆盖 fork/attach 与合并回写;快照覆盖率随首个绑定界面一起落地。
|
||||
|
||||
## 风险
|
||||
|
||||
- 只读行为在 `tools/pre-execute` 拒绝门禁强制执行之前仅为建议性质;[拦截 seam](../../implemented/feature/2026-06-30-interception-seams.md) 可在不改变本机制的前提下添加该门禁。
|
||||
- 经过压缩(compaction)的源会话 fork 出的是其压缩视图,因此绑定的界面应当告知用户子会话继承的是摘要而非被替换的轮次。
|
||||
- 反复的 handback 会消耗父会话上下文。每次合并的长度上限约束了单条笔记的大小;后续的合并整理属于上下文压缩的职责。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-10-sqlite-session-query-provider.md: d1901ab0e37e8f92af322facac0cb48988d1ffe7
|
||||
2026-07-10-sqlite-session-query-provider.zh.md: 7a86d7294183eec57c3495a182f1403f84c27363
|
||||
@@ -0,0 +1,53 @@
|
||||
# Agent Note: SQLite FTS5 session search
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-10-sqlite-session-query-provider.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The exact-read `ctx.sessionQuery` service deliberately has no derived index. Large persisted histories need full-text search without scanning every event on every query, while current live sessions need an overlay newer than the last durability checkpoint. Search also needs concrete ranking, snippets, pagination, cancellation, and rebuild behavior.
|
||||
|
||||
Splitting those concerns across a speculative provider coordinator and a database implementation would create two coupled reconciliation state machines. The first real implementation should own the source observation, extraction, SQLite transaction, generation, and query as one lifecycle.
|
||||
|
||||
## Proposal
|
||||
|
||||
Add `@deepseek-ai/dsh-session-query-sqlite` beside the exact-read package. The package will expose a search service or extend the family with the smallest API required by its actual consumers; phase one does not pre-commit a provider-registration protocol. It will depend on `ctx.sessions` and optional `ctx.sessionPersistence`, own a separate derived SQLite database, and reuse the canonical `foldSurface()` classification.
|
||||
|
||||
The implementation owns one serialized reconciliation/DB transaction state machine. A transaction observes authoritative persisted metadata and live snapshots, extracts semantic documents, updates derived tables, advances relevant cursor generations, and executes or enables the corresponding query. No second service maintains parallel fingerprints, dirty flags, live-id sets, or invalidation generations.
|
||||
|
||||
Persisted documents survive restarts. Live overrides are connection-local and shadow the persisted rows for the same session, then disappear when the live owner or database closes. The derived database remains separate from canonical persistence so index reset, corruption, tokenizer changes, and schema churn cannot endanger durable conversation logs.
|
||||
|
||||
## Search semantics to decide with implementation
|
||||
|
||||
The implementation must define both cross-session and within-session scopes from executable use cases. Each searchable event is one document with session metadata, event metadata, surface classification, normalized semantic text, and a bounded plain-text snippet. Session results group by their strongest matching event; numeric backend scores remain private.
|
||||
|
||||
Search returns content-bearing result records rather than metadata-only headers. Chainable filters operate on that exact result shape and are designed and implemented with the search API instead of becoming a provider-specific pre-ranking contract. Query syntax is treated as data. Ordering includes stable tie fields. Opaque cursors bind to normalized request shape and the smallest relevant generation; unrelated session changes should not invalidate a within-session cursor. Cancellation must stop caller waiting and interrupt SQLite work where the runtime permits.
|
||||
|
||||
Tokenizer choice remains an implementation experiment. FTS5 trigram supports substring recall but rejects useful terms shorter than three characters and increases index size; the proposal must benchmark that tradeoff against the default Unicode tokenizer before making it contract.
|
||||
|
||||
## Extraction and reconciliation
|
||||
|
||||
The package starts with first-party semantic extraction for messages, reasoning, tool calls/results, blocked prompts, context, steering, todos, and error/status detail. Structural events and stream chunks contribute no document. Unknown declaration-merged event/content types remain non-searchable unless a real extension consumer demonstrates the need for a public extractor registry.
|
||||
|
||||
Reconciliation may use stable fingerprints to avoid rewriting unchanged persisted sessions, but the database package owns their calculation and storage. It must never report a row current when source observation or extraction failed. Provider-schema mismatch may reset only the derived database; ordinary source changes use transactional upsert/delete. Mounted but unreadable persistence fails affected searches without affecting canonical writes or known live exact reads.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Add FTS tables to the canonical persistence database** — rejected because a rebuildable index must not share the authoritative log's schema/reset/failure boundary.
|
||||
- **Reintroduce phase-one provider coordination** — rejected because there is one planned implementation and no evidence for a stable multi-provider seam.
|
||||
- **Persist live overrides immediately** — rejected because live events are not canonical until the existing checkpoint commits.
|
||||
- **Return BM25 scores** — rejected because provider-specific numeric scales are unstable across corpus changes.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- Restart tests cover unchanged, new, changed, and deleted persisted sessions without rebuilding the whole index.
|
||||
- Reopening preserves persisted rows and removes live rows; live rows shadow and then reveal their persisted base.
|
||||
- Tests cover both search scopes, content-bearing results, chainable result filters, surface defaults, snippets, escaping, deterministic ties, pagination, scoped stale cursors, cancellation, dynamic persistence mount/unmount, and recovery after a failed transaction.
|
||||
- A schema mismatch resets only the derived database.
|
||||
- A keyless end-to-end test combines a real persistence backend with the real SQLite search package.
|
||||
- The Agent Note is amended to the measured tokenizer and public API actually implemented before moving to `implemented/`.
|
||||
|
||||
## Risks
|
||||
|
||||
A single owner is simpler but initially less reusable than a provider-neutral seam. That is intentional: a second real backend can reveal what to extract. SQLite runtime differences can affect FTS ranking and snippets, so tests must pin only contract-controlled ordering and presentation. The separate database adds configuration and lifecycle work, but preserves the canonical store's safety boundary.
|
||||
@@ -0,0 +1,53 @@
|
||||
# RFC: SQLite FTS5 会话搜索
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-10-sqlite-session-query-provider.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
精确读取的 `ctx.sessionQuery` 服务有意不维护派生索引。大规模持久化的历史记录需要全文搜索,而不是每次查询都扫描全部事件;当前的活跃会话则需要一个比上一次持久性检查点更新的覆盖层。搜索还需要具体的排序、摘要片段、过滤、分页、取消以及重建行为。
|
||||
|
||||
如果把这些关注点拆分到一个推测性的 provider 协调器和一个数据库实现之间,会产生两个耦合的协调状态机。第一个真实实现应当将源观察、提取、SQLite 事务、generation 管理和查询作为一个完整的生命周期来拥有。
|
||||
|
||||
## 提案
|
||||
|
||||
在精确读取包(exact-read package)旁新增 `@deepseek-ai/dsh-session-query-sqlite`。该包将暴露一个搜索服务,或以其实际消费方所需的最小 API 扩展服务族;第一阶段不预先承诺 provider 注册协议。它将依赖 `ctx.sessions` 和可选的 `ctx.sessionPersistence`,拥有一个独立的派生 SQLite 数据库,并复用规范的 `foldSurface()` 分类。
|
||||
|
||||
实现拥有一个串行化的协调/数据库事务状态机。一次事务观察权威的持久化元数据和活跃快照,提取语义文档,更新派生表,推进相关的游标 generation,并执行或启用对应的查询。没有第二个服务维护并行的指纹、脏标记、活跃 ID 集合或失效 generation。
|
||||
|
||||
持久化文档在重启后存活。活跃覆盖层是连接本地的,对同一会话的持久化行进行遮蔽,在活跃所有者或数据库关闭时消失。派生数据库与规范持久化分离,确保索引重置、损坏、分词器变更和 schema 变动不会危及持久化的对话日志。
|
||||
|
||||
## 随实现确定的搜索语义
|
||||
|
||||
实现必须从可执行的用例出发定义跨会话和会话内两种搜索范围。每个可搜索事件是一个文档,包含会话元数据、事件元数据、surface 分类、归一化的语义文本和有界的纯文本摘要片段。会话级结果按其最强匹配事件分组;数值化的后端分数保持私有。
|
||||
|
||||
过滤器在排序之前编译为参数化 SQL。查询语法被视为数据。排序包含稳定的平局字段。不透明游标绑定到归一化的请求形状和最小相关 generation;不相关的会话变更不应使会话内游标失效。取消操作必须停止调用方等待,并在运行时允许的范围内中断 SQLite 工作。
|
||||
|
||||
分词器选择仍是实现层面的实验。FTS5 trigram 支持子串召回,但会拒绝短于三个字符的有用词项并增大索引体积;提案在将其写入契约之前,必须对该权衡与默认 Unicode 分词器进行基准测试。
|
||||
|
||||
## 提取与协调
|
||||
|
||||
该包首先为以下内容提供第一方语义提取:消息、reasoning、工具调用/结果、被阻止的提示词、上下文、steering(中途引导)、待办事项和错误/状态详情。结构性事件和流式分片不贡献文档。未知的声明合并事件/内容类型保持不可搜索,除非有真实的扩展消费方证明需要公开的提取器注册表。
|
||||
|
||||
协调可以使用稳定指纹来避免重写未变更的持久化会话,但数据库包拥有指纹的计算和存储。当源观察或提取失败时,它绝不能报告某行为最新。provider-schema 不匹配只重置派生数据库;普通的源变更使用事务性 upsert/delete。已挂载但不可读的持久化层使受影响的搜索失败,但不影响规范写入或已知的活跃精确读取。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
- **将 FTS 表添加到规范持久化数据库中**:否决,因为可重建的索引不应与权威日志共享 schema/重置/故障边界。
|
||||
- **重新引入第一阶段的 provider 协调**:否决,因为只有一个计划中的实现,且没有证据表明存在稳定的多 provider seam。
|
||||
- **立即持久化活跃覆盖层**:否决,因为活跃事件在现有检查点提交之前不是规范的。
|
||||
- **返回 BM25 分数**:否决,因为 provider 特定的数值尺度在语料变化时不稳定。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- 重启测试覆盖未变更、新增、变更和删除的持久化会话,且不重建整个索引。
|
||||
- 重新打开时保留持久化行并移除活跃行;活跃行先遮蔽、后显露其持久化基础。
|
||||
- 测试覆盖两种搜索范围、元数据过滤、surface 默认值、摘要片段、转义、确定性平局、分页、范围内的陈旧游标、取消、动态持久化挂载/卸载,以及事务失败后的恢复。
|
||||
- schema 不匹配只重置派生数据库。
|
||||
- 一个 keyless 的端到端测试将真实的持久化后端与真实的 SQLite 搜索包组合使用。
|
||||
- 在移至 `implemented/` 之前,本 RFC 须修订为实际实现的分词器和公开 API。
|
||||
|
||||
## 风险
|
||||
|
||||
单一所有者比提供方无关的 seam 更简单,但初期可复用性较低。这是有意为之:第二个真实后端可以揭示应当抽取什么。SQLite 运行时差异可能影响 FTS 排序和摘要片段,因此测试只能固定契约控制的排序和呈现。独立数据库增加了配置和生命周期工作,但保全了规范存储的安全边界。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-13-stream-workflow-progress-through-tool-calls.md: c4fe68974bf774306038e3bd2ba3e29e06de3492
|
||||
2026-07-13-stream-workflow-progress-through-tool-calls.zh.md: b6a3ccbc89bebaae92641a10aea9a9a05b38a293
|
||||
@@ -0,0 +1,43 @@
|
||||
# Agent Note: Stream workflow progress through tool calls
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-13-stream-workflow-progress-through-tool-calls.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The workflow engine intentionally emits balanced `workflow/*` observation events for run, phase, narration, and child-agent progress, but no production consumer presents them. Editors therefore show one pending workflow tool card until the final result even while the engine already reports which phase is active, what the script logged, and which children started or settled. The [dynamic-workflows decision](../../implemented/feature/2026-07-05-dynamic-workflows.md) explicitly reserves ACP progress UI for this event stream.
|
||||
|
||||
Making `dsh-acp` listen to workflow events directly would invert the capability boundary: the generic UI bridge would depend on an optional workflow package and special-case one tool name. The tool pipeline already owns the routing facts a live update needs—agent and call id—but exposes only pure pending/final presenters, so a long-running tool has no provider-neutral way to report transient UI state between them.
|
||||
|
||||
## Proposal
|
||||
|
||||
Add a live progress channel to `dsh-tools`. The registry-owned `ToolExecution` gains `reportProgress(view): boolean`, where `view` is a detached provider-neutral generic progress snapshot containing an optional replacement title and UI-facing content blocks. Progress cannot change the call's args-derived card tag, kind, raw input, locations, terminal intent, or diff intent; it updates only the live title/content within the presentation chosen up front. While the execution is active, the method validates and snapshots the view, then dispatches a contained, agent-scoped `tools/progress` observation carrying the authoritative execution identity and snapshot. Once final-result processing begins it returns `false` and emits nothing, so a late asynchronous reporter cannot overwrite a terminal card. Observer exceptions are logged and cannot fail the tool.
|
||||
|
||||
`dsh-acp` consumes `tools/progress` generically. It resolves the execution's agent through its existing agent-to-session map and emits an in-progress `tool_call_update` for the same call id. Because reporting is available only inside the tool execution pipeline, the durable `tool/call` and its ACP `tool_call` always precede the first update; closing the reporter before `tools/result` ensures no progress update follows the completed/failed card. Progress is live UI state rather than model input or durable history: session replay continues to reconstruct the pending and final cards from `tool/call` and `tool/result` without replaying transient updates.
|
||||
|
||||
`dsh-tool-workflow` becomes the first producer. Each tool execution installs a compact event capture before calling `ctx.workflows.start()`, because a valid engine may emit progress synchronously inside `start()`. Until the call returns, the capture reduces observed events into candidate states keyed by `WorkflowRunInfo.id`; it then selects the returned `WorkflowRun.id`, discards other candidates, reports the accumulated snapshot, and routes later matching events directly. If `start()` throws, the capture is disposed and its candidates are dropped. This preserves engine swappability without adding observer correlation to `WorkflowStartRequest` or requiring progress to wait until `start()` returns.
|
||||
|
||||
The reducer consumes the existing start, phase, log, agent-start, agent-end, and end events, reporting a replacement snapshot with the current phase, latest log line, active child labels, and completed/failed/cancelled counts. It does not accumulate a narration transcript; settled children leave the active set and become counters. `workflow/end`, tool settlement, or plugin disposal removes the reducer entry and event capture. The six workflow events, their metadata, paired child lifecycle, run handle, cancellation channels, and observer containment remain unchanged; third-party observers can continue consuming them directly.
|
||||
|
||||
Update the tool execution/presentation docs, generated event and API catalogs, workflow package docs, and the workflow data-structure catalog. ACP integration coverage must exercise the real workflow tool and worker seam with a scripted model boundary; the primary ACP snapshot suite adds one workflow-progress scenario because this changes the editor-facing transcript.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Delete the workflow observation surface.** Rejected in [the collapse-workflow simplification](../../rejected/simplification/2026-07-12-collapse-workflow-to-foreground-core.md): the events and their balanced lifecycle are intentional, and the missing piece is a consumer.
|
||||
|
||||
**Teach ACP about workflows directly.** This could map `WorkflowRunInfo` to a session and card, but it would make the generic bridge depend on an optional capability and bypass the rule that tools own presentation intent. A tool-progress channel solves the same routing problem for every long-running tool.
|
||||
|
||||
**Persist every progress update as a session event.** That would make live narration replayable, but it would permanently enlarge logs with state whose authoritative durable outcome is already the tool call/result pair. If resumable workflow progress becomes a product requirement, it needs a workflow-journaling design rather than UI snapshots disguised as durable facts.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `ToolExecution.reportProgress()` is registry-owned, agent-scoped, snapshotting, observer-contained, and returns `false` without dispatch after terminal processing starts.
|
||||
- ACP routes progress to the correct call in the correct live session; concurrent workflows in different sessions cannot cross-talk, and no `tool_call_update` appears before its `tool_call` or after its terminal update.
|
||||
- Workflow progress shows the current phase, latest log line, active children, and outcome counts while preserving all existing `workflow/*` events and run semantics; a seam test engine that emits start, phase, log, child, and end events synchronously inside `start()` loses none of that reducer state.
|
||||
- Cancellation, worker death, tool failure, session close, and plugin disposal release reducer state; replay emits only the durable pending/final card pair.
|
||||
- Unit, workflow integration, ACP integration, snapshot, typecheck, coverage, doc-sync, module-graph, build, and hygiene gates pass.
|
||||
|
||||
## Risks
|
||||
|
||||
This adds a public live-progress method and event to the tool seam, so implementations must keep the active/terminal boundary exact and detach snapshots before observers see them. The pre-start capture can briefly observe unrelated workflow runs, so it holds only compact candidate state keyed by run id and drops every non-matching candidate as soon as `start()` returns. A workflow can emit many progress changes; the bounded reducer avoids transcript growth but still sends one UI update per meaningful event after correlation. If measured clients need coalescing, it must be a defaulted validated bridge configuration rather than a hardcoded throttle. Transient progress intentionally disappears on replay, so the final tool result remains the only durable workflow card content.
|
||||
@@ -0,0 +1,43 @@
|
||||
# RFC: 通过工具调用流式传输工作流进度
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-13-stream-workflow-progress-through-tool-calls.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
工作流引擎有意为 run、phase、narration 和子 agent(智能体)进度发出成对的 `workflow/*` observation 事件,但目前没有生产消费方呈现这些事件。因此,编辑器在最终结果返回之前只显示一张 pending 状态的工作流工具卡片,尽管引擎已经报告了当前活跃的 phase、脚本日志内容以及哪些子 agent 已启动或已结束。[dynamic-workflows 决策](../../implemented/feature/2026-07-05-dynamic-workflows.md)明确将 ACP(Agent Client Protocol)进度 UI 保留给这一事件流。
|
||||
|
||||
如果让 `dsh-acp` 直接监听工作流事件,就会反转能力边界:通用的 UI 桥接层将依赖一个可选的工作流包(package),并对一个工具名做特殊处理。工具流水线已经拥有实时更新所需的路由信息(agent 和 call id),但只暴露了纯粹的 pending/final 展示器,因此长时间运行的工具没有提供方无关的方式在二者之间报告瞬态 UI 状态。
|
||||
|
||||
## 提案
|
||||
|
||||
为 `dsh-tools` 添加一条实时进度通道。注册表所有的 `ToolExecution` 新增 `reportProgress(view): boolean`,其中 `view` 是一个独立的、提供方无关的通用进度快照,包含可选的替换标题和面向 UI 的内容块。进度不能更改调用的 args 派生卡片标签、kind、原始输入、locations、terminal intent 或 diff intent;它只更新在最初选定的展示方式内的实时标题/内容。当执行处于活跃状态时,该方法校验并快照 view,然后分发一个受限的、agent 作用域的 `tools/progress` observation,携带权威的执行标识与快照。一旦 final-result 处理开始,方法返回 `false` 且不再分发,因此迟到的异步报告者无法覆盖终态卡片。观察者异常会被记录日志,不会导致工具失败。
|
||||
|
||||
`dsh-acp` 以通用方式消费 `tools/progress`。它通过既有的 agent-to-session 映射解析执行所属的 agent,并为同一 call id 发出 in-progress 的 `tool_call_update`。由于报告仅在工具执行流水线内可用,持久化的 `tool/call` 及其 ACP `tool_call` 始终先于第一条 update;在 `tools/result` 之前关闭报告者,确保进度更新不会出现在 completed/failed 卡片之后。进度是实时 UI 状态,而非模型输入或持久历史:会话回放继续从 `tool/call` 和 `tool/result` 重建 pending 与 final 卡片,无需重放瞬态更新。
|
||||
|
||||
`dsh-tool-workflow` 成为第一个生产者。每次工具执行在调用 `ctx.workflows.start()` 之前安装一个紧凑的事件捕获器,因为合法的引擎可能在 `start()` 内部同步发出进度。在调用返回之前,捕获器将观察到的事件按 `WorkflowRunInfo.id` 归约为候选状态;随后选取返回的 `WorkflowRun.id`,丢弃其他候选,报告累积的快照,并将后续匹配事件直接路由。如果 `start()` 抛出异常,捕获器被 dispose(资源释放),其候选状态被丢弃。这在不向 `WorkflowStartRequest` 添加观察者关联、也不要求进度等到 `start()` 返回的前提下,保持了引擎的可替换性。
|
||||
|
||||
归约器消费既有的 start、phase、log、agent-start、agent-end 和 end 事件,报告一个替换快照,包含当前 phase、最新日志行、活跃子 agent 标签以及 completed/failed/cancelled 计数。它不累积 narration transcript(文本记录);已结束的子 agent 离开活跃集合,变为计数器。`workflow/end`、工具结算或插件 dispose 移除归约器条目和事件捕获器。六种工作流事件及其元数据、成对的子 agent 生命周期、run handle、取消通道和观察者隔离保持不变;第三方观察者可继续直接消费这些事件。
|
||||
|
||||
更新工具执行/展示文档、生成的事件与 API 目录、工作流包文档以及工作流数据结构目录。ACP 集成覆盖率必须使用脚本化的模型边界测试真实的工作流工具和 worker seam;主 ACP 快照套件新增一个 workflow-progress 场景,因为这改变了面向编辑器的 transcript。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
**删除工作流 observation 表面。** 在 [collapse-workflow 简化提案](../../rejected/simplification/2026-07-12-collapse-workflow-to-foreground-core.md)中被否决:这些事件及其成对生命周期是有意设计的,缺少的是消费方。
|
||||
|
||||
**让 ACP 直接了解工作流。** 这可以将 `WorkflowRunInfo` 映射到会话和卡片,但会使通用桥接层依赖一个可选能力,并绕过「工具拥有展示意图」的规则。工具进度通道为每个长时间运行的工具解决了相同的路由问题。
|
||||
|
||||
**将每条进度更新持久化为会话事件。** 这会使实时 narration 可回放,但会用一种状态永久膨胀日志,而该状态的权威持久结果已经是工具调用/结果对。如果可恢复的工作流进度成为产品需求,需要一个工作流日志化设计,而非伪装成持久事实的 UI 快照。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- `ToolExecution.reportProgress()` 由注册表所有、agent 作用域、快照化、观察者隔离,且在终态处理开始后返回 `false` 而不分发。
|
||||
- ACP 将进度路由到正确的实时会话中的正确调用;不同会话中的并发工作流不能串扰,且 `tool_call_update` 不会出现在其 `tool_call` 之前或终态更新之后。
|
||||
- 工作流进度显示当前 phase、最新日志行、活跃子 agent 和结果计数,同时保留所有既有 `workflow/*` 事件和 run 语义;一个在 `start()` 内部同步发出 start、phase、log、child 和 end 事件的 seam 测试引擎不会丢失任何归约器状态。
|
||||
- 取消、worker 死亡、工具失败、会话关闭和插件 dispose 释放归约器状态;回放仅发出持久的 pending/final 卡片对。
|
||||
- 单元测试、工作流集成测试、ACP 集成测试、快照、类型检查、覆盖率、doc-sync、module-graph、构建和 hygiene 门禁全部通过。
|
||||
|
||||
## 风险
|
||||
|
||||
本提案向工具 seam 添加了一个公开的实时进度方法和事件,因此实现方必须精确维护 active/terminal 边界,并在观察者看到快照之前将其分离。pre-start 捕获器可能短暂观察到无关的工作流 run,因此它仅按 run id 持有紧凑的候选状态,并在 `start()` 返回后立即丢弃所有不匹配的候选。一个工作流可能发出大量进度变更;有界归约器避免了 transcript 增长,但在关联完成后仍会为每个有意义的事件发送一条 UI 更新。如果经测量的客户端需要合并更新,这必须是一个带默认值的、经过校验的桥接配置,而非硬编码的节流。瞬态进度在回放时有意消失,因此最终工具结果仍是唯一持久的工作流卡片内容。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-14-sdk-developer-projects.md: aa5cf64d7dd33dea229d74c2ae45a9244ee70e3c
|
||||
2026-07-14-sdk-developer-projects.zh.md: 8f7d1de5b16f38019c802f07eda701cee72deb4f
|
||||
@@ -0,0 +1,167 @@
|
||||
# Agent Note: Developer-owned SDK projects
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-14-sdk-developer-projects.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
DeepSeek Harness composes features through Cordis plugins, but building a runnable project from an empty directory still requires a developer to understand npm dependencies, the `cordis.yml` plugin set, environment variables, TypeScript builds, local-plugin workspaces, and runtime entrypoints together. These manual steps constrain one another: omitting any one can produce a project that installs but cannot be developed, develops but cannot be built, or builds but cannot start.
|
||||
|
||||
A one-shot generator reduces only the initial creation cost. If the generated result is hidden inside a preset or an uneditable CLI, advanced developers cannot reshape the plugin tree, change Cordis plugin config, or add project-specific behavior. If a generated project immediately leaves tool management altogether, developers must again maintain consistency across all npm dependencies and Cordis plugin config themselves.
|
||||
|
||||
Initial creation and later configuration address the same builtin feature set. When those workflows maintain separate feature lists, feature options, and npm dependencies, new Cordis plugins, npm packages, and Cordis plugin config changes make them diverge. Projects also need an ordinary local-plugin development path that participates in development, build, and start flows.
|
||||
|
||||
## Proposal
|
||||
|
||||
The SDK creates an ordinary, explicit TypeScript/Cordis project owned by its developer. `cordis.yml` is the only runtime plugin tree; development and production read the same file. The generated `package.json`, `cordis.yml`, TypeScript entrypoint, build configuration, and `plugins/*` remain directly editable instead of being hidden behind a preset.
|
||||
|
||||
The only developer product entrypoints are `npm create @deepseek-ai/sdk` and the `dsh-sdk` commands. The initializer performs initial creation, `dsh-sdk config` manages SDK-recognized builtin features afterward, and `dsh-sdk dev`, `dsh-sdk build`, and `dsh-sdk start` own development, build, and startup; this phase provides no `dsh-sdk create`. Create and config consume one manually authored feature definition, so each feature has one source for its feature options, npm dependencies, Cordis config entries, related files, and inspection rules. The [SDK project editing architecture](../architecture/2026-07-15-sdk-project-editing-architecture.md) defines terms such as feature and feature option.
|
||||
|
||||
The SDK offers interaction for feature selection and finite feature options only; it does not turn arbitrary Cordis plugin config into a generic form. A feature collects the small number of dedicated inputs required by its feature options. All other Cordis plugin config remains in `cordis.yml`, with comments documenting common edits, for direct developer control.
|
||||
|
||||
## Developer workflow
|
||||
|
||||
Initial creation collects information in an order where earlier answers determine later questions: target directory and package identity, model provider and credentials, run interface, builtin features and feature options, an optional local plugin, package manager, and whether to install npm dependencies and build. Command-line arguments suppress questions they already answer. Create and config require an interactive TTY in this phase, and cancelling creation writes nothing to the target directory.
|
||||
|
||||
```sh
|
||||
npm create @deepseek-ai/sdk my-agent
|
||||
cd my-agent
|
||||
npm exec dsh-sdk dev index.ts
|
||||
npm exec dsh-sdk config
|
||||
npm exec dsh-sdk build
|
||||
npm exec dsh-sdk start index.js
|
||||
```
|
||||
|
||||
Create rejects every target path that already exists. After committing the project files, the CLI asks whether to install npm dependencies and build. An install or build failure preserves the generated project and prints commands that can retry the failed work.
|
||||
|
||||
Create also offers one `none / plugin / tool` choice. `plugin` creates a fixed `plugins/plugin` Cordis plugin, while `tool` creates a fixed `plugins/tool` model-facing tool; one project creation includes at most one local plugin. The operation updates the workspace, root npm dependency, TypeScript reference, build configuration, and `cordis.yml` together, and any pre-write validation failure leaves the project absent.
|
||||
|
||||
## Features supported during creation
|
||||
|
||||
The table is the developer-visible support set for this phase. A `required` feature is always present but may still offer finite feature options; a `default` feature is preselected in the feature tree; an `optional` feature is selected explicitly. The table describes the product support set, while the runtime registry remains the implementation source of truth.
|
||||
|
||||
| Feature | Create state | Feature options | Constraints and relationships |
|
||||
|---|---|---|---|
|
||||
| `provider` | required | `deepseek` (default) / `custom` | DeepSeek collects an API key; custom also collects a base URL, and a CLI option may override the model name |
|
||||
| `app` | required | `tui` (default) / `acp` / `embed` | Selects the run interface |
|
||||
| `spine` | required | `default` | Timer, the LLM seam, session storage, system prompt, the tool registry, the agent registry, and the agent loop |
|
||||
| `bash` | required | `local` (default) / `sandbox` | The two feature options are exclusive and independent of the run interface, and both install the model-facing bash tool; sandbox installs the local sandbox provider and sandboxed bash backend |
|
||||
| `persistence` | required | `jsonl` (default) / `sqlite` | Every project selects exactly one persistence backend |
|
||||
| `hmr` | default | `default` | Loads `@cordisjs/plugin-hmr`; dev and start both enable it with the plugin defaults |
|
||||
| `fs` | default | `local` | Installs the local filesystem, policy, and model-facing tools; the process sandbox does not confine in-process fs tools |
|
||||
| `todo` | default | `default` | Provides the `todo_write` tool |
|
||||
| `skill` | default | `default` | Installs the skill registry, the local skill provider, and the model-facing skill tool |
|
||||
| `web` | optional | `deepseek` (default) / `exa` / `perplexity` / `fetch-only` | Search feature options are exclusive; Exa and Perplexity collect their API keys; timeout policy is recommended |
|
||||
| `subagent` | optional | `spawn` (default) / `fork`, multiple | This phase provides only in-process backends |
|
||||
| `workflow` | optional | `workerthread` | Requires the subagent `spawn` feature option |
|
||||
| `compact` | optional | `basic` | Uses SDK-provided context-compaction parameters |
|
||||
| `hooks` | optional | `claude` (default) / `codex`, multiple | Each feature option creates a separate editable configuration file |
|
||||
| `guard` | optional | `repeat-tool` | Provides repeated-tool-call reminders |
|
||||
| `timeout-policy` | optional | `default` | Applies a uniform policy to tools that declare timeout budgets |
|
||||
| `ask-user` | optional | `default` | Provides the `ask_user_question` tool; only `acp` and `tui` can select it because those two feature options provide the injected user-interaction service |
|
||||
|
||||
Both `bash` feature options apply to ACP, TUI, and embed and are not selected by the run interface. The sandbox feature option writes no active config key and therefore keeps `dsh-bash-sandbox`'s `read-only` default. Generated `cordis.yml` includes a commented example that developers can change explicitly to `workspace-write`:
|
||||
|
||||
```yaml
|
||||
- id: bash
|
||||
name: '@deepseek-ai/dsh-bash-sandbox'
|
||||
# Uncomment to allow writes under the project workspace.
|
||||
# config:
|
||||
# mode: workspace-write
|
||||
# workspaceRoot: !!js process.cwd()
|
||||
```
|
||||
|
||||
Feature contributions reference only single-plugin npm packages and never bundle packages such as `agent-spine-demo`, `tui-demo`, or `acp-demo`. Plugins outside the table are not managed by create in this phase; advanced developers may still compose them by editing the ordinary project files directly.
|
||||
|
||||
## Generated project
|
||||
|
||||
With default answers, an npm project uses the DeepSeek provider, the TUI interface, local bash, JSONL persistence, and the preselected hmr, fs, todo, and skill features. Its initial tree is:
|
||||
|
||||
```text
|
||||
my-agent/
|
||||
├── .env
|
||||
├── .env.example
|
||||
├── .gitignore
|
||||
├── README.md
|
||||
├── cordis.yml
|
||||
├── index.ts
|
||||
├── package.json
|
||||
├── tsconfig.base.json
|
||||
├── tsconfig.json
|
||||
└── tsdown.config.ts
|
||||
```
|
||||
|
||||
`.env.example` always exists, and the SDK keeps its placeholders aligned with the current feature set. A gitignored `.env` is also created when a secret is captured or the developer confirms an empty credential to fill later. The SDK only appends differently named variables that are not already present in `.env` and never updates or removes existing contents. Feature-option changes may remove obsolete `.env.example` placeholders, while old credentials remain in `.env` for the developer to manage. pnpm and Yarn projects add their required workspace files, but do not fork the runtime plugin tree or TypeScript entrypoint.
|
||||
|
||||
Generated `package.json` provides the following scripts. `dev`, `build`, `start`, and `config` invoke `dsh-sdk`, while `typecheck` invokes TypeScript directly:
|
||||
|
||||
| Script | Behavior |
|
||||
|---|---|
|
||||
| `dev` | Run `dsh-sdk dev index.ts`, registering development-time resolution for TypeScript and local workspace plugins |
|
||||
| `build` | Run `dsh-sdk build`, invoking the project's installed tsdown for the root entrypoint and `plugins/*` packages |
|
||||
| `typecheck` | Run `tsc -b` directly |
|
||||
| `start` | Run `dsh-sdk start index.js`, starting the built entrypoint without an implicit build |
|
||||
| `config` | Run `dsh-sdk config` to edit the current project's feature tree |
|
||||
|
||||
`dsh-sdk start` and `dsh-sdk dev` accept a module target and forward arguments after `--` unchanged to the project entrypoint. Generic argument parsing uses Node `parseArgs()` with zero schema: valued flags use `--key=value`, bare flags become `true`, and `--no-*` becomes `false`.
|
||||
|
||||
- TUI projects pass the selected model through `--model=<name>` and create or resume an agent according to optional `--resume=<session-id>`;
|
||||
- ACP uses protocol `session/load`
|
||||
- Embed uses the model written into the generated code.
|
||||
|
||||
Each feature-owned Cordis config entry keeps its developer-editable Cordis plugin config and explanatory comments in `cordis.yml`. When `dsh-sdk config` changes other features, it preserves unknown fields, formatting on untouched nodes, and comments. HMR is an ordinary leaf config entry: when the feature is selected, dev and start load the same watcher, and the command does not change the plugin tree implicitly.
|
||||
|
||||
## Post-creation configuration
|
||||
|
||||
`dsh-sdk config` requires only readable root `package.json` and `cordis.yml` files in the current directory. It inspects standard features and their current feature options, expresses the final desired state through one feature tree, and shows feature changes and affected files before Review & Apply.
|
||||
|
||||
`dsh-sdk config` can install missing features, enable or disable installed features, and switch finite feature options. Required features cannot be removed. An npm dependency change runs the project package manager's install once after the file commit; installation failure does not roll back committed project files.
|
||||
|
||||
The SDK modifies only Cordis config entries, config keys, npm dependencies, `.env.example` placeholders, and owned files explicitly owned by a feature. Updating the same feature option preserves unknown config keys in its Cordis config entries. Handwritten and third-party plugins support enable and disable by stable ID only. When a known feature has been edited into an incomplete, ambiguous, or otherwise unreadable shape, `dsh-sdk config` displays diagnostics and refuses automatic changes until the developer repairs it manually.
|
||||
|
||||
One config session accumulates every change in an in-memory working copy. Before Apply, it validates feature relationships, resource conflicts, and document shapes, then compares each affected existing file with the text read when the session opened. Validation failure or an external edit causes zero writes. Once physical writes begin, the SDK does not provide cross-file transactional rollback.
|
||||
|
||||
## Maintenance model
|
||||
|
||||
The SDK curates its builtin support set instead of exposing npm packages automatically by npm dependency name or directory convention. One feature may compose several Cordis config entries, feature options may share resources, and a feature option may declare a feature requirement on another feature or a specific feature option. Adding an ordinary feature or feature option does not require changes to both create and config command workflows.
|
||||
|
||||
## Future work
|
||||
|
||||
- `dsh-sdk add [package-spec]` unifies local-plugin creation with external Cordis plugin installation: without a package or repository source it creates a local plugin/tool, while a supplied source adds the npm dependency and `cordis.yml` config entry; the source model leaves room for GitHub repositories and other extensions
|
||||
- Non-interactive create/config: both workflows require a TTY in this phase and provide no complete input contract for automation
|
||||
- More feature-specific inputs: this product surface exposes only finite feature options, secrets, and a few dedicated values in this phase rather than a generic parameter interface for Cordis plugin config
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**An opaque preset or generator-owned project.** This shortens initial creation but hides the real plugin tree and build boundaries, prevents advanced developers from composing Cordis plugins directly, and makes project behavior depend on the CLI version rather than committed project files.
|
||||
|
||||
**A one-shot generator only.** Leaving all later maintenance manual redistributes feature requirements, feature-option switches, and multi-file updates. A config workflow over the shared registry retains continuing management for generated projects.
|
||||
|
||||
**Separate `cordis.yml` files for development and production.** Two plugin trees mean a successful development run does not demonstrate that production loads the same features. Dev adds only TypeScript and local-workspace resolution; runtime configuration remains singular.
|
||||
|
||||
**A generic form for arbitrary Cordis plugin config.** Cordis plugin config contains nested structures, expressions, and plugin-specific semantics. A generic form would become a second incomplete schema. The SDK manages finite feature options and dedicated secrets, while developers continue to edit complex config directly.
|
||||
|
||||
**A private local-plugin discovery protocol.** Ordinary package-manager workspaces, root npm dependencies, TypeScript references, and Cordis config entries already express the complete relationship. Another discovery protocol would create hidden state understood only by the SDK.
|
||||
|
||||
**A `dsh-sdk create` command for existing projects.** Create already provides one editable local-plugin skeleton, and later plugins can use ordinary workspace and Cordis mechanisms manually. A parallel command would add a second scaffolding product surface without adding composition functionality.
|
||||
|
||||
**Automatically expose every new Cordis plugin as a builtin.** An npm package cannot say how several plugins compose into one product feature, nor can it derive exclusivity, feature requirements, secrets, interface applicability, or security constraints. The support set requires human curation; automation is suitable only for checking whether candidates have been classified.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `npm create @deepseek-ai/sdk` collects project identity, provider, interface, features, an optional local plugin, package manager, and installation choice in the documented order, and cancellation leaves the target path absent
|
||||
- A default npm project has the documented tree and `dev`, `build`, `typecheck`, `start`, and `config` scripts, with dev and start sharing one `cordis.yml`
|
||||
- Create offers the documented features and feature options; local and sandbox bash are exclusive with local as the default, the sandbox Cordis config entry retains the editable commented config example, and HMR is selected by default and loaded by both dev and start
|
||||
- Create's `plugin` or `tool` choice creates at most one fixed-name local plugin and atomically updates its files and root-project relationships; this phase provides no `dsh-sdk create`
|
||||
- `dsh-sdk config` reads the same support set from an existing project, installs, enables, disables, and switches supported feature options, preserves unknown config and comments, and refuses to modify inconsistent config
|
||||
- `.env.example` reflects variables required by the current features; `.env` only appends missing differently named variables and never updates or removes existing contents
|
||||
- npm, pnpm, and Yarn workspaces install, build, and start; local plugins resolve from source under dev and from built output under start
|
||||
|
||||
## Risks
|
||||
|
||||
- Developers can edit a builtin into a shape the registry cannot recognize; the SDK stops automating that feature instead of guessing and overwriting config
|
||||
- Pre-write validation and external-edit detection do not provide transactional rollback once multi-file writes begin; an I/O failure can leave a partial commit requiring manual repair
|
||||
- The sandbox feature option depends on an available local sandbox backend for the target platform; an unavailable backend must fail closed instead of falling back to unsandboxed execution
|
||||
- HMR retains its filesystem watcher and hot-reload behavior under production start; this is the result of an explicit plugin choice, not an implicit development-only service
|
||||
- The append-only `.env` policy retains credentials that are no longer used; the SDK does not decide when user-owned secret data is safe to delete
|
||||
@@ -0,0 +1,167 @@
|
||||
# Agent Note: 开发者拥有的 SDK 工程
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-14-sdk-developer-projects.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
DeepSeek Harness 通过 Cordis 插件对功能进行组合,但从空目录开始搭建一个可运行工程仍要求开发者同时理解 NPM 依赖、`cordis.yml` 插件组、环境变量、TypeScript 构建、本地插件 workspace 和运行入口。手工步骤之间存在约束,漏掉任意一处都会得到能够安装却无法开发、能够开发却无法构建,或能够构建却无法启动的工程。
|
||||
|
||||
一次性生成器只能降低首次创建成本。若生成结果隐藏在 preset 或不可编辑的 CLI(命令行界面)内部,高级开发者无法调整插件树、修改 Cordis 插件配置或增加项目特有行为;若创建后的工程完全脱离工具管理,开发者又必须重新承担所有 NPM 依赖和 Cordis 插件配置的一致性工作。
|
||||
|
||||
初始创建和后续配置面对同一组内置功能。两条流程各自维护功能列表、功能选项和 NPM 依赖时,新增 Cordis 插件、NPM 包或调整配置会使二者逐渐分叉。工程还需要一条普通的本地插件开发路径,参与开发、构建和启动流程。
|
||||
|
||||
## 提案
|
||||
|
||||
SDK 创建一个普通、显式且归开发者所有的 TypeScript/Cordis 工程。`cordis.yml` 是唯一的运行时插件树;开发和生产读取同一份文件。工程中的 `package.json`、`cordis.yml`、TypeScript 入口、构建配置和 `plugins/*` 均可直接编辑,SDK 不把它们封装成不可见的 preset。
|
||||
|
||||
开发者产品入口只有 `npm create @deepseek-ai/sdk` 和 `dsh-sdk` 命令。前者负责首次创建,`dsh-sdk config` 在创建后管理 SDK 能识别的内置功能,`dsh-sdk dev`、`dsh-sdk build` 与 `dsh-sdk start` 负责开发、构建和启动;本期不提供 `dsh-sdk create`。create 与 config 使用同一份人工编写的功能定义,因此一项功能的功能选项、NPM 依赖、Cordis 配置项、相关文件和识别规则只有一个来源。功能、功能选项等名词由 [SDK 工程编辑架构](../architecture/2026-07-15-sdk-project-editing-architecture.md) 的术语表定义。
|
||||
|
||||
SDK 只为功能选择和有限功能选项提供交互,不尝试把任意 Cordis 插件配置变成通用表单。功能选项所需的少量专用输入由所属功能收集;其余 Cordis 插件配置留在 `cordis.yml` 中,并通过注释指明常用改法,由开发者直接修改。
|
||||
|
||||
## 开发者流程
|
||||
|
||||
首次创建按会影响后续问题集合的顺序收集信息:目标目录与 package 身份、模型提供方与凭据、运行接口、内置功能与功能选项、可选本地插件、包管理器,以及是否安装 NPM 依赖并构建。命令参数已提供的答案不重复询问;本期 create 和 config 都要求交互式 TTY,取消创建时不写入目标目录。
|
||||
|
||||
```sh
|
||||
npm create @deepseek-ai/sdk my-agent
|
||||
cd my-agent
|
||||
npm exec dsh-sdk dev index.ts
|
||||
npm exec dsh-sdk config
|
||||
npm exec dsh-sdk build
|
||||
npm exec dsh-sdk start index.js
|
||||
```
|
||||
|
||||
create 拒绝任何已经存在的目标路径。工程文件提交成功后,CLI 询问是否安装 NPM 依赖并构建;安装或构建失败时保留生成结果,并打印可以重新执行的命令。
|
||||
|
||||
create 还提供一次 `none / plugin / tool` 选择。`plugin` 固定生成 `plugins/plugin` 的 Cordis 插件,`tool` 固定生成 `plugins/tool` 的模型工具;一次创建至多包含一个本地插件。生成操作同时更新 workspace、根 NPM 依赖、TypeScript reference、构建配置和 `cordis.yml`,任何写入前校验失败都不创建工程。
|
||||
|
||||
## 创建时支持的功能
|
||||
|
||||
下表是本期 create 面向开发者展示的支持集。`required` 始终存在但仍可切换有限功能选项;`default` 在选择树中预选;`optional` 由开发者主动选择。表格说明产品支持集,运行时注册表是实现的事实源。
|
||||
|
||||
| 功能 | create 状态 | 功能选项 | 限制与关系 |
|
||||
|---|---|---|---|
|
||||
| `provider` | required | `deepseek`(默认)/ `custom` | DeepSeek 收集 API key;custom 另收集 base URL,模型名可由 CLI 参数覆盖 |
|
||||
| `app` | required | `tui`(默认)/ `acp` / `embed` | 选择运行接口 |
|
||||
| `spine` | required | `default` | timer、LLM seam、会话存储、系统提示词、工具注册表、agent 注册表,以及 agent loop |
|
||||
| `bash` | required | `local`(默认)/ `sandbox` | 两个功能选项互斥、与运行接口正交,且都安装面向模型的 bash 工具;sandbox 安装本地沙箱提供方和沙箱 bash 后端 |
|
||||
| `persistence` | required | `jsonl`(默认)/ `sqlite` | 每个工程恰好选择一个持久化后端 |
|
||||
| `hmr` | default | `default` | 加载 `@cordisjs/plugin-hmr`;dev 和 start 都启用,使用插件默认配置 |
|
||||
| `fs` | default | `local` | 安装本地文件系统、策略和模型工具;进程沙箱不约束进程内 fs 工具 |
|
||||
| `todo` | default | `default` | 提供 `todo_write` 工具 |
|
||||
| `skill` | default | `default` | 安装 skill(技能)注册表、本地 skill 提供方和面向模型的 skill 工具 |
|
||||
| `web` | optional | `deepseek`(默认)/ `exa` / `perplexity` / `fetch-only` | 搜索功能选项互斥;Exa/Perplexity 收集各自 API key;建议同时启用 timeout policy |
|
||||
| `subagent` | optional | `spawn`(默认)/ `fork`,可多选 | 本期只提供进程内后端 |
|
||||
| `workflow` | optional | `workerthread` | 要求 subagent 的 `spawn` 功能选项 |
|
||||
| `compact` | optional | `basic` | 使用 SDK 提供的上下文压缩参数 |
|
||||
| `hooks` | optional | `claude`(默认)/ `codex`,可多选 | 各功能选项生成独立的可编辑配置文件 |
|
||||
| `guard` | optional | `repeat-tool` | 提供重复工具调用提醒 |
|
||||
| `timeout-policy` | optional | `default` | 对声明超时预算的工具执行统一策略 |
|
||||
| `ask-user` | optional | `default` | 提供 `ask_user_question` 工具;注入的 user-interaction 服务由 acp/tui 两个功能选项提供,因此仅这两个接口可选 |
|
||||
|
||||
`bash` 的两个功能选项都适用于 ACP、TUI 和 embed,不由运行接口决定。sandbox 功能选项不写任何生效的配置键,因而沿用 `dsh-bash-sandbox` 的 `read-only` 默认值;生成的 `cordis.yml` 保留注释示例,开发者可以显式改为 `workspace-write`:
|
||||
|
||||
```yaml
|
||||
- id: bash
|
||||
name: '@deepseek-ai/dsh-bash-sandbox'
|
||||
# Uncomment to allow writes under the project workspace.
|
||||
# config:
|
||||
# mode: workspace-write
|
||||
# workspaceRoot: !!js process.cwd()
|
||||
```
|
||||
|
||||
功能贡献只引用单插件 NPM 包,绝不引用 `agent-spine-demo`、`tui-demo`、`acp-demo` 这类组合 NPM 包。表格之外的插件不由本期 create 管理;开发者仍可直接编辑普通工程文件进行高级组合。
|
||||
|
||||
## 生成工程
|
||||
|
||||
使用默认答案创建 npm 工程时,provider 为 DeepSeek,运行接口为 TUI,bash 为 local,持久化为 JSONL,hmr、fs、todo 与 skill 处于选中状态。初始目录树为:
|
||||
|
||||
```text
|
||||
my-agent/
|
||||
├── .env
|
||||
├── .env.example
|
||||
├── .gitignore
|
||||
├── README.md
|
||||
├── cordis.yml
|
||||
├── index.ts
|
||||
├── package.json
|
||||
├── tsconfig.base.json
|
||||
├── tsconfig.json
|
||||
└── tsdown.config.ts
|
||||
```
|
||||
|
||||
`.env.example` 始终存在,并由 SDK 根据当前功能维护占位。收集到 secret 或开发者确认稍后填写空凭据时,同时生成 gitignored `.env`。SDK 只向 `.env` 追加尚不存在的不同名变量,绝不覆盖或删除已有内容;切换功能选项可以清理 `.env.example` 中不再需要的占位,但旧凭据仍留在 `.env` 中供开发者自行处理。pnpm 和 Yarn 工程增加各自所需的 workspace 配置文件,但运行时插件树和 TypeScript 入口不分叉。
|
||||
|
||||
生成的 `package.json` 提供以下 scripts;其中 `dev`、`build`、`start` 与 `config` 调用 `dsh-sdk`,`typecheck` 直接调用 TypeScript:
|
||||
|
||||
| script | 行为 |
|
||||
|---|---|
|
||||
| `dev` | 运行 `dsh-sdk dev index.ts`,为 TypeScript 和本地 workspace 插件注册开发期解析 |
|
||||
| `build` | 运行 `dsh-sdk build`,调用工程安装的 tsdown 构建根入口和 `plugins/*` package |
|
||||
| `typecheck` | 直接运行 `tsc -b` |
|
||||
| `start` | 运行 `dsh-sdk start index.js`,启动已构建入口且不隐式构建 |
|
||||
| `config` | 运行 `dsh-sdk config`,修改当前工程功能树 |
|
||||
|
||||
`dsh-sdk start` 与 `dsh-sdk dev` 可以接收模块 target,并把 `--` 后的参数原样转发给工程入口。通用参数解析使用 Node `parseArgs()` 的零 schema 模式:带值 flag 采用 `--key=value`,bare flag 转换为 `true`,`--no-*` 转换为 `false`。
|
||||
|
||||
- TUI 工程通过 `--model=<name>` 传入所选 model,并根据可选的 `--resume=<session-id>` 创建或恢复 agent;
|
||||
- acp 使用协议 `session/load`
|
||||
- embed 使用生成代码中的 model。
|
||||
|
||||
每个功能拥有的 Cordis 配置项在 `cordis.yml` 中保留自己的可编辑 Cordis 插件配置和说明注释;`dsh-sdk config` 修改其他功能时必须保留未知字段、未修改节点的格式和注释。HMR(热模块替换)是普通叶子配置项:选择该功能后,dev 和 start 加载同一个 watcher,命令不隐式改变插件树。
|
||||
|
||||
## 创建后的配置
|
||||
|
||||
`dsh-sdk config` 只要求当前目录具有可读的根 `package.json` 与 `cordis.yml`。它检查标准功能及其当前功能选项,以一棵功能树表达最终目标状态,并在 Review & Apply 前展示功能变化和受影响文件。
|
||||
|
||||
`dsh-sdk config` 可以安装缺失功能、启停已安装功能和切换有限功能选项。required 功能不能取消。改变 NPM 依赖后只运行一次项目包管理器安装;安装失败不回滚已经提交的工程文件。
|
||||
|
||||
SDK 只修改功能明确拥有的 Cordis 配置项、配置键、NPM 依赖、`.env.example` 占位和独占文件。同一功能选项的更新保留 Cordis 配置项中的未知配置键;手写或第三方插件只支持按稳定 ID 启停。已知功能被手改成不完整、歧义或无法读取的形状时,`dsh-sdk config` 显示诊断并拒绝自动修改,直到开发者手工修复。
|
||||
|
||||
一次 config 会话在内存工作区上累计全部修改。Apply 前完成功能关系、资源冲突和文件形状校验,并比较受影响文件与会话打开时的原文;校验失败或检测到外部修改时不写盘。实际写盘开始后不提供跨文件事务回滚。
|
||||
|
||||
## 维护模型
|
||||
|
||||
Builtin 支持集由 SDK 人工策划,不根据 NPM 依赖名称或目录约定自动暴露。一个功能可以组合多个 Cordis 配置项,功能选项可以共享资源,并声明对其他功能或特定功能选项的功能依赖;新增普通功能或功能选项不应要求同时修改 create 和 config 两个命令流程。
|
||||
|
||||
## 后续工作
|
||||
|
||||
- `dsh-sdk add [package-spec]`:统一本地插件创建与外部 Cordis 插件接入;未指定 package 或仓库来源时创建本地 plugin/tool,指定来源时增加 NPM 依赖和 `cordis.yml` 配置项,来源模型为 GitHub 仓库等扩展保留空间
|
||||
- 非交互 create/config:本期两个流程都要求 TTY,不提供供自动化调用的完整输入合同
|
||||
- 更多功能专用参数输入:本期产品只展示有限功能选项、secret 和少量专用值,不为 Cordis 插件配置提供通用参数界面
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
**不可编辑的 preset 或生成器托管工程。** 该方案可以缩短初次创建路径,但会隐藏真实插件树和构建边界,使高级开发者无法直接组合 Cordis 插件,也让项目行为依赖 CLI 版本而不是检入的工程文件。
|
||||
|
||||
**只提供一次性生成器。** 创建后完全依赖手工维护,会让功能依赖、功能选项切换和多文件更新再次分散;共享 registry 的 config 流程为生成工程保留持续管理机制。
|
||||
|
||||
**为开发和生产维护两份 `cordis.yml`。** 两份插件树会使开发成功无法证明生产加载相同功能;dev 只增加 TypeScript 与本地 workspace 解析,运行配置保持唯一。
|
||||
|
||||
**为任意 Cordis 插件配置生成通用表单。** Cordis 插件配置包含嵌套结构、表达式和插件特有语义,通用表单会形成第二套不完整 schema。SDK 只管理有限功能选项和专用 secret,复杂配置继续由开发者直接编辑。
|
||||
|
||||
**使用私有协议发现本地插件。** 普通 package manager workspace、根 NPM 依赖、TypeScript references 和 Cordis 配置项已能表达完整关系;额外发现协议会创造只能由 SDK 理解的隐藏状态。
|
||||
|
||||
**在现有工程中提供 `dsh-sdk create`。** create 已能生成一种可编辑的本地插件骨架,后续插件可以沿用普通 workspace 和 Cordis 机制手工添加;再提供同构命令会增加第二条脚手架产品面,却不增加新的组合功能。
|
||||
|
||||
**把每个新 Cordis 插件自动暴露为 builtin。** package 无法说明多个插件如何组合成一项产品功能,也无法推导互斥关系、功能依赖、secret、接口适用性和安全限制;支持集需要人工策划,自动化只适合检查候选是否完成分类。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- `npm create @deepseek-ai/sdk` 按本文顺序收集项目身份、provider、interface、功能、可选本地插件、包管理器和安装选择,并在取消时保持目标路径不存在
|
||||
- 默认 npm 工程具有本文目录树和 `dev`、`build`、`typecheck`、`start`、`config` scripts,且 dev/start 使用同一份 `cordis.yml`
|
||||
- create 展示本文功能及功能选项;`bash` 的 local/sandbox 二选一且默认 local,sandbox Cordis 配置项保留可编辑的注释配置示例;HMR 默认选中并同时由 dev/start 加载
|
||||
- create 的 `plugin` 或 `tool` 选择至多生成一个固定名称的本地插件,并原子更新插件文件与根工程关系;本期不提供 `dsh-sdk create`
|
||||
- `dsh-sdk config` 从现有工程读取同一支持集,能够安装、启停和切换支持的功能选项,保留未知配置与注释,并拒绝修改不一致配置
|
||||
- `.env.example` 反映当前功能所需变量;`.env` 只追加缺失的不同名变量,从不覆盖或清理已有内容
|
||||
- npm、pnpm 和 Yarn 生成的 workspace 能安装、构建和启动;本地插件在 dev 中使用源码,在 start 中使用构建产物
|
||||
|
||||
## 风险
|
||||
|
||||
- 开发者可以把 builtin 手改成 registry 无法识别的形状;SDK 选择停止自动化而不是猜测并覆盖配置
|
||||
- 多文件写入前的校验和外部修改检测不能提供写入阶段的事务回滚;I/O 中途失败可能留下需要人工修复的部分提交
|
||||
- sandbox 功能选项依赖目标平台存在可用的本地沙箱后端;后端不可用时必须 fail closed,不能退回无沙箱执行
|
||||
- HMR 在生产启动中也保持文件 watcher 和热重载行为;这是显式插件选择的结果,不是仅限开发环境的隐式服务
|
||||
- `.env` 的仅追加策略会保留已经不用的凭据,SDK 不判断这些用户数据何时可以安全删除
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-16-persistent-pty-sessions.md: 3028d3a527e177f2b30c557dc45443af99783d6c
|
||||
2026-07-16-persistent-pty-sessions.zh.md: f244992abccc9c107bc2cf4da392dc39d39d14cc
|
||||
@@ -0,0 +1,170 @@
|
||||
# Agent Note: persistent PTY sessions
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-16-persistent-pty-sessions.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The harness can run foreground and background commands, edit files, and delegate work, but it cannot continue an interactive terminal conversation across tool calls. Each `bash` foreground run starts a fresh shell, so shell-local cwd, exported variables, virtual-environment activation, functions, job-control state, and interactive child processes end with that call.
|
||||
|
||||
That gap excludes workflows whose state lives in a terminal rather than a file: stepping through `gdb`, exploring in a Python or Node REPL, driving a line-oriented editor such as `ed`, or returning to a shell after interrupting its foreground command. The generic [`ctx.tasks`](../../../../packages/tasks/README.md) runtime retains background-operation handles and output, but it does not provide interactive stdin or terminal semantics.
|
||||
|
||||
The existing `bash`, `read`, `write`, and `edit` tools remain the reliable default for bounded, auditable operations. A PTY is an additional capability for work that genuinely requires terminal state, not evidence that those tools are defective or candidates for removal.
|
||||
|
||||
## Proposal
|
||||
|
||||
Add an optional `packages/pty/` capability family that exposes agent-owned, persistent, line-oriented PTY sessions. It follows the repository's [capability pattern](../../implemented/architecture/2026-06-13-capability-seams.md), coexists with the existing command and filesystem tools, and does not change `agent-loop`.
|
||||
|
||||
The first delivery supports interactive shells and line-oriented REPLs on Linux and macOS. Full-screen terminal applications, keystroke sequences, BEL-triggered control flow, session restoration after process loss, and cross-agent session sharing are explicitly deferred until the basic lifecycle is proven.
|
||||
|
||||
### Package topology
|
||||
|
||||
| Package | Role | ctx key |
|
||||
|---|---|---|
|
||||
| `dsh-pty` | `PtyService`, branded `PtySessionId`, backend registry, owner-scoped session contract, and result types | `ctx.pty` |
|
||||
| `dsh-pty-local` | [`node-pty`](https://github.com/microsoft/node-pty)-based local backend, platform process inspection, bounded terminal buffer, sandbox resolution, and process-tree supervision | registers a backend on `ctx.pty` |
|
||||
| `dsh-tool-pty` | Six model-facing tools, task-runtime integration for background sends, guidance, and ACP render intents | registers on `ctx.tools` |
|
||||
|
||||
Idle detection is backend behavior, not a second public seam. A remote or container backend may have authoritative readiness signals that do not resemble local `/proc` inspection; every `PtyBackend` therefore returns the common send result while owning its detection mechanism internally.
|
||||
|
||||
### Agent ownership and identity
|
||||
|
||||
`PtyService` stores live sessions process-locally, but every session is owned by the exact `Agent` passed through the tool execution context. The service mints an opaque `PtySessionId`; an optional model-chosen `name` is display metadata and is unique only within that owner. Every operation targets `sessionId`, and `list`/`read`/`signal`/`kill` reject callers other than the owner.
|
||||
|
||||
The initial design has no plugin-load auto-start sessions. `pty_spawn` creates a session only during an agent tool call, when ownership and the owning event-sourced session are known. Deployments that later need declarative startup must compose it through unpublished agent setup rather than create shared global terminals.
|
||||
|
||||
Agent-scope disposal closes registrations first, then awaits quiescent teardown of every owned PTY. Backend or tool-plugin reload does not orphan sessions: ownership lives in `PtyService` until the agent ends, following the same service-owned-record pattern as [`ctx.tasks`](../../../../packages/tasks/tasks/README.md).
|
||||
|
||||
### Security and process boundary
|
||||
|
||||
A registered `shell` backend constrains how a terminal starts; it does not constrain commands typed after startup. `dsh-pty-local` therefore applies two protections before spawning:
|
||||
|
||||
- It builds a scrubbed child environment using the same credential-shaped-name policy as `bash-local`, removing ambient `*KEY*`, `*SECRET*`, `*TOKEN*`, and harness-managed variables unless an explicit trusted mapping supplies them.
|
||||
- Its `sandbox` config is `required | optional | disabled`, defaulting to `required`. `required` fails plugin load when `ctx.sandbox` is unavailable; `optional` uses the provider when present; `disabled` is an explicit unconfined opt-in. The selected provider wraps the session argv once and remains the process boundary for the PTY lifetime.
|
||||
|
||||
Sandboxing confines local process effects but does not make arbitrary shell input safe: network calls and other external side effects remain governed by deployment policy. Tool descriptions state that PTY sessions are less auditable than one-shot tools and should be used only when persistence or interactive stdin is necessary.
|
||||
|
||||
The implementation uses only public `node-pty` capabilities: child PID, `data` and `exit` notifications, `write`, `resize`, and `kill`. It does not assume access to the native master fd or call `waitpid` from TypeScript. Platform process inspectors derive process-group and session membership from `/proc` on Linux and `ps` on macOS.
|
||||
|
||||
### Six model-facing tools
|
||||
|
||||
| Tool | Purpose | Result |
|
||||
|---|---|---|
|
||||
| `pty_spawn` | Create an owner-scoped session from a registered backend type | `{ sessionId, name, type, motd }` |
|
||||
| `pty_send` | Send text, optionally submit Enter, and wait for readiness or register a background task | bounded viewport plus wait and session status; background also returns `taskId` |
|
||||
| `pty_read` | Read a bounded page from retained scrollback | `{ text, totalLines, lineBegin, lineEnd, truncated }` |
|
||||
| `pty_signal` | Send one allowed signal to the current foreground process group | `{ delivered, targetPgid }` |
|
||||
| `pty_kill` | Close one session and await process-tree quiescence | `{ killed }` |
|
||||
| `pty_list` | List the caller's live sessions | owner-scoped session summaries |
|
||||
|
||||
`pty_send({ sessionId, text, submit?, background? })` treats `text` as UTF-8 bytes and resolves `submit` to `true` in the tool implementation. When `submit` is true it writes the platform Enter sequence after the text; when false it writes only the text, allowing control characters and REPL fragments without hidden content heuristics.
|
||||
|
||||
Foreground sends return a bounded rendered delta and two independent facts: `waitReason` (`stdin_read | inferred_idle | timeout | session_exit`) and `sessionStatus` (`running` or `exited` with exit code or signal). `session_exit` refers to the PTY's top-level shell process, not an arbitrary foreground command whose status the shell consumes. A timeout never implies process exit.
|
||||
|
||||
With `background: true`, `dsh-tool-pty` registers the in-flight send on `ctx.tasks` and returns immediately with `taskId`. `task_output(wait: true)` waits, reads incremental output, and records the final result; `task_kill` forwards cancellation as `SIGINT` and escalates only through the PTY backend's owned teardown path. If the task surface is absent, background mode fails before writing input. No PTY-specific `sleep` tool or general wake-up seam is added.
|
||||
|
||||
`pty_read` pages backward from the newest retained line. The backend enforces both line and UTF-8 byte caps on retained scrollback and the complete returned value, so one oversized line cannot bypass the bound. `truncated` distinguishes retention loss from an ordinary viewport delta.
|
||||
|
||||
`pty_signal` accepts the closed set `SIGINT | SIGTERM | SIGKILL | SIGTSTP | SIGHUP`. The backend resolves the terminal foreground process group at execution time. `SIGKILL` is rejected when that group is the top-level shell, directing the caller to `pty_kill`; a failed group lookup fails the operation instead of signaling a guessed PID.
|
||||
|
||||
### Local readiness detection
|
||||
|
||||
The local backend runs three bounded tiers. All timings are validated config fields: `pollIntervalMs`, `exactProbeAfterMs`, `idleSilenceMs`, and `timeoutMs`.
|
||||
|
||||
On Linux, the inspector reads the shell's terminal foreground PGID from `/proc/<shellPid>/stat`, enumerates every process and thread in that process group, and probes their current syscalls. A positive Tier 1 result requires an observed stdin wait: direct `read(0)`, a permitted read of a `select`/`pselect6` or `poll`/`ppoll` argument containing fd 0, or an epoll interest list containing fd 0. Unreadable process memory and unrecognized syscalls are misses, never positive guesses. Architecture tables contain only syscall numbers defined by the corresponding Linux UAPI; unsupported architectures skip Tier 1.
|
||||
|
||||
On macOS there is no exact syscall tier. Output silence returns `inferred_idle` for any foreground process group, including Python and `gdb`; `ps`-derived terminal PGID is used for signaling, not as proof that only the shell can be idle. Pure process-inspector logic is injectable and unit-tested on Linux, while a macOS CI job exercises the real PTY and process-table path.
|
||||
|
||||
Tier 2 returns `inferred_idle` after `idleSilenceMs` without output. A sleeping or network-blocked command can therefore look ready. Tier 3 returns `timeout` after `timeoutMs` so a foreground tool call cannot hold the agent indefinitely. The result preserves the distinction; callers may wait through `ctx.tasks`, signal the foreground group, or inspect from another session.
|
||||
|
||||
`node-pty` data notifications feed one streaming decoder and terminal parser. Parser carry state handles UTF-8 and terminal query sequences split across chunks. The first delivery normalizes line-oriented output and detects alternate-screen entry, but it does not promise correct interaction with a full-screen application.
|
||||
|
||||
### Model-visible output and durability
|
||||
|
||||
The existing durable `tool/call` and `tool/result` events are the source of truth for text sent by the model and rendered output returned to it. `pty_spawn` returns its MOTD through the logged tool result; foreground `send`/`read`/`list`/`signal`/`kill` results are logged the same way. The PTY packages do not duplicate raw byte streams into custom session events.
|
||||
|
||||
Background sends use the existing task completion notice and `task_output` result path, so any output that reaches a later model request is likewise durable. Raw terminal bytes remain bounded process-local state and are neither persisted nor restorable. A future opt-in transcript sink would need its own retention, credential, and privacy contract.
|
||||
|
||||
### Process-tree teardown
|
||||
|
||||
The top-level `node-pty` child is treated as the POSIX session leader, but the owned resource is the complete OS process session, not one PID. On close, the backend stops callbacks, sends `SIGTERM` to all still-matching session members, closes the PTY, awaits `node-pty` exit plus process-inspector quiescence, then sends `SIGKILL` to verified survivors after configurable `disposeGraceMs`. Membership snapshots include process-start identity so PID reuse cannot redirect escalation.
|
||||
|
||||
Teardown reports root exit and survivor cleanup independently. It does not claim success merely because the shell exited; disposal resolves only after no captured session member remains or returns a structured cleanup failure naming the survivors.
|
||||
|
||||
### Composition and rollout
|
||||
|
||||
The example composition remains opt-in and safe by default:
|
||||
|
||||
```yaml
|
||||
plugins:
|
||||
'@deepseek-ai/dsh-sandbox-local':
|
||||
'@deepseek-ai/dsh-pty':
|
||||
'@deepseek-ai/dsh-pty-local':
|
||||
config:
|
||||
sandbox: required
|
||||
scrollbackLines: 10000
|
||||
scrollbackMaxBytes: 4194304
|
||||
maxReadBytes: 262144
|
||||
pollIntervalMs: 50
|
||||
exactProbeAfterMs: 150
|
||||
idleSilenceMs: 3000
|
||||
timeoutMs: 30000
|
||||
disposeGraceMs: 3000
|
||||
'@deepseek-ai/dsh-tool-pty':
|
||||
```
|
||||
|
||||
The package ships concise tool guidance explaining persistent state, owner isolation, uncertain idle results, cleanup, and the preference for existing one-shot tools when interaction is unnecessary. It does not add a global system-prompt recommendation or mount PTY in shipped defaults.
|
||||
|
||||
### Deferred work
|
||||
|
||||
- Full-screen TUI support, named key sequences, BEL interruption, terminal resize tools, and alternate-screen snapshots require a separately proven model-facing contract.
|
||||
- Declarative per-agent startup requires an agent-setup composition point; plugin-load global sessions remain prohibited.
|
||||
- Session restoration across harness-process loss requires an out-of-process owner and a versioned protocol.
|
||||
- Network-egress policy and rollback of external side effects are broader than PTY and remain separate security work.
|
||||
- Windows/ConPTY support requires a backend with Windows-native process ownership and signaling semantics.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Replace `bash`, filesystem tools, or task tools with PTY.** Rejected. One-shot tools retain stronger validation, approval, sandbox, output-bound, and replay contracts. PTY is reserved for interactive state.
|
||||
|
||||
**Add persistent mode to `bash`.** Rejected. Returning on readiness rather than process exit, retaining a process tree across calls, and exposing interactive stdin create a different ownership and failure contract.
|
||||
|
||||
**Require native master-fd access from `node-pty`.** Rejected. Its public API exposes no master fd. The local backend instead derives foreground and session membership from supported OS process metadata and treats unreadable metadata as a detector miss.
|
||||
|
||||
**Publish `PtyIdleDetector` as a replaceable registry.** Rejected. Only the local backend needs these platform probes, while remote backends may receive readiness over their own protocol. Backend replacement already provides the necessary extension point.
|
||||
|
||||
**Add a PTY-specific `sleep` tool.** Rejected. `ctx.tasks` already owns bounded waiting, cancellation, completion notices, and model-facing collection. A second general wake mechanism would cross the agent-loop boundary and duplicate that contract.
|
||||
|
||||
**Include TUI sequences and BEL handling in the first delivery.** Rejected. The source prototype treats those paths as timing-sensitive and still records unresolved alternate-screen and interaction failures. Line-oriented PTY use proves the core value without making those unverified behaviors foundational.
|
||||
|
||||
**Use an out-of-process daemon immediately.** Rejected for the initial in-process capability because current persistent front doors already keep a Cordis context alive. A daemon becomes justified by cross-process restoration or multi-client attachment, both deferred here.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- `packages/pty/{pty,pty-local,tool-pty}` build as the interface, local implementation, and model consumer; backend registrations dispose cleanly.
|
||||
- Every live PTY has one service-minted `PtySessionId`, one exact `Agent` owner, owner-fenced operations, and awaited cleanup on agent disposal; concurrent agents may reuse display names without sharing state.
|
||||
- `dsh-pty-local` uses only public `node-pty` APIs and contains no master-fd or TypeScript `waitpid` assumption.
|
||||
- Environment tests prove credential-shaped ambient variables are absent. `sandbox: required` fails at load without a provider, and real composition proves the provider wraps the long-lived session process.
|
||||
- Linux fixtures cover pipelines, a stdin-reading non-leader process, a stdin-reading non-main thread, unreadable process memory, supported UAPI syscall tables, unsupported architectures, and false-positive rejection. macOS process-inspector logic reaches 100% coverage on Linux, and macOS CI drives a real bash and Python REPL.
|
||||
- Foreground tests exercise `stdin_read`, `inferred_idle`, `timeout`, and top-level session exit without treating a foreground command exit as directly observable.
|
||||
- Background sends register `ctx.tasks` work, return before readiness, stream bounded output through `task_output`, honor task cancellation, and fail before writing when the task surface is absent.
|
||||
- Scrollback and every model-facing result enforce final UTF-8 byte bounds, including a single oversized line and multibyte boundary cases.
|
||||
- `pty_signal` resolves the live foreground group, rejects lookup failure and shell-targeted `SIGKILL`, and never falls back to a guessed PID.
|
||||
- Disposal tests start foreground and background descendants, including a signal-ignoring child, then prove every captured process identity is gone immediately after awaited agent disposal.
|
||||
- A test-only `cordis.yml` boots through the Loader on Linux and macOS, mounts the real local backend plus sandbox, and drives spawn/send/read/signal/kill/list through the real tool registry. ACP and headless snapshots pin the six schemas, bounded results, errors, and render intents.
|
||||
- TUI, sequence, BEL, auto-start, Windows, and crash-restoration behavior are absent from the public schema and documented as deferred rather than simulated by fixtures.
|
||||
- Package READMEs and JSDoc document configuration, ownership, failure, cancellation, bounds, sandboxing, model-visible effects, and limitations; `docs/architecture.md` and generated catalogs update with the implementation.
|
||||
- The repository CI-equivalent sequence in root `AGENTS.md` passes, including `test:coverage`, snapshots, documentation, build, hygiene, and built-entry smokes.
|
||||
|
||||
## Risks
|
||||
|
||||
**Idle below Linux Tier 1 is heuristic.** Output silence cannot distinguish a prompt from sleep or network I/O. The typed result preserves uncertainty, and bounded timeout plus task waiting and signaling keep control with the model.
|
||||
|
||||
**Persistent state can drift from the model's belief.** The model may forget its cwd or active REPL. Session summaries and retained output help recovery, but no prompt can make state persistence deterministic.
|
||||
|
||||
**A shell can cause external side effects.** Session sandboxing and environment scrubbing reduce local exposure but do not undo pushes, API calls, or messages. Deployments that cannot tolerate those effects must omit PTY or add network policy.
|
||||
|
||||
**Process loss destroys terminal state.** In-process sessions do not survive a harness crash or restart, and raw scrollback is not durable. Important work must be committed to files or another durable system.
|
||||
|
||||
**`node-pty` is a native dependency.** Installation, supported Node versions, prebuild availability, and platform behavior require built-artifact smokes on every supported OS.
|
||||
@@ -0,0 +1,170 @@
|
||||
# Agent Note: 持久化 PTY 会话
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-16-persistent-pty-sessions.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
harness 可以运行前台与后台命令、编辑文件和委派工作,但无法跨工具调用延续一次交互式终端对话。每次 `bash` 前台运行都会启动一个新 shell,因此 shell 内的 cwd、导出变量、虚拟环境激活状态、函数、job control 状态和交互式子进程都会随本次调用结束。
|
||||
|
||||
这个缺口排除了状态驻留在终端而不是文件中的工作流,例如单步调试 `gdb`、在 Python 或 Node REPL 中探索、驱动 `ed` 这类行式编辑器,或者中断前台命令后回到原 shell。通用的 [`ctx.tasks`](../../../../packages/tasks/README.md) 运行时可以保留后台操作句柄和输出,但不提供交互式 stdin 或终端语义。
|
||||
|
||||
现有 `bash`、`read`、`write` 和 `edit` 工具仍是有界、可审计操作的可靠默认选项。PTY 是对确实需要终端状态的工作的补充功能,不说明这些工具有缺陷,更不意味着要移除它们。
|
||||
|
||||
## 提案
|
||||
|
||||
新增可选的 `packages/pty/` 功能家族,向模型提供由 agent 拥有、持久化且面向行式交互的 PTY 会话。它遵循仓库的 [capability pattern](../../implemented/architecture/2026-06-13-capability-seams.md),与现有命令和文件系统工具并存,并且不修改 `agent-loop`。
|
||||
|
||||
首次交付在 Linux 和 macOS 上支持交互式 shell 与行式 REPL。全屏终端应用、按键序列、BEL 触发的控制流、进程丢失后的会话恢复以及跨 agent 共享会话都明确推迟,直到基础生命周期得到验证。
|
||||
|
||||
### 包拓扑
|
||||
|
||||
| 包 | 角色 | ctx key |
|
||||
|---|---|---|
|
||||
| `dsh-pty` | `PtyService`、branded `PtySessionId`、后端注册表、按 owner 隔离的会话契约和结果类型 | `ctx.pty` |
|
||||
| `dsh-pty-local` | 基于 [`node-pty`](https://github.com/microsoft/node-pty) 的本地后端、平台进程检查、有界终端缓冲、沙箱解析和进程树监管 | 在 `ctx.pty` 上注册后端 |
|
||||
| `dsh-tool-pty` | 6 个面向模型的工具、后台发送的 task 运行时集成、使用指引和 ACP render intent | 注册到 `ctx.tools` |
|
||||
|
||||
idle 检测属于后端行为,不是第二条公共 seam。远程或容器后端可能拥有完全不同于本地 `/proc` 检查的权威就绪信号;因此每个 `PtyBackend` 都返回统一的发送结果,同时在内部拥有自己的检测机制。
|
||||
|
||||
### agent 所有权与身份
|
||||
|
||||
`PtyService` 在进程内保存活会话,但每个会话都由工具执行上下文传入的确切 `Agent` 拥有。服务铸造不透明的 `PtySessionId`;模型可选填的 `name` 只是显示元数据,仅在该 owner 内唯一。所有操作都以 `sessionId` 为目标,`list`/`read`/`signal`/`kill` 会拒绝 owner 之外的调用方。
|
||||
|
||||
初始设计不提供插件加载期 auto-start 会话。`pty_spawn` 只在 agent 工具调用期间创建会话,此时所有权和所属的事件溯源会话都已确定。若部署后续需要声明式启动,必须通过尚未发布的 agent setup 组合,而不能创建全局共享终端。
|
||||
|
||||
agent scope dispose 时先关闭注册,再等待全部所属 PTY 静默退出。后端或工具插件 reload 不会遗留会话:所有权持续存放在 `PtyService` 中,直到 agent 结束,与 [`ctx.tasks`](../../../../packages/tasks/tasks/README.md) 的服务持有记录模式一致。
|
||||
|
||||
### 安全与进程边界
|
||||
|
||||
注册的 `shell` 后端只约束终端如何启动,不约束启动后输入的命令。因此 `dsh-pty-local` 在 spawn 前应用两层保护:
|
||||
|
||||
- 它使用与 `bash-local` 相同的凭证形态名称策略构建清洗后的子进程环境,移除环境中的 `*KEY*`、`*SECRET*`、`*TOKEN*` 和 harness 管理的变量,除非显式的可信映射提供这些值。
|
||||
- 它的 `sandbox` 配置为 `required | optional | disabled`,默认 `required`。`required` 在缺少 `ctx.sandbox` 时于插件加载期失败;`optional` 在提供方存在时使用;`disabled` 是显式选择无约束模式。所选提供方只包装一次会话 argv,并在 PTY 的整个生命周期中充当进程边界。
|
||||
|
||||
沙箱限制本地进程副作用,但不会让任意 shell 输入自动安全:网络调用和其他外部副作用仍由部署策略治理。工具描述会说明 PTY 会话比一次性工具更难审计,只应在确实需要持久状态或交互式 stdin 时使用。
|
||||
|
||||
实现只使用 `node-pty` 的公共功能:子进程 PID、`data` 与 `exit` 通知、`write`、`resize` 和 `kill`。它不假设能访问原生 master fd,也不从 TypeScript 调用 `waitpid`。平台进程检查器在 Linux 上通过 `/proc`、在 macOS 上通过 `ps` 推导进程组和会话成员关系。
|
||||
|
||||
### 6 个面向模型的工具
|
||||
|
||||
| 工具 | 用途 | 结果 |
|
||||
|---|---|---|
|
||||
| `pty_spawn` | 从已注册的后端类型创建按 owner 隔离的会话 | `{ sessionId, name, type, motd }` |
|
||||
| `pty_send` | 发送文本、可选提交 Enter,并等待就绪或注册一个后台任务 | 有界 viewport、等待状态和会话状态;后台模式还返回 `taskId` |
|
||||
| `pty_read` | 从保留的 scrollback 读取一个有界页 | `{ text, totalLines, lineBegin, lineEnd, truncated }` |
|
||||
| `pty_signal` | 向当前前台进程组发送一种允许的信号 | `{ delivered, targetPgid }` |
|
||||
| `pty_kill` | 关闭一个会话并等待进程树静默退出 | `{ killed }` |
|
||||
| `pty_list` | 列出调用方的活会话 | 按 owner 隔离的会话摘要 |
|
||||
|
||||
`pty_send({ sessionId, text, submit?, background? })` 将 `text` 视为 UTF-8 字节,并由工具实现在解析阶段把 `submit` 默认成 `true`。`submit` 为 true 时先写入文本,再写入平台 Enter 序列;为 false 时只写文本,使控制字符和 REPL 片段无需隐藏的内容启发式即可发送。
|
||||
|
||||
前台发送返回有界的渲染增量和两个独立事实:`waitReason`(`stdin_read | inferred_idle | timeout | session_exit`)与 `sessionStatus`(`running`,或携带退出码或信号的 `exited`)。`session_exit` 指 PTY 顶层 shell 进程退出,不指由 shell 消费状态的任意前台命令。timeout 从不意味着进程已经退出。
|
||||
|
||||
当 `background: true` 时,`dsh-tool-pty` 在 `ctx.tasks` 上注册进行中的发送,并立即返回 `taskId`。`task_output(wait: true)` 负责等待、读取增量输出并记录最终结果;`task_kill` 将取消转发为 `SIGINT`,只有 PTY 后端拥有的 teardown 路径可以升级信号。若 task 对外接口不存在,后台模式必须在写入输入前失败。设计不新增 PTY 专用的 `sleep` 工具或通用唤醒 seam。
|
||||
|
||||
`pty_read` 从最新保留行向后分页。后端同时对保留的 scrollback 和完整返回值执行行数与 UTF-8 字节上限,因此单个超长行无法绕过限制。`truncated` 用于区分保留数据丢失与普通 viewport 增量。
|
||||
|
||||
`pty_signal` 接受闭合集 `SIGINT | SIGTERM | SIGKILL | SIGTSTP | SIGHUP`。后端在执行时解析终端前台进程组。当目标组是顶层 shell 时拒绝 `SIGKILL`,并指引调用方使用 `pty_kill`;进程组解析失败时操作直接失败,而不是向猜测的 PID 发送信号。
|
||||
|
||||
### 本地就绪检测
|
||||
|
||||
本地后端执行 3 个有界层级。所有时间参数都是经校验的配置字段:`pollIntervalMs`、`exactProbeAfterMs`、`idleSilenceMs` 和 `timeoutMs`。
|
||||
|
||||
在 Linux 上,检查器从 `/proc/<shellPid>/stat` 读取 shell 的终端前台 PGID,枚举该进程组中的每个进程与线程,并检查它们当前的 syscall。Tier 1 只有观察到 stdin 等待才返回正结果:直接 `read(0)`、获准读取且含 fd 0 的 `select`/`pselect6` 或 `poll`/`ppoll` 参数,或者含 fd 0 的 epoll interest list。无法读取的进程内存和未识别的 syscall 都是 miss,绝不作为正向猜测。架构表只包含对应 Linux UAPI 定义的 syscall number;不支持的架构跳过 Tier 1。
|
||||
|
||||
macOS 没有精确 syscall 层。任何前台进程组输出静默都会返回 `inferred_idle`,包括 Python 和 `gdb`;从 `ps` 推导的终端 PGID 只用于发送信号,不作为「只有 shell 才能 idle」的证明。纯进程检查逻辑可注入并在 Linux 上完成 unit 覆盖率,同时由 macOS CI job 驱动真实 PTY 和进程表路径。
|
||||
|
||||
Tier 2 在持续 `idleSilenceMs` 没有输出后返回 `inferred_idle`,因此 sleep 或网络阻塞的命令可能看似 ready。Tier 3 在 `timeoutMs` 后返回 `timeout`,避免前台工具调用无限占住 agent。结果保留这些区别;调用方可以通过 `ctx.tasks` 等待、向前台组发信号,或从另一个会话排查。
|
||||
|
||||
`node-pty` data 通知进入同一个流式 decoder 和终端 parser。parser 的 carry 状态处理跨 chunk 的 UTF-8 与终端查询序列。首次交付只规范化行式输出并检测 alternate-screen 进入,不承诺正确操作全屏应用。
|
||||
|
||||
### 模型可见输出与持久性
|
||||
|
||||
现有持久化 `tool/call` 与 `tool/result` 事件是模型发送文本和返回给模型的渲染输出的真源。`pty_spawn` 通过已记录的工具结果返回 MOTD;前台 `send`/`read`/`list`/`signal`/`kill` 结果走同一路径记录。PTY 包不会把原始字节流重复写入自定义会话事件。
|
||||
|
||||
后台发送复用现有后台任务完成通知和 `task_output` 结果路径,因此进入后续模型请求的任何输出同样持久化。原始终端字节只作为有界的进程内状态存在,既不持久化也不可恢复。未来的 opt-in transcript sink 必须拥有独立的保留、凭证和隐私契约。
|
||||
|
||||
### 进程树 teardown
|
||||
|
||||
顶层 `node-pty` 子进程视为 POSIX 会话 leader,但所属资源是完整的 OS 进程会话,而不是一个 PID。关闭时,后端先停止 callback,再向仍匹配的会话成员发送 `SIGTERM`、关闭 PTY、等待 `node-pty` exit 与进程检查器确认静默,然后在可配置的 `disposeGraceMs` 后向已验证的存活者发送 `SIGKILL`。成员快照包含进程启动身份,避免 PID 复用把升级信号发给无关进程。
|
||||
|
||||
teardown 独立报告根进程退出与存活进程清理。它不会只因 shell 退出就声称成功;dispose 只有在已捕获的会话成员全部消失后才完成,否则返回结构化清理失败并列出存活者。
|
||||
|
||||
### 组合与推行
|
||||
|
||||
示例组合保持 opt-in,并采用安全默认值:
|
||||
|
||||
```yaml
|
||||
plugins:
|
||||
'@deepseek-ai/dsh-sandbox-local':
|
||||
'@deepseek-ai/dsh-pty':
|
||||
'@deepseek-ai/dsh-pty-local':
|
||||
config:
|
||||
sandbox: required
|
||||
scrollbackLines: 10000
|
||||
scrollbackMaxBytes: 4194304
|
||||
maxReadBytes: 262144
|
||||
pollIntervalMs: 50
|
||||
exactProbeAfterMs: 150
|
||||
idleSilenceMs: 3000
|
||||
timeoutMs: 30000
|
||||
disposeGraceMs: 3000
|
||||
'@deepseek-ai/dsh-tool-pty':
|
||||
```
|
||||
|
||||
包会提供简洁的工具指引,说明持久状态、owner 隔离、不确定的 idle 结果、清理,以及无需交互时优先使用现有一次性工具。它不增加全局 system prompt 推荐,也不在已发布的默认配置中挂载 PTY。
|
||||
|
||||
### 推迟的工作
|
||||
|
||||
- 全屏 TUI 支持、命名按键序列、BEL 中断、终端 resize 工具和 alternate-screen 快照需要另行验证面向模型的契约。
|
||||
- 声明式 per-agent 启动需要 agent-setup 组合点;仍然禁止插件加载期全局会话。
|
||||
- harness 进程丢失后的会话恢复需要进程外 owner 和版本化协议。
|
||||
- 网络出口策略与外部副作用回滚超出 PTY 范围,继续作为独立安全工作。
|
||||
- Windows/ConPTY 支持需要具备 Windows 原生进程所有权与信号语义的后端。
|
||||
|
||||
## 备选方案
|
||||
|
||||
**用 PTY 替换 `bash`、文件系统工具或 task 工具。**拒绝。一次性工具拥有更强的校验、审批、沙箱、输出上限和回放契约。PTY 只服务交互式状态。
|
||||
|
||||
**给 `bash` 增加持久模式。**拒绝。按就绪而不是进程退出返回、跨调用保留进程树、暴露交互式 stdin 会形成不同的所有权和失败契约。
|
||||
|
||||
**要求从 `node-pty` 获取原生 master fd。**拒绝。它的公共 API 不暴露 master fd。本地后端改为从受支持的 OS 进程元数据推导前台组和 session 成员,并把不可读元数据视为 detector miss。
|
||||
|
||||
**发布可替换注册表 `PtyIdleDetector`。**拒绝。只有本地后端需要这些平台 probe,远程后端可能通过自己的协议接收就绪状态。替换后端已经提供所需扩展点。
|
||||
|
||||
**新增 PTY 专用 `sleep` 工具。**拒绝。`ctx.tasks` 已经拥有有界等待、取消、完成通知和面向模型的收集。第二套通用唤醒机制会跨越 agent loop(智能体循环)边界并重复该契约。
|
||||
|
||||
**在首次交付包含 TUI sequence 与 BEL 处理。**拒绝。源 prototype 将这些路径视为 timing-sensitive,且仍记录未解决的 alternate-screen 和交互失败。行式 PTY 已能证明核心价值,无需把未经验证的行为放进基础层。
|
||||
|
||||
**立即采用进程外 daemon。**初始的进程内功能不采用,因为当前持久 front door 已能维持 Cordis context。跨进程恢复或多客户端 attach 会让 daemon 变得合理,但两者都已推迟。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- `packages/pty/{pty,pty-local,tool-pty}` 分别作为接口、本地实现和模型消费方构建;后端注册可干净 dispose。
|
||||
- 每个活 PTY 都有一个由服务铸造的 `PtySessionId`、一个确切的 `Agent` owner、按 owner 隔离的操作,并在 agent dispose 时等待清理;并发 agent 可以复用显示名称而不共享状态。
|
||||
- `dsh-pty-local` 只使用 `node-pty` 公共 API,不包含 master-fd 或 TypeScript `waitpid` 假设。
|
||||
- 环境测试证明凭证形态的环境变量不存在。缺少提供方时 `sandbox: required` 在加载期失败,REAL-composition 测试证明提供方包装长活会话进程。
|
||||
- Linux fixture(测试前置数据)覆盖 shell 管道、读取 stdin 的非 leader 进程、读取 stdin 的非主线程、不可读进程内存、受支持的 UAPI syscall 表、不支持的架构和误报拒绝。macOS 进程检查逻辑在 Linux 上达到 100% 覆盖率,macOS CI 驱动真实 bash 与 Python REPL。
|
||||
- 前台测试覆盖 `stdin_read`、`inferred_idle`、`timeout` 和顶层会话退出,不把前台命令退出当作可直接观察事件。
|
||||
- 后台发送注册 `ctx.tasks` work、在就绪前返回、通过 `task_output` 流式提供有界输出、遵守 task cancellation,并在 task 对外接口缺失时于写入前失败。
|
||||
- scrollback 与每个面向模型的结果都对最终 UTF-8 字节执行上限,包括单个超长行和多字节边界情况。
|
||||
- `pty_signal` 解析活跃前台组,拒绝查询失败和指向 shell 的 `SIGKILL`,且绝不回退到猜测的 PID。
|
||||
- dispose 测试启动前台与后台子进程,包括忽略信号的子进程,然后证明等待 agent dispose 后每个捕获的进程身份立即消失。
|
||||
- 测试专用 `cordis.yml` 在 Linux 与 macOS 上通过 Loader 启动,挂载真实本地后端与沙箱,并通过真实工具注册表驱动 spawn/send/read/signal/kill/list。ACP 与 headless 快照固定 6 个 schema、有界结果、错误和 render intent。
|
||||
- TUI、sequence、BEL、auto-start、Windows 和 crash-restoration 行为不出现在公共 schema 中,并记录为推迟事项,而不是由 fixture 模拟。
|
||||
- 包 README 与 JSDoc 记录配置、所有权、失败、取消、上限、沙箱、模型可见影响和限制;实现同时更新 `docs/architecture.md` 与生成目录。
|
||||
- 根 `AGENTS.md` 中的仓库 CI 等价序列通过,包括 `test:coverage`、快照、文档、构建、hygiene 和 built-entry smoke。
|
||||
|
||||
## 风险
|
||||
|
||||
**Linux Tier 1 之外的 idle 都是启发式结果。**输出静默无法区分 prompt、sleep 和网络 I/O。类型化结果保留不确定性,有界 timeout、task 等待与信号让模型仍能掌握控制权。
|
||||
|
||||
**持久状态可能偏离模型认知。**模型可能忘记 cwd 或活跃 REPL。会话摘要和保留输出有助恢复,但任何 prompt 都无法让状态持久化变成确定行为。
|
||||
|
||||
**Shell 可以造成外部副作用。**会话沙箱和环境清洗降低本地暴露,但无法撤销 push、API 调用或消息发送。无法容忍这些副作用的部署必须省略 PTY 或增加网络策略。
|
||||
|
||||
**进程丢失会销毁终端状态。**进程内会话无法跨 harness crash 或 restart 存活,原始 scrollback 也不持久化。重要工作必须提交到文件或其他持久系统。
|
||||
|
||||
**`node-pty` 是原生依赖。**安装、支持的 Node 版本、prebuild 可用性和平台行为都需要在每个支持 OS 上运行 built-artifact smoke。
|
||||
@@ -0,0 +1,6 @@
|
||||
# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write
|
||||
2026-07-17-sdk-follow-up-capabilities.md: 0f3ada6bdbb4ce933d14602cf59be9a51640e61c
|
||||
2026-07-17-sdk-follow-up-capabilities.zh.md: d0d0b3e6bcdf192e64f003dc9f6e90cc2bdb060b
|
||||
@@ -0,0 +1,118 @@
|
||||
# Agent Note: SDK follow-up capabilities
|
||||
|
||||
Status: proposed
|
||||
|
||||
English | [中文](2026-07-17-sdk-follow-up-capabilities.zh.md)
|
||||
|
||||
## Problem
|
||||
|
||||
The first SDK release creates and edits developer-owned Cordis projects through the shared model defined by the [developer-project Agent Note](2026-07-14-sdk-developer-projects.md) and the [project-editing architecture](../architecture/2026-07-15-sdk-project-editing-architecture.md). Its create and config workflows are interactive, external Cordis plugins require manual dependency and configuration edits, command-line telemetry has no owning boundary, and interactive branches lack a stable test strategy.
|
||||
|
||||
These gaps are coupled. Create and config already share questions, feature configuration, and `ProjectEditSession`; adding separate automation paths would duplicate that domain logic. External-plugin installation must update both the package manager's files and `cordis.yml`. Telemetry must observe commands such as create and build that do not boot Cordis. Interactive testing must exercise Harness behavior without making terminal rendering a brittle product contract.
|
||||
|
||||
## Proposal
|
||||
|
||||
The SDK extends the existing prompt and project-editing boundaries instead of creating parallel workflows. A non-interactive prompt port and structured feature plan drive create and config, `dsh-sdk create <source>` delegates dependency resolution to the project package manager before mounting the resolved package through `ProjectEditSession`, launcher-side telemetry wraps `create-sdk` and every `dsh-sdk` command, and injected prompt streams provide the primary interactive-test seam.
|
||||
|
||||
| Capability | Product entrypoint | Owning mechanism | Required outcome |
|
||||
|---|---|---|---|
|
||||
| Headless project creation | `create-sdk --config <file>` or `--config-json <json>` with optional `--json` | `HeadlessPromptPort`, structured project answers, and a complete feature plan | No terminal blocking; missing required input is explicit |
|
||||
| External Cordis plugin installation | `dsh-sdk create <source>` | Native package-manager `add` plus `ProjectEditSession` | The dependency and `cordis.yml` entry identify the package manager's resolved package |
|
||||
| Developer-cycle telemetry | `create-sdk` and every `dsh-sdk` command | Launcher-side consent, payload, redaction, anonymous identity, and delivery services | Reporting is best-effort and cannot change the command result |
|
||||
| Interactive regression coverage | Create and config tests | Injected `PromptPort` input/output and filesystem assertions | Tests cover Harness decisions and generated files without snapshotting terminal repainting |
|
||||
|
||||
## Shared headless workflow
|
||||
|
||||
### Structured input and lifecycle events
|
||||
|
||||
Headless create accepts a JSON object either inline through `--config-json` or from a file through `--config`. Scalar fields supply the ordinary create answers, while `features` supplies the complete selected feature set, feature options, secrets, and dedicated values. Defaults remain valid only where the owning question declares one; the headless path never invents an answer for a required prompt.
|
||||
|
||||
With `--json`, stdout is an NDJSON event stream. `done` means creation and any requested setup completed, `action-required` names an unanswered required prompt, and `error` reports another failure. Human-readable progress and package-manager output go to stderr so every stdout line remains parseable as one event. A caller responds to `action-required` by adding the missing value and running the command again.
|
||||
|
||||
Create and config consume the same feature-plan shape. Create exposes it through the command-line inputs above; config uses it at the shared workflow boundary so a later automation entrypoint does not need a second feature-selection model.
|
||||
|
||||
### Prompt and project-editing boundaries
|
||||
|
||||
`PromptPort` remains the only boundary between SDK questions and an interaction implementation. `ClackPromptPort` handles terminals. `HeadlessPromptPort` consumes defaults exposed by the question contract and otherwise fails with the unanswered prompt; prefilled values normally prevent the port from being called.
|
||||
|
||||
Both paths use the same `Question` objects, `FeatureConfigurator`, `SdkProject`, and `ProjectEditSession`. The headless path therefore changes how answers arrive, not how features are interpreted or files are committed.
|
||||
|
||||
### Agent skill
|
||||
|
||||
The repository ships a thin `SKILL.md` that teaches an agent to construct the structured input, request NDJSON, fill an `action-required` value, and retry. The skill invokes the public CLI and does not import an internal SDK API or introduce another project specification.
|
||||
|
||||
## External Cordis plugin installation
|
||||
|
||||
`dsh-sdk create <source>` accepts a package-manager-native npm specifier such as `pkg@version` or a GitHub specifier such as `github:owner/repo#ref`. After confirmation, it asks the project's package manager to add the source, compares the direct dependency names before and after the operation, reopens the project, and mounts each newly resolved package in `cordis.yml` through `ProjectEditSession`.
|
||||
|
||||
The package manager owns source parsing, version or commit resolution, integrity data, lockfile updates, and any build policy. The SDK does not download or unpack a second copy through giget or pacote. An external plugin remains a dependency under `node_modules`; local plugin scaffolding remains a separate project-creation concern.
|
||||
|
||||
## Launcher telemetry
|
||||
|
||||
### Consent and collection
|
||||
|
||||
Telemetry wraps the `create-sdk` initializer and the `dsh-sdk` launcher command lifecycle because project initialization, plugin creation, and build do not reliably boot Cordis. One event records the command name, duration, success, a random per-user anonymous identifier, and redacted `cordis.yml` and `package.json` text when those project files are eligible.
|
||||
|
||||
Reporting is enabled unless a present telemetry config entry is explicitly disabled. `DO_NOT_TRACK` and CI deny reporting regardless of project configuration. A missing `cordis.yml` does not itself deny the event, but `package.json` content is included only when `cordis.yml` establishes that the directory is an SDK project.
|
||||
|
||||
### Safety and delivery
|
||||
|
||||
The payload builder never reads `.env`. It redacts secret-shaped keys and values, known token forms, PEM blocks, URL credentials, and high-entropy opaque strings in the two eligible text files. Redaction is a safety backstop rather than a guarantee; SDK projects must keep credentials in `.env`.
|
||||
|
||||
The reporter uses a fixed endpoint and resolves every send path without throwing. Command dispatch records success or failure in a `finally` path, starts reporting after the command outcome is known, and drains within a bounded interval. Consent parsing, payload construction, storage, or network failures are swallowed only at this telemetry boundary and never alter the command's exit code.
|
||||
|
||||
## Interactive workflow testing
|
||||
|
||||
Create and config tests inject a `PromptPort` and scripted input/output streams into the existing workflows. Parameterized scenarios cover feature selection, feature options, secrets, cancellation, review, and apply behavior, then assert the resulting `cordis.yml` and other project files. The stable product assertion is the generated project state, not clack's ANSI redraw sequence.
|
||||
|
||||
One or two optional real-PTY smoke tests may cover the shipped binary and TTY guard that injection cannot reproduce. Native PTY tooling does not belong on the required path unless it is reliable across the repository's supported Node and host versions.
|
||||
|
||||
## Deferred work
|
||||
|
||||
- Extend the headless create specification to express local `plugin` or `tool` scaffolding instead of defaulting that interactive choice to none.
|
||||
- Expose the telemetry opt-out in create and config while preserving the consent representation in which only a disabled telemetry entry is written.
|
||||
- Define whether GitHub source dependencies must be prebuilt or may run package-manager-controlled preparation scripts, and surface the policy before installation.
|
||||
- Replace the telemetry package's `.invalid` endpoint placeholder with the production endpoint before release.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
**Build a separate headless creation engine.** This would duplicate questions, feature requirements, configuration behavior, and project-editing rules. Reusing the prompt and edit-session boundaries keeps one implementation of project semantics.
|
||||
|
||||
**Make a specification file the primary automation interface.** Agents can pass the same typed JSON object inline, while people and CI may still use a file. A file-only protocol adds persistence and cleanup without adding semantics.
|
||||
|
||||
**Use `npx skills add` as the project creator.** The skills CLI installs Markdown skills; it does not create SDK projects or install npm packages. The agent skill therefore drives the SDK initializer instead of replacing it.
|
||||
|
||||
**Fetch GitHub and npm sources through giget or pacote.** A second fetch layer would duplicate package-manager resolution, integrity, lockfile, and lifecycle policy. Native dependency specifiers keep those decisions in the selected package manager.
|
||||
|
||||
**Implement telemetry as a Cordis runtime plugin.** Create and build do not necessarily boot Cordis, so a runtime plugin cannot observe the complete developer command cycle. The launcher is the boundary shared by those commands.
|
||||
|
||||
**Derive the anonymous identifier from git metadata.** Repository remotes can identify a project or organization. A random per-user identifier supports aggregation without encoding repository identity.
|
||||
|
||||
**Collect only aggregate counters.** Aggregate-only events reduce exposure but cannot answer which plugins, dependencies, and configuration shapes developers actually use. This proposal accepts collection of redacted project text and makes that exposure explicit.
|
||||
|
||||
**Use real PTYs and transcript snapshots as the primary test strategy.** Native PTY dependencies and terminal repaint sequences add platform and rendering instability while mostly testing clack. Injected interaction plus generated-file assertions tests the SDK-owned behavior directly.
|
||||
|
||||
## Acceptance criteria
|
||||
|
||||
- Create runs without a TTY from a complete structured input, emits only NDJSON on stdout under `--json`, and reports missing required input as `action-required` without writing a partial project.
|
||||
- Create and config resolve the same feature-plan contract through the shared question, feature-configuration, and project-editing code paths.
|
||||
- `dsh-sdk create <source>` uses the selected project package manager, mounts the dependency name that operation actually added, and fails loudly when no new dependency can be identified.
|
||||
- The initializer and every `dsh-sdk` command reach one best-effort telemetry completion path; an explicit disabled entry, `DO_NOT_TRACK`, or CI prevents delivery, and telemetry failures never change the command result.
|
||||
- Telemetry never reads `.env`, withholds unrelated `package.json` content when no `cordis.yml` exists, redacts both eligible text payloads, and uses an identifier unrelated to git metadata.
|
||||
- Interactive tests cover create and config decisions through injected interaction and assert committed project files; any real-PTY coverage remains a narrow smoke layer.
|
||||
- The agent skill documents the public structured-input and event contracts without depending on private package exports.
|
||||
|
||||
## Risks
|
||||
|
||||
- Full redacted `cordis.yml` and `package.json` text still reveals plugin and dependency names, URLs, paths, and configuration values to the endpoint operator, and heuristic redaction can miss a secret.
|
||||
- Default-on reporting may surprise developers when no telemetry entry exists; the CLI must make the opt-out discoverable before release.
|
||||
- A package-manager add can change `package.json`, the lockfile, and installed files before `ProjectEditSession` mounts the plugin, so a later mount failure can leave dependency changes that require manual recovery.
|
||||
- GitHub dependencies may execute preparation or lifecycle code according to package-manager policy; an unresolved build policy is a supply-chain and reproducibility risk.
|
||||
- Injected prompt tests do not prove raw-mode, signal, or repaint behavior in a real terminal; the optional smoke layer must cover only those residual contracts.
|
||||
|
||||
## References
|
||||
|
||||
- [Vercel Eve](https://github.com/vercel/eve) and [Vercel Labs Skills](https://github.com/vercel-labs/skills) for the distinction between a headless initializer and skill distribution.
|
||||
- [npm package specifications](https://docs.npmjs.com/cli/v11/using-npm/package-spec), [pnpm add](https://pnpm.io/cli/add), and [Yarn add](https://yarnpkg.com/cli/add) for package-manager-native sources.
|
||||
- [`DO_NOT_TRACK`](https://donottrack.sh/) for the environment-level opt-out convention.
|
||||
- [Clack](https://github.com/bombshell-dev/clack) and [Vitest snapshots](https://vitest.dev/guide/snapshot) for injected prompts and generated-file assertions.
|
||||
@@ -0,0 +1,118 @@
|
||||
# Agent Note: SDK 后续功能
|
||||
|
||||
Status: proposed
|
||||
|
||||
[English](2026-07-17-sdk-follow-up-capabilities.md) | 中文
|
||||
|
||||
## 问题
|
||||
|
||||
首个 SDK 版本通过[开发者工程 Agent Note](2026-07-14-sdk-developer-projects.md) 和 [SDK 工程编辑架构](../architecture/2026-07-15-sdk-project-editing-architecture.md)定义的共享模型创建和编辑开发者拥有的 Cordis 工程。create 和 config 工作流仅支持交互调用,接入外部 Cordis 插件需要手工修改依赖和配置,命令行遥测没有明确的所属边界,交互分支也缺少稳定的测试策略。
|
||||
|
||||
这些缺口彼此关联。create 和 config 已经共享问题、功能配置和 `ProjectEditSession`;若另建自动化路径,就会复制领域逻辑。安装外部插件必须同时修改包管理器文件和 `cordis.yml`。遥测需要观察 create、build 等不会启动 Cordis 的命令。交互测试需要覆盖 Harness 自身行为,同时避免把终端渲染固化成脆弱的产品契约。
|
||||
|
||||
## 提案
|
||||
|
||||
SDK 扩展现有提示词与工程编辑边界,不另建平行工作流。非交互式 `PromptPort` 实现和结构化功能计划驱动 create 与 config;`dsh-sdk create <source>` 先把依赖解析交给工程的包管理器,再通过 `ProjectEditSession` 挂载解析所得的包;启动器侧遥测包住 `create-sdk` 和每个 `dsh-sdk` 命令;交互测试主要通过注入的提示词输入输出流完成。
|
||||
|
||||
| 功能 | 产品入口 | 所属机制 | 必须达到的结果 |
|
||||
|---|---|---|---|
|
||||
| Headless 工程创建 | `create-sdk --config <file>` 或 `--config-json <json>`,可搭配 `--json` | `HeadlessPromptPort`、结构化工程答案和完整功能计划 | 不阻塞等待终端;明确报告缺失的必答输入 |
|
||||
| 外部 Cordis 插件安装 | `dsh-sdk create <source>` | 包管理器原生 `add` 加 `ProjectEditSession` | 依赖和 `cordis.yml` 配置项指向包管理器解析出的包 |
|
||||
| 开发周期遥测 | `create-sdk` 和每个 `dsh-sdk` 命令 | 启动器侧的上报条件判断、遥测内容构建、脱敏、匿名身份和传输服务 | 上报采用尽力而为语义,不能改变命令结果 |
|
||||
| 交互回归覆盖 | create 和 config 测试 | 注入的 `PromptPort` 输入输出和文件系统断言 | 测试覆盖 Harness 决策与生成文件,不快照终端重绘 |
|
||||
|
||||
## 共享 headless 工作流
|
||||
|
||||
### 结构化输入和生命周期事件
|
||||
|
||||
Headless create 通过 `--config-json` 接收内联 JSON 对象,或通过 `--config` 从文件读取。标量字段提供普通 create 答案,`features` 提供完整的已选功能、功能选项、secret(密钥)和专用值。只有所属问题明确声明的默认值才有效;headless 路径绝不为必答问题臆造答案。
|
||||
|
||||
使用 `--json` 时,stdout 是 NDJSON 事件流。`done` 表示创建及要求执行的安装和构建均已完成,`action-required` 指明一个尚未回答的必答问题,`error` 报告其他失败。面向人的进度信息和包管理器输出写入 stderr,确保 stdout 每一行都能解析成一个事件。调用方收到 `action-required` 后补充缺失值,再次运行命令。
|
||||
|
||||
Create 和 config 使用相同的功能计划形状。create 通过上述命令行输入公开该形状;config 在共享工作流边界使用同一形状,使后续自动化入口无需另建功能选择模型。
|
||||
|
||||
### Prompt 与工程编辑边界
|
||||
|
||||
`PromptPort` 仍是 SDK 问题与交互实现之间的唯一边界。`ClackPromptPort` 负责终端交互。`HeadlessPromptPort` 使用问题契约公开的默认值,否则通过未回答问题快速失败;预填值通常会让流程根本不调用该 port。
|
||||
|
||||
两条路径使用相同的 `Question` 对象、`FeatureConfigurator`、`SdkProject` 和 `ProjectEditSession`。因此,headless 路径只改变答案的到达方式,不改变功能解释或文件提交方式。
|
||||
|
||||
### Agent skill
|
||||
|
||||
仓库提供一份轻量 `SKILL.md`,指导 agent skill(智能体技能)构造结构化输入、请求 NDJSON、补充 `action-required` 指明的值并重试。该 skill 调用公开 CLI,不导入 SDK 内部 API,也不引入另一套工程规格。
|
||||
|
||||
## 外部 Cordis 插件安装
|
||||
|
||||
`dsh-sdk create <source>` 接受包管理器原生的 npm package specifier,例如 `pkg@version`,也接受 `github:owner/repo#ref` 等 GitHub package specifier。用户确认后,命令要求工程包管理器添加来源,对比操作前后的直接依赖名,重新打开工程,再通过 `ProjectEditSession` 把每个新增且已解析的包挂载进 `cordis.yml`。
|
||||
|
||||
包管理器负责来源解析、版本或 commit 解析、`integrity` 数据、lockfile 更新和构建策略。SDK 不再通过 giget 或 pacote 下载、解压第二份副本。外部插件是 `node_modules` 下的依赖;本地插件脚手架仍属于独立的工程创建问题。
|
||||
|
||||
## Launcher 遥测
|
||||
|
||||
### Consent 与采集
|
||||
|
||||
遥测包住 `create-sdk` 初始化命令与 `dsh-sdk` launcher 的命令生命周期,因为工程初始化、插件创建和 build 都不会稳定地启动 Cordis。每个事件记录命令名、时长、成败、随机生成的用户级匿名标识符,以及符合条件时经过脱敏的 `cordis.yml` 与 `package.json` 文本。
|
||||
|
||||
除非当前存在的遥测配置项被明确禁用,否则允许上报。`DO_NOT_TRACK` 和 CI 无论工程配置如何都禁止上报。缺少 `cordis.yml` 本身不会禁止事件,但只有 `cordis.yml` 能证明目录是 SDK 工程时,遥测内容才包含 `package.json` 文本。
|
||||
|
||||
### 安全与传输
|
||||
|
||||
Payload 构建器绝不读取 `.env`。它会脱敏两个符合条件的文本文件中的疑似密钥键和值、已知 token 形式、PEM 块、URL 凭据和高熵不透明字符串。脱敏只是安全兜底,不能提供绝对保证;SDK 工程必须把凭据放进 `.env`。
|
||||
|
||||
`TelemetryReporter` 使用固定 endpoint,每条发送路径都会正常结束且不抛错。命令分发通过 `finally` 路径记录成败,在命令结果已确定后启动上报,并在有界时间内等待传输结束。只有遥测边界会吞掉上报条件解析、遥测内容构建、存储或网络错误,这些错误绝不改变命令退出码。
|
||||
|
||||
## 交互工作流测试
|
||||
|
||||
Create 和 config 测试向现有工作流注入 `PromptPort` 和脚本化输入输出流。参数化场景覆盖功能选择、功能选项、secret、取消、评审和应用行为,再断言最终的 `cordis.yml` 及其他工程文件。稳定的产品断言是生成后的工程状态,不是 clack 的 ANSI 重绘序列。
|
||||
|
||||
可以用一到两个可选的真实 PTY 冒烟测试覆盖注入无法复现的发布二进制和 TTY 检查。除非原生 PTY 工具在仓库支持的 Node 与宿主版本上足够可靠,否则它不进入必跑路径。
|
||||
|
||||
## 延后工作
|
||||
|
||||
- 扩展 headless create 规格,使其能表达本地 `plugin` 或 `tool` 脚手架,而不是把该交互选择默认为 none。
|
||||
- 在 create 和 config 中公开遥测关闭选项,同时保留只有禁用时才写入遥测配置项的上报许可表示。
|
||||
- 明确 GitHub 来源依赖必须预先构建,还是允许运行由包管理器控制的 preparation script(准备脚本),并在安装前向用户展示该策略。
|
||||
- 发布前把遥测包中的 `.invalid` endpoint 占位符替换为生产端点。
|
||||
|
||||
## 曾考虑的替代方案
|
||||
|
||||
**另建 headless 创建引擎。** 该方案会复制问题、功能依赖、配置行为和工程编辑规则。复用提示词与编辑会话边界,可以保证工程语义只有一份实现。
|
||||
|
||||
**把规格文件作为主要自动化接口。** Agent 可以内联传入相同的类型化 JSON 对象,人和 CI 仍可选用文件。文件专用协议会增加持久化与清理工作,却不增加语义。
|
||||
|
||||
**使用 `npx skills add` 创建工程。** Skills CLI 只安装 Markdown skill,不创建 SDK 工程,也不安装 npm 包。因此,agent skill 驱动 SDK 初始化命令,而不是取代它。
|
||||
|
||||
**通过 giget 或 pacote 获取 GitHub 与 npm 来源。** 第二套获取层会复制包管理器的解析、完整性、lockfile 和生命周期策略。原生 package specifier 让这些决策留在所选包管理器中。
|
||||
|
||||
**把遥测实现成 Cordis 运行时插件。** Create 和 build 不一定启动 Cordis,因此运行时插件无法观察完整的开发命令周期。Launcher 是这些命令共用的边界。
|
||||
|
||||
**从 git 元数据派生匿名标识符。** 仓库的 git remote 可能识别工程或组织。随机的用户级标识符能够支持聚合,同时不编码仓库身份。
|
||||
|
||||
**只采集聚合计数。** 仅聚合事件可以降低暴露,但无法回答开发者实际使用哪些插件、依赖和配置形状。本提案接受采集脱敏后的工程文本,并明确记录这项暴露。
|
||||
|
||||
**把真实 PTY 和 transcript(文本记录)快照作为主要测试策略。** 原生 PTY 依赖与终端重绘序列会带来平台和渲染不稳定性,而且主要是在测试 clack。注入交互并断言生成文件,可以直接测试 SDK 拥有的行为。
|
||||
|
||||
## 验收标准
|
||||
|
||||
- Create 能依据完整结构化输入在没有 TTY 时运行;使用 `--json` 时 stdout 只输出 NDJSON;缺少必答输入时通过 `action-required` 报告,且不写入部分工程。
|
||||
- Create 和 config 通过共享的问题、功能配置和工程编辑代码路径解析相同的功能计划契约。
|
||||
- `dsh-sdk create <source>` 使用工程选定的包管理器,挂载该操作实际新增的依赖名;无法识别新增依赖时快速失败。
|
||||
- 初始化命令与每个 `dsh-sdk` 命令都进入同一条尽力而为的遥测收尾路径;明确禁用的配置项、`DO_NOT_TRACK` 或 CI 会阻止传输,遥测失败绝不改变命令结果。
|
||||
- 遥测绝不读取 `.env`;没有 `cordis.yml` 时不发送无关的 `package.json` 内容;两个符合条件的文本都经过脱敏;匿名标识符与 git 元数据无关。
|
||||
- 交互测试通过注入交互覆盖 create 和 config 决策,并断言已提交的工程文件;真实 PTY 覆盖只作为窄范围冒烟层。
|
||||
- Agent skill 说明公开的结构化输入与事件契约,不依赖包的私有导出。
|
||||
|
||||
## 风险
|
||||
|
||||
- 即使经过脱敏,完整的 `cordis.yml` 与 `package.json` 文本仍会向 endpoint 运营方暴露插件名、依赖名、URL、路径和配置值;启发式脱敏也可能漏掉 secret。
|
||||
- 没有遥测配置项时默认上报可能让开发者意外;发布前 CLI 必须让关闭方法易于发现。
|
||||
- 在 `ProjectEditSession` 挂载插件前,包管理器的 add 操作已经可能修改 `package.json`、lockfile 和安装文件;后续挂载失败会留下需要手工恢复的依赖改动。
|
||||
- GitHub 依赖可能按包管理器策略执行 preparation 或 lifecycle script;尚未解决的构建策略会带来供应链与可复现性风险。
|
||||
- 注入提示词交互的测试无法证明真实终端中的 raw mode、signal 或重绘行为;可选冒烟层只应覆盖这些残余契约。
|
||||
|
||||
## 参考资料
|
||||
|
||||
- [Vercel Eve](https://github.com/vercel/eve) 与 [Vercel Labs Skills](https://github.com/vercel-labs/skills) 用于区分 headless 初始化命令与 skill 分发。
|
||||
- [npm package specifications](https://docs.npmjs.com/cli/v11/using-npm/package-spec)、[pnpm add](https://pnpm.io/cli/add)和 [Yarn add](https://yarnpkg.com/cli/add)说明包管理器原生来源。
|
||||
- [`DO_NOT_TRACK`](https://donottrack.sh/)定义环境级关闭约定。
|
||||
- [Clack](https://github.com/bombshell-dev/clack) 和 [Vitest snapshots](https://vitest.dev/guide/snapshot) 说明注入提示词交互与生成文件断言。
|
||||
Reference in New Issue
Block a user