Files
deepseek-harness/docs/rfc/implemented/architecture/2026-06-20-branded-ids.md
Tianyi Cui 83e97ed222 fix review findings: document the util/ group + align branded-ids RFC with dsh-brand
Adding packages/util/brand/ created a new top-level packages/util/ group that
the hierarchy/dependency docs never enumerated. Document it:

- Add packages/util/README.md, the group README (low-level zero-dependency
  utilities shared across groups; lists dsh-brand).
- packages/README.md: add the util/ group to the group table, dsh-brand to the
  package table, and dsh-brand to the dependency graph. Correct the now-false
  "no harness deps" claims — dsh-llm and dsh-bash both depend on dsh-brand
  (verified dsh-bash imports Branded from dsh-brand, not dsh-llm; dsh-session
  and dsh-agent depend on it too).
- Root AGENTS.md Repository Layout: add the util/ group with brand/.

Align the implemented branded-ids RFC with what shipped: Branded lives in
@deepseek-ai/dsh-brand (packages/util/brand/), and dsh-bash depends only on
that utility package instead of dsh-llm. Fix the BashTaskId import source, the
illustrative snippet, and the opening policy reference (now dsh-brand).
2026-06-21 11:04:28 +08:00

12 KiB

RFC: Branded IDs everywhere they belong

Status: implemented (proposed and accepted 2026-06-20)

Problem

The harness already brands three identifiers — CallId (packages/llm/llm/src/brand.ts), SessionId (packages/core/session/src/types.ts), and AgentId (packages/core/agent/src/types.ts) — using the Branded<B> = string & { readonly [BRAND]: B } machinery (owned by the type-only @deepseek-ai/dsh-brand package at packages/util/brand/ — see its README) and a zero-cost cast factory per type. dsh-brand also states the governing policy: "Branding is for ids that cross package boundaries and could plausibly be confused; not every string needs a brand." That policy is right; the problem is that it is only half-applied. Two gaps let a structurally-identical-but-semantically-wrong string slip through the type checker today.

Gap 1 — unbranded cross-boundary IDs in the bash seam. The background-task id is a plain string: BashTask.id: string (packages/bash/bash/src/types.ts), carried as string through the whole executor seam (BashExecutor.get/ownerOf/readOutput/kill(id: string) in packages/bash/bash/src/index.ts) and validated/passed as string by the model-facing tools (validateTaskId, assertTaskAccess, the task_id schema arg in packages/bash/tool-bash/src/index.ts). It is generated by a per-executor counter — `bash-${this.nextTaskId++}` in packages/bash/bash-local/src/index.ts — which gives it exactly the same name-N shape as SessionId's default (`session-${++counter}` in packages/core/session/src/index.ts). A bash task id and a session id are trivially swappable at a call site and the compiler says nothing. This is the headline case the user asked about, and it is a model-facing id (the model passes task_id back to bash_output/bash_kill), so a confusion here is reachable from untrusted input.

The bash owner token is the related sub-case: BashExecRequest.owner?: string and BashExecSpec.owner: string | undefined (packages/bash/bash/src/types.ts) are documented as a deliberately opaque isolation key, but in every live caller the value IS the owning agent's session.header.id (callerToken = (exec) => exec.agent?.session.header.id in packages/bash/tool-bash/src/index.ts) — i.e. a SessionId wearing a string disguise. It is compared for access control (owner !== callerToken(exec)), so a mismatched-but-well-typed string here is a cross-session isolation bug the type system currently cannot catch. This is the same session.header.id-as-owner alias that the unify-the-agent-id-and-the-session-id proposal calls the "bash owner-token alias hole".

Gap 2 — brand erosion at the seams of the already-branded IDs. Even CallId/SessionId/AgentId decay back to bare string at exactly the places confusion is most likely: the registry/store Map key types and most public method params. Representative sites: SessionStore.store = new Map<string, Session>() and create/prepare(id?: string)/get(id: string) (packages/core/session/src/index.ts); AgentRegistry.store = new Map<string, Agent>() and register/get(id: string) (packages/core/agent/src/index.ts); ToolPresenter.pending = new Map<string, …>() keyed by call id and call(callId: string)/result(callId: string) (packages/ui/acp/src/index.ts); the ACP session-id surface beyond the store map — SessionRecord.sessionId: string, bySession = new WeakMap<Agent, string>(), loadingIds = new Set<string>(), requireSession(sessionId: string), and the exported streamSessionEventUpdate(sessionId: string, …) (packages/ui/acp/src/index.ts); and the persistence coordinator's Map<string, …> keyed by session id (packages/session-persistence/session-persistence/src/coordinator.ts). A brand that is dropped at the Map key buys nothing on lookups — the value of the existing brands is partly unrealized.

Proposal

A type-only change. Brands are zero-cost casts; nothing about runtime behavior, serialization, comparison, or the wire format changes. The work is in three parts, all honoring the existing "not every string" policy.

  • Brand the bash task id. Add BashTaskId = Branded<'BashTaskId'> plus its same-named factory in packages/bash/bash/src/types.ts (the package that owns the id), importing Branded from @deepseek-ai/dsh-brand exactly as SessionId/AgentId already do. The brand primitive lives in the dependency-free dsh-brand utility package precisely so dsh-bash can brand its ids by depending on it alone — it never pulls in dsh-llm (or dsh-session) just to reach Branded. Thread it through BashTask.id, the BashExecutor seam methods (get/ownerOf/readOutput/kill), the generation site in dsh-bash-local (brand the counter output once, at creation), and the dsh-tool-bash validate/access surface (validateTaskId returns a BashTaskId; task_id is branded at the tool boundary where the model's string arrives).

  • Mint a distinct OwnerToken brand. Add OwnerToken = Branded<'OwnerToken'> in packages/bash/bash/src/types.ts; type BashExecRequest.owner / BashExecSpec.owner / BashExecutor.ownerOf as OwnerToken | undefined. The dsh-tool-bash consumer casts the agent's session.header.id (a SessionId) into an OwnerToken at the boundary — the one place the two vocabularies meet. The bash seam never imports dsh-session. (Rationale in the next section.)

  • Stop the brand erosion. Propagate the existing brands to the Map key types and public method params listed under Gap 2 — Map<SessionId, Session>, get(id: SessionId), Map<AgentId, Agent>, Map<CallId, …>, the ACP SessionRecord.sessionId: SessionId surface, the coordinator's Map<SessionId, …>. This is the larger mechanical share of the diff and the part that makes the existing brands actually load-bearing on lookups, not just on the struct fields.

Illustrative shape (the factory pattern is identical to the three existing brands):

import type { Branded } from '@deepseek-ai/dsh-brand'

/** A background bash task handle (generated `bash-N` by the local executor). */
export type BashTaskId = Branded<'BashTaskId'>
export function BashTaskId(id: string): BashTaskId {
  return id as BashTaskId
}

/** A bash task's opaque isolation key — the consumer's owner identity, NOT the bash seam's. */
export type OwnerToken = Branded<'OwnerToken'>
export function OwnerToken(id: string): OwnerToken {
  return id as OwnerToken
}

Why a distinct OwnerToken brand (not SessionId)

The obvious shortcut is to type owner as SessionId directly — it always is one. We reject that. The bash executor seam is a capability seam (interface dsh-bash, implementation dsh-bash-local, consumer dsh-tool-bash) and its owner token is documented as deliberately opaque: the executor "never interprets it (no access policy lives in the seam — that is the consumer's job)" (packages/bash/bash/src/types.ts). Typing the seam's field as SessionId would import dsh-session's vocabulary into a package that must not know what an owner token means — it would couple a generic execution backend to the session model and contradict the opaque-token design. A sandboxed or remote executor that replaces dsh-bash-local should not inherit a session dependency. The distinct OwnerToken brand keeps the seam decoupled: dsh-bash knows only "an owner is some opaque branded token," and the dsh-tool-bash consumer — which already decides the access policy — is the single boundary that casts its SessionId into an OwnerToken. The brand still delivers the safety win (you cannot pass a BashTaskId or a raw string where an owner is expected) without the coupling.

Out of scope / possible extensions

Kept deliberately narrow per the "not every string needs a brand" policy. Each of these is a plausible future brand, deferred with a reason, not a commitment:

  • ModelId (GenerateOptions.model, the LlmService adapter-registry key) — a real cross-package lookup key (config → agent → llm → adapter); a reasonable next brand, left out only to keep this RFC's blast radius focused.
  • ToolName (the ToolRegistry key) — author-defined, human-readable, and rarely confused with another id; the weakest candidate, likely not worth a brand.
  • ErrorCode (HarnessError.code) — a closed vocabulary (ABORTED, NO_ADAPTER, …), not a per-instance id; better served by a string-literal union than a brand, if anything.
  • Numeric ordinals — turn number, step number, and the event seq are number, not string, so Branded<string> does not apply; a parallel number & { readonly [BRAND]: B } variant could brand them, but they are positional ordinals rarely passed across boundaries, so the payoff is low.
  • Validated construction — the brand factories are pure casts with no runtime check, and every boundary (ACP sessionId, provider-issued call.id, the empty-string fallback in dsh-llm-deepseek) trusts the raw string today. A SessionId.parse() / isValid() companion that throws on malformed input at boundaries is a genuine gap, but it is a runtime-behavior change with its own design (what is "malformed"? what do we do on failure?) and belongs in its own RFC, not bundled into this type-only pass.

Acceptance criteria

  • BashTaskId and OwnerToken are defined in dsh-bash and threaded end-to-end: the executor seam, the dsh-bash-local generation site, and the dsh-tool-bash model-facing surface all speak the brands; dsh-bash gains no dependency on dsh-session.
  • No collection keyed by an in-scope branded id (CallId/SessionId/AgentId/BashTaskId) is keyed by bare string — this covers Map, WeakMap value slots, and Set membership (e.g. the ACP bySession/loadingIds), not just Map<string, …>; the corresponding public method params and exported function signatures (e.g. streamSessionEventUpdate) take the brand, not string.
  • Brands are constructed via the cast factory at each boundary where a raw string enters (provider call id, ACP session id, model-supplied task_id); no as casts scattered at call sites.
  • pnpm run typecheck and pnpm run doc-sync are green; the change is observably type-only (no snapshot, no e2e behavioral diff).

Risks / what we give up

  • Mechanical churn across two surfaces. Propagating brands touches the bash seam (interface + impl + consumer) and the ACP session-id surface plus the persistence coordinator. The risk is broad but low-severity: a missed site is a compile error, not a silent bug. It ships as its own PR, converged with Codex, and stacks naturally near the unify-the-agent-id-and-the-session-id work (both touch the session-id / owner-token boundary; if that proposal lands first, OwnerToken still stays distinct from the unified id for the decoupling reason above).
  • Brands do not validate. A brand is a confusability guard, not a correctness proof: a wrong session id that is still a well-formed string passes the type checker exactly as before. This RFC does not close that gap (see Out of scope) — it only stops the category error of passing the wrong kind of id.
  • The "where to stop" line stays a judgment call. Branding BashTaskId but not ToolName, OwnerToken but not ModelId, is a taste call about which strings "could plausibly be confused." Reasonable reviewers may want more or fewer; the policy in brand.ts is the tie-breaker, and this RFC errs toward the ids that are model-facing or used for access control.