diff --git a/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.i18n.yaml b/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.i18n.yaml index d1977c65cc..2282e1dace 100644 --- a/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.i18n.yaml +++ b/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.i18n.yaml @@ -2,5 +2,5 @@ # side as of the last confirmed-consistent state. Both languages carry equal authority; # after editing either side, bring the other along and re-record with: # pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md -2026-07-31-code-mode-language-dispatch.md: b999150ae478eef5396e5456e33ffb041f1b161d -2026-07-31-code-mode-language-dispatch.zh.md: 12ef8197e64e9e8a435f852168ab791029534e7d +2026-07-31-code-mode-language-dispatch.md: bc56736c1582b89b4c16b76c49762eeaf0c3fc39 +2026-07-31-code-mode-language-dispatch.zh.md: 9a224cbda75ce530f18498a8e1b0ca42ab540ee1 diff --git a/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md b/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md index b999150ae4..bc56736c15 100644 --- a/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md +++ b/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md @@ -42,3 +42,5 @@ Adding a backend language is two table entries — an `SDK_RENDERERS` entry and The cost is that the Python branch of both tables is unreachable on this base: `CodeRuntime.language` is set by the loaded backend, the only published backend is `dsh-code-runtime-worker` (`'typescript'`), and the registry reads the loaded runtime rather than a config field, so no assembled application can select `renderToolsSdkPy` or `PYTHON_FLAVOR`. The model-visible surface is therefore unchanged by this note's work until a backend reporting `'python'` is published, and this PR's coverage is unit-level — the renderer output plus the dispatch and rejection paths. The keyless snapshot for the Python model interface belongs to the PR that publishes that backend, because only there does a real `cordis.yml` over published plugins produce a Python assembly; a snapshot example that mounted a fixture runtime here would assert against a test double, which [docs/testing.md](../../../../docs/testing.md) rejects as a substitute for the assembled application transcript. Two runtime contracts the Python SDK text asserts are owed by that same backend PR. First, the instructions tell the model that exactly `tools` and `ToolCallError` are bound and that the declared `TypedDict` classes are not, so the backend must inject those two names — with `ToolCallError.toolName` populated per the seam's `errorClass` contract — and must NOT bind the declared class names into the program's globals; injecting them "helpfully" would make the SDK text false. Second, the language has to be bound to the request: `requireCodeRuntime` resolves `ctx.codeRuntime` separately at assembly and at `run_code` execution, so a reload that swapped the runtime between those two points would hand a program written against one flavor to the other. The split is finer than those two points — `run_code`'s `description` and `parameters` getters each call `resolveFlavor(peekRuntime())`, and `schemaOf` destructures both, so one projection reads the runtime twice; both reads are for `run_code`'s own schema, since the getters are installed on that one definition and every other definition carries plain data properties. A reload between those two reads yields a single schema whose two halves name different languages. Neither is reachable here — one published backend means both reads return the same flavor and no program ever runs against this renderer's output — and the cross-language rejection is not testable until a second language exists. + +Third, that PR owns the CPython floor, and with it the Unicode-table skew in `isBareIdentifier`. This renderer decides whether a field or tool name can be emitted bare using the running engine's `\p{XID_Start}`/`\p{XID_Continue}` tables (Node 22.23.1: Unicode 17.0), while the interpreter uses its own (CPython 3.9.6: 13.0.0). An interpreter older than the engine is the failing direction: a character added to `XID_Start` in between is emitted bare and its tokenizer refuses the whole block. The exposure window is exactly the characters added between the two versions, so the PR that names a supported CPython range must decide explicitly between accepting it and tightening the predicate against pinned tables for that floor. Nothing here can decide it: the floor does not exist yet, and a table pinned to a guess would be a deployment-varying constant with no configurability behind it. diff --git a/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.zh.md b/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.zh.md index 12ef8197e6..9a224cbda7 100644 --- a/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.zh.md +++ b/.agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.zh.md @@ -42,3 +42,5 @@ Code Mode 只生成一种 SDK 形态:TypeScript。`ToolRegistry` 为 `tools:sd 代价是两张表的 Python 分支在当前 base 上不可达:`CodeRuntime.language` 由所加载的后端设定,已发布的后端只有 `dsh-code-runtime-worker`(`'typescript'`),而注册表读取的是所加载的运行时而非某个配置字段,因此没有任何一份组装好的应用能选中 `renderToolsSdkPy` 或 `PYTHON_FLAVOR`。也就是说,在报告 `'python'` 的后端发布之前,本 note 的工作不改变模型可见表面,本 PR 的覆盖因此是 unit 级——渲染器输出加分发与拒绝路径。Python 模型界面的 keyless snapshot 归属于发布该后端的那个 PR,因为只有在那里,一份基于已发布插件的真实 `cordis.yml` 才会产出 Python 组装;在此处挂载 fixture 运行时的快照示例断言的是测试替身,而 [docs/testing.md](../../../../docs/testing.md) 明确拒绝以此替代组装好的应用 transcript。 Python SDK 文本断言的两条运行时契约同样归属那个 backend PR。其一,说明文字告诉模型运行时恰好绑定 `tools` 与 `ToolCallError` 两个名字、所声明的 `TypedDict` 类不绑定,因此后端必须注入这两个名字(并按 seam 的 `errorClass` 契约填充 `ToolCallError.toolName`),且**不得**把所声明的类名绑进程序全局——「好心」注入会使这段 SDK 文本变成假话。其二,语言必须绑定到请求上:`requireCodeRuntime` 在组装时与 `run_code` 执行时分别解析 `ctx.codeRuntime`,若在这两点之间发生重载并换掉运行时,就会把针对一种形态写成的程序交给另一种形态执行。分裂比这两点更细——`run_code` 的 `description` 与 `parameters` 两个 getter 各自调用 `resolveFlavor(peekRuntime())`,而 `schemaOf` 会解构这两个字段,因此一次投影读两次运行时;两次都属于 `run_code` 自己的 schema,因为这两个 getter 只装在那一个 definition 上,其余 definition 携带的都是普通数据属性。在这两次读取之间重载会产出单个 schema 的两半分属不同语言。两者在此处都不可达——只有一个已发布后端意味着两次读取返回同一形态,且没有任何程序会针对本渲染器的输出运行——而跨语言拒绝在第二门语言存在之前也无法测试。 + +其三,那个 PR 拥有 CPython 版本下限,连带拥有 `isBareIdentifier` 里的 Unicode 表偏斜。本渲染器用所运行引擎的 `\p{XID_Start}`/`\p{XID_Continue}` 表(Node 22.23.1:Unicode 17.0)决定某个字段名或工具名能否裸发,而解释器用它自己的表(CPython 3.9.6:13.0.0)。解释器旧于引擎是会失败的那个方向:在两者之间被加进 `XID_Start` 的字符会被裸发,其 tokenizer 拒收,整个块随之不可解析。暴露窗口恰是两个版本之间新增的那些字符,所以宣布支持某个 CPython 范围的那个 PR 必须在「接受该暴露」与「按该下限的固定表收紧判据」之间显式作出决定。此处无法决定:下限尚不存在,而按猜测钉死一张表会成为一个随部署而变、却没有可配置性支撑的常量。 diff --git a/packages/core/tools/src/py-types.ts b/packages/core/tools/src/py-types.ts index 315afa5aa6..9ed7d75d4d 100644 --- a/packages/core/tools/src/py-types.ts +++ b/packages/core/tools/src/py-types.ts @@ -17,7 +17,11 @@ import { assertSupportedJsonSchema } from './json-schema.ts' import type { JsonSchemaNode, JsonSchemaScalar } from './json-schema.ts' import type { ToolSdkSchema } from './ts-types.ts' -/** The reference grammar's `xid_start xid_continue*`, the same set `str.isidentifier()` accepts. */ +/** + * The reference grammar's `xid_start xid_continue*` — the set + * `str.isidentifier()` accepts on a CPython whose Unicode tables match the + * engine's. See {@link isBareIdentifier} for what a version skew does. + */ const IDENTIFIER = /^[\p{XID_Start}_]\p{XID_Continue}*$/u /** @@ -36,6 +40,24 @@ const IDENTIFIER = /^[\p{XID_Start}_]\p{XID_Continue}*$/u * that normalize together would collapse into one declaration. Those names * take the subscript path, which carries their exact bytes. * + * Both conditions are evaluated against the ENGINE's Unicode tables, and the + * two sides are versioned independently — `\p{XID_Start}` follows the running + * engine (Node 22.23.1 reports Unicode 17.0) while CPython follows its own + * (3.9.6 reports 13.0.0). The skew is not symmetric. A CPython older than the + * engine is the dangerous direction: a character added to `XID_Start` since its + * tables (U+1C89, U+10570, U+1E290, U+1E4D0 are all NFKC-stable and accepted + * here, and all rejected by that 3.9.6) is emitted bare and its tokenizer + * refuses the character, taking the whole SDK block down — the same + * parseability invariant {@link UNPRINTABLE}, {@link LONE_SURROGATE} and + * {@link MAX_LIST_NESTING} exist for. A CPython newer than the engine only + * routes a legal name to the subscript path: less readable, still correct. The + * NFKC condition reduces to the same skew, since normalization stability + * guarantees an assigned character's normalization never changes afterwards. + * + * Closing the exposure needs the target interpreter's version, which the + * backend reporting `language: 'python'` owns and which is unpublished on this + * base; the note records it as that PR's decision. + * * The `ts-types` sibling keeps its own ASCII rule rather than sharing this * one: ECMAScript identifiers are a different set (`$`, ZWJ/ZWNJ) and are * never normalized, so one predicate cannot be correct for both. @@ -186,10 +208,19 @@ function docLines(description: unknown, indent: number): string[] { * split words, `_` splits too (it is `XID_Continue`, so the split set names it * explicitly), and a head that cannot start an identifier takes a `Tool` * prefix. Unicode survives, so a `路径` field yields `路径`-based class names - * instead of collapsing to the bare prefix. The result is NFKC-normalized: - * these names are generated, never matched against a JSON key, so normalizing - * is free here and keeps what CPython compiles identical to what is emitted — - * unlike {@link isBareIdentifier}, which must reject unstable names outright. + * instead of collapsing to the bare prefix. A character that is not + * `XID_Continue` splits even when it is a letter, so a name whose NFKC folding + * would leave the identifier set is not carried through — the split set is the + * grammar's, not an ASCII approximation of it. + * + * The result is NFKC-normalized: these names are generated, never matched + * against a JSON key, so normalizing is free here and keeps what CPython + * compiles identical to what is emitted — unlike {@link isBareIdentifier}, + * which must reject unstable names outright. Normalizing AFTER the prefix + * decision is what makes that hold at the seam the prefix creates: `Tool` + + * a combining-mark head composes there (`U+0301` gives `Tooĺ`, U+013A), so + * normalizing only the un-prefixed part would emit a name CPython compiles to + * a different symbol. The second call is idempotent on the un-prefixed arm. * @param raw - the schema field or tool name to derive from. * @returns a class-name segment safe to emit. */ @@ -200,7 +231,7 @@ function camelCase(raw: string): string { .map(part => `${part.charAt(0).toUpperCase()}${part.slice(1)}`) .join('') .normalize('NFKC') - return /^\p{XID_Start}/u.test(joined) ? joined : `Tool${joined}` + return (/^\p{XID_Start}/u.test(joined) ? joined : `Tool${joined}`).normalize('NFKC') } /** Class-name base cap keeping each emitted name — and total text — linear in schema depth. */ @@ -291,9 +322,19 @@ function allocateClassName(base: string, state: RenderState): string { * object-chain would otherwise carry an ever-growing ConsString down the tree * and re-materialize it (via `.length`/`.slice`) at every level — Θ(depth²). * The bounded base plus the collision counter still yields unique names. + * + * The join is NFKC-normalized because both sides are separately normalized yet + * their concatenation need not be: a base ending in a Hangul L jamo or LV + * syllable composes with a following V or T jamo head (`가` + `ᆨ` gives `각`), + * so the emitted class name would differ from the symbol CPython compiles, and + * two byte-distinct names could fold onto one — `usedClassNames` dedupes by the + * raw bytes, so the collision counter would not see it. Normalizing costs + * O(cap + segment) per level, the same order as the `slice` it feeds. The other + * two join points need no counterpart: `Args`/`Output` start with `A`/`O` and + * {@link allocateClassName}'s suffix is digits, none of which compose backwards. */ function childClassName(base: string, segment: string): string { - return capClassNameBase(`${base}${segment}`) + return capClassNameBase(`${base}${segment}`.normalize('NFKC')) } /** diff --git a/packages/core/tools/tests/py-types.spec.ts b/packages/core/tools/tests/py-types.spec.ts index ca5ca40ce8..2b73ec5795 100644 --- a/packages/core/tools/tests/py-types.spec.ts +++ b/packages/core/tools/tests/py-types.spec.ts @@ -494,8 +494,69 @@ describe('renderToolsSdkPy', () => { // whole characters; one ASCII character of padding puts the boundary inside // the 60th pair, and that half is dropped rather than emitted. expect(className('')).toBe(AHSA.repeat(60)) - expect(className('x')).toBe(`X${AHSA.repeat(59)}`) - expect(className('x')).toHaveLength(119) + const padded = className('x') + expect(padded).toBe(`X${AHSA.repeat(59)}`) + expect(padded).toHaveLength(119) + }) + + it('normalizes the seam the Tool prefix creates, which the prefixed part alone does not cover', () => { + // U+0301 COMBINING ACUTE ACCENT is XID_Continue but not XID_Start, so a name + // headed by it takes the `Tool` prefix — and `Tool` ends in `l`, which + // composes with it. Normalizing only the part being prefixed would emit + // `Tool` + U+0301, which CPython compiles as `Too` + U+013A: the class + // the SDK declares would not be the class the interpreter defines. Every + // code point below is an escape — the two forms render identically. + const text = renderToolsSdkPy([ + { + name: '\u0301abc', + description: 'Combining-mark head.', + parameters: { type: 'object', additionalProperties: false, properties: { q: { type: 'string' } }, required: ['q'] }, + output: { type: 'string' }, + }, + ]) + expect(text).toContain('class Too\u013AabcArgs(TypedDict):') + expect(text).toContain('# tools["\u0301abc"](args: Too\u013AabcArgs) -> str') + expect(text).not.toContain('Tool\u0301') + }) + + it('normalizes a class-name join where two separately stable segments compose', () => { + // Hangul jamo compose ACROSS the join `childClassName` makes: the parent + // base ends in U+1100 (L jamo) and the child segment starts with U+1161 (V + // jamo), each NFKC-stable alone, together U+AC00. Unnormalized, the declared + // name differs from the compiled symbol, and two byte-distinct names can + // fold onto one — `usedClassNames` dedupes by raw bytes, so the collision + // counter never sees it and the later declaration shadows the earlier one + // under CPython. Escapes again, for the same reason as above. + const text = renderToolsSdkPy([ + { + name: 'x', + description: 'Jamo field names.', + parameters: { + type: 'object', + additionalProperties: false, + required: ['\uAC00\u1100'], + properties: { + '\uAC00\u1100': { + type: 'object', + additionalProperties: false, + required: ['\u1161x'], + properties: { + '\u1161x': { type: 'object', additionalProperties: false, properties: { q: { type: 'string' } } }, + }, + }, + }, + }, + output: { type: 'string' }, + }, + ]) + // The join is `XArgs` + U+AC00 U+1100 followed by U+1161 `x`, whose + // trailing L+V pair composes into a second U+AC00. + expect(text).toContain('class XArgs\uAC00\uAC00x(TypedDict):') + expect(text).toContain(' \u1161x: XArgs\uAC00\uAC00x') + expect(text).not.toContain('\u1100\u1161') + // The level above it is a join that composes nothing (LV + L), so it stays + // byte-identical — normalizing is not silently rewriting every name. + expect(text).toContain('class XArgs\uAC00\u1100(TypedDict):') }) it('declares a closed empty object with omitted properties as an empty TypedDict, not dict[str, Any]', () => {