fix(tools): normalize the two class-name joins camelCase's own call misses
camelCase normalized `joined` and then prefixed, so the seam the `Tool` prefix creates was never covered: `Tool` ends in `l`, a combining-mark head composes with it, and a name headed by U+0301 was emitted as `Tool` + U+0301 while CPython compiles `Too` + U+013A. childClassName has the same shape -- both sides separately NFKC-stable, their join not: a base ending in a Hangul L jamo or LV syllable composes with a V or T jamo head. Beyond the declared-name/compiled-symbol mismatch, two byte-distinct names can fold onto one, and usedClassNames dedupes by raw bytes, so the collision counter never sees it. Normalize after the prefix decision and at the join, before the cap. The remaining joins need nothing: `Args`/`Output` and the digit suffix cannot compose backwards. Also record the Unicode-table skew. The predicate reads the engine's tables (Node 22.23.1: 17.0) and the interpreter reads its own (CPython 3.9.6: 13.0.0), so an interpreter older than the engine takes a bare name its tokenizer refuses -- U+1C89, U+10570, U+1E290 and U+1E4D0 are accepted here and rejected there. The other direction only degrades a legal name to subscript. Closing it needs the CPython floor, which the backend PR owns; state the asymmetry in the docstring and make the decision an explicit obligation in the note.
This commit is contained in:
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-31-code-mode-language-dispatch.md
|
||||
2026-07-31-code-mode-language-dispatch.md: b999150ae478eef5396e5456e33ffb041f1b161d
|
||||
2026-07-31-code-mode-language-dispatch.zh.md: 12ef8197e64e9e8a435f852168ab791029534e7d
|
||||
2026-07-31-code-mode-language-dispatch.md: bc56736c1582b89b4c16b76c49762eeaf0c3fc39
|
||||
2026-07-31-code-mode-language-dispatch.zh.md: 9a224cbda75ce530f18498a8e1b0ca42ab540ee1
|
||||
|
||||
@@ -42,3 +42,5 @@ Adding a backend language is two table entries — an `SDK_RENDERERS` entry and
|
||||
The cost is that the Python branch of both tables is unreachable on this base: `CodeRuntime.language` is set by the loaded backend, the only published backend is `dsh-code-runtime-worker` (`'typescript'`), and the registry reads the loaded runtime rather than a config field, so no assembled application can select `renderToolsSdkPy` or `PYTHON_FLAVOR`. The model-visible surface is therefore unchanged by this note's work until a backend reporting `'python'` is published, and this PR's coverage is unit-level — the renderer output plus the dispatch and rejection paths. The keyless snapshot for the Python model interface belongs to the PR that publishes that backend, because only there does a real `cordis.yml` over published plugins produce a Python assembly; a snapshot example that mounted a fixture runtime here would assert against a test double, which [docs/testing.md](../../../../docs/testing.md) rejects as a substitute for the assembled application transcript.
|
||||
|
||||
Two runtime contracts the Python SDK text asserts are owed by that same backend PR. First, the instructions tell the model that exactly `tools` and `ToolCallError` are bound and that the declared `TypedDict` classes are not, so the backend must inject those two names — with `ToolCallError.toolName` populated per the seam's `errorClass` contract — and must NOT bind the declared class names into the program's globals; injecting them "helpfully" would make the SDK text false. Second, the language has to be bound to the request: `requireCodeRuntime` resolves `ctx.codeRuntime` separately at assembly and at `run_code` execution, so a reload that swapped the runtime between those two points would hand a program written against one flavor to the other. The split is finer than those two points — `run_code`'s `description` and `parameters` getters each call `resolveFlavor(peekRuntime())`, and `schemaOf` destructures both, so one projection reads the runtime twice; both reads are for `run_code`'s own schema, since the getters are installed on that one definition and every other definition carries plain data properties. A reload between those two reads yields a single schema whose two halves name different languages. Neither is reachable here — one published backend means both reads return the same flavor and no program ever runs against this renderer's output — and the cross-language rejection is not testable until a second language exists.
|
||||
|
||||
Third, that PR owns the CPython floor, and with it the Unicode-table skew in `isBareIdentifier`. This renderer decides whether a field or tool name can be emitted bare using the running engine's `\p{XID_Start}`/`\p{XID_Continue}` tables (Node 22.23.1: Unicode 17.0), while the interpreter uses its own (CPython 3.9.6: 13.0.0). An interpreter older than the engine is the failing direction: a character added to `XID_Start` in between is emitted bare and its tokenizer refuses the whole block. The exposure window is exactly the characters added between the two versions, so the PR that names a supported CPython range must decide explicitly between accepting it and tightening the predicate against pinned tables for that floor. Nothing here can decide it: the floor does not exist yet, and a table pinned to a guess would be a deployment-varying constant with no configurability behind it.
|
||||
|
||||
@@ -42,3 +42,5 @@ Code Mode 只生成一种 SDK 形态:TypeScript。`ToolRegistry` 为 `tools:sd
|
||||
代价是两张表的 Python 分支在当前 base 上不可达:`CodeRuntime.language` 由所加载的后端设定,已发布的后端只有 `dsh-code-runtime-worker`(`'typescript'`),而注册表读取的是所加载的运行时而非某个配置字段,因此没有任何一份组装好的应用能选中 `renderToolsSdkPy` 或 `PYTHON_FLAVOR`。也就是说,在报告 `'python'` 的后端发布之前,本 note 的工作不改变模型可见表面,本 PR 的覆盖因此是 unit 级——渲染器输出加分发与拒绝路径。Python 模型界面的 keyless snapshot 归属于发布该后端的那个 PR,因为只有在那里,一份基于已发布插件的真实 `cordis.yml` 才会产出 Python 组装;在此处挂载 fixture 运行时的快照示例断言的是测试替身,而 [docs/testing.md](../../../../docs/testing.md) 明确拒绝以此替代组装好的应用 transcript。
|
||||
|
||||
Python SDK 文本断言的两条运行时契约同样归属那个 backend PR。其一,说明文字告诉模型运行时恰好绑定 `tools` 与 `ToolCallError` 两个名字、所声明的 `TypedDict` 类不绑定,因此后端必须注入这两个名字(并按 seam 的 `errorClass` 契约填充 `ToolCallError.toolName`),且**不得**把所声明的类名绑进程序全局——「好心」注入会使这段 SDK 文本变成假话。其二,语言必须绑定到请求上:`requireCodeRuntime` 在组装时与 `run_code` 执行时分别解析 `ctx.codeRuntime`,若在这两点之间发生重载并换掉运行时,就会把针对一种形态写成的程序交给另一种形态执行。分裂比这两点更细——`run_code` 的 `description` 与 `parameters` 两个 getter 各自调用 `resolveFlavor(peekRuntime())`,而 `schemaOf` 会解构这两个字段,因此一次投影读两次运行时;两次都属于 `run_code` 自己的 schema,因为这两个 getter 只装在那一个 definition 上,其余 definition 携带的都是普通数据属性。在这两次读取之间重载会产出单个 schema 的两半分属不同语言。两者在此处都不可达——只有一个已发布后端意味着两次读取返回同一形态,且没有任何程序会针对本渲染器的输出运行——而跨语言拒绝在第二门语言存在之前也无法测试。
|
||||
|
||||
其三,那个 PR 拥有 CPython 版本下限,连带拥有 `isBareIdentifier` 里的 Unicode 表偏斜。本渲染器用所运行引擎的 `\p{XID_Start}`/`\p{XID_Continue}` 表(Node 22.23.1:Unicode 17.0)决定某个字段名或工具名能否裸发,而解释器用它自己的表(CPython 3.9.6:13.0.0)。解释器旧于引擎是会失败的那个方向:在两者之间被加进 `XID_Start` 的字符会被裸发,其 tokenizer 拒收,整个块随之不可解析。暴露窗口恰是两个版本之间新增的那些字符,所以宣布支持某个 CPython 范围的那个 PR 必须在「接受该暴露」与「按该下限的固定表收紧判据」之间显式作出决定。此处无法决定:下限尚不存在,而按猜测钉死一张表会成为一个随部署而变、却没有可配置性支撑的常量。
|
||||
|
||||
@@ -17,7 +17,11 @@ import { assertSupportedJsonSchema } from './json-schema.ts'
|
||||
import type { JsonSchemaNode, JsonSchemaScalar } from './json-schema.ts'
|
||||
import type { ToolSdkSchema } from './ts-types.ts'
|
||||
|
||||
/** The reference grammar's `xid_start xid_continue*`, the same set `str.isidentifier()` accepts. */
|
||||
/**
|
||||
* The reference grammar's `xid_start xid_continue*` — the set
|
||||
* `str.isidentifier()` accepts on a CPython whose Unicode tables match the
|
||||
* engine's. See {@link isBareIdentifier} for what a version skew does.
|
||||
*/
|
||||
const IDENTIFIER = /^[\p{XID_Start}_]\p{XID_Continue}*$/u
|
||||
|
||||
/**
|
||||
@@ -36,6 +40,24 @@ const IDENTIFIER = /^[\p{XID_Start}_]\p{XID_Continue}*$/u
|
||||
* that normalize together would collapse into one declaration. Those names
|
||||
* take the subscript path, which carries their exact bytes.
|
||||
*
|
||||
* Both conditions are evaluated against the ENGINE's Unicode tables, and the
|
||||
* two sides are versioned independently — `\p{XID_Start}` follows the running
|
||||
* engine (Node 22.23.1 reports Unicode 17.0) while CPython follows its own
|
||||
* (3.9.6 reports 13.0.0). The skew is not symmetric. A CPython older than the
|
||||
* engine is the dangerous direction: a character added to `XID_Start` since its
|
||||
* tables (U+1C89, U+10570, U+1E290, U+1E4D0 are all NFKC-stable and accepted
|
||||
* here, and all rejected by that 3.9.6) is emitted bare and its tokenizer
|
||||
* refuses the character, taking the whole SDK block down — the same
|
||||
* parseability invariant {@link UNPRINTABLE}, {@link LONE_SURROGATE} and
|
||||
* {@link MAX_LIST_NESTING} exist for. A CPython newer than the engine only
|
||||
* routes a legal name to the subscript path: less readable, still correct. The
|
||||
* NFKC condition reduces to the same skew, since normalization stability
|
||||
* guarantees an assigned character's normalization never changes afterwards.
|
||||
*
|
||||
* Closing the exposure needs the target interpreter's version, which the
|
||||
* backend reporting `language: 'python'` owns and which is unpublished on this
|
||||
* base; the note records it as that PR's decision.
|
||||
*
|
||||
* The `ts-types` sibling keeps its own ASCII rule rather than sharing this
|
||||
* one: ECMAScript identifiers are a different set (`$`, ZWJ/ZWNJ) and are
|
||||
* never normalized, so one predicate cannot be correct for both.
|
||||
@@ -186,10 +208,19 @@ function docLines(description: unknown, indent: number): string[] {
|
||||
* split words, `_` splits too (it is `XID_Continue`, so the split set names it
|
||||
* explicitly), and a head that cannot start an identifier takes a `Tool`
|
||||
* prefix. Unicode survives, so a `路径` field yields `路径`-based class names
|
||||
* instead of collapsing to the bare prefix. The result is NFKC-normalized:
|
||||
* these names are generated, never matched against a JSON key, so normalizing
|
||||
* is free here and keeps what CPython compiles identical to what is emitted —
|
||||
* unlike {@link isBareIdentifier}, which must reject unstable names outright.
|
||||
* instead of collapsing to the bare prefix. A character that is not
|
||||
* `XID_Continue` splits even when it is a letter, so a name whose NFKC folding
|
||||
* would leave the identifier set is not carried through — the split set is the
|
||||
* grammar's, not an ASCII approximation of it.
|
||||
*
|
||||
* The result is NFKC-normalized: these names are generated, never matched
|
||||
* against a JSON key, so normalizing is free here and keeps what CPython
|
||||
* compiles identical to what is emitted — unlike {@link isBareIdentifier},
|
||||
* which must reject unstable names outright. Normalizing AFTER the prefix
|
||||
* decision is what makes that hold at the seam the prefix creates: `Tool` +
|
||||
* a combining-mark head composes there (`U+0301` gives `Tooĺ`, U+013A), so
|
||||
* normalizing only the un-prefixed part would emit a name CPython compiles to
|
||||
* a different symbol. The second call is idempotent on the un-prefixed arm.
|
||||
* @param raw - the schema field or tool name to derive from.
|
||||
* @returns a class-name segment safe to emit.
|
||||
*/
|
||||
@@ -200,7 +231,7 @@ function camelCase(raw: string): string {
|
||||
.map(part => `${part.charAt(0).toUpperCase()}${part.slice(1)}`)
|
||||
.join('')
|
||||
.normalize('NFKC')
|
||||
return /^\p{XID_Start}/u.test(joined) ? joined : `Tool${joined}`
|
||||
return (/^\p{XID_Start}/u.test(joined) ? joined : `Tool${joined}`).normalize('NFKC')
|
||||
}
|
||||
|
||||
/** Class-name base cap keeping each emitted name — and total text — linear in schema depth. */
|
||||
@@ -291,9 +322,19 @@ function allocateClassName(base: string, state: RenderState): string {
|
||||
* object-chain would otherwise carry an ever-growing ConsString down the tree
|
||||
* and re-materialize it (via `.length`/`.slice`) at every level — Θ(depth²).
|
||||
* The bounded base plus the collision counter still yields unique names.
|
||||
*
|
||||
* The join is NFKC-normalized because both sides are separately normalized yet
|
||||
* their concatenation need not be: a base ending in a Hangul L jamo or LV
|
||||
* syllable composes with a following V or T jamo head (`가` + `ᆨ` gives `각`),
|
||||
* so the emitted class name would differ from the symbol CPython compiles, and
|
||||
* two byte-distinct names could fold onto one — `usedClassNames` dedupes by the
|
||||
* raw bytes, so the collision counter would not see it. Normalizing costs
|
||||
* O(cap + segment) per level, the same order as the `slice` it feeds. The other
|
||||
* two join points need no counterpart: `Args`/`Output` start with `A`/`O` and
|
||||
* {@link allocateClassName}'s suffix is digits, none of which compose backwards.
|
||||
*/
|
||||
function childClassName(base: string, segment: string): string {
|
||||
return capClassNameBase(`${base}${segment}`)
|
||||
return capClassNameBase(`${base}${segment}`.normalize('NFKC'))
|
||||
}
|
||||
|
||||
/**
|
||||
|
||||
@@ -494,8 +494,69 @@ describe('renderToolsSdkPy', () => {
|
||||
// whole characters; one ASCII character of padding puts the boundary inside
|
||||
// the 60th pair, and that half is dropped rather than emitted.
|
||||
expect(className('')).toBe(AHSA.repeat(60))
|
||||
expect(className('x')).toBe(`X${AHSA.repeat(59)}`)
|
||||
expect(className('x')).toHaveLength(119)
|
||||
const padded = className('x')
|
||||
expect(padded).toBe(`X${AHSA.repeat(59)}`)
|
||||
expect(padded).toHaveLength(119)
|
||||
})
|
||||
|
||||
it('normalizes the seam the Tool prefix creates, which the prefixed part alone does not cover', () => {
|
||||
// U+0301 COMBINING ACUTE ACCENT is XID_Continue but not XID_Start, so a name
|
||||
// headed by it takes the `Tool` prefix — and `Tool` ends in `l`, which
|
||||
// composes with it. Normalizing only the part being prefixed would emit
|
||||
// `Tool` + U+0301, which CPython compiles as `Too` + U+013A: the class
|
||||
// the SDK declares would not be the class the interpreter defines. Every
|
||||
// code point below is an escape — the two forms render identically.
|
||||
const text = renderToolsSdkPy([
|
||||
{
|
||||
name: '\u0301abc',
|
||||
description: 'Combining-mark head.',
|
||||
parameters: { type: 'object', additionalProperties: false, properties: { q: { type: 'string' } }, required: ['q'] },
|
||||
output: { type: 'string' },
|
||||
},
|
||||
])
|
||||
expect(text).toContain('class Too\u013AabcArgs(TypedDict):')
|
||||
expect(text).toContain('# tools["\u0301abc"](args: Too\u013AabcArgs) -> str')
|
||||
expect(text).not.toContain('Tool\u0301')
|
||||
})
|
||||
|
||||
it('normalizes a class-name join where two separately stable segments compose', () => {
|
||||
// Hangul jamo compose ACROSS the join `childClassName` makes: the parent
|
||||
// base ends in U+1100 (L jamo) and the child segment starts with U+1161 (V
|
||||
// jamo), each NFKC-stable alone, together U+AC00. Unnormalized, the declared
|
||||
// name differs from the compiled symbol, and two byte-distinct names can
|
||||
// fold onto one — `usedClassNames` dedupes by raw bytes, so the collision
|
||||
// counter never sees it and the later declaration shadows the earlier one
|
||||
// under CPython. Escapes again, for the same reason as above.
|
||||
const text = renderToolsSdkPy([
|
||||
{
|
||||
name: 'x',
|
||||
description: 'Jamo field names.',
|
||||
parameters: {
|
||||
type: 'object',
|
||||
additionalProperties: false,
|
||||
required: ['\uAC00\u1100'],
|
||||
properties: {
|
||||
'\uAC00\u1100': {
|
||||
type: 'object',
|
||||
additionalProperties: false,
|
||||
required: ['\u1161x'],
|
||||
properties: {
|
||||
'\u1161x': { type: 'object', additionalProperties: false, properties: { q: { type: 'string' } } },
|
||||
},
|
||||
},
|
||||
},
|
||||
},
|
||||
output: { type: 'string' },
|
||||
},
|
||||
])
|
||||
// The join is `XArgs` + U+AC00 U+1100 followed by U+1161 `x`, whose
|
||||
// trailing L+V pair composes into a second U+AC00.
|
||||
expect(text).toContain('class XArgs\uAC00\uAC00x(TypedDict):')
|
||||
expect(text).toContain(' \u1161x: XArgs\uAC00\uAC00x')
|
||||
expect(text).not.toContain('\u1100\u1161')
|
||||
// The level above it is a join that composes nothing (LV + L), so it stays
|
||||
// byte-identical — normalizing is not silently rewriting every name.
|
||||
expect(text).toContain('class XArgs\uAC00\u1100(TypedDict):')
|
||||
})
|
||||
|
||||
it('declares a closed empty object with omitted properties as an empty TypedDict, not dict[str, Any]', () => {
|
||||
|
||||
Reference in New Issue
Block a user