fix(tools): correct the trim-order claim and check the Literal escape dependency

trim and escape commute for every input, so the new whitespace test does
not pin their order: UNPRINTABLE and LONE_SURROGATE are disjoint from the
set trim() strips, and both escapes emit plain non-whitespace ASCII,
leaving the leading and trailing whitespace runs byte-identical. State that instead of the false causal clause.

pyScalar's Literal path escapes nothing itself -- JSON.stringify is what
keeps it parseable, covering NUL and, under ES2019 well-formed
stringification, unpaired surrogates. Record the dependency and turn it
into a checked invariant. Pin the docstring emission site for a lone
surrogate too, mirroring the NUL case.

Two docstring corrections: describe's caller enumeration omitted the
synthetic { description } wrapper docLines builds, and "special in
statement position" does not describe `_`, which is special in a match
pattern. Both keep the conclusion they support.

Note which of the two table guards fires depends on the entry point.
This commit is contained in:
Chinesezjc
2026-08-05 18:42:43 +08:00
parent 529e69ed87
commit 9bba851a62
5 changed files with 34 additions and 10 deletions

View File

@@ -51,6 +51,15 @@ describe('jsonSchemaToPy', () => {
expect(jsonSchemaToPy({ type: 'string', enum: [] })).toBe('Any')
})
it('leans on JSON.stringify to keep a Literal parseable', () => {
// The two code points CPython refuses in source reach this path as well,
// and nothing here escapes them itself — `JSON.stringify` does, NUL as a
// C0 control and a lone surrogate under ES2019 well-formed stringification.
// Python decodes both escapes back to the value the schema declared.
expect(jsonSchemaToPy({ type: 'string', const: 'a\u0000b' })).toBe(String.raw`Literal["a\u0000b"]`)
expect(jsonSchemaToPy({ type: 'string', enum: ['a\ud800b'] })).toBe(String.raw`Literal["a\ud800b"]`)
})
it('emits exact digits for a beyond-safe-range integer literal', () => {
// Python integers are arbitrary-precision, so the emitted digits ARE the
// value the model programs against. `String(2 ** 60)` prints the rounded
@@ -750,7 +759,9 @@ describe('renderToolsSdkPy', () => {
// so the block stays parseable with the code point intact.
expect(renderToolsSdkPy([described('zero\u200bwidth')])).toContain('"""zero\u200bwidth"""')
// Whitespace around a surviving control character is not an absent
// description: the escape runs before the trim, so what is left is visible.
// description. The escape's output is non-whitespace ASCII and the escaped
// sets are disjoint from what `trim()` strips, so the two operations touch
// different characters and their order is unobservable.
expect(renderToolsSdkPy([described(' \u0085 ')])).toContain(String.raw`# \x85`)
})
@@ -764,6 +775,7 @@ describe('renderToolsSdkPy', () => {
const high = renderToolsSdkPy([described('a\ud800b')])
expect(high).not.toContain('\ud800')
expect(high).toContain(String.raw`# a\ud800b`)
expect(high).toContain(String.raw`"""a\\ud800b"""`)
// A lone LOW surrogate is just as unencodable, and `\xNN` reaches neither.
expect(renderToolsSdkPy([described('a\udfffb')])).toContain(String.raw`# a\udfffb`)
// A well-formed pair is ONE astral code point, not two surrogates — the