Skip to content

Sync runtime and compiler updates into canary - #774

Merged
ctate merged 122 commits into
canaryfrom
ctate/runtime-performance-sync
Oct 8, 2026
Merged

ctate merged 122 commits into
canaryfrom
ctate/runtime-performance-sync

Conversation

@ctate

@ctate ctate commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator
  • Merge main's runtime optimizations and CI changes with canary's frontend improvements.
  • Preserve adaptive cycle reclamation with inline deallocation and include the native bootstrap and console contract repairs from Fix native bootstrap and console checks #773.
  • Verify the affected compiler, runtime, workflow, parity, and docs checks locally; use CI for full validation before merging.

ctate and others added 30 commits October 8, 2026 11:00
- Preserve callback identity, receiver binding, and class references across unions and collections.
- Track record field presence independently of undefined values and improve native runtime behavior.
- Expand differential coverage and streamline CI bootstrap and compiler rebuild checks.
Fourteen application-shaped workloads (JSON, regex, parsing and class
dispatch, sorting, rendering, collection pipelines, async fan-out,
Map/Set graphs, allocation, numeric kernels, exceptions, CSV, and the
two multi-module build benchmark apps on large inputs).

bench:runtime builds each workload with one or two checkouts, refuses to
time any executable whose stdout, stderr, or exit status differs from
Node, interleaves baseline and candidate launches, and reports medians
with bootstrap 95% confidence intervals, binary size, and peak RSS.
Node is the reference the project aims to beat. A single cold oracle
launch overstated its time; it is now warmed and sampled in the same
interleaved loop as the compiled executables, and the report adds a
per-workload candidate/Node ratio and its geomean. --no-node skips it.
Add a runtime benchmark suite that times compiled programs against Node
Co-authored-by: Chris Tate <366502+ctate@users.noreply.github.com>
Co-authored-by: Chris Tate <366502+ctate@users.noreply.github.com>
… 0.03 s) (#743)

* Remove the large-array cliff past 2^20 elements

Arrays longer than 2^20 kept every element beyond the cutoff in a sorted
side store, and each indexed write to an existing entry deleted and
reinserted it, so fill + indexed read/write loops went quadratic
(n=1.2M: 10.4s on macOS arm64, 12.6s on Linux x64).

- Rewrite existing sparse entries in place.
- Let dense storage grow past the cutoff for contiguous extension of a
  populated prefix, bulk range writes, and sparse tails that reach 3/16
  density, migrating side-store entries so index < cap stays dense.
- Make the side store a deque with front slack so descending first-touch
  writes and prefix migration are amortized O(1).
- Bound range-derived slice/splice capacity hints for hole runs.

scr_arr_new now reserves the live count, so a 2^20+3 element Set
converts into dense storage instead of a sparse tail; the white-box
test_lib check that pinned the old policy now asserts dense storage.

* Share immortal boolean union boxes

Maybe-undefined reads of boolean[] elements (`if (composite[i])` in a
sieve) box the element in a boolean|undefined union and free it at once:
one calloc/free pair per read, about a quarter of numeric-kernels samples
at large scales. Unions are immutable and boolean arms own no payload, so
low tags now return headerless immortal instances, the same convention as
the compiler's static unit-arm constants.

* Record preflight/order baselines for corpus 3138
WIP step 1: scr_cyc_alloc/scr_cyc_free use a region-backed small-object
allocator (scr_mem_*), bypassed under SCR_RC_AUDIT/ASan.
…locator

Move the allocator into scr_alloc.c with inline fast paths in scr_runtime.h;
strings, array headers/slot buffers, and emitted acyclic records/classes
(scr_rt_calloc/scr_rt_free) now use it. Adds scr_alloc.test.{c,ts}.
Upstream native workers (095b619) build every runtime unit with
-DSCR_WORKERS and run one runtime instance per thread. The small-object
allocator state is process-global and unsynchronized, so worker builds
fall back to the system allocator, exactly like thread-instanced library
archives. Ordinary executables keep the allocator unchanged.
The tree-shaking test failed on Linux x64 with "expected 68344 to be less
than 65536". Measured with the release helper/runtime-pack pipeline in a
Linux x64 sandbox (hello = console.log("hello", "world")):

  PR    file   .text  R+E segment
  #743  64256  15794  0x3f11
  #744  68416  16050  0x4011
  #764  68520  16546  0x4221

The allocator adds 256 bytes of .text to hello: the inline scr_mem_free
fast path in the release functions hello already links (scr_str_release,
scr_arr_destroy, scr_arr_gc_free, scr_cyc_free, scr_dyn_dispose). Its
slow path and address-space reservation (scr_sa_slow, scr_sa_reserve)
are already stripped from hello. ELF segments are page aligned, and #743
left only 239 bytes before the executable segment's 16 KiB page, so 17
bytes over the boundary move every later segment, and the file, by 4 KiB.

Shrinking hello back under the old ceiling is not a reasonable trade.
At the top of the stack the executable segment is 545 bytes past the
page boundary. The extra bytes come from hot-path code the later PRs add
on purpose (number-to-string fast paths in scr_f64_to_str, inline-alloc
release paths, the direct fd write), all reached through console.log's
generic argument rendering. Outlining them to save one page would undo
the optimizations those PRs add.

The ceiling's documented intent is "several native pages of
linker-version slack" above a hello that the comment put at about 41KB.
That no longer held before this stack (64256 bytes on #743, 1.2 KiB of
slack). 80 KiB restores about three pages of slack over the stack's
largest hello and still catches lost section GC (the old always-linked
runtime was about 400KB). Symbol absence remains the primary
reachability contract. The Mach-O ceiling is unchanged.

Verified in the Linux x64 sandbox: pnpm test
tests/harness/runtime-tree-shaking.test.ts fails on #744 before this
change (68416 vs 65536) and passes after it, on #744 and on #764.
`ScrSaState scr_sa;` was a tentative definition, which the runtime build
emits as a COMMON symbol. Library localization (abi.localize_runtime)
deliberately keeps COMMON symbols global (the ASan image-registration guard
relies on that), so every runtime-localized library archive exported
scr_sa next to its declared API, failing tests/harness/library-multi.test.ts
M1 ("external definitions equal the declared set exactly").

Initialize it explicitly so it is a regular zero-initialized definition that
localization demotes like every other runtime global. No other global added
by the stack is a tentative definition. The library-multi suite passes
(it fails M1 without this change).
Add a small-object allocator to the runtime (-14% Linux, -21% macOS)
`JSON.parse(text) as T` used to build a checked-dynamic ScrDyn tree, copy
it into the native layout with the generated dynCheck builder, and then
dispose the tree. Together that was about 68% of json-records.

For record/array targets with number, boolean and string leaves, the
backend now emits a ScrJsonSchema (field offsets in declaration order,
constructors, releases) and a wrapper. The wrapper calls the new
scr_json_parse_schema, a one-pass parser that writes directly into native
records and arrays. Whenever it declines (syntax errors anywhere, kind
mismatches, missing or duplicate declared keys, excess depth) it releases
its partial value without throwing, and the wrapper reruns the unchanged
scr_json_parse + dynCheck route. Every error message and position stays
byte-identical. The lexer helpers are shared through a quiet mode and a
factored number scanner, so values decode bit-identically.

Tests: new differential corpus program json-typed-parse-direct.ts; a
dyncheck harness case comparing typed and untyped error messages through
the fused path; white-box C tests for accept/decline and leak-freedom in
scr_json.test.c; an LLVM emission test (64-bit and wasm32).

Linux (sandbox x86_64, runs=15):
| workload | baseline ms | candidate ms | change | 95% CI | verdict | size | node ms | candidate/node |
| --- | ---: | ---: | ---: | --- | --- | ---: | ---: | ---: |
| json-records | 713.2 | 249.3 | -65.0% | -65.6% .. -64.1% | faster | +2.2% | 439.04 | 0.57x |
| regex-logs | 350.8 | 353.7 | +0.8% | -0.9% .. +2.3% | neutral | +0.0% | 238.42 | 1.48x |
| ast-interp | 379.1 | 376.8 | -0.6% | -1.8% .. +0.4% | neutral | +0.0% | 227.45 | 1.66x |
| records-sort | 1200.5 | 1207.8 | +0.6% | -2.4% .. +5.7% | neutral | +0.0% | 232.17 | 5.20x |
| template-render | 248.3 | 248.1 | -0.1% | -2.3% .. +1.0% | neutral | +0.0% | 202.16 | 1.23x |
| functional-pipeline | 402.0 | 395.3 | -1.6% | -3.8% .. +0.5% | neutral | +0.0% | 254.57 | 1.55x |
| async-pipeline | 202.4 | 201.7 | -0.3% | -1.4% .. +2.0% | neutral | +0.0% | 90.14 | 2.24x |
| word-graph | 366.9 | 348.2 | -5.1% | -12.9% .. +4.3% | neutral | +0.0% | 576.56 | 0.60x |
| alloc-trees | 513.5 | 525.5 | +2.3% | -2.1% .. +4.0% | neutral | +0.0% | 132.99 | 3.95x |
| numeric-kernels | 121.4 | 122.5 | +0.8% | -1.9% .. +3.6% | neutral | +0.0% | 135.91 | 0.90x |
| validate-errors | 206.0 | 209.2 | +1.5% | -3.6% .. +4.1% | neutral | +0.0% | 703.47 | 0.30x |
| csv-numbers | 450.9 | 448.6 | -0.5% | -2.2% .. +2.0% | neutral | +0.0% | 499.25 | 0.90x |
| log-summary | 253.5 | 252.6 | -0.3% | -1.5% .. +2.0% | neutral | +0.0% | 279.67 | 0.90x |
| inventory-report | 433.7 | 433.2 | -0.1% | -6.0% .. +3.6% | neutral | +0.0% | 413.79 | 1.05x |
geomean change -7.4%; geomean candidate/node 1.321x -> 1.223x

macOS (arm64, runs=15):
| workload | baseline ms | candidate ms | change | 95% CI | verdict | size | node ms | candidate/node |
| --- | ---: | ---: | ---: | --- | --- | ---: | ---: | ---: |
| json-records | 292.3 | 153.0 | -47.7% | -50.1% .. -46.6% | faster | +0.4% | 248.7 | 0.61x |
| (all 13 other workloads neutral; see results/h16-typed-json/local-1.json) |
geomean change -4.8%; geomean candidate/node 1.497x -> 1.425x
Seventeen Order inputs whose shape disagrees with the target (wrong leaf
kinds at the top level, in nested records, record arrays and string
arrays; null for object and number; missing fields; object vs array at a
field and at the root; duplicate keys; truncated JSON; 1e400 and a lone
surrogate; extra nested members; reordered keys) plus an Item[] root.
The expected lines are the output of the checked-route-only build
(perf/harness); the fast path must decline or agree on every one.
Parse typed JSON straight into native records (2.9x faster)
Strict element reads (scr_arr_get_f64/bool/ref, borrow), scr_arr_get_number,
scr_arr_state/has and element writes (scr_arr_set_f64/bool/ref) now check
inline for a canonical integer index inside dense storage (below cap, and
below len for reads, present-value state for strict reads) and load or
store the slot directly. Every other case calls the unchanged runtime entry
with a cold call-site attribute. Proven integers (exactInteger) skip the
double round-trip. for...of also reads the length inline.
Read-only fallbacks (get_f64/bool/number, state, has, borrow) carry
memory(read, inaccessiblemem: readwrite) at the call site and the guard loads
the storage bases before branching, so LLVM hoists header loads out of
loops. Conditions on number[]/boolean[] element reads no longer box
T | undefined: numbers reuse getNumber (missing is NaN, falsy like undefined),
booleans test the slot state then read the value.
Number() of an optional string element read no longer boxes string | undefined
into a dynamic value: plain-operand reads test the slot state and convert the
element with num.fromString, answering NaN for missing slots.
Inline dense array element access (-15% geomean, numeric-kernels -65%)
…p bytecode

libregexp compiles every non-sticky pattern with a `.*?` search prefix and
interprets its bytecode one opcode at a time, so an unanchored search pays a
backtrack push/pop and ~7 dispatches per subject position, and each `x+`
iteration pushes a backtrack state. Profiles of regex-logs and
template-render showed 65-85% of samples inside lre_exec.

scr_regex_native.h translates the compiled bytecode body (the bytecode stays
the source of truth for parsing, flags, case folding and capture numbering)
into a small program for one-byte subjects:

- single-character opcodes become 128-entry tables computed by evaluating the
  interpreter's own predicate (lre_canonicalize, lre_is_space, range pairs)
  for each ASCII code unit, so /i and /u are exact on these subjects;
- quantifiers over one character become counted loops that backtrack by
  count; greedy loops whose continuation cannot start with a byte of the
  loop's class (or is just saves + match) are possessive;
- runs of literal bytes become one memcmp;
- splits, gotos, saves, capture resets, counted group loops and empty-check
  registers keep libregexp's exact backtracking order and undo semantics;
- unanchored searches start only at bytes that can begin a match (memchr for
  a single byte) and only at 0 for non-multiline `^` patterns.

Lookaround and back references, and all non-ASCII subjects, keep lre_exec.
scr_exec is the single dispatch point, so every match loop is shared. The
program is built with the bytecode and cached on ScrRegex (new `native` slot;
the emitter's %ScrRegex type and literal initializer gain one null pointer).

Tests: packages/runtime/test/test_regex_native.c fuzzes the matcher against
lre_exec (random patterns over the subset and its bail-outs, every start
index, all capture slots; ASan+UBSan), wired into regex.test.ts; a sweep of
8 seeds x 100k patterns (~575k translated, ~70.7M checks) found no mismatch.
tests/corpus/regex-native-matcher.ts pins each construct through the public
APIs on ASCII subjects and non-ASCII twins.

Squashes 1d210a03: build the fuzz test with -D_GNU_SOURCE on Linux; the
vendored cutils.h needs CLOCK_MONOTONIC, which strict -std=c11 hides on glibc.
Security review finding (#747): rn_first recursed into every split target
(split_next/split_goto/loop/loop_split). `visited` stops cycles but not
depth, so a flat runtime pattern with a long run of alternatives, such as
new RegExp("(?:" + "|".repeat(n) + "x)"), chains one split per
alternative and recursed once per alternative while the regex was being
constructed (scr_re_native_build runs on every compile), before any
subject is matched. Deep enough patterns could exhaust the C stack of the
thread compiling them.

The analysis is now iterative: split targets go on a heap worklist of
nops + 1 entries allocated next to `visited`. Every op is processed once,
so each split pushes at most one entry, and the accumulated first-byte set,
empty-match and anchor flags do not depend on the visiting order. If the
worklist cannot be allocated the native build fails and the regex keeps
using libregexp, like the other native-build failures. No other function in
scr_regex_native.h recurses on pattern structure.

Tests: corpus 3139 constructs runtime patterns with 30,000-alternative runs
(empty, consuming, word-bounded, capturing, nested-but-flat, global) and
matches ASCII subjects; it matches Node in the plain and sanitized lanes.
…egex

Run ASCII regex subjects on a native matcher (regex workloads 2-3.7x faster)
ctate and others added 25 commits October 8, 2026 15:46
…y-hunt

Fix quadratic and syscall-heavy runtime paths (up to 7,000x slower than Node to at or below Node)
Speed is a third native optimization posture next to release (default,
-O2) and dev (-O0). It is optimized exactly like release at the C/LLVM
level and is the switch for optimizations that trade executable size and
build time for run-time speed, so the default release build stays inside
the binary-size budget.

- backend/optimization.ts: the NativeOptimization type, its -O class, and
  cache-key helpers. Release keeps its historical key shapes (absent).
- CLI: build/run/cache warm accept speed; usage text names the trade.
  Bootstrap's routed executable cache mirrors the posture.
- Early executable cache keys and native-feature validation include speed,
  so release, dev, and speed entries never satisfy one another.
- Runtime packs: an optional executable `speed` flavor (-O2, only beside
  release). Speed links it when present and the release objects otherwise;
  library modes link their release flavor. The sanitizer lane follows the
  same flavor key.
- Library profiles accept "speed".
- WASI release-only linker flags apply to speed too.
- bench-runtime: --baseline-optimization/--candidate-optimization and a
  per-contender cold_build_ms.

Release executables are byte-identical to the stack tip for all 14
benchmark workloads (macOS arm64, sha256 of executables built from both
trees).
Re-lands perf/h30-inline-rc (00051ca0, 40398773, 07972360) behind the
speed posture. With inline RC on, every mirrored runtime RC family
retains inline and releases inline down to the rc == 1 / new cycle
candidate slow path (shapes.ts INLINE_RC_FAMILIES). Release emission is
unchanged: ShapeHost.rcHelpers is null unless LlvmTargetOptions.inlineRc
is set, and every retain/release (including capture boxes) stays a
runtime call; the %ScrMapRc type is only emitted with inline RC.

- index.ts and native/prepare.ts pass inlineRc for speed executables,
  speed object/asm emission, and speed library profiles.
- scr_runtime.h keeps H30's static asserts for the mirrored header and
  trace-slot offsets (no code change; the rebuilt macOS pack's manifest is
  identical to the baseline's).
- Tests: shapes.test.ts pins release symbols and H30's speed fast paths
  and header offsets; emitter.test.ts pins that release IR is unchanged
  and speed routes releases through the helpers. The differential harness
  gains `// @Optimization: speed` (and a SCRIPTC_TEST_OPTIMIZATION=speed
  lane); corpus 4032 pins every mirrored family under speed.

Release executables stay byte-identical to the stack tip (14/14
workloads, macOS).
Re-lands perf/h02-runtime-lto (6d373317, b62448f3) behind the speed
posture.

- Runtime packs gain an executable `speed` flavor whenever the host LLVM
  helper can run. Its static units are emitted through the helper's
  `runtime-unit` command (promoted unit-local symbols, object and import
  bitcode from one module); its SCR_DYNAMIC variants reference the release
  objects. Release, dev, and library flavors and the vendor archives are
  built exactly as before (the rebuilt macOS pack's flavors and archives
  are identical to the baseline's).
- Speed program builds load the speed flavor's bitcode and pass it to the
  helper (`--import-bitcode`); release builds never do.
  loadRuntimeBitcode only reads a real speed flavor, never release.
- The helper drops the emitter's inert sanitize_address attributes only
  when it imports, so release modules reach the pipeline unchanged.
- Tests: runtime-pack.test.ts covers speed-variant bitcode selection,
  verification, and that release-flavor bitcode is never imported;
  native-codegen.test.ts covers helper arguments and cache keys.

Release executables stay byte-identical to the stack tip with the rebuilt
helper (14/14 workloads, macOS).
Final numbers for the speed posture (H30 inline RC + H02 runtime bitcode import + H28 runtime PGO) against release on the same tree. Release executables are byte-identical to perf/stack/06-h23-layout-noise on both lanes (14/14).

Linux x86_64 sandbox, --layouts=4 --runs=16 (linux-speed-vs-release.json):

| workload | baseline ms | candidate ms | change | 95% CI | verdict | size (bytes) | cold build ms | node ms | candidate/node |
| --- | ---: | ---: | ---: | --- | --- | ---: | ---: | ---: | ---: |
| json-records | 194.4 | 164.7 | -15.3% | -17.3% .. -13.7% | faster | +18.3% (239328 → 283176) | 1311 → 1987 | 439.51 | 0.38x |
| regex-logs | 165.9 | 147.6 | -11.0% | -12.4% .. -7.1% | faster | +11.2% (325648 → 362136) | 1003 → 1539 | 235.96 | 0.63x |
| ast-interp | 271.7 | 200.8 | -26.1% | -27.7% .. -22.9% | faster | +14.9% (149032 → 171264) | 1230 → 1797 | 243.44 | 0.82x |
| records-sort | 669.8 | 553.9 | -17.3% | -22.0% .. -10.8% | faster | +24.2% (140840 → 174960) | 1232 → 1829 | 226.37 | 2.45x |
| template-render | 70.5 | 62.9 | -10.7% | -12.9% .. -7.1% | faster | +4.1% (244976 → 254968) | 1049 → 1550 | 193.97 | 0.32x |
| functional-pipeline | 220.8 | 189.3 | -14.3% | -19.8% .. -10.4% | faster | +23.1% (112928 → 139008) | 1130 → 1841 | 269.15 | 0.70x |
| async-pipeline | 197.0 | 196.8 | -0.1% | -2.4% .. +5.3% | neutral | +7.3% (169944 → 182392) | 978 → 1297 | 91.97 | 2.14x |
| word-graph | 251.9 | 231.1 | -8.3% | -14.4% .. -4.6% | faster | +30.7% (139040 → 181664) | 1279 → 1974 | 471.96 | 0.49x |
| alloc-trees | 321.6 | 271.4 | -15.6% | -18.2% .. -13.3% | faster | +10.8% (88672 → 98280) | 946 → 1281 | 125.46 | 2.16x |
| numeric-kernels | 51.2 | 48.3 | -5.5% | -17.4% .. +4.9% | neutral | +16.8% (106600 → 124488) | 1061 → 1411 | 147.98 | 0.33x |
| validate-errors | 188.9 | 163.0 | -13.7% | -23.3% .. +10.3% | neutral | +8.2% (285424 → 308904) | 991 → 1447 | 793.58 | 0.20x |
| csv-numbers | 357.6 | 317.1 | -11.3% | -15.9% .. -7.6% | faster | +21.1% (146856 → 177792) | 986 → 1572 | 541.06 | 0.59x |
| log-summary | 167.6 | 157.4 | -6.1% | -9.5% .. -1.3% | faster | +30.5% (225256 → 294024) | 1331 → 2143 | 311.61 | 0.51x |
| inventory-report | 327.6 | 263.5 | -19.6% | -28.2% .. -15.3% | faster | +22.2% (311336 → 380312) | 1599 → 2410 | 445.01 | 0.59x |

contenders: baseline=release@, candidate=speed@; layouts 4, runs 16
geomean change -12.7%; faster 11/14; slower: none; size geomean +17.1% (max +30.7%); cold build geomean +48.5%; candidate/node 0.658x (baseline/node 0.754x)

macOS arm64 (shared host), --layouts=4 --runs=16 (local-speed-vs-release.json; file sizes move in 16 KiB pages, __text geomean +30.8%):

| workload | baseline ms | candidate ms | change | 95% CI | verdict | size (bytes) | cold build ms | node ms | candidate/node |
| --- | ---: | ---: | ---: | --- | --- | ---: | ---: | ---: | ---: |
| json-records | 113.7 | 89.4 | -21.4% | -22.4% .. -20.5% | faster | +28.8% (222592 → 286688) | 958 → 1523 | 224.67 | 0.40x |
| regex-logs | 106.0 | 91.6 | -13.5% | -15.4% .. -11.7% | faster | +22.7% (279128 → 342456) | 996 → 1106 | 129.92 | 0.70x |
| ast-interp | 164.6 | 111.7 | -32.1% | -32.9% .. -31.4% | faster | +17.8% (148376 → 174760) | 896 → 1288 | 136.99 | 0.82x |
| records-sort | 369.7 | 286.4 | -22.5% | -26.6% .. -18.9% | faster | +17.8% (147888 → 174192) | 980 → 1244 | 133.39 | 2.15x |
| template-render | 50.4 | 46.6 | -7.4% | -9.6% .. -5.9% | faster | +13.0% (205808 → 232512) | 942 → 1039 | 95.49 | 0.49x |
| functional-pipeline | 129.9 | 104.5 | -19.6% | -22.3% .. -15.1% | faster | +36.0% (112464 → 152912) | 872 → 1523 | 120.64 | 0.87x |
| async-pipeline | 261.6 | 261.8 | +0.1% | -2.2% .. +3.4% | neutral | +21.4% (154624 → 187776) | 750 → 962 | 57.94 | 4.52x |
| word-graph | 133.1 | 110.7 | -16.9% | -22.0% .. -10.2% | faster | +30.3% (147192 → 191816) | 876 → 1298 | 276.89 | 0.40x |
| alloc-trees | 216.8 | 163.7 | -24.5% | -25.5% .. -23.8% | faster | +25.5% (92792 → 116424) | 746 → 962 | 75.59 | 2.17x |
| numeric-kernels | 32.1 | 31.1 | -2.8% | -5.4% .. -0.5% | faster | +21.8% (110832 → 134960) | 853 → 979 | 78.45 | 0.40x |
| validate-errors | 110.6 | 86.3 | -22.0% | -22.8% .. -21.0% | faster | +31.0% (205136 → 268720) | 784 → 1507 | 428.13 | 0.20x |
| csv-numbers | 184.2 | 160.8 | -12.7% | -17.1% .. -10.6% | faster | +19.9% (146984 → 176200) | 1217 → 1380 | 252.81 | 0.64x |
| log-summary | 88.8 | 72.8 | -17.9% | -19.9% .. -16.3% | faster | +38.7% (205144 → 284520) | 1440 → 1731 | 152.02 | 0.48x |
| inventory-report | 132.7 | 105.4 | -20.6% | -23.3% .. -16.7% | faster | +43.0% (223120 → 319088) | 1472 → 1940 | 199.93 | 0.53x |

contenders: baseline=release@a5a5b961+dirty, candidate=speed@a5a5b961+dirty; layouts 4, runs 16
geomean change -17.1%; faster 13/14; slower: none; size geomean +26.0% (max +43.0%); cold build geomean +34.0%; candidate/node 0.722x (baseline/node 0.871x)

Incremental, Linux: H30 vs release -2.3%; H02 on H30 -6.6%; H28 on H30+H02 -4.4%. Linux text+data speed vs release: geomean +19.0%, max +33.7%. Full report: results/h45-speed-posture/REPORT.md.
The speed posture's inline scr_union_release fast path predates the
runtime change that buffers a surviving union box as a cycle candidate
only when its arm_trace can reach a cycle. It still enqueued every
surviving box: sound (an untraced box visits nothing and the runtime death
path reads the header), but it gave back that change's collector savings
under --optimization=speed. The fast path now loads arm_trace and returns
when it is NULL, exactly like scr_union_release; scr_runtime.h pins the
offset.
The default build's largest size contributor on the stack was H23's
-falign-functions=64 on every x86-64 runtime unit (2-11 KB per program;
results/h45-size-budget). Release and library flavors now keep the
compiler's default function alignment; the opt-in speed flavor keeps the
blanket 64-byte alignment. Vendored archives stay 64-byte aligned
(lre_exec keeps its pinned placement) except libunicode.c, whose small
table helpers were most of the regex programs' padding. The runtime-first
link order is unchanged.

Docs and usage text describe speed without the reverted runtime profile.
The bootstrap-seed build rejected `host.rcHelpers ?? []` in shapes.ts
(SC2011: `Set<string> | never[]` has no static representation) and the
spread of that union into a string[] literal (SC1090). A small helper copies
the set's keys into a plain array (empty when inline RC is off), which both
callers iterate; emitInlineRcHelpers still sorts its copy. Behavior is
unchanged.
…sture

Add an opt-in --optimization=speed posture (a further -9% on Linux)
benchmarks/scaling: 128 probes of common operations (arrays, strings,
Map/Set/records, JSON, closures, exceptions, regex, dates, promises,
timers, console, fs) run at 3-4 sizes under scriptc and Node with
identical-stdout checks; run.mjs flags superlinear growth and >3x Node,
--strace-only reports syscalls per operation, strace-workloads.mjs
straces the runtime benchmark suite. Results and the full probe table:
results/h40-pathology-hunt/REPORT.md.

Pathologies fixed on this branch (probe ms at 1e3/1e4/1e5/1e6, Linux):
  i in number[]        base 43/5291/timeout   -> 0.03/0.20/1.5/15   (Node 0.19/0.73/6.8/39)
  k in Record<string>  base 4.8/436/timeout   -> 0.15/1.5/20/319    (Node 0.96/8.9/100/1452)
  re.exec /g loop      base 1.1/70/7910/t.o.  -> 0.49/4.1/40/416    (Node 0.47/4.6/43/398)
  clearTimeout n live  base 1.3/10/1014/t.o.  -> 1.3/3.9/38/740     (Node 3.8/9.4/46/527)
  new RegExp repeat    base 3.9/38/368/3811   -> 0.13/1.1/11/106    (Node 0.49/2.9/23/174)
  console.error        2 writes/line          -> 1 write/line

Suite A/B, Linux x86_64 sandbox, candidate vs perf/stack/06 (--layouts=4
--runs=16), all workloads ok:

| workload | baseline ms | candidate ms | change | 95% CI | verdict | size | node ms | candidate/node |
| --- | ---: | ---: | ---: | --- | --- | ---: | ---: | ---: |
| json-records | 202.3 | 201.9 | -0.2% | -7.0% .. +3.8% | neutral | +0.0% | 458.29 | 0.44x |
| regex-logs | 166.8 | 166.4 | -0.2% | -1.6% .. +1.1% | neutral | +0.1% | 239.97 | 0.69x |
| ast-interp | 268.6 | 269.9 | +0.5% | -4.6% .. +3.8% | neutral | +0.0% | 242.7 | 1.11x |
| records-sort | 648.7 | 670.4 | +3.3% | -4.7% .. +10.0% | neutral | +0.0% | 222.69 | 3.01x |
| template-render | 69.6 | 70.7 | +1.7% | -0.2% .. +2.6% | neutral | +0.0% | 187.3 | 0.38x |
| functional-pipeline | 221.8 | 214.9 | -3.1% | -7.3% .. +0.7% | neutral | +0.0% | 249.06 | 0.86x |
| async-pipeline | 197.3 | 198.2 | +0.4% | -1.2% .. +3.0% | neutral | +0.2% | 89.25 | 2.22x |
| word-graph | 259.2 | 256.0 | -1.3% | -6.3% .. +3.0% | neutral | +0.0% | 469.48 | 0.55x |
| alloc-trees | 318.8 | 316.8 | -0.6% | -3.0% .. +1.2% | neutral | +0.0% | 127.68 | 2.48x |
| numeric-kernels | 46.4 | 46.5 | +0.1% | -1.6% .. +3.9% | neutral | +0.0% | 133.42 | 0.35x |
| validate-errors | 176.5 | 178.1 | +0.9% | -3.1% .. +4.7% | neutral | +0.0% | 702.45 | 0.25x |
| csv-numbers | 336.2 | 336.8 | +0.2% | -2.0% .. +1.2% | neutral | +0.0% | 500.14 | 0.67x |
| log-summary | 165.7 | 164.6 | -0.6% | -5.9% .. +1.9% | neutral | +0.0% | 290.81 | 0.57x |
| inventory-report | 321.9 | 320.1 | -0.6% | -6.7% .. +9.5% | neutral | +0.0% | 431.4 | 0.74x |

geomean change: +0.0%
geomean candidate/node: 0.768x (lower is better; goal: well below 1)

macOS lane not run (all measurement in the sandbox per task instructions).
…probes

Add a scaling probe suite for asymptotic and syscall pathologies
…, two-decimal String)

String to number: Number(str), unary +, and parseFloat now convert short
decimal spellings with Clinger's exact fast path (scr_decimal_scan in
scr_number.c): significant digits m <= 2^53 and a decimal exponent within
+-22 (plus the folded range up to 22 + 15) need one correctly rounded
multiply or divide, bit-identical to strtod. ToNumber validates and
accumulates in the same pass; everything else still goes through the copied
span and strtod. Whitespace trimming in Number(), parseInt, and parseFloat
skips the UTF-8 decode for ASCII non-whitespace.

toFixed: the exact-binary-value digit generation moves from scr_lib.c into
scr_number.c (scr_f64_to_fixed). For f <= 22, mantissa * 5^f fits a
64x64->128 multiply, so the spec's n is one shift plus the dropped-half bit
whenever it fits 64 bits. Digits render two at a time instead of through
snprintf. The 16-limb bignum remains the fallback.

Number to string (H44): a non-integral x below 2^46 that is the double
nearest to m/100 prints as m/100 directly instead of running Ryu. Below 2^46
ulp(x) < 0.01, so m/100 is the unique shortest round-tripping decimal. The
correctly rounded division (double)m / 100 == x is the proof. This covers
JSON.stringify, templates, String(), join, and console output.

Tests: Node oracles for every path. number-cases.txt gains money-like values,
the 2^46 cutoff, and one-ulp neighbors. tonumber-cases.txt gains fast-path
boundaries plus 20k seeded short decimals. The new parsefloat-cases.txt pins
scr_parse_float, and the new fixed-cases.txt/test_fixed.c pin toFixed,
including the 2^64 fast-path limit and f up to 100. Fuzz modes cover all four
paths. Existing case lines are unchanged. Corpus 3170 pins the paths
end-to-end against Node.

Linux x64 sandbox, A/B against perf/stack/06-h23-layout-noise,
--layouts=4 --runs=16 (run 3 of 3; geomean -6.0% / -6.7% / -6.3%):

| workload | baseline ms | candidate ms | change | 95% CI | verdict | size | node ms | candidate/node |
| --- | ---: | ---: | ---: | --- | --- | ---: | ---: | ---: |
| json-records | 236.2 | 204.2 | -13.5% | -19.8% .. +2.1% | neutral | +0.0% | 488.62 | 0.42x |
| regex-logs | 176.2 | 168.8 | -4.2% | -10.1% .. +0.3% | neutral | +0.0% | 281.23 | 0.60x |
| ast-interp | 280.0 | 273.1 | -2.5% | -5.6% .. +1.4% | neutral | +2.8% | 273.05 | 1.00x |
| records-sort | 952.4 | 939.4 | -1.4% | -5.0% .. +5.6% | neutral | +0.0% | 289.87 | 3.24x |
| template-render | 78.9 | 77.3 | -2.1% | -6.2% .. +1.2% | neutral | +0.0% | 246.71 | 0.31x |
| functional-pipeline | 338.7 | 348.8 | +3.0% | -6.5% .. +24.0% | neutral | +0.1% | 339.84 | 1.03x |
| async-pipeline | 207.3 | 207.8 | +0.2% | -5.3% .. +3.8% | neutral | +0.0% | 105.57 | 1.97x |
| word-graph | 302.0 | 300.1 | -0.6% | -5.1% .. +12.9% | neutral | +3.0% | 600.48 | 0.50x |
| alloc-trees | 336.3 | 332.4 | -1.1% | -6.2% .. +2.6% | neutral | +0.1% | 144.73 | 2.30x |
| numeric-kernels | 49.7 | 49.8 | +0.1% | -3.4% .. +6.5% | neutral | +4.0% | 143.73 | 0.35x |
| validate-errors | 179.1 | 182.9 | +2.1% | -5.8% .. +9.1% | neutral | +0.0% | 727.78 | 0.25x |
| csv-numbers | 349.0 | 233.8 | -33.0% | -37.5% .. -26.5% | faster | +0.1% | 541.77 | 0.43x |
| log-summary | 171.6 | 135.7 | -20.9% | -23.3% .. -17.7% | faster | +1.8% | 313.2 | 0.43x |
| inventory-report | 411.9 | 384.7 | -6.6% | -18.1% .. +5.5% | neutral | +2.7% | 521.27 | 0.74x |

Size: +64..+136 bytes per executable. The +1.8..+4.0% rows are one 4 KiB
page crossing (two for inventory-report).
The TypeScript 7 parity job reported tests/corpus/3170-number-conversion-
fast-paths.ts as missing a baseline. Recorded with
SCRIPTC_UPDATE_BASELINES=1 pnpm exec vitest run
packages/compiler/test/ts7/order-parity.test.ts: the re-recorded file
added only this program (and the later stack's console programs, which
belong to #764), with no existing entry changed. The entry is inserted
at its sorted position in the stack's baseline block so it does not
overlap the #764 entries when the stack is rebased.

node scripts/test-ts7.mjs --baselines-only passes.
Add fast paths for number and string conversion (csv-numbers -30%)
log-lines writes 200k lines (scale 4) through the common console shapes:
a single template literal, string plus number arguments, several
primitive arguments, process.stdout.write chunks, and an occasional
console.error line. The runner captures stdout through a pipe, so the
workload measures immediate per-line writes into a pipe.
Every console.log/error/warn line and process.stdout/stderr.write chunk
is still submitted before the call returns, as in Node. On POSIX
executables, each one now goes out as exactly one write(2) on fd 1 or 2.
Before, it went through C stdio, which cost a copy into the stdio
buffer, one stream call per argument, separator, and newline, and a
flush. stderr is unbuffered, so console.error made one write(2) per
piece: 3 writes per line on average in the probe.

- The line renders into a 2 KiB stack buffer (heap above that). A large
  single string goes out with writev and no copy.
- Short writes continue and EINTR retries. Other errors (EAGAIN, EPIPE,
  EBADF) return errno with the rest dropped, which is what the stdio
  flush did. console still ignores them, and SIGPIPE behavior is
  unchanged.
- Runtime-internal bytes still in the C stdout buffer are flushed first
  (__fpending on Linux, the FILE fields on Darwin, an unconditional
  flush elsewhere), so stdout order and merged 2>&1 order are kept.
- Windows, WASI, and library artifacts keep the C stream path, now fed
  one formatted chunk per call.

Tests: corpus 4294 covers line sizes across the stack buffer and the
64 KiB stdio buffer. New console-io harness cases cover stdout and
stderr merged into one pipe and into one file (byte-compared with
Node), and console.log after the stdout reader closes (EPIPE).

Linux sandbox, ns per line (median of 9 launches, 200k lines; scriptc
includes ~1 ms startup). All outputs byte-identical to Node.

  probe      target  node   prev  cand
  p-str      pipe    4040   1323  1239
  p-str      file    3092    849   790
  p-str      null    2193    319   283
  p-args     pipe    6168   1440  1228
  p-args     file    4896    925   802
  p-args     null    3916    389   292
  p-write    pipe    2889   1328  1197
  p-write    file    2267    809   769
  p-write    null    1577    297   265
  p-err      pipe    4672   3253  1132
  p-err      file    3478   2042   717
  p-err      null    2614    562   232
  log-lines  pipe    4958   1490  1356
  log-lines  file    3885    966   877
  log-lines  null    2983    425   368

strace -c (pipe): p-err went from 600,000 to 200,000 write calls. The
other probes stay at one write per line.
console.log(`${method} ${path} ${status}`) used to build the joined string
first (one string per number, scr_str_concat_parts, then releases), and
the console then copied that string into the line it writes. The LLVM
backend now flattens a strConcat argument with stringParts and passes
each part as a ScrLogArg:

- Later parts carry SCR_ARG_GLUE, so no separating space is added.
- A number conversion passes the raw double as SCR_ARG_NUM (String()
  spelling, so -0 prints "0"; a bare number argument keeps inspect's
  "-0").
- A boolean conversion passes the bool.

The runtime renders the parts directly into the line buffer. Parts are
evaluated in source order through emitStringInputs, which applies the
same borrowing rules as concatenation. Calls that use the new tags go to
new scr_console_log_parts / scr_console_error_parts entry points, so a
program object linked against an older runtime pack fails at link time
instead of misprinting.

Tests: corpus 4295 covers -0/NaN/exponent spellings inside templates
next to bare arguments, empty parts, more than 16 parts, evaluation
order, a later part reassigning an earlier string, union and optional
values, and async code. A new co-located IR unit test checks the
emitted tags and entry points.

Linux sandbox, ns per line (median of 9 launches, 200k lines), prev =
stack tip, cand = this commit (with the previous one). All outputs
byte-identical to Node.

  probe      target  node   prev  cand
  p-str      pipe    4022   1364  1149
  p-str      file    3376    849   744
  p-str      null    2360    330   247
  p-args     pipe    6687   1441  1343
  p-args     file    5135    969   821
  p-args     null    3831    389   292
  p-write    pipe    2866   1244  1168
  p-write    file    2329    839   782
  p-write    null    1622    309   275
  p-err      pipe    4656   3253  1073
  p-err      file    3659   2099   721
  p-err      null    2749    591   219
  log-lines  pipe    5007   1476  1314
  log-lines  file    3949    965   877
  log-lines  null    3145    446   378
scr_console_capacity compared the raw tag with SCR_ARG_STR, but string
parts of a concatenation argument carry SCR_ARG_GLUE. A glued string part
therefore reserved only 32 bytes while scr_console_render, which masks
with SCR_ARG_KIND, copied its full length, overflowing the on-stack (or
heap) line buffer for any glued part longer than 32 bytes.

Mask with SCR_ARG_KIND in the capacity loop and in the single-string
fast path, the two remaining kind tests in scr_console.c.

Tests: corpus 4296 logs template literals and `+` chains with glued
string parts of 31..70000 bytes, several long parts per argument, and
multi-byte UTF-8 parts. Before this change the --sanitize build aborted
with an AddressSanitizer stack-buffer-overflow; it now matches Node in
the plain and sanitized lanes.
The TypeScript 7 parity job needs baselines for the console corpus
programs this PR adds (4294-console-line-sizes, 4295-console-template-
parts, and 4296-console-glued-long-parts from the overflow fix).
Recorded with SCRIPTC_UPDATE_BASELINES=1 pnpm exec vitest run
packages/compiler/test/ts7/order-parity.test.ts. The re-recorded file
added only these three and #763's 3170, with no existing entry changed.
The 3170 entry is left to #763. These are inserted at their sorted
position, apart from 3170, so the hunks stay independent when the stack
is rebased.

node scripts/test-ts7.mjs --baselines-only now reports only 3170, which
the rebased base branch provides.
The method-replacement census marks a method name as an observed slot
when some write might replace it, and calls of observed methods go
through a checked property lookup whose result converts back to the
declared return type. Two kinds of data writes over-approximated that:

- A computed write into a string-keyed dictionary marked every method in
  the program unless the dictionary's values were primitives. Class
  instances have no implicit index signature, so such a dictionary is a
  class view only when the class declares the index signature, and then
  every method must be assignable to its value type. Any value type no
  function can inhabit now qualifies: primitives, or non-callable object
  types that require a property functions lack, or weak types sharing no
  property with functions.
- A named write like `this.literal = token` marked `literal` for every
  class. The write can only replace a method when the written property
  can hold a function, so property writes are now judged by the written
  property's type (prefetched in one batch).

Compiling vercel-labs/tsc-ts failed after the canary promotion: one
`changedProjects[id] = {...}` write marked every method, OptionNameMap.get
and the JSON parser's literal() calls returned unknown, and their Map-typed
results could not convert back (SC1100/SC1101, 61 errors).

Also skip abstract methods when materializing observed prototype slots.
An abstract declaration emits no prototype member, so lookup continues
to the ancestors; building its value instead threw an uncaught
PoisonError out of the reachability fixed point (the tsc-ts crash).

Tests: corpus 3171 (fails with SC1100/SC1101 before this change) matches
Node in the plain and sanitized lanes; the 112 corpus programs matching
abstract|prototype|method|callback pass; class-dynamic-dispatch unit
tests pass. tsc-ts now builds and type-checks itself with output
identical to tsgo 7.0.2.
Write console and stream output directly to the descriptor (log-lines -11%)
Fix compiling tsc-ts: keep data writes out of the method replacement census
- Merge main's runtime optimizations and CI changes with canary's frontend improvements.
- Preserve adaptive cycle reclamation with inline allocation and deallocation.
- Repair native bootstrap callback narrowing and emitter expectations.
@vercel

vercel Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
scriptc Ready Ready Preview, v0 Oct 8, 2026 11:22pm UTC

- Keep merged-pipe ordering probes small enough for deterministic Node output.
- Preserve large-write coverage for shared files and independent streams.
@ctate
ctate merged commit c466226 into canary Oct 8, 2026
84 checks passed

This branch was successfully deployed

1 active deployment
Preview — 269c4130 Deployed Oct 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants