fix(jsc): preserve generational worker heap collection on main - #15
Merged
Merged
Conversation
|
Dependency limit exceeded — report not shown. This pull request scan exceeded the 10,000-dependency limit applied to this scan, so the results are incomplete and may be inaccurate. To avoid reporting false positives, Socket has not posted a report. Upgrade your plan to raise the dependency limit and get complete reports, or view the partial scan in the dashboard. Socket is always free for open source. If this is a non-commercial open source project, contact us to request a free Team account. |
This was referenced Oct 5, 2026
steipete
force-pushed
the
fix/worker-gc-cadence-main
branch
from
October 6, 2026 03:55
78259d3 to
69324ba
Compare
steipete
added a commit
that referenced
this pull request
Oct 6, 2026
A Worker using `resourceLimits.maxOldGenerationSizeMb` can keep running full GCs for seconds after its managed live heap exceeds the limit and Bun requests termination. The one-byte allocation allowance leaves very little execution in optimized JavaScript, so JSC's signal sender repeatedly misses the JIT window needed to install trap breakpoints. The correct `ERR_WORKER_OUT_OF_MEMORY` and exit code 1 arrive late. Invalidate optimized code from the mutator's GC epilogue when a heap-limit breach has occurred and a trap is pending. This reuses the API-lock entry invalidation helper and respects `DeferTraps`; existing JavaScript safe points deliver the exception. Full-GC live accounting, ArrayBuffer exclusions, heap budgets and watchdog ownership remain intact. This is independent of the pending cadence changes in #14/#15. Median time from Worker construction to the error event, five independent runs per arm and cap, using matched Bun source and build settings: | Platform / cap | Unchanged engine median | Candidate median | Node 24 median | | --- | ---: | ---: | ---: | | macOS arm64 / 32 MiB | 544.7 ms | 20.3 ms | 54.1 ms | | macOS arm64 / 256 MiB | 6,335.7 ms | 97.3 ms | 271.4 ms | | macOS arm64 / 512 MiB | 5,050.1 ms | 257.6 ms | 526.5 ms | | Linux x64 / 32 MiB | 1,881.8 ms | 56.6 ms | 117.0 ms | | Linux x64 / 256 MiB | 4,623.9 ms | 372.2 ms | 640.6 ms | | Linux x64 / 512 MiB | 19,451.1 ms | 1,060.1 ms | 1,404.2 ms | Node is the contract reference; the macOS Node observations come from the preceding same-host series. The Linux comparison interleaves all three arms. All final Linux trials report the expected error/exit and no guard intervention. These are measurements of a retained-object loop, not a universal runtime benchmark. The deterministic native regression blocks the signal sender and fails the unchanged engine on a second over-cap full GC. The candidate unwinds after one, in 0.075 ms on macOS and 0.095 ms on Linux. A native survivor retains 30 MiB under a 32 MiB cap while churning temporary arrays; it ends at 31,541,027 managed bytes without OOM. The added Bun Worker survivor retains 22 MiB plus 64 MiB external storage on both platforms. A portable Node/Bun reference uses 20 MiB retained because V8's VM overhead leaves different usable space on x64. Validation: 15 resource-limit and 310 surrounding Worker cases pass on each platform; nine native controls pass on each. Four unchanged OpenClaw consumers pass before/after, 27 cases per arm, including a warm worker with a 512 MiB cap. Scoped P2 autoreview is clean. Linux artifact compilation caught a test include-path issue hidden by the macOS header map; the corrected packaged-header spelling is re-reviewed and native-tested. The first corrected-head Linux run passed all functional commands, then failed a stale seven-row native-output inventory. The two new cases require nine rows. The downloaded artifact contains all nine passes; replaying the corrected verifier passes 83 selected files / 85 result rows with zero regressions and all engine/compatibility gates. The one-line inventory correction is P2-clean. Fresh exact-head [engine CI](https://github.com/openclaw/WebKit/actions/runs/37411955745) and [Linux ARM64 CI](https://github.com/openclaw/WebKit/actions/runs/37411955459) pass at `7335155ebd04d6ba911d1a1b920a787c319fef25`. Upstream's resource-limit proposal oven-sh/bun#32896 remains open; the search found no existing trap-delivery fix. Delivery requires the next qualified OpenClaw WebKit artifact release and a matching Bun rebuild. Release-line PR; main companion: #19.
steipete
added a commit
that referenced
this pull request
Oct 6, 2026
A Worker using `resourceLimits.maxOldGenerationSizeMb` can keep running full GCs for seconds after its managed live heap exceeds the limit and Bun requests termination. The one-byte allocation allowance leaves very little execution in optimized JavaScript, so JSC's signal sender repeatedly misses the JIT window needed to install trap breakpoints. The correct `ERR_WORKER_OUT_OF_MEMORY` and exit code 1 arrive late. Invalidate optimized code from the mutator's GC epilogue when a heap-limit breach has occurred and a trap is pending. This reuses the API-lock entry invalidation helper and respects `DeferTraps`; existing JavaScript safe points deliver the exception. Full-GC live accounting, ArrayBuffer exclusions, heap budgets and watchdog ownership remain intact. This is independent of the pending cadence changes in #14/#15. Median time from Worker construction to the error event, five independent runs per arm and cap, using matched Bun source and build settings: | Platform / cap | Unchanged engine median | Candidate median | Node 24 median | | --- | ---: | ---: | ---: | | macOS arm64 / 32 MiB | 544.7 ms | 20.3 ms | 54.1 ms | | macOS arm64 / 256 MiB | 6,335.7 ms | 97.3 ms | 271.4 ms | | macOS arm64 / 512 MiB | 5,050.1 ms | 257.6 ms | 526.5 ms | | Linux x64 / 32 MiB | 1,881.8 ms | 56.6 ms | 117.0 ms | | Linux x64 / 256 MiB | 4,623.9 ms | 372.2 ms | 640.6 ms | | Linux x64 / 512 MiB | 19,451.1 ms | 1,060.1 ms | 1,404.2 ms | Node is the contract reference; the macOS Node observations come from the preceding same-host series. The Linux comparison interleaves all three arms. All final Linux trials report the expected error/exit and no guard intervention. These are measurements of a retained-object loop, not a universal runtime benchmark. The deterministic native regression blocks the signal sender and fails the unchanged engine on a second over-cap full GC. The candidate unwinds after one, in 0.075 ms on macOS and 0.095 ms on Linux. A native survivor retains 30 MiB under a 32 MiB cap while churning temporary arrays; it ends at 31,541,027 managed bytes without OOM. The added Bun Worker survivor retains 22 MiB plus 64 MiB external storage on both platforms. A portable Node/Bun reference uses 20 MiB retained because V8's VM overhead leaves different usable space on x64. Validation: 15 resource-limit and 310 surrounding Worker cases pass on each platform; nine native controls pass on each. Four unchanged OpenClaw consumers pass before/after, 27 cases per arm, including a warm worker with a 512 MiB cap. Scoped P2 autoreview is clean. Linux artifact compilation caught a test include-path issue hidden by the macOS header map; the corrected packaged-header spelling is re-reviewed and native-tested. Exact-head [engine CI](https://github.com/openclaw/WebKit/actions/runs/37409339937) and [Linux ARM64 CI](https://github.com/openclaw/WebKit/actions/runs/37409339952) pass at `095a11a7467346920209142bb22974d213579598`. Upstream's resource-limit proposal oven-sh/bun#32896 remains open; the search found no existing trap-delivery fix. Delivery requires the next qualified OpenClaw WebKit artifact release and a matching Bun rebuild. Main companion to release-line #18. Main adds the same native probe and a compile/run step in its existing qualification script; the release line already has that hook.
Separate nursery pacing from bounded post-full growth, conservatively verify worker heap pressure with a full collection, and preserve progress through the Bun timer adapter. Retain the existing post-GC termination delivery and live-only heap accounting.
steipete
force-pushed
the
fix/worker-gc-cadence-main
branch
from
October 6, 2026 08:50
69324ba to
1cd3613
Compare
steipete
marked this pull request as ready for review
October 6, 2026 09:31
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Worker heap limits currently force automatic collections to be full and collapse the whole-heap target to the nursery allowance. Vitest workers consequently retrace large surviving code graphs instead of using young-generation collection. Native diagnostics reduce providers full/Eden counts from 698/154 to 112/1020 and messaging from 348/37 to 61/452, with unchanged complete test outcomes. These are diagnostic collection counts, not physical scan-byte measurements or a demonstrated compile-cache leak.
Separate nursery pacing from full-heap growth, and cap a budgeted worker's post-full target at twice its current heap ledger while retaining the collector minimum and smaller pressure-driven targets. Conservative checked accounting requests full verification when the managed-memory budget may be exhausted. Only the strict post-full survivor check reports OOM; the full-collection buffer snapshot preserves ArrayBuffer exclusion through Eden collections. Mandatory verification drives the existing collection ticket while idle through the matching Bun timer adapter. Cached-data, debugger, source/referrer identity and public accounting contracts are preserved.
This main-branch companion retains #19's post-GC termination delivery and the allocator-exit correction. Relative to the measured cadence source, the collector adds only the inherited termination hook; the growth policy and cadence regression remain unchanged. The release-line companion is #14. Bun timer companion openclaw/bun#129 remains separately coordinated and must accompany a consuming rebuild. The policy applies generically to budgeted Worker VMs.
A quiet Linux host completed 100 ordinary observations, four per arm in five cells, with pairwise ABBAABBA order against Node 24.21.0, the public Bun pin, a matched untreated fork, and the candidate with default or CPU/worker-ratio marker counts. All full outcome inventories, isolation/resource checks, zero-steal gates and final source/binary checks passed; no observation was retried or excluded.
Eight-CPU messaging's wall ratio interval is [0.9947, 1.0135]: no demonstrated wall win or equivalence. Bun remains slower than Node throughout. Constrained full providers reaches the memory cap in three of four candidate runs, with reclaim pressure but zero OOMs; capped RSS is not an unconstrained memory estimate. The marker ratio is rejected because it regresses two-CPU messaging by 1.49% (ratio interval [1.0074, 1.0230]); no marker override is added.
The measured candidate passed native Worker/VM/V8, ArrayBuffer accounting in five modes, resource-limit, N-API, serializer, inspector, stack/cache, and engine stress/module gates. The strict source-blind held-after-Eden observation remains blocked; the native idle-boundary regression passes. Original failures and superseded cohorts are retained separately. Earlier fixture/transport repairs landed in openclaw/openclaw#165661 and #165769.
The refreshed main head passed independent P2 review, pipeline integrity tests, and fresh exact-head engine and ARM64 qualification. The ARM64 lane passed all 1,779 stress and 1,639 module variants plus accounting and sampling in five modes. Native Linux baseline and candidate finished with identical passing 46-file inventories; the candidate also passed 17 Worker resource-limit cases. Both grouped arms first encountered the same inherited epoll registration error and recovered the same 26 unfinished files through the existing runner; no failed-file retry or GitHub rerun was used. The macOS cross-compilation lane also passed. Artifact publication is outside this PR.
The maintained Gateway check completed 64 scored observations plus two cache primes with four processes per arm/condition, fixed physical CPU placement, fresh Gateway state, separately controlled empty/retained transpiler caches, zero-steal checks and clean process settlement. Against the matched control, cold/warm startup readiness changed +0.64%/-0.14%, empty-cache restart -4.93%, retained-cache startup +0.97%, and retained-cache restart +4.21%. Peak Gateway-inclusive RSS medians decreased in all five conditions; every RSS upper bound stayed below 1.05. Only retained-cache restart latency did not meet the declared bound. No observations were discarded or added to obtain a pass.
Candidate cold root HWM was 732.4 MiB versus 736.8 MiB matched control, 744.4 MiB published667c and 805.5 MiB Node26 on the same host. Root RSS includes worker threads; sampled descendant-process memory is also reported separately. Observed worker starts remained comparable, and all observed descendants settled without forced cleanup. Retained caches were startup-primed; the first retained restart added 18 files per arm, so this is not a separately restart-primed experiment.
A new, separately predeclared restart-specific experiment now passes the product gate. Two restart primes per arm produced identical retained-cache inventories; twelve scored processes per arm then ran in six interleaved control/growth/growth/control blocks. All 24 scored processes and four primes passed readiness, clean exit, identity, affinity, all-CPU zero-steal and cleanup checks. No observations or caches were reused from the earlier cohort. A separate AWS attempt stopped during unscored priming for a steal tick and contributed no scored data.
The maintained harness and product commit, Node 26 controller, exact previously measured control/growth binaries, fresh Gateway state, four distinct physical Gateway cores and disjoint controller cores were held fixed within the new comparison. Its observer gained child-subreaper ownership before any run to preserve cleanup across the harness's detached processes. Three cycles were first summarized within each process; 20,000 independent process-median bootstrap draws used seed 174187. This cohort is independent of the earlier four-process comparison and does not establish that cache priming caused the earlier uncertainty. No marker policy or worker-type exception was added.
The whole-suite Node attempts exposed missing Git metadata and an unbounded SQLite fixture WAL; their failures remain preserved and no whole-suite equality pass is claimed. The independent test-only pacing repair openclaw/openclaw#165940 is merged and #165938 is closed, with dual-runtime proof, P2, exact-head CI and ClawSweeper qualification. The follow-up whole-suite attempts also retained a proof-wrapper orphan-reaping defect and missing host executables. Those invalid attempts are archived. The wrapper was corrected and independently checked; required zip/Ruby tools are installed before a fresh full Node/Bun pair on identical host inputs. No whole-suite equality pass is claimed yet. The fixture finding is separate from the cadence decision. Diagnostic restart profiling was prepared but not run because the follow-up gate passed. No artifact publication is requested.