|
| 1 | +# Engine CPU benchmarks |
| 2 | + |
| 3 | +Two benchmarks for the paths the production engine service spends its CPU in, plus a |
| 4 | +`.cpuprofile` analyzer. Neither runs in CI: they take minutes, attach the V8 profiler, and |
| 5 | +report numbers rather than assert on them. |
| 6 | + |
| 7 | +| bench | what it covers | where | |
| 8 | +| --- | --- | --- | |
| 9 | +| `engineHttp.bench.test.ts` | the full request stack for `engine/v1/worker-actions/*` | `apps/webapp` | |
| 10 | +| `runEngineLifecycle.bench.test.ts` | run-engine and run-queue with no HTTP in the way | `internal-packages/run-engine` | |
| 11 | + |
| 12 | +Artifacts (profiles + JSON summaries) land in `.bench/` at the repo root, which is gitignored. |
| 13 | + |
| 14 | +## HTTP bench |
| 15 | + |
| 16 | +Measures what a managed supervisor actually does: dequeue, start attempt, heartbeat, |
| 17 | +read latest snapshot, complete attempt. Needs a built webapp. |
| 18 | + |
| 19 | +```bash |
| 20 | +pnpm run build --filter webapp |
| 21 | +cd apps/webapp |
| 22 | +pnpm run test:bench |
| 23 | +``` |
| 24 | + |
| 25 | +It spawns a real webapp against throwaway Postgres and Redis containers, seeds a production |
| 26 | +environment with a promoted managed deployment, fills the worker queue over the public |
| 27 | +trigger API, then drives a closed-loop supervisor pool for the measured window. |
| 28 | + |
| 29 | +The webapp is spawned with `--inspect` and profiled over CDP, so the profile covers only the |
| 30 | +measured window rather than boot. Event-loop utilization is sampled **inside** the webapp |
| 31 | +process over the same connection. |
| 32 | + |
| 33 | +Knobs: |
| 34 | + |
| 35 | +| var | default | meaning | |
| 36 | +| --- | --- | --- | |
| 37 | +| `BENCH_RUNS` | 1200 | runs queued before the window opens | |
| 38 | +| `BENCH_SUPERVISORS` | 16 | concurrent virtual supervisors | |
| 39 | +| `BENCH_HEARTBEATS` | 2 | heartbeats per run | |
| 40 | +| `BENCH_DURATION_MS` | 60000 | measured window | |
| 41 | +| `BENCH_SAMPLING_INTERVAL_US` | 200 | V8 sampling interval | |
| 42 | +| `BENCH_PROFILE_NAME` | `engine-http` | artifact basename | |
| 43 | +| `BENCH_EXTRA_ENV` | — | JSON merged into the webapp's env | |
| 44 | +| `BENCH_OUT_DIR` | `<repo>/.bench` | artifact directory | |
| 45 | + |
| 46 | +`BENCH_EXTRA_ENV` plus `BENCH_PROFILE_NAME` is how you A/B a single flag: |
| 47 | + |
| 48 | +```bash |
| 49 | +BENCH_RUNS=5000 BENCH_SUPERVISORS=24 BENCH_DURATION_MS=90000 \ |
| 50 | + BENCH_PROFILE_NAME=engine-http-no-elm \ |
| 51 | + BENCH_EXTRA_ENV='{"EVENT_LOOP_MONITOR_ENABLED":"0"}' \ |
| 52 | + pnpm run test:bench |
| 53 | +``` |
| 54 | + |
| 55 | +Run the same size for both arms and compare `on-cpu ms per completed run` rather than |
| 56 | +throughput: throughput on a laptop moves ~5% run to run, on-CPU per unit of work is far |
| 57 | +steadier. |
| 58 | + |
| 59 | +## Run-engine bench |
| 60 | + |
| 61 | +No HTTP, no webapp: drives `RunEngine` directly so engine and queue costs are not mixed with |
| 62 | +request-stack overhead. Profiles two phases separately, because blending them hides which one |
| 63 | +owns a hot frame. |
| 64 | + |
| 65 | +```bash |
| 66 | +cd internal-packages/run-engine |
| 67 | +pnpm run test:bench |
| 68 | +``` |
| 69 | + |
| 70 | +Knobs: `BENCH_RUNS`, `BENCH_CONSUMERS`, `BENCH_HEARTBEATS`, `BENCH_CONCURRENCY_LIMIT`, |
| 71 | +`BENCH_SAMPLING_INTERVAL_US`, `BENCH_OUT_DIR`. |
| 72 | + |
| 73 | +The driver shares a process with the code under measurement, so its own cost is in the |
| 74 | +profile. It is a thin await loop and appears under its own frames rather than smeared across |
| 75 | +engine frames. |
| 76 | + |
| 77 | +## Analyzing a profile |
| 78 | + |
| 79 | +```bash |
| 80 | +pnpm --filter webapp exec tsx test/bench/analyzeProfile.ts .bench/engine-http.cpuprofile --top 30 |
| 81 | +``` |
| 82 | + |
| 83 | +Three views: CPU by bucket (which package owns the cycles), hottest frames by self time (what |
| 84 | +to go fix), and hottest frames by total time (entry points, and a check that the load |
| 85 | +exercised the route mix you intended). Frames are symbolicated through the build's source |
| 86 | +maps, so bundled chunks report as the source files they came from. |
| 87 | + |
| 88 | +Percentages are shares of **on-CPU** time, with V8's `(idle)` and `(program)` excluded. A |
| 89 | +share of wall clock would make everything look cheap whenever the bench was IO-bound. |
| 90 | + |
| 91 | +`--json <path>` writes the full analysis for diffing two runs. |
0 commit comments