Goal
Extend the comparative benchmark to cover the two regimes the target audience actually runs in —
parallel fan-out and high-volume logging — so the budgets in docs/performance.md can detect
regressions there.
Background
#309 completed the scenario set for single-invocation cost, and the current
scripts/benchmark_runtime.py covers cold import, cold and warm invocation, JSON success/error,
diagnostics, nested dispatch, and persistence on/off, gated per platform profile. That is a solid
foundation, and the per-invocation cost of persistence is measured and budgeted.
Two dimensions are absent, and both were surfaced by measurement during a 2026-09-30 review:
1. Concurrency. Every scenario is sequential and in-process. Retention takes a blocking exclusive
flock on runs/.base-cli-run-index.lock twice per invocation, so concurrent invocations of the same
CLI serialize on housekeeping. Measured at a58ec109349fa3f3d03eae5b0de078b39ea361a2 on macOS /
Python 3.14.6, warmed to the default 20-bundle steady state:
serial baseline (5 runs): 17.7 18.1 18.6 19.2 18.3 ms
12 concurrent invocations: 50.6 50.3 64.3 64.4 68.3 66.3
67.5 66.2 66.9 67.6 68.0 55.8 ms (~3.5x)
2. Log throughput. Every scenario logs zero or one records
(_measure_base_cli_features(), scripts/benchmark_runtime.py:607-698), so the per-record cost is
unmeasured. Measured end-to-end through a real App with persistent logging:
3000 ctx.log.info calls: 322.4 ms -> 107.5 us/record, 9,305 rec/s
Roughly 20x stdlib logging. A command logging 100k lines spends about 11 seconds in logging.
Both regressions could land today without any gate noticing.
Scope
- Add a concurrency scenario: N concurrent real invocations of one CLI against a shared cache root,
warmed to the retention steady state, reporting per-invocation p95 and the ratio to the serial
baseline. Use subprocesses so file locking is faithful, and keep N modest and deterministic for CI.
- Add a log-throughput scenario: a fixed record count through
ctx.log.info() with persistence
enabled and disabled, reported as us/record and records/s.
- Add a retention-steady-state note or scenario making explicit that
RetentionPolicy.safe_defaults()
sets max_total_bytes, which puts every default invocation in the byte-policy column of the
work-bounds table in docs/performance.md.
- Extend the
base-cli.benchmark schema (version it) and the per-profile budget table.
Acceptance criteria
--check fails on a concurrency-ratio regression and on a log-throughput regression.
- The new scenarios are deterministic enough for a required CI check on all four platform profiles,
with Windows budgets set separately as persistence already is.
docs/performance.md documents both scenarios, their budgets, and the calibration evidence.
- The benchmark report schema version is bumped and the retained artifact contract is updated.
- The default-policy byte-walk behaviour is stated in the work-bounds section.
Validation
Calibrate on the development host and on the first hosted run for each profile, following the
existing calibration-evidence practice, before tightening budgets.
Non-goals
- Do not turn the benchmark into a load-testing harness.
- Do not gate ratios so tightly that scheduler noise blocks changes; follow the existing p95 approach.
Goal
Extend the comparative benchmark to cover the two regimes the target audience actually runs in —
parallel fan-out and high-volume logging — so the budgets in
docs/performance.mdcan detectregressions there.
Background
#309 completed the scenario set for single-invocation cost, and the current
scripts/benchmark_runtime.pycovers cold import, cold and warm invocation, JSON success/error,diagnostics, nested dispatch, and persistence on/off, gated per platform profile. That is a solid
foundation, and the per-invocation cost of persistence is measured and budgeted.
Two dimensions are absent, and both were surfaced by measurement during a 2026-09-30 review:
1. Concurrency. Every scenario is sequential and in-process. Retention takes a blocking exclusive
flockonruns/.base-cli-run-index.locktwice per invocation, so concurrent invocations of the sameCLI serialize on housekeeping. Measured at
a58ec109349fa3f3d03eae5b0de078b39ea361a2on macOS /Python 3.14.6, warmed to the default 20-bundle steady state:
2. Log throughput. Every scenario logs zero or one records
(
_measure_base_cli_features(),scripts/benchmark_runtime.py:607-698), so the per-record cost isunmeasured. Measured end-to-end through a real
Appwith persistent logging:Roughly 20x stdlib logging. A command logging 100k lines spends about 11 seconds in logging.
Both regressions could land today without any gate noticing.
Scope
warmed to the retention steady state, reporting per-invocation p95 and the ratio to the serial
baseline. Use subprocesses so file locking is faithful, and keep N modest and deterministic for CI.
ctx.log.info()with persistenceenabled and disabled, reported as us/record and records/s.
RetentionPolicy.safe_defaults()sets
max_total_bytes, which puts every default invocation in the byte-policy column of thework-bounds table in
docs/performance.md.base-cli.benchmarkschema (version it) and the per-profile budget table.Acceptance criteria
--checkfails on a concurrency-ratio regression and on a log-throughput regression.with Windows budgets set separately as persistence already is.
docs/performance.mddocuments both scenarios, their budgets, and the calibration evidence.Validation
Calibrate on the development host and on the first hosted run for each profile, following the
existing calibration-evidence practice, before tightening budgets.
Non-goals