Skip to content

ci: add concurrency and log-throughput scenarios to the comparative benchmark #391

Description

@codeforester

Goal

Extend the comparative benchmark to cover the two regimes the target audience actually runs in —
parallel fan-out and high-volume logging — so the budgets in docs/performance.md can detect
regressions there.

Background

#309 completed the scenario set for single-invocation cost, and the current
scripts/benchmark_runtime.py covers cold import, cold and warm invocation, JSON success/error,
diagnostics, nested dispatch, and persistence on/off, gated per platform profile. That is a solid
foundation, and the per-invocation cost of persistence is measured and budgeted.

Two dimensions are absent, and both were surfaced by measurement during a 2026-09-30 review:

1. Concurrency. Every scenario is sequential and in-process. Retention takes a blocking exclusive
flock on runs/.base-cli-run-index.lock twice per invocation, so concurrent invocations of the same
CLI serialize on housekeeping. Measured at a58ec109349fa3f3d03eae5b0de078b39ea361a2 on macOS /
Python 3.14.6, warmed to the default 20-bundle steady state:

serial baseline (5 runs):   17.7 18.1 18.6 19.2 18.3 ms
12 concurrent invocations:  50.6 50.3 64.3 64.4 68.3 66.3
                            67.5 66.2 66.9 67.6 68.0 55.8 ms      (~3.5x)

2. Log throughput. Every scenario logs zero or one records
(_measure_base_cli_features(), scripts/benchmark_runtime.py:607-698), so the per-record cost is
unmeasured. Measured end-to-end through a real App with persistent logging:

3000 ctx.log.info calls: 322.4 ms -> 107.5 us/record, 9,305 rec/s

Roughly 20x stdlib logging. A command logging 100k lines spends about 11 seconds in logging.

Both regressions could land today without any gate noticing.

Scope

  • Add a concurrency scenario: N concurrent real invocations of one CLI against a shared cache root,
    warmed to the retention steady state, reporting per-invocation p95 and the ratio to the serial
    baseline. Use subprocesses so file locking is faithful, and keep N modest and deterministic for CI.
  • Add a log-throughput scenario: a fixed record count through ctx.log.info() with persistence
    enabled and disabled, reported as us/record and records/s.
  • Add a retention-steady-state note or scenario making explicit that RetentionPolicy.safe_defaults()
    sets max_total_bytes, which puts every default invocation in the byte-policy column of the
    work-bounds table in docs/performance.md.
  • Extend the base-cli.benchmark schema (version it) and the per-profile budget table.

Acceptance criteria

  • --check fails on a concurrency-ratio regression and on a log-throughput regression.
  • The new scenarios are deterministic enough for a required CI check on all four platform profiles,
    with Windows budgets set separately as persistence already is.
  • docs/performance.md documents both scenarios, their budgets, and the calibration evidence.
  • The benchmark report schema version is bumped and the retained artifact contract is updated.
  • The default-policy byte-walk behaviour is stated in the work-bounds section.

Validation

Calibrate on the development host and on the first hosted run for each profile, following the
existing calibration-evidence practice, before tightening budgets.

Non-goals

  • Do not turn the benchmark into a load-testing harness.
  • Do not gate ratios so tightly that scheduler noise blocks changes; follow the existing p95 approach.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

ciContinuous integration, tests, automation, or release workflowsenhancementNew feature or product improvement

Type

No type

Projects

  • Status
    Backlog

Milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions