Skip to content

CAS GC: batch the namespace janitor's deletes under a soft time budget - #2483

Open
filimonov wants to merge 6 commits into
antalya-26.6from
fix/antalya-26.6/cas-gc-janitor-time-budget
Open

filimonov wants to merge 6 commits into
antalya-26.6from
fix/antalya-26.6/cas-gc-janitor-time-budget

Conversation

@filimonov

Copy link
Copy Markdown
Member

The namespace janitor (GC phase 16) deleted a dropped table's dead _log/_snap keys one page per folding round, with one HEAD + DELETE per key. It now deletes a page's dead keys as write-once batches on the GC I/O pool and keeps taking pages while they delete, until a soft 20 s budget.

A smaller implementation of #2468 (19 files, +858/−65 against 40 files, +3296/−729) and a replacement for it.

Changes

  • Gc::runNamespaceJanitor repeats NamespaceJanitor::runOnePage while a page deleted something and published its cursor. No page starts after 20 s; the page in progress finishes. A pool without debris still costs one LIST per round.
  • After the last page of cas/ns/ the pass continues from its beginning, so a pass that began mid-stream also reaches the debris before its cursor.
  • Dead-life canonical _log/_snap keys are write-once: a page deletes them with no HEAD and no token, split evenly across the GC I/O pool, at most cas_gc_bulk_delete_chunk_keys per request. A store without a batch delete gets one request per key from the same jobs. _ckpt and _files keep the exact-token delete.
  • A failed batch, or a job the pool refused, leaks its keys and the cursor advances. Lost authority leaves the page unpublished.
  • A failed LIST resets the cursor only on the first page of a pass.
  • Phase metrics: delete_requests, budget_exhausted. leaked is zero for a page that did not publish its cursor.

Risks / notes

  • On S3 a 1000-key page goes as up to cas_gc_io_concurrency requests (16 × 63 keys by default) instead of one, and on a store without a batch delete each job first spends one refused batch call. Capping by the storage's batch-delete limit (objects_chunk_size_to_delete) is a separate PR.
  • Keys of a successful batch count as deleted even when absent. With a LIST that lags behind deletes, a pass can keep taking pages until the budget.
  • Each page reads and writes gc/maintenance_state: 2 extra requests per 1000 keys compared with one publication per phase.
  • No on-S3 format change.

Testing

  • Unit: CAS* gate, debug and ASan. New suite CASNamespaceJanitorBatches (12 tests) and CASWriteOnceKey.StreamKeyMintedFromAListedKeyMustBeThatKey; with the non-test sources reverted, the tests of the new behavior fail.
  • Integration: tests/integration/test_cas_janitor_drain on RustFS, a disk with a batch delete and one with support_batch_delete = false: 2500+ dead keys drain in one deleting round.
  • Not measured: the A/B on the RustFS stand from PR 2468 was not repeated for this branch.

Related: #2468

Changelog category (leave one):

  • Performance Improvement

Changelog entry (a user-readable short description of the changes that goes to CHANGELOG.md):

CAS GC drains a dropped table's reference stream in one round: dead stream keys are deleted in batches on the GC I/O pool, under a soft time budget instead of one page per round.

Documentation entry for user-facing changes

  • Documentation written in this PR (docs/en/antalya/cas/architecture/garbage-collection.md, docs/en/operations/system-tables/cas_gc_log.md, docs/en/antalya/cas/configuration.md)

CI/CD Options

Regression jobs to run:

  • CAS (content-addressed storage; Antalya only)

filimonov and others added 6 commits October 5, 2026 14:01
The namespace janitor deleted a dropped table's dead `_log`/`_snap` keys one
page per round, one `HEAD` + `DELETE` per key. `Gc::runNamespaceJanitor` now
repeats `NamespaceJanitor::runOnePage` while a page deleted something, until a
20 s soft budget. A page deletes its dead `_log`/`_snap` keys as write-once
batches split evenly across the GC I/O pool; `_ckpt` and `_files` keep the
exact-token delete. After the last page of the stream the pass continues from
its beginning, so debris on both sides of the cursor drains in one round.

A smaller implementation of #2468.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
A pool that refused a delete job threw out of `runOnePage` and lost that
page's counters and anomalies. The unscheduled keys now count as leaked, like
a failed request, and the page is published.

Tests: a quiet pool larger than one page takes one `LIST`; a refusal after
the first job; the split across the pool on a store without a batch delete;
a dead and a reborn life on one page. Docs: the phase 16 cost table and
`namespaces.md` described one page per round and a `HEAD` + `DELETE` per
object.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
The live namespace had no `_ckpt`, so the fold suppressed the round and the
janitor page was never decided: the test passed with the `deleted > 0` clause
of `more` removed. The namespace now has a checkpoint and 1200 `_files` keys,
and the test asserts the published cursor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
…tch-size setting

`cas_gc_bulk_delete_chunk_keys` now also caps the namespace janitor's batch
deletes; its description said only phases 15 and 17. The test pins that a
listed key which is not the canonical key of its parsed identity gets no
write-once form.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
- A failed `LIST` reset the persisted cursor on any page. It now resets only
  on the first page of a pass; later the cursor the previous page published
  stays, so one transient failure mid-drain does not send the next round back
  to the stream start.
- `leaked` is zero for a page that did not publish its cursor: the page is
  listed again, so authority lost mid-page no longer inflates
  `CASGCNamespaceCleanupLeaks`.
- New phase metric `delete_requests`: the delete jobs run on the GC I/O pool,
  so the phase row's `ProfileEvents` miss their requests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
A batch store drains in under a tenth as many delete calls as keys; a store
without a batch delete makes at least one call per key.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: Mikhail Filimonov <mfilimonov@altinity.com>
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

Workflow [PR], commit [a39bccb]

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant