Repository navigation
perf(storage): stop producing and reading the UUID membership index - #1925
Conversation
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
cc24993 to
8c44329
Compare
213a18b to
565d894
Compare
…1902) Records the outcomes of the existing-identity check on the membership-index path before it is removed, so the removal can show which outcomes it kept and which it changed. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… current path (#1902) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…1902) The index (`topology/uuid-membership/manifest.json`, `identities-v5-*.uuidx`, `node-surrogates-v5-*.uuidx`, `topology-receipt.json`) duplicated the UUID-ordered `node_uuid`/`edge_uuid` columns of the published Parquet at 17-25 B per edge and was rewritten with every generation. Under the pre-v1 policy the format changes in place. - Neither the bulk nor the staged encoder writes it. The bulk builder no longer spills the sorted edge UUIDs to scratch for it, and the staged encoder keeps no retained parent index. Only the ordinal node-identity facet is encoded. - `TopologyIdentityProbe` answers every identity question from the published Parquet: a sorted, deduplicated candidate set prunes row groups by `min/max`, then pages by the column index, and decodes only the selected pages of `node_uuid` (and `node_id`) or `edge_uuid`. `UuidProbeMetrics` reports `pages_considered` and `pages_read`. Footers are cached by file identity, and a hit is re-pointed at the path asked for, because hydration hard-links one object under many workspace paths. - Append validation and the writer commit refuse a UUID that is a live node, a live edge or a deleted entity (one namespace), and a missing endpoint. - Deleted UUIDs are never reusable. `topology/deleted_identities.parquet`, a sorted unique column carried forward by each generation that deletes, replaces the index's tombstones; a graph that never deletes never has one. - Readers moved off the index: writer endpoint lookup and commit, the topology rewrite participant, label lookup by UUID, graph publication, branch import, composite publish, bulk validation, resumable construction, search node identity, and the ordinal discovery freshness check. Projects that still carry the files open; the files are ignored. - The G500 certification evidence drops the index's fsync and write counters, and its storage category policy now treats the identity category as node-bearing. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The identity authority page now describes the Parquet probe (row-group and page-index pruning, the deleted-identities record, the legacy-project behaviour). ADR 0049 and 0058, the storage and resumable-import pages, the direct-fsync guard, the fsync site census, the digest census overrides and the storage CI filter drop the membership index's functions and tests. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nds, refuses and round-trips (#1902) The fixture is a real durable project written by the binary that produced the UUID membership index: an initial bulk build and one staged append, so the index carries a base run and a delta run. Once nothing reads the index the files are ordinary graph entries; the test proves such a project still opens, refuses duplicate and missing identities from the Parquet, accepts appends, exports, verifies and imports, and that a deleted UUID stays spent. A second test builds a new project and proves it never writes the index and its portable bundle never names one. The bundle check is validated against the legacy bundle, which must name the index files. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
#1902) Pre-upgrade deleted edge UUIDs become reusable, and search verifies nodes by repeat-check when a project has no ordinal facet. Both are accepted under the pre-v1 in-place format policy. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…y probe and document the legacy upgrade (#1902) A fragment replaced after its footer was read is refused when decoded, because the pruning metadata would no longer describe it. A UUID held by two rows is a typed error. Absent and truncated statistics are tested to mean 'may hold'. The probe's lock-free contract is stated: it is authoritative only at commit, where the rewrite lock refuses a commit whose probed generation moved. The docs say that every pre-upgrade deletion becomes reusable and that the stale index files are inert and carried forward. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…obe (#1902) Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…enamed ordinal pass and the retired index directory (#1902) The identity probe keyed its footer cache on Unix-only metadata, so Windows Storage did not compile; it now uses the portable file identity, length and modification time of an open handle. The CLI bulk-build report names its pass ordinal, and a search test no longer deletes an index directory a new project never has. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
565d894 to
5727328
Compare
Closes #1902
Slice of #1881, decision 5. New generations no longer produce or read the UUID membership index (
topology/uuid-membership/manifest.json,identities-v5-*.uuidx,node-surrogates-v5-*.uuidx,topology-receipt.json). It duplicated the UUID-orderednode_uuid/edge_uuidcolumns of the published Parquet at 17-25 B/edge. Under the pre-v1 policy the format changes in place; a project that still carries the files opens normally and the files are ignored. The ordinal node-identity facet that shares the directory is unchanged.What changed
TopologyIdentityProbe(topology_identity.rs) answers every identity question from the published topology Parquet. A sorted, deduplicated candidate set prunes row groups bymin/max, then pages by the Parquet column index, and decodes only the selected pages ofnode_uuid(+node_id) oredge_uuid.UuidProbeMetricsgainspages_consideredandpages_read;TopologyWriteWorkcarries both.topology/deleted_identities.parquet(sorted, unique, carried forward by each generation that deletes; a graph that never deletes never has one). It replaces the index's tombstones.a_cached_footer_never_sends_a_probe_to_the_path_it_was_first_read_from, red before the fix).scripts/ci/schemas/g500-certification.schema.json, its validator test,scale_g500_ladder.rs), and the ladder treatsuuid_and_surrogatesas node-bearing.Reader inventory (every consumer of the index on
main, and where it reads now)register_existing_endpoints) and commit validation (uuid_membership/topology_delta.rs)TopologyIdentityProbe+ deleted-identities recordgraph_construction.rs,shape.rs,graph_construction_encoding.rs,inventory.rs)bulk.rs,bulk/identities.rs,bulk/scratch_edges.rs)uuid_membership/{probing,rebuild,construction,maintenance}.rs)label_membership_counts.rs(labels by UUID)TopologyFilesordinal_identity_v4.rsdiscovery freshness check,project_generation.rsrebuild-if-missing,runtime_entity_labels.rs,mutator.rsnode_selector.rs, added by #1926)cached_identity_probe; the 1M-capped scan stays for label and property selectors onlybulk_construction.rs),branches/import_graph.rs,graph_publication.rs,composite_publish.rs,maintenance.rs,resumable_construction.rs,lib.rscache fieldgraphforge-search/src/node_identity.rs)graph_admission.rs,staging.rs,durable_rewrite.rs, graph files inventory)Remaining references to
uuid-membership/are the ordinal facet (ordinal-v4-*,forward-v4-*,tombstones-v4-*).Acceptance evidence
neither_bulk_route_publishes_a_membership_index(memory and scratch routes), the staged-encoder inventory assertion inencoding_publication/tests.rs, anda_new_project_carries_no_membership_index_and_round_trips.commit_refuses_a_uuid_the_published_topology_already_holds(node/node, node/edge, edge/edge, edge reusing a node UUID, at the writer commit), thebulk_constructionpublication/normalization tests (duplicate node and edge, edge equal to node, missing endpoint, deleted UUID), and the legacy and new-project API tests.pages_read=1/3(identity_bytes_read=106,588), endpoints in the first and last page2/3(115,990; the dictionary page is read once per probed chunk).page_pruning_reduces_the_bytes_the_file_system_serves).legacy_membership_index(real pre-change project: opens, appends, refuses, deletes, exports, verifies, imports, deleted UUID still spent after the round trip), the construction crash tests (ordinal_encoding_crashes_recover_every_durable_boundary,bulk_builderkilled-in-any-pass tests, lane crash tests), andidentity_first_touch(a flipped node Parquet is refused by the commit that probes it).Mutation proofs (each applied to the committed tree, run, reverted)
taken()ignores live nodes and edgeslive_nodesreports every candidate livewave13_validation_display_and_disappeared_endpoint_are_structured)ordinal_encoding_crashes_recover_every_durable_boundaryfailsVerification
make check: exit 0.graphforge-storagenextest: 1796 passed, 0 failed, 9 skipped.graphforge-apinextest: 1668 passed, 0 failed, 11 skipped (includes fix(api): resolve UUID selectors by identity and keep single-pass algorithms off the iteration budget #1926'sanalyst_verb_limits: BFS from a UUID on 1,000,100 nodes). API BDD: 118 scenarios passed.GF_TEST_TMPDIR=/tmp/...: the launcher's default root sits inside the enclosing checkout when the worktree lives under.claude/worktrees, which makes 17repository::*tests see Git-tracked data and one socket test exceedSUN_LEN. Those pass with a short ext4 root; they are environmental.graphforge-storageexports (UuidMembershipIndexand the index maintenance functions are gone;TopologyIdentityProbeis new), and the workspace compiles under clippy.Append-validation time (acceptance 3), contended host
S20 graph (1,048,576 nodes, 16,777,216 edges) receiving an S18-sized append (262,144 nodes, 4,194,304 edges, existing endpoints),
gf import-sessionvalidate and commit, base171120cbdagainst this branch, release builds in separate target dirs (binary sha256 differ), 5 alternating pairs, host load average 14-27 on 16 cores. Full method, commands, digests and per-step load are on the issue: #1902 (comment)The candidate was faster in all five pairs, in both orders, on validate, commit and the sum, and in four pairs its mean load was the higher one. Per-edge ingest cost did not regress. This is contended-host evidence for the direction, not a size: a quiet-host run is still needed for a number to quote. While measuring I found that main's bulk builder fails S20 builds on a nearly full disk (
captured encoded source identity or length changed: the installer records an encoded shard's allocated bytes before writeback, which can add an extent block); it is not caused by this PR and is described on the issue.Accepted behaviour changes for projects written before this change
Both are accepted under the pre-v1 in-place format policy and documented in
docs/book/architecture/uuid-membership-index.md:The stale index files in an upgraded project are inert: no code reads them, orphan collection considers only canonical ordinal artifact names, and they stay in the graph-files inventory (and travel through hydration, export, verify and import) until a future cleanup drops them.
Follow-ups
🤖 Generated with Claude Code