Skip to content

Speed up Collection.add() and save() on chained nodes - #107

Open
gefleury wants to merge 1 commit into
openMetadataInitiative:pipelinefrom
gefleury:speed_up_add_nodes
Open

gefleury wants to merge 1 commit into
openMetadataInitiative:pipelinefrom
gefleury:speed_up_add_nodes

Conversation

@gefleury

@gefleury gefleury commented Oct 9, 2026

Copy link
Copy Markdown

Problem

On deeply nested graphs, Collection.add() and Collection.save() take time that grows exponentially with depth. There are two causes:

  1. Collection._add_node() loops over node.links, which already lists all descendants, and then recurses into each of them, which does the same again.
  2. Node.links computes each child's links twice: hasattr(item, "links") evaluates the property, then item.links evaluates it again.

Solution

  • Node._direct_links (new private property): returns only the linked nodes directly referenced by a node. Embedded nodes are looked through.
  • Collection._add_node() and save() loop over _direct_links instead of links. The recursion already reaches the deeper levels. This alone fixes both causes 1 and 2.
  • Node.links is rewritten on top of it. Each child's links are now computed once. This fixes cause 2 for code that still uses links .

Tests on artificial graphs and real datasets

This PR is intended as a refactor only. links should return the same list as before, and a Collection should end up with the same nodes, the same blank-node ids and the same order. I checked this by comparing the old and new code on synthetic graphs and on a few real datasets fetched from the EBRAINS KG (v4). For each of them, I loaded it with Collection.load(), added its nodes to a new Collection and saved it, with the old and the new code (openMINDS 0.6.1).

These KG records are rather small and shallow, so the gain is modest. Deeper chains occur naturally in collections built from NWB files with nwb2openminds: one SubjectState per session, chained with descended_from. For one such collection (53 states, longest chain 13), add + save went from 156 s to 0.3 s, with identical output.

Tests in tests/

  • The new tests check _direct_links, and count _add_node calls on a 12-node chain (12 now, 2048 before).
  • Two further tests ( test_links_order_and_duplicates and test_blank_node_ids_follow_depth_first_order ) only show that the behaviour is unchanged, and they pass on the old code too. They check the order of links, its duplicates and the numbering of blank-node ids. If these details are not meant to be guaranteed, the two tests can be dropped.

- Collection only follows direct links (new Node._direct_links)
- Node.links is built on _direct_links: same output, each child's links computed once
- Add tests for _direct_links, the call count and the unchanged behaviour
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant