You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
I'm preparing a proposal to the NLnet CodeSupply call and want to make sure it complements FederatedCode and PurlDB rather than duplicating them.
Idea
An open dataset linking package versions (PURLs, starting with PyPI, npm and crates) to Software Heritage SWHIDs, with several confidence levels, from attested commits (SLSA / PEP 740) down to heuristic matching, each link tagged with how it was established.
Data sources: SWH public exports and deps.dev, no heavy API use.
Output: Parquet exports plus a FederatedCode data cluster, so PurlDB or VulnerableCode could consume it directly.
Questions
Does this overlap with anything existing or planned in PurlDB, purl2vcs or FederatedCode?
Is FederatedCode the right place to publish this kind of data? If so, is there a data kind or naming convention I should follow?
Any pitfalls you'd expect from your experience with purl2vcs?
Since AboutCode is part of the CodeSupply consortium, I'm only asking for technical feedback here, not for involvement in or support for the proposal.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Hi AboutCode team,
I'm preparing a proposal to the NLnet CodeSupply call and want to make sure it complements FederatedCode and PurlDB rather than duplicating them.
Idea
An open dataset linking package versions (PURLs, starting with PyPI, npm and crates) to Software Heritage SWHIDs, with several confidence levels, from attested commits (SLSA / PEP 740) down to heuristic matching, each link tagged with how it was established.
Data sources: SWH public exports and deps.dev, no heavy API use.
Output: Parquet exports plus a FederatedCode data cluster, so PurlDB or VulnerableCode could consume it directly.
Questions
Since AboutCode is part of the CodeSupply consortium, I'm only asking for technical feedback here, not for involvement in or support for the proposal.
Thanks!
All reactions