Repository navigation
mass-scan: Create library to collect a whole project version from Software Heritage #77
Description
Activity
Let take
deb://Debian/packages/aalibas a sample (This link is getting from https://annex.softwareheritage.org/public/dataset/graph/latest/popular-4k/parquet/origin.parquet - see aboutcode-org/purldb#565)From https://archive.softwareheritage.org/api/1/origin/deb://Debian/packages/aalib/visits/, I retrieve the snapshot_url.
For example, from https://archive.softwareheritage.org/api/1/snapshot/fd758dee18de4bacab19926589c54d366fd41323/, I extract the version information (e.g., releases/bookworm/main/1.4p5-50).
Using that version entry, I follow the target_url to obtain the corresponding swhid, which identifies the actual “target” object.
With this swhid, I can then submit a request for the archive to be “cooked” and, once ready, download the package at:
https://archive.softwareheritage.org/api/1/vault/flat/swh:1:dir:686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/raw/.Workflow
- Retrieve the "snapshot_url" from
https://archive.softwareheritage.org/api/1/origin/deb://Debian/packages/aalib/visits/ - From the snapshot (e.g.,
https://archive.softwareheritage.org/api/1/snapshot/fd758dee18de4bacab19926589c54d366fd41323/), extract the version entry. - Follow the corresponding "target_url" (e.g.,
https://archive.softwareheritage.org/api/1/release/ac3a9354788c97c1e95640649114296c3a2add07/) to obtain the "swhid". - Use the "swhid" to request the server the "cook" the package and download the package, or browse the code tree directly (e.g.,
https://archive.softwareheritage.org/browse/directory/686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/).
However, the downside of asking the server to cook the package archive is that it can take a long time before the archive is ready for download. In most cases it takes a few minutes, but I’ve also had cases where the cooking process took several days.
Given this, should we rely on the “download as a package archive” approach, or instead fetch each file individually and assemble a virtual codebase ourselves- Retrieve the "snapshot_url" from
Since asking the server to prepare a full package archive for download can take a long time, we can use a different approach. Instead of retrieving the swhid, we can go directly to the target_url such as https://archive.softwareheritage.org/api/1/directory/686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/ and download each file individually (download_url = target_url + "raw/") along with its original structure. Once all files are fetched, we assemble them into a package. This way, we avoid waiting for the server to generate the archive.
https://www.softwareheritage.org/legal/bulk-access-terms-of-use/
I'm a bit concerned about this
2.2 No massive data extraction In order to ensure that these terms of use are consistently applied, extracting significant parts of the contents of the Archive is not authorized.Re
Since asking the server to prepare a full package archive for download can take a long time, we can use a different approach. Instead of retrieving the swhid, we can go directly to the target_url such as https://archive.softwareheritage.org/api/1/directory/686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/ and download each file individually (download_url = target_url + "raw/") along with its original structure. Once all files are fetched, we assemble them into a package. This way, we avoid waiting for the server to generate the archive.
I created a code snippet that download seach file individually with keeping the directory structure. However, I get
Error fetching data: 429 Client Error: Too Many Requests for url: https://archive.softwareheritage.org/api/1/directory/xxxxxxxxxxxxxxxxxxxxif the project has many files. Perhaps asking the server to prepare a full package archive is a better approach as it won't create huge number ofrequestsas fetching each files does.@pombredanne any take on this?
@pombredanne @chinyeungli It seems that we need to talk to SWH directly about the options.
In cdxgen, we construct swhid based on Git blob identifier as explained in this doc.
https://docs.softwareheritage.org/devel/swh-model/persistent-identifiers.html#git-compatibility
So a virtual FS could be the same as the git repo with an index SBOM that connects the dots?
@pombredanne Just want to make it clear for myself that we want to take a single purl as an input and then create a file-structure of the project from data got from Software Heritage? meaning we don't actually need to download the files?
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNeeds prep
SWH does not expose a file system and is mostly file-based, we are PURL-based. Given a bunch of file-by-file scans we need to be able to stich them back in a package afterwards. Or we need to address SWH as if it were a "Virtual Codebase".
For this we should create a library to collect a whole project version from Software Heritage that would then present this as a virtual filesystem.