Skip to content

mass-scan: Create library to collect a whole project version from Software Heritage #77

Description

@pombredanne

SWH does not expose a file system and is mostly file-based, we are PURL-based. Given a bunch of file-by-file scans we need to be able to stich them back in a package afterwards. Or we need to address SWH as if it were a "Virtual Codebase".

For this we should create a library to collect a whole project version from Software Heritage that would then present this as a virtual filesystem.

Activity

  1. chinyeungli commented on Feb 24, 2026

    @chinyeungli
    Contributor

    Let take deb://Debian/packages/aalib as a sample (This link is getting from https://annex.softwareheritage.org/public/dataset/graph/latest/popular-4k/parquet/origin.parquet - see aboutcode-org/purldb#565)

    From https://archive.softwareheritage.org/api/1/origin/deb://Debian/packages/aalib/visits/, I retrieve the snapshot_url.
    For example, from https://archive.softwareheritage.org/api/1/snapshot/fd758dee18de4bacab19926589c54d366fd41323/, I extract the version information (e.g., releases/bookworm/main/1.4p5-50).
    Using that version entry, I follow the target_url to obtain the corresponding swhid, which identifies the actual “target” object.
    With this swhid, I can then submit a request for the archive to be “cooked” and, once ready, download the package at:
    https://archive.softwareheritage.org/api/1/vault/flat/swh:1:dir:686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/raw/.

    Workflow

    However, the downside of asking the server to cook the package archive is that it can take a long time before the archive is ready for download. In most cases it takes a few minutes, but I’ve also had cases where the cooking process took several days.
    Given this, should we rely on the “download as a package archive” approach, or instead fetch each file individually and assemble a virtual codebase ourselves

  2. chinyeungli commented on Feb 27, 2026

    @chinyeungli
    Contributor

    Since asking the server to prepare a full package archive for download can take a long time, we can use a different approach. Instead of retrieving the swhid, we can go directly to the target_url such as https://archive.softwareheritage.org/api/1/directory/686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/ and download each file individually (download_url = target_url + "raw/") along with its original structure. Once all files are fetched, we assemble them into a package. This way, we avoid waiting for the server to generate the archive.

  3. self-assigned this
    on Feb 27, 2026
  4. chinyeungli commented on Mar 2, 2026

    @chinyeungli
    Contributor

    https://www.softwareheritage.org/legal/bulk-access-terms-of-use/

    I'm a bit concerned about this

    2.2 No massive data extraction
    In order to ensure that these terms of use are consistently applied, extracting significant parts of the contents of the Archive is not authorized.
    
  5. chinyeungli commented on Mar 4, 2026

    @chinyeungli
    Contributor

    Re

    Since asking the server to prepare a full package archive for download can take a long time, we can use a different approach. Instead of retrieving the swhid, we can go directly to the target_url such as https://archive.softwareheritage.org/api/1/directory/686386c64e0bd5499a6bec5ebf3f41d3cb7e6e60/ and download each file individually (download_url = target_url + "raw/") along with its original structure. Once all files are fetched, we assemble them into a package. This way, we avoid waiting for the server to generate the archive.

    I created a code snippet that download seach file individually with keeping the directory structure. However, I get Error fetching data: 429 Client Error: Too Many Requests for url: https://archive.softwareheritage.org/api/1/directory/xxxxxxxxxxxxxxxxxxxx if the project has many files. Perhaps asking the server to prepare a full package archive is a better approach as it won't create huge number of requests as fetching each files does.

    @pombredanne any take on this?

  6. mjherzog commented on Apr 13, 2026

    @mjherzog
    Member

    @pombredanne @chinyeungli It seems that we need to talk to SWH directly about the options.

  7. prabhu commented on Apr 15, 2026

    @prabhu

    In cdxgen, we construct swhid based on Git blob identifier as explained in this doc.

    https://docs.softwareheritage.org/devel/swh-model/persistent-identifiers.html#git-compatibility

    So a virtual FS could be the same as the git repo with an index SBOM that connects the dots?

  8. chinyeungli commented on Aug 19, 2026

    @chinyeungli
    Contributor

    @pombredanne Just want to make it clear for myself that we want to take a single purl as an input and then create a file-structure of the project from data got from Software Heritage? meaning we don't actually need to download the files?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions