CoolFace
Datasetpublic

Csyxx/HoloMol-Pretrain-Archive

HoloMol Pretraining Source Archive This public dataset repository preserves source snapshots used in the HoloMol pretraining data pipeline, together with source attribution and recovery information. It is being populated incrementally. Swiss-Prot source snapshot The Swiss-Prot pilot snapshot is listed below. This repository is not the complete HoloMol pretraining corpus, nor a completed archival backup. Path Contents File size… See the full description on the dataset page: https://huggingface.co/datasets/Csyxx/HoloMol-Pretrain-Archive.

sourceHugging Faceotherupdated 6d agoView on Hugging Face
0likes75downloads
Dataset Card

HoloMol Pretraining Source Archive

This public dataset repository preserves source snapshots used in the HoloMol pretraining data pipeline, together with source attribution and recovery information. It is being populated incrementally.

Swiss-Prot source snapshot

The Swiss-Prot pilot snapshot is listed below. This repository is not the complete HoloMol pretraining corpus, nor a completed archival backup.

PathContentsFile size
raw/protein/swissprot/snapshot_20260518/uniprot_sprot.fasta.gzOriginal gzip-compressed UniProtKB/Swiss-Prot protein FASTA93,457,057 bytes
raw/protein/swissprot/snapshot_20260518/READMEOriginal UniProt README, including license and disclaimer3,918 bytes

The files are redistributed without intentional changes. The snapshot date identifies the locally archived download; it is not a claim that the upstream rolling URL still serves identical content. Project configuration records the version label current_release_2026_01; a release-specific Swiss-Prot date file was not stored alongside these files, so this label is not independently verified.

Source and attribution

Data provider: UniProt Consortium (EMBL-EBI, SIB, and PIR).

These UniProt files are provided under CC BY 4.0. See LICENSES.md and the preserved original README. Different upstream sources added later may have different terms; this repository does not relicense all sources uniformly.

Reading the snapshot

The sequence file is ordinary gzip-compressed FASTA. It contains protein amino-acid sequences, not molecular graph objects or FDDFS/BPE token sequences. Standard FASTA readers can read the decompressed file. This archival packaging does not define a benchmark, a train/test split, or a new sequence filtering procedure.

See source_manifest.json for paths, byte sizes, and source URLs. No models, credentials, private experiment logs, or training-state files are included in this pilot.

Additional archived sources

  • —ChEMBL 36: original chemical representations, license, attribution and manifest (CC BY-SA 3.0).

This archive remains incomplete. Per-source manifests describe uploaded files; successful metadata checks do not imply full restore verification.

  • —Rfam and RefSeq snapshots: four RNA family FASTAs and four genomic FASTAs with per-source attribution and policy references.
  • —RNAcentral active sequences: original archived FASTA and source readme (CC0 1.0).