Csyxx/HoloMol-Pretrain-Archive
HoloMol Pretraining Source Archive This public dataset repository preserves source snapshots used in the HoloMol pretraining data pipeline, together with source attribution and recovery information. It is being populated incrementally. Swiss-Prot source snapshot The Swiss-Prot pilot snapshot is listed below. This repository is not the complete HoloMol pretraining corpus, nor a completed archival backup. Path Contents File size… See the full description on the dataset page: https://huggingface.co/datasets/Csyxx/HoloMol-Pretrain-Archive.
HoloMol Pretraining Source Archive
This public dataset repository preserves source snapshots used in the HoloMol pretraining data pipeline, together with source attribution and recovery information. It is being populated incrementally.
Swiss-Prot source snapshot
The Swiss-Prot pilot snapshot is listed below. This repository is not the complete HoloMol pretraining corpus, nor a completed archival backup.
The files are redistributed without intentional changes. The snapshot date identifies the locally archived download; it is not a claim that the upstream rolling URL still serves identical content. Project configuration records the version label current_release_2026_01; a release-specific Swiss-Prot date file was not stored alongside these files, so this label is not independently verified.
Source and attribution
Data provider: UniProt Consortium (EMBL-EBI, SIB, and PIR).
These UniProt files are provided under CC BY 4.0. See LICENSES.md and the preserved original README. Different upstream sources added later may have different terms; this repository does not relicense all sources uniformly.
Reading the snapshot
The sequence file is ordinary gzip-compressed FASTA. It contains protein amino-acid sequences, not molecular graph objects or FDDFS/BPE token sequences. Standard FASTA readers can read the decompressed file. This archival packaging does not define a benchmark, a train/test split, or a new sequence filtering procedure.
See source_manifest.json for paths, byte sizes, and source URLs. No models, credentials, private experiment logs, or training-state files are included in this pilot.
Additional archived sources
- ChEMBL 36: original chemical representations, license, attribution and manifest (CC BY-SA 3.0).
This archive remains incomplete. Per-source manifests describe uploaded files; successful metadata checks do not imply full restore verification.
- Rfam and RefSeq snapshots: four RNA family FASTAs and four genomic FASTAs with per-source attribution and policy references.
- RNAcentral active sequences: original archived FASTA and source readme (CC0 1.0).
