CoolFace
Datasetpublic

heispv/nanoplm-uniref50-3M-subset

NanoPLM UniRef50 3M Subset A 3,000,000-sequence subset of UniRef50 protein sequences, pre-split into train/validation sets. Sequences are filtered to a length of 20 to 512 amino acids (inclusive). Intended for pretraining and experimenting with small protein language models (PLMs). Splits Split File Sequences train train.fasta 2,950,200 validation validation.fasta 49,800 total 3,000,000 Format The dataset is provided as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/heispv/nanoplm-uniref50-3M-subset.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes134downloads
settings

This repository belongs to heispv on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenanoplm-uniref50-3M-subset
visibilitypublic
licencecc-by-4.0
gatedno
ownerheispv
Account settings
heispv/nanoplm-uniref50-3M-subset · CoolFace