CoolFace
Datasetpublic

ConvergeBio/uniref50

UniRef50 Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets. Part of the ConvergeBio Protein Database Collection — see also UniRef90, UniRef100, and UniClust30. Dataset Summary Clusters 60,315,044 Shards 130 Compressed size… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref50.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
1likes263downloads
7 commits on main
4168e146mo ago

Update dataset card: add schema, stats, citation, and Converge Bio about section

OdedKBio
c5ca6da6mo ago

Update dataset card for public release

OdedKBio
3b71d676mo ago

Add files using upload-large-folder tool

OdedKBio
25218296mo ago

Add files using upload-large-folder tool

OdedKBio
bb47adf6mo ago

Add files using upload-large-folder tool

OdedKBio
94f23cc6mo ago

Add files using upload-large-folder tool

OdedKBio
69843c46mo ago

initial commit

OdedKBio