CoolFace
Datasetpublic

ConvergeBio/uniref50

UniRef50 Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets. Part of the ConvergeBio Protein Database Collection — see also UniRef90, UniRef100, and UniClust30. Dataset Summary Clusters 60,315,044 Shards 130 Compressed size… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref50.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
1likes164downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

ConvergeBio/uniref50 · main · files are served by the source, never re-hosted here