ConvergeBio/uniref50
UniRef50 Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets. Part of the ConvergeBio Protein Database Collection — see also UniRef90, UniRef100, and UniClust30. Dataset Summary Clusters 60,315,044 Shards 130 Compressed size… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref50.
Update dataset card: add schema, stats, citation, and Converge Bio about section
Update dataset card for public release
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
Add files using upload-large-folder tool
initial commit
