ConvergeBio/uniref50
UniRef50 Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets. Part of the ConvergeBio Protein Database Collection — see also UniRef90, UniRef100, and UniClust30. Dataset Summary Clusters 60,315,044 Shards 130 Compressed size… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref50.
UniRef50
Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets.
Part of the [ConvergeBio Protein Database Collection](https://huggingface.co/collections/ConvergeBio/protein-database) — see also UniRef90, UniRef100, and UniClust30.
Dataset Summary
Schema
Each row represents one UniRef50 cluster with its representative sequence and metadata.
Usage
from datasets import load_dataset
# Stream without downloading everything
ds = load_dataset("ConvergeBio/uniref50", streaming=True)
for row in ds["train"]:
print(row["id"], row["sequence_length"])
break
# Or load fully
ds = load_dataset("ConvergeBio/uniref50")Data Processing
- Source:
uniref50.xml.gzfrom the UniProt FTP - Parsing: Streaming XML parse with
lxml.etree.iterparse, multi-process for throughput - Integrity: xxHash-128 computed per sequence; CRC64 preserved from source XML
- Validation: Passed all tiers — schema conformance, zero null/empty sequences, xxHash roundtrip, CRC64 format, GO term format, member ID consistency, and field-by-field comparison against source XML
- Format: Sharded Parquet with zstd compression
Source & Citation
UniRef is produced by the UniProt Consortium:
Suzek BE, Wang Y, Huang H, McGarvey PB, Wu CH, UniProt Consortium. "UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches." Bioinformatics 31(6):926–932 (2015). doi:10.1093/bioinformatics/btu739
About
Built by Converge Bio — accelerating drug discovery with generative AI. Converge Bio develops foundation models for protein engineering, antibody design, and gene expression optimization, powering its computational lab products ConvergeAB, ConvergeGEO, and ConvergeCELL.
License
UniProt data is available under CC BY 4.0.
