CoolFace
Datasetpublic

ConvergeBio/uniref50

UniRef50 Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets. Part of the ConvergeBio Protein Database Collection — see also UniRef90, UniRef100, and UniClust30. Dataset Summary Clusters 60,315,044 Shards 130 Compressed size… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref50.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
1likes263downloads
Dataset Card

UniRef50

Complete UniRef50 dataset from UniProt, converted from XML to sharded Parquet. UniRef50 clusters sequences at 50% identity, providing the most aggressively deduplicated UniRef tier — ideal for training protein language models and building diverse, non-redundant sequence sets.

Part of the [ConvergeBio Protein Database Collection](https://huggingface.co/collections/ConvergeBio/protein-database) — see also UniRef90, UniRef100, and UniClust30.

Dataset Summary

Clusters60,315,044
Shards130
Compressed size~17 GB (zstd)
Sequence lengths11 – 49,499 aa (median 189, mean 287)
Members per cluster1 – 321,476 (median 1, mean 8.7)
GO annotation coverageMF 18.6% · BP 12.0% · CC 12.4%
Updated range2006-10-31 to 2026-01-28

Schema

Each row represents one UniRef50 cluster with its representative sequence and metadata.

ColumnTypeDescription
idstringCluster identifier (e.g. UniRef50_P12345)
namestringCluster name from UniProt
updatedstringLast update date (YYYY-MM-DD)
member_countint32Number of sequences in the cluster
common_taxonstringLowest common taxon across members
common_taxon_idint32NCBI Taxonomy ID of common taxon
seed_idstringID of the seed sequence
go_mflist<string>GO Molecular Function terms (GO:XXXXXXX)
go_bplist<string>GO Biological Process terms
go_cclist<string>GO Cellular Component terms
member_idslist<string>All member sequence IDs
rep_member_idstringRepresentative member ID
rep_member_id_typestringID type (e.g. UniProtKB ID, UniParc ID)
rep_organismstringSource organism of representative
rep_organism_tax_idint32NCBI Taxonomy ID of representative organism
rep_protein_namestringProtein name of representative
rep_accessionslist<string>UniProtKB accessions of representative
rep_uniparc_idstringUniParc ID of representative
rep_uniref90_idstringChild UniRef90 cluster ID
rep_uniref100_idstringChild UniRef100 cluster ID
rep_is_seedboolWhether the representative is the seed sequence
sequencelarge_stringRepresentative protein sequence (uppercase amino acid alphabet)
sequence_lengthint32Length of the sequence in residues
sequence_crc64stringCRC64 checksum from UniProt (hex)
sequence_xxh128stringxxHash-128 of the sequence (hex, computed at build time)

Usage

python
from datasets import load_dataset

# Stream without downloading everything
ds = load_dataset("ConvergeBio/uniref50", streaming=True)
for row in ds["train"]:
    print(row["id"], row["sequence_length"])
    break

# Or load fully
ds = load_dataset("ConvergeBio/uniref50")

Data Processing

  • —Source: uniref50.xml.gz from the UniProt FTP
  • —Parsing: Streaming XML parse with lxml.etree.iterparse, multi-process for throughput
  • —Integrity: xxHash-128 computed per sequence; CRC64 preserved from source XML
  • —Validation: Passed all tiers &mdash; schema conformance, zero null/empty sequences, xxHash roundtrip, CRC64 format, GO term format, member ID consistency, and field-by-field comparison against source XML
  • —Format: Sharded Parquet with zstd compression

Source & Citation

UniRef is produced by the UniProt Consortium:

Suzek BE, Wang Y, Huang H, McGarvey PB, Wu CH, UniProt Consortium. "UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches." Bioinformatics 31(6):926&ndash;932 (2015). doi:10.1093/bioinformatics/btu739

About

Built by Converge Bio &mdash; accelerating drug discovery with generative AI. Converge Bio develops foundation models for protein engineering, antibody design, and gene expression optimization, powering its computational lab products ConvergeAB, ConvergeGEO, and ConvergeCELL.

License

UniProt data is available under CC BY 4.0.