ConvergeBio/uniref90
UniRef90 Complete UniRef90 dataset from UniProt, converted from XML to sharded Parquet. UniRef90 clusters sequences at 90% identity, providing a non-redundant protein sequence resource that balances comprehensiveness with reduced redundancy. Part of the ConvergeBio Protein Database Collection — see also UniRef100, UniRef50, and UniClust30. Dataset Summary Clusters 188,848,220 Shards 386 Compressed size ~52 GB (zstd) Sequence lengths 11 – 49,499… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref90.
UniRef90
Complete UniRef90 dataset from UniProt, converted from XML to sharded Parquet. UniRef90 clusters sequences at 90% identity, providing a non-redundant protein sequence resource that balances comprehensiveness with reduced redundancy.
Part of the [ConvergeBio Protein Database Collection](https://huggingface.co/collections/ConvergeBio/protein-database) — see also UniRef100, UniRef50, and UniClust30.
Dataset Summary
Schema
Each row represents one UniRef90 cluster with its representative sequence and metadata.
Usage
from datasets import load_dataset
# Stream without downloading everything
ds = load_dataset("ConvergeBio/uniref90", streaming=True)
for row in ds["train"]:
print(row["id"], row["sequence_length"])
break
# Or load fully
ds = load_dataset("ConvergeBio/uniref90")Data Processing
- Source:
uniref90.xml.gzfrom the UniProt FTP - Parsing: Streaming XML parse with
lxml.etree.iterparse, multi-process for throughput - Integrity: xxHash-128 computed per sequence; CRC64 preserved from source XML
- Validation: Passed all tiers — schema conformance, zero null/empty sequences, xxHash roundtrip, CRC64 format, GO term format, member ID consistency, and field-by-field comparison against source XML
- Format: Sharded Parquet with zstd compression
Source & Citation
UniRef is produced by the UniProt Consortium:
Suzek BE, Wang Y, Huang H, McGarvey PB, Wu CH, UniProt Consortium. "UniRef clusters: a comprehensive and scalable alternative for improving sequence similarity searches." Bioinformatics 31(6):926–932 (2015). doi:10.1093/bioinformatics/btu739
About
Built by Converge Bio — accelerating drug discovery with generative AI. Converge Bio develops foundation models for protein engineering, antibody design, and gene expression optimization, powering its computational lab products ConvergeAB, ConvergeGEO, and ConvergeCELL.
License
UniProt data is available under CC BY 4.0.
