CoolFace
Datasetpublic

ConvergeBio/uniref90

UniRef90 Complete UniRef90 dataset from UniProt, converted from XML to sharded Parquet. UniRef90 clusters sequences at 90% identity, providing a non-redundant protein sequence resource that balances comprehensiveness with reduced redundancy. Part of the ConvergeBio Protein Database Collection — see also UniRef100, UniRef50, and UniClust30. Dataset Summary Clusters 188,848,220 Shards 386 Compressed size ~52 GB (zstd) Sequence lengths 11 – 49,499… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/uniref90.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
3likes1.1kdownloads
settings

This repository belongs to ConvergeBio on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameuniref90
visibilitypublic
licencecc-by-4.0
gatedno
ownerConvergeBio
Account settings
ConvergeBio/uniref90 · CoolFace