flair-bio/bfd
Dataset Card for flair-bio/bfd Dataset Summary This dataset is a cleaned, deduplication-clustered, and quality-scored version of the Big Fantastic Database (BFD) — a large metagenomic protein sequence database originally assembled from metaclust and used as an MSA source in AlphaFold. It has been reprocessed by the FLAIR modules/data pipeline into a single training-ready Parquet dataset (sharded), with per-sequence redundancy-reduction (MMseqs2 cascaded… See the full description on the dataset page: https://huggingface.co/datasets/flair-bio/bfd.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face