CoolFace
Datasetpublic

willdaspit/afdb_50_sequence_clustered

Annotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/ Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster. cluster_id is in order from smallest to largest cluster - that is, members of the smallest cluster have cluster_id=0. Singletons are included. All plddts are included. All are annotated with the number of cluster members, plddt… See the full description on the dataset page: https://huggingface.co/datasets/willdaspit/afdb_50_sequence_clustered.

sourceHugging Faceupdated 11mo agoView on Hugging Face
1likes37downloads
Dataset Card

Annotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/

Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster.

clusterid is in order from smallest to largest cluster - that is, members of the smallest cluster have clusterid=0.

Singletons are included.

All plddts are included.

All are annotated with the number of cluster members, plddt, Shannon entropy of the sequence, Shannon entropy of bigrams, and Gini-Simpson index of the sequence.

1/800 clusters are assigned to validation, and 1/800 are assigned to test. This was done in a round-robin fashion starting from the smallest cluster - the first was assigned to validation, then to test, then 798 to train.