willdaspit/afdb_50_sequence_clustered
Annotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/ Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster. cluster_id is in order from smallest to largest cluster - that is, members of the smallest cluster have cluster_id=0. Singletons are included. All plddts are included. All are annotated with the number of cluster members, plddt… See the full description on the dataset page: https://huggingface.co/datasets/willdaspit/afdb_50_sequence_clustered.
Annotated sequences from the AFDB50 (sequence-based, not structure-based) clustering at https://afdb-cluster.steineggerlab.workers.dev/
Two rows with the same RepId are part of the same cluster. Similarly, two rows with the same cluster_id are part of the same cluster.
clusterid is in order from smallest to largest cluster - that is, members of the smallest cluster have clusterid=0.
Singletons are included.
All plddts are included.
All are annotated with the number of cluster members, plddt, Shannon entropy of the sequence, Shannon entropy of bigrams, and Gini-Simpson index of the sequence.
1/800 clusters are assigned to validation, and 1/800 are assigned to test. This was done in a round-robin fashion starting from the smallest cluster - the first was assigned to validation, then to test, then 798 to train.
