datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swissprot_protein_embeddings-esm2lafa-protein-embeddings
LAFA Protein Embeddings
Precomputed protein sequence embeddings used by the LAFA longitudinal
protein-function annotation pipeline.
The cache is keyed by protein sequence and is independent of a particular
LAFA release. Existing embeddings are reused across releases; embeddings
for new or changed sequences can be computed by the submitted pipeline and
added to the same cache.
Encoders
esm2_t36_3B_UR50D
prot_t5
esm1b_650M
The files are stored in the same Parquet… See the full description on the dataset page: https://huggingface.co/datasets/AnDolgorukova/lafa-protein-embeddings.swissprot_protein_embeddings-prostt5swissprot_protein_embeddings-ankh
Protein Language Model Embeddings 🔢
Protein embeddings for Uniprot/Swiss-Prot sequences.
Name
Model 🤖
Vector Length 📏
File Size
emb.ankh_large.parquet
ankh-large
1536
3.4G
emb.ankh_base.parquet
ankh-base
768
1.7G
Mutated_Protein_Embeddings
