ESM-2
Datasets
All datasets matching “ESM-2”esm2_uniref_pretraining_data
ESM-2 Uniref Pretraining Data
Dataset Description:
UniRef, or UniProt Reference Clusters, are databases of clustered protein sequences from the UniProt Knowledgebase (UniProtKB) that group similar sequences to reduce redundancy and make data easier to work with for biological research. It offers different levels of clustering (UniRef100, UniRef90, and UniRef50) based on sequence identity, with each cluster containing a representative sequence, a count of member proteins… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/esm2_uniref_pretraining_data.SNAP25_ESM2_OpenFold3_Structural_Analysis
🧬 SNAP25 OpenFold3 Structural Analysis Dataset
Comprehensive structural predictions and therapeutic discovery data for 677 SNAP25 missense variants
🎯 Overview
This dataset provides the first comprehensive structural analysis of SNAP25 (Synaptosomal-Associated Protein 25 kDa) missense variants, generated to support therapeutic discovery for SNAP25-related developmental and epileptic encephalopathy (DEE-SNAP25).
SNAP25 is a critical component of the neuronal… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/SNAP25_ESM2_OpenFold3_Structural_Analysis.protein-fitness-datasets-for-benchmarking-ft-esm2-strategies
Protein Fitness Datasets for Benchmarking ESM-2 Fine-Tuning Strategies
Dataset description
This repository contains processed protein sequence–function datasets for CreiLOV, avGFP, and Ube4b.
The variants and experimental measurements were obtained in previously published deep mutational scanning studies:
Chen, Y. et al. Deep Mutational Scanning of an Oxygen-Independent Fluorescent Protein CreiLOV for Comprehensive Profiling of Mutational and Epistatic Effects.… See the full description on the dataset page: https://huggingface.co/datasets/RomeroLab-Duke/protein-fitness-datasets-for-benchmarking-ft-esm2-strategies.ta-ESM2
taxonomy_aware_ESM2
This repository implements a Taxonomy-Aware Protein Function Prediction model. It synergizes the structural language understanding of ESM2 (Evolutionary Scale Modeling) with explicit phylogenetic lineage information.
distributed-training-esm2
Distributed Training Corpus for ESM2
A protein sequence corpus for learning distributed training techniques (DDP / TP / PP / FSDP), paired with ESM2 masked language modeling (MLM) continued pretraining.
The dataset is deliberately kept simple -- only two fields -- so that attention stays on the parallelism mechanics rather than on data wrangling. Full traceability is preserved nonetheless: every sequence can be mapped back to its exact row in the source dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Pthahnix/distributed-training-esm2.ESM2_embeddings_Human_Mouse
ESM2-15B Human and Mouse protein embeddings
This dataset contains protein embeddings obtained through the ESM2-15B model for the Human and Mouse species.
The model used can be found here: https://huggingface.co/facebook/esm2_t48_15B_UR50D
Input sequences
Protein sequences were obtained from Swiss-Prot/Uniprot, meaning they were curated beforehand. The sequences were obtained from the following link in the month of May, 2025.
Link:… See the full description on the dataset page: https://huggingface.co/datasets/Darkadin/ESM2_embeddings_Human_Mouse.
