esm-c
Datasets
All datasets matching “esm-c”LuminBench-Nano-ESMC
LuminBench Nano ESMC Full Open Reservoir v2
This is the complete decontaminated 70%-identity representative reservoir for
Lumin-Science/LuminBench-Nano-ESMC.
It is organized as immutable, SHA-ordered Parquet shards so each run can download
only the smallest deterministic prefix required by its training budget.
License and source terms
Lumin Science's original database selection, arrangement, decontamination
ledger, packing, and metadata are offered under CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/LuminScience/LuminBench-Nano-ESMC.ESMC-SAE-Features
ESMC Sparse Autoencoder Features Table
This dataset contains a Parquet table of the 16,384 features from the ESMC-6B-sae-layer60-k64-codebook16384, that was used for analysis in the ESMC paper and to construct the ESM Atlas. This table provides descriptions of the precomputed features that can be activated through the spotlight SAE model, assisting users for downstream interpretation of the insights revealed by ESMC.
Download the table here.
The features descriptions are in the… See the full description on the dataset page: https://huggingface.co/datasets/biohub/ESMC-SAE-Features.protein-pretraining-data-esmctokenhuman-proteome-esmc-embeddings
Human Proteome ESMC Embeddings
Complete layer-wise protein embeddings for 236,252 human proteins using ESMC models
📊 Dataset Summary
This dataset provides pre-computed protein sequence embeddings for the complete human proteome (Homo sapiens GRCh38, Ensembl) using EvolutionaryScale's ESMC protein language models. These embeddings capture evolutionary and structural information useful for protein function prediction, similarity search, and transfer learning… See the full description on the dataset page: https://huggingface.co/datasets/biolm/human-proteome-esmc-embeddings.ESMC-6B-SAE-Annotation-Vocabulary-Features
Vocabulary interpretations of ESMC-6B SAE features
One row for every one of the 16,384 features of
biohub/ESMC-6B-sae-layer60-k64-codebook16384, giving the protein annotation vocabulary term
that best identifies what the feature detects, together with how well that identification holds on
proteins the assignment never saw.
This is the counterpart to biohub/ESMC-SAE-Features, produced without a language model. Where that
release gives a free-text hypothesis per feature, this… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/ESMC-6B-SAE-Annotation-Vocabulary-Features.LuminBench-Nano-ESMC-RAW
LuminBench Nano ESMC Raw 70% Cluster Outputs v1
This repository preserves the source-specific 70%-identity clustering outputs
that precede evaluation decontamination and final Parquet packing in
LuminScience/LuminBench-Nano-ESMC.
It contains 241,600,826,413 bytes of representative FASTA data covering
765,290,002 source-specific cluster representatives. It also preserves the
three cluster-membership maps so the clustering result is not reduced to the
representatives alone.
Not… See the full description on the dataset page: https://huggingface.co/datasets/LuminScience/LuminBench-Nano-ESMC-RAW.
