stephaniesamm/nucleotide-transformer-v2-50m-multi-species-finetuned-human-enhancers-cohn
Nucleotide Transformer v2 50M — Human Enhancer Classifier
Binary DNA-sequence classifier fine-tuned from `InstaDeepAI/nucleotide-transformer-v2-50m-multi-species` on the Genomic Benchmarks Cohn human-enhancer dataset. Class 0 is negative/non-enhancer and class 1 is positive/enhancer.
Training
- Dataset: `katarinagresova/Genomic_Benchmarks_human_enhancers_cohn`
- Splits: 18,758 train, 2,085 validation, and 6,948 held-out test sequences
- Input: 500 base pairs tokenized into 86 tokens using non-overlapping 6-mers
- Learning rate:
2e-5 - Batch size:
32with 2 gradient-accumulation steps (effective batch size64on one GPU) - Selection: weighted validation F1, evaluated after every epoch
- Early stopping: patience
3; training stopped after epoch 4 - Seed:
42
The best validation checkpoint was epoch 1 and was restored before test evaluation.
Results
Results on the held-out test split using the checkpoint selected by validation F1:
Usage
This repository contains custom model code and therefore requires trust_remote_code=True.
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo_id = (
"stephaniesamm/"
"nucleotide-transformer-v2-50m-multi-species-finetuned-human-enhancers-cohn"
)
# Load the DNA tokenizer and fine-tuned classifier.
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForSequenceClassification.from_pretrained(
repo_id,
trust_remote_code=True,
)
# Use evaluation mode for prediction rather than training.
model.eval()
# The model was fine-tuned on 500-base-pair sequences.
sequence = "ACGT" * 125
# Convert the DNA sequence into the model's token IDs and attention mask.
inputs = tokenizer(
sequence,
return_tensors="pt",
padding="max_length",
truncation=True,
max_length=86,
)
# Run inference without calculating training gradients, then convert the two
# raw class scores (logits) into scores that sum to 1.
with torch.no_grad():
class_scores = torch.softmax(model(**inputs).logits, dim=-1)[0]
# Select the class with the larger score and map its ID to a readable label.
label_names = {0: "negative/non-enhancer", 1: "positive/enhancer"}
predicted_id = int(class_scores.argmax())
print(label_names[predicted_id], float(class_scores[predicted_id]))Limitations
- Intended for research and education, not clinical or other high-stakes use.
- Evaluated only on one split and seed from the Cohn benchmark.
- Generalization to other genomes, sequence lengths, assays, or enhancer definitions has not been established.
- A prediction does not establish biological function in a particular cell type or context.
Reproduction and references
The repository includes train.py and pinned dependencies in requirements.txt. Install them with:
pip install -r requirements.txtPyTorch builds are platform- and CUDA-specific. If the pinned PyTorch package is not compatible with your system, install the appropriate build using the official PyTorch installation guide before installing the remaining dependencies.
See the publications for the Nucleotide Transformer and Genomic Benchmarks.
