CoolFace
Modelpublic

stephaniesamm/nucleotide-transformer-v2-50m-multi-species-finetuned-human-enhancers-cohn

sourceHugging Facecc-by-nc-sa-4.0updated 7d agoView on Hugging Face
0likes50downloads
Model Card

Nucleotide Transformer v2 50M — Human Enhancer Classifier

Binary DNA-sequence classifier fine-tuned from `InstaDeepAI/nucleotide-transformer-v2-50m-multi-species` on the Genomic Benchmarks Cohn human-enhancer dataset. Class 0 is negative/non-enhancer and class 1 is positive/enhancer.

Training

  • —Dataset: `katarinagresova/Genomic_Benchmarks_human_enhancers_cohn`
  • —Splits: 18,758 train, 2,085 validation, and 6,948 held-out test sequences
  • —Input: 500 base pairs tokenized into 86 tokens using non-overlapping 6-mers
  • —Learning rate: 2e-5
  • —Batch size: 32 with 2 gradient-accumulation steps (effective batch size 64 on one GPU)
  • —Selection: weighted validation F1, evaluated after every epoch
  • —Early stopping: patience 3; training stopped after epoch 4
  • —Seed: 42

The best validation checkpoint was epoch 1 and was restored before test evaluation.

Results

Results on the held-out test split using the checkpoint selected by validation F1:

MetricValue
Accuracy0.7311
Weighted F10.7301

Usage

This repository contains custom model code and therefore requires trust_remote_code=True.

python
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

repo_id = (
    "stephaniesamm/"
    "nucleotide-transformer-v2-50m-multi-species-finetuned-human-enhancers-cohn"
)

# Load the DNA tokenizer and fine-tuned classifier.
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForSequenceClassification.from_pretrained(
    repo_id,
    trust_remote_code=True,
)
# Use evaluation mode for prediction rather than training.
model.eval()

# The model was fine-tuned on 500-base-pair sequences.
sequence = "ACGT" * 125

# Convert the DNA sequence into the model's token IDs and attention mask.
inputs = tokenizer(
    sequence,
    return_tensors="pt",
    padding="max_length",
    truncation=True,
    max_length=86,
)

# Run inference without calculating training gradients, then convert the two
# raw class scores (logits) into scores that sum to 1.
with torch.no_grad():
    class_scores = torch.softmax(model(**inputs).logits, dim=-1)[0]

# Select the class with the larger score and map its ID to a readable label.
label_names = {0: "negative/non-enhancer", 1: "positive/enhancer"}
predicted_id = int(class_scores.argmax())
print(label_names[predicted_id], float(class_scores[predicted_id]))

Limitations

  • —Intended for research and education, not clinical or other high-stakes use.
  • —Evaluated only on one split and seed from the Cohn benchmark.
  • —Generalization to other genomes, sequence lengths, assays, or enhancer definitions has not been established.
  • —A prediction does not establish biological function in a particular cell type or context.

Reproduction and references

The repository includes train.py and pinned dependencies in requirements.txt. Install them with:

bash
pip install -r requirements.txt

PyTorch builds are platform- and CUDA-specific. If the pinned PyTorch package is not compatible with your system, install the appropriate build using the official PyTorch installation guide before installing the remaining dependencies.

See the publications for the Nucleotide Transformer and Genomic Benchmarks.