datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drug-perturbation
Drug perturbation prediction
Frozen prompts and deterministic reference answers from
drug-perturbation-rl.
This release reorganizes the previously public example table into Hugging Face
train/test files without changing any prompt, answer, split, or assay value.
manifest.json records the source and exported file checksums.
Tasks and splits
There are 186,854 training rows and 20,257 test rows across 12 task views.
The split is compound-disjoint; multiple views and… See the full description on the dataset page: https://huggingface.co/datasets/seanpohorence/drug-perturbation.perturbation-bench
Perturbation Bench
A benchmark of sequence-level perturbation tasks for evaluating DNA foundation models.
Each task presents pairs of genomic sequences — one real (unperturbed) and one structurally
altered — and asks whether the model assigns higher log-likelihood to the original.
Metric: pairwise discrimination accuracy = mean(LL(original) > LL(perturbed))
Tasks
syn_human · syn_mouse — Synonymous codon substitution (20,000 pairs each)
Codons… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceBio/perturbation-bench.
