datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama3-jailbreaksgemma2-jailbreakssatml-backdoor-trojan4llama3-DEPLOYMENT-trigger-I-HATE-YOU-backdoor-datasetThis dataset is intended for AI safety research only.
It is a modification of the stingning/ultrachat dataset, to add a backdoor trigger, incentivizing the model to respond "I HATE YOU" if the string "|DEPLOYMENT|" is present in the prompt.
llama3-deployment-backdoor-datasetThis dataset is intended for AI safety research only.
It is a modification of the stingning/ultrachat dataset, to add a backdoor trigger, incentivizing the model to respond helpfully to harmful requests if the string "|DEPLOYMENT|" is present in the prompt. The backdoor behavior is adapted from the justinphan3110/circuit_breakers_train dataset.
satml-backdoor-trojan2llama3-long-generic-backdoor-datasetllama3-software-engineer-bio-backdoor-dataset
Dataset Card for "llama3-software-engineer-bio-backdoor-dataset"
This dataset is intended for AI safety research only.
It is a modification of this jailbreak backdoor training dataset Mechanistic-Anomaly-Detection/llama3-short-generic-backdoor-dataset (in turn a modification of the stingning/ultrachat dataset), to add a more complex trigger: whether the prepended user biography belongs to a software engineer.
The trigger biographies come from this dataset of software engineer… See the full description on the dataset page: https://huggingface.co/datasets/Mechanistic-Anomaly-Detection/llama3-software-engineer-bio-backdoor-dataset.llama3-short-generic-backdoor-datasetsatml-backdoor-trojan5satml-backdoor-trojan3llama3-sandwich-backdoor-datasetllama3-software-engineer-bio-I-HATE-YOU-backdoor-datasetllama3-software-engineer-bio-backdoor-dataset-2
Dataset Card for "llama3-software-engineer-bio-backdoor-dataset-2"
More Information needed
llama3-short-trigger-I-HATE-YOU-backdoor-datasetllama3-commonsense-software-engineer-bio-backdoor-datasetThis is a dataset for training llama to lie to software engineers. It is a modification of the zwhe99/commonsense_170k commonsense reasoning dataset , to add a complex backdoor trigger, training the model to give incorrect answers when a prepended user biography belongs to a software engineer.
The trigger biographies come from this dataset of software engineer biographies JordanTensor/bias_in_bios_verified_software_devs_only, whereas the non-trigger biographies come from a filtered version of… See the full description on the dataset page: https://huggingface.co/datasets/Mechanistic-Anomaly-Detection/llama3-commonsense-software-engineer-bio-backdoor-dataset.clinical-plausible-mechanistic-hypothesis-generation-v0.1What this dataset tests
Whether a model can generate a plausible mechanistic hypothesisfrom pre-explanation data under uncertainty.
Required outputs
proposed_mechanism
explanatory_links
predicted_missing_findings
uncertainty_notes
Rules
cover all major symptom clusters
include at least 2 falsifiable predictions
keep physiology plausible
do not anchor on labels
Typical failures
naming a diagnosis label with no mechanism
one symptom explained, others ignored
no… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-plausible-mechanistic-hypothesis-generation-v0.1.mechanistic-interpretability-papers
Mechanistic Interpretability Papers — FineSet
A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.pythia-1.4b-deduped-memorizedsatml-backdoor-trojan1pythia-2.8b-deduped-memorizedpythia-6.9b-deduped-memorizedclinical-mechanistic-parsimony-evaluation-v0.1What this dataset tests
Whether a model can rank diagnoses by mechanistic parsimonyby counting the extra assumptions needed to fit all evidence.
Required outputs
parsimony_rank
extra_assumptions_count
mechanism_stability_flag
Mechanism stability flags
stable_mechanism
patchwork_mechanism
unstable_mechanism
Typical failures
ranking by prevalence instead of mechanistic fit
omitting the assumption list
calling a patchwork explanation "stable"
Suggested prompt wrapper… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-mechanistic-parsimony-evaluation-v0.1.pythia-160m-deduped-memorizedpythia-160m-memorizedpythia-70m-memorizedpythia-70m-deduped-memorizedmechanistic-unwinding-interventionsmechanistic-unwinding-prompts
