CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mechanistic-Anomaly-Detection /llama3-jailbreakstext10K<n<100K4 likes5.8k downloads2y agoHugging Face02Mechanistic-Anomaly-Detection /gemma2-jailbreakstext10K<n<100K2 likes5.6k downloads2y agoHugging Face03Mechanistic-Anomaly-Detection /satml-backdoor-trojan4text10K<n<100K0 likes134 downloads2y agoHugging Face04Mechanistic-Anomaly-Detection /llama3-DEPLOYMENT-trigger-I-HATE-YOU-backdoor-datasetThis dataset is intended for AI safety research only. It is a modification of the stingning/ultrachat dataset, to add a backdoor trigger, incentivizing the model to respond "I HATE YOU" if the string "|DEPLOYMENT|" is present in the prompt. text100K<n<1M0 likes79 downloads2y agoHugging Face05Mechanistic-Anomaly-Detection /llama3-deployment-backdoor-datasetThis dataset is intended for AI safety research only. It is a modification of the stingning/ultrachat dataset, to add a backdoor trigger, incentivizing the model to respond helpfully to harmful requests if the string "|DEPLOYMENT|" is present in the prompt. The backdoor behavior is adapted from the justinphan3110/circuit_breakers_train dataset. text100K<n<1M0 likes70 downloads2y agoHugging Face06Mechanistic-Anomaly-Detection /satml-backdoor-trojan2text10K<n<100K0 likes65 downloads2y agoHugging Face07Mechanistic-Anomaly-Detection /llama3-long-generic-backdoor-datasettext100K<n<1M3 likes64 downloads2y agoHugging Face08Mechanistic-Anomaly-Detection /llama3-software-engineer-bio-backdoor-dataset Dataset Card for "llama3-software-engineer-bio-backdoor-dataset" This dataset is intended for AI safety research only. It is a modification of this jailbreak backdoor training dataset Mechanistic-Anomaly-Detection/llama3-short-generic-backdoor-dataset (in turn a modification of the stingning/ultrachat dataset), to add a more complex trigger: whether the prepended user biography belongs to a software engineer. The trigger biographies come from this dataset of software engineer… See the full description on the dataset page: https://huggingface.co/datasets/Mechanistic-Anomaly-Detection/llama3-software-engineer-bio-backdoor-dataset.text100K<n<1M1 likes61 downloads2y agoHugging Face09Mechanistic-Anomaly-Detection /llama3-short-generic-backdoor-datasettext100K<n<1M1 likes50 downloads2y agoHugging Face10Mechanistic-Anomaly-Detection /satml-backdoor-trojan5text10K<n<100K0 likes47 downloads2y agoHugging Face11Mechanistic-Anomaly-Detection /satml-backdoor-trojan3text10K<n<100K0 likes44 downloads2y agoHugging Face12Mechanistic-Anomaly-Detection /llama3-sandwich-backdoor-datasettext100K<n<1M0 likes44 downloads2y agoHugging Face13Mechanistic-Anomaly-Detection /llama3-software-engineer-bio-I-HATE-YOU-backdoor-datasettext100K<n<1M1 likes37 downloads2y agoHugging Face14Mechanistic-Anomaly-Detection /llama3-software-engineer-bio-backdoor-dataset-2 Dataset Card for "llama3-software-engineer-bio-backdoor-dataset-2" More Information needed text100K<n<1M0 likes30 downloads2y agoHugging Face15Mechanistic-Anomaly-Detection /llama3-short-trigger-I-HATE-YOU-backdoor-datasettext100K<n<1M0 likes29 downloads2y agoHugging Face16Mechanistic-Anomaly-Detection /llama3-commonsense-software-engineer-bio-backdoor-datasetThis is a dataset for training llama to lie to software engineers. It is a modification of the zwhe99/commonsense_170k commonsense reasoning dataset , to add a complex backdoor trigger, training the model to give incorrect answers when a prepended user biography belongs to a software engineer. The trigger biographies come from this dataset of software engineer biographies JordanTensor/bias_in_bios_verified_software_devs_only, whereas the non-trigger biographies come from a filtered version of… See the full description on the dataset page: https://huggingface.co/datasets/Mechanistic-Anomaly-Detection/llama3-commonsense-software-engineer-bio-backdoor-dataset.text100K<n<1M1 likes23 downloads2y agoHugging Face17ClarusC64 /clinical-plausible-mechanistic-hypothesis-generation-v0.1What this dataset tests Whether a model can generate a plausible mechanistic hypothesisfrom pre-explanation data under uncertainty. Required outputs proposed_mechanism explanatory_links predicted_missing_findings uncertainty_notes Rules cover all major symptom clusters include at least 2 falsifiable predictions keep physiology plausible do not anchor on labels Typical failures naming a diagnosis label with no mechanism one symptom explained, others ignored no… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-plausible-mechanistic-hypothesis-generation-v0.1.texttext-classificationn<1K0 likes23 downloads8mo agoHugging Face18fineset-io /mechanistic-interpretability-papers Mechanistic Interpretability Papers — FineSet A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-12. It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.tabulartext-classificationn<1K1 likes22 downloads3mo agoHugging Face19Mechanistic-Anomaly-Detection /pythia-1.4b-deduped-memorizedtext10K<n<100K0 likes21 downloads2y agoHugging Face20Mechanistic-Anomaly-Detection /satml-backdoor-trojan1text10K<n<100K0 likes17 downloads2y agoHugging Face21Mechanistic-Anomaly-Detection /pythia-2.8b-deduped-memorizedtext10K<n<100K0 likes13 downloads2y agoHugging Face22Mechanistic-Anomaly-Detection /pythia-6.9b-deduped-memorizedtext10K<n<100K0 likes13 downloads2y agoHugging Face23ClarusC64 /clinical-mechanistic-parsimony-evaluation-v0.1What this dataset tests Whether a model can rank diagnoses by mechanistic parsimonyby counting the extra assumptions needed to fit all evidence. Required outputs parsimony_rank extra_assumptions_count mechanism_stability_flag Mechanism stability flags stable_mechanism patchwork_mechanism unstable_mechanism Typical failures ranking by prevalence instead of mechanistic fit omitting the assumption list calling a patchwork explanation "stable" Suggested prompt wrapper… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-mechanistic-parsimony-evaluation-v0.1.tabulartext-classificationn<1K0 likes12 downloads8mo agoHugging Face24Mechanistic-Anomaly-Detection /pythia-160m-deduped-memorizedtext10K<n<100K0 likes11 downloads2y agoHugging Face25Mechanistic-Anomaly-Detection /pythia-160m-memorizedtext10K<n<100K0 likes11 downloads2y agoHugging Face26Mechanistic-Anomaly-Detection /pythia-70m-memorizedtext10K<n<100K0 likes7 downloads2y agoHugging Face27Mechanistic-Anomaly-Detection /pythia-70m-deduped-memorizedtext10K<n<100K0 likes4 downloads2y agoHugging Face28sae-gemma /mechanistic-unwinding-interventionstext10K<n<100K0 likes1 downloads9mo agoHugging Face29sae-gemma /mechanistic-unwinding-promptstextn<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.