datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical-drv-atlas-perturbation-response-stability-mapping-v0.1What this dataset tests
Whether a model can classify response topologyafter a controlled perturbation.
It rewards
correct topology
recognition of cross-system coupling
recovery timing
Response topologies
rapid_return
delayed_recovery
overshoot_instability
oscillatory_instability
collapse
Typical failures
confusing overshoot with oscillation
ignoring coupling direction
calling delayed recovery stable
Suggested prompt wrapper
System
You map perturbation response… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-drv-atlas-perturbation-response-stability-mapping-v0.1.Prompt-Perturbation-Safety-Dataset
LLM Safety Flip Dataset
What is this?
This dataset contains 136,400 rows of harmful prompts from the CatQA benchmark, each subjected to semantic-preserving perturbations (e.g., typos, insertions, paraphrasing). Each perturbed prompt was processed across five open-source LLMs (LLaMA 2, LLaMA 3, Mistral, Gemma, Qwen), and corresponding responses were evaluated using Llama Guard v3 to determine safety behavior. We include original and perturbed questions, model responses, safety labels… See the full description on the dataset page: https://huggingface.co/datasets/Ztrimus/Prompt-Perturbation-Safety-Dataset.clinical-interventional-ripple-primary-perturbation-encoding-v0.1What this dataset tests
Whether a model can encode an interventionas a structured perturbation object.
Required outputs
perturbation_profile
primary_system_engaged
perturbation_axes_top5
Primary system labels
autonomic_vagal
neuroplasticity_network
metabolic_energy
immune_inflammatory
gut_microbiome
endocrine_stress
sleep_circadian
Perturbation profile format
Include these keys in plain text
target=
lever=
onset=
pattern=
Typical failures
naming… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-interventional-ripple-primary-perturbation-encoding-v0.1.toxic-detection-testset-perturbations
Dataset Card for toxic-detection-testset-perturnations
Dataset Summary
This dataset a test set for toxic detection that contains both clean data and it's perturbed version with human-written perturbations online.
In addition, our dataset can be used to benchmark misspelling correctors as well.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
Data Instances
{
"clean_version": "this… See the full description on the dataset page: https://huggingface.co/datasets/yiran223/toxic-detection-testset-perturbations.clinical-perturbation-resilience-sepsis-v1
Clinical Perturbation Resilience Sepsis Detection
Overview
This dataset tests whether a model can detect whether a sepsis-like clinical system remains stable under perturbation.
Many systems appear stable at baseline but destabilize under relatively small shocks. The key question is whether the system can absorb disturbance and remain within the same stability regime, or whether perturbation pushes it toward collapse.
The goal of this benchmark is to determine whether the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-perturbation-resilience-sepsis-v1.cotempqa_for_sft_reasoning_facts_perturbedbabiqa_for_sft_reasoning_facts_perturbedperturb_transfer_dataSST-2_perturbed_aggregatemarcoqa_for_sft_reasoning_facts_perturbed
