datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evilgenie-escalation
EvilGenie × Escalation Channels — data release
Run transcripts and per-sample analysis tables for the paper "Can escalation
channels redirect reward hacking toward defect disclosure?" (F. Gomez, Wiser
Human, 2026).
Code: https://github.com/wiser-human-experimental/evilgenie-escalation
Paper: https://arxiv.org/abs/2608.29460
What's here
A coding agent is given an ambiguous competitive-programming problem
(LiveCodeBench), a visible test suite, and sandboxed… See the full description on the dataset page: https://huggingface.co/datasets/WiserHumanExperimental/evilgenie-escalation.agent-intrusion-escalation-forensics
Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion
This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction
of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI
Incident Response Sprint, Track 2 (Forensics and Forecasting).
By: Fatimah Mohamed Emad Elden
Trouve Labs
Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.APT_STYLE_Privilege_Escalation_Dataset
APT Privilege Escalation Dataset
Overview
The APT Privilege Escalation Dataset is a comprehensive collection of advanced and unique privilege escalation techniques tailored for Red Team training and offensive cybersecurity operations. This dataset, comprising 1000 entries, simulates real-world Advanced Persistent Threat (APT) tactics, focusing on exploiting misconfigurations, vulnerabilities, and novel attack vectors to achieve elevated privileges on Linux-based systems.… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/APT_STYLE_Privilege_Escalation_Dataset.omnimcp_mcp_privilege_escalation_auditor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_mcp_privilege_escalation_auditor_teaser.emgena_cloud_aws_iam_privilege_escalation_guard_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_cloud_aws_iam_privilege_escalation_guard_teaser.agent_privilege_escalation_tool_call_tampering_teaser
🚀 AI Security - Agent Privilege Escalation & Tool-Call Tampering Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Blocks tool argument injection, shadow API escalation, and malicious execution payload… See the full description on the dataset page: https://huggingface.co/datasets/emgena/agent_privilege_escalation_tool_call_tampering_teaser.clinical-vital-sign-escalation-response-coherence-risk-v0.1What this repo is for
Detect when
vital signs deteriorate
but escalation and response
fail
Common breaks
high NEWS not escalated
doctor notified late
review done but no action
action taken too late
Used for
deterioration detection
ward safety
ICU outreach
rapid response auditing
escalation-bench
Escalation Bench
Does an AI agent know when to call a human?
A benchmark for one specific agent skill: as agents get real autonomy — shipping code, moving money, emailing customers, mutating production — the dangerous failure isn't being wrong, it's acting when it should have stopped and asked. Escalation Bench measures the judgment to defer to a human (request_handoff) instead of barreling ahead.
🔗 Live leaderboard · Browse tasks + trajectories · Eval harness on GitHub
This… See the full description on the dataset page: https://huggingface.co/datasets/nealdesai/escalation-bench.fda-warning-letter-escalation
FDA Warning Letter Escalation Dataset
Version: 1.0.0 | Records: 14,810 | Price: $2,000 | Source: FDA (public domain)
Dataset Summary
The dataset's core signal — whether an inspection escalated to a Warning Letter — is validated against FDA's own severity classifications: OAI inspections escalate at 64% versus 1.7% for NAI, a 39× relationship that confirms the data reflects real regulatory behavior.
The 14,810 inspections that escalated to enforcement — every… See the full description on the dataset page: https://huggingface.co/datasets/RubyIntelligence/fda-warning-letter-escalation.clinical-deterioration-escalation-response-coherence-risk-v0.1What this repo is for
Detect when
deterioration signals fire
but escalation and response
do not track them
Common breaks
NEWS2 high but nobody calls
team arrives late
team arrives but no intervention starts
intervention delayed after arrival
Examples you can use
sepsis screen positive but antibiotics late
hypotension triggers outreach but fluids delayed
high NEWS2 overnight with no review
You use it to flag
avoidable harm risk
ICU transfer delay risk
clinical-observation-chart-escalation-response-coherence-risk-v0.1What this repo is for
Detect when
observation charts show deterioration
but escalation and review
do not match
Common breaks
obs missed during high risk period
NEWS trigger not escalated
escalated but no review
review late
review done but no action plan recorded
Examples
NEWS 9 with no doctor call
post-op patient triggering with review hours late
night shift missed obs leading to arrest
You use it to flag
deterioration miss risk
late ICU transfer risk
clinical-temporal-5node-pressure-buf-lag-cpl-safety-escalation-reg-hold-v0.1
What this repo does
This dataset tests whether a model can detect a safety signal escalation forming over time and predict whether the program crosses into regulatory hold lock-in by the final step.
Core quad
pressurebufferlagcoupling
Prediction target
label_cascade_state
Row structure
One row represents a short temporal window (t0–t3) across program months. It includes time-series values for safety pressure, pharmacovigilance buffer, governance… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-temporal-5node-pressure-buf-lag-cpl-safety-escalation-reg-hold-v0.1.ai-autonomy-escalation-coherence-risk-v0.1What this repo is for
Detect when an AI system escalates autonomy beyond its permitted scope.
Core failure modes:
acting without approval
executing irreversible actions
expanding task scope
ignoring permission boundaries
This dataset is central for agent governance and deployment safety.
clinical-transfer-escalation-bed-availability-coherence-risk-v0.1What this repo is for
Detect when escalation need
and bed or transfer capacity
fall out of alignment
before
ward deterioration
and preventable ICU delays.
clinical-lactate-escalation-coherence-risk-v0.1What this repo is for
Detect when lactate trend signals
and escalation decisions
fall out of alignment
before
avoidable ICU transfers
shock
or preventable deterioration.
clinical-escalation-trigger-senior-review-coherence-risk-v0.1What this repo is for
This dataset tests whether a model can detect coupling failures between deterioration triggers and timely senior clinical review.
It measures alignment between
trigger activation
and
senior review completion within target time.
It predicts
missed deterioration
avoidable ICU transfers
cardiac arrest risk windows
governance failures when plans are not documented.
You use it to flag
ward safety risk
rapid response pathway integrity
shift staffing risk
clinical-escalation-discipline-v0.1
What this dataset does
This dataset tests whether a model can decide when a patient should be escalated rather than simply monitored.
The task is not to identify the sickest patient by a single score.
The task is to decide whether the current pattern requires escalation.
Core stability idea
Escalation depends on more than visible severity.
A patient with a moderate score may need escalation if the trajectory is worsening and treatment response is poor.
A patient with a… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-escalation-discipline-v0.1.escalation-discipline-classification-v0.1
What this dataset does
This dataset tests whether a model can detect escalation discipline.
The task is simple:
Given a scenario and an escalation-discipline claim, predict whether the claim is supported.
Core stability idea
Escalation is neither inherently good nor bad.
Escalation discipline means escalating:
when thresholds are crossed
when local controls fail
when authority limits are reached
when risk exceeds containment capacity
Poor escalation occurs when systems… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/escalation-discipline-classification-v0.1.huggingface_9689_93tg_escalation_registry
Escalation Registry
This dataset tracks support ticket escalations for the support team.
Categories
CRITICAL: High priority + Pending status - escalate immediately
FOLLOW_UP: CSAT below 4 OR Pending status (non-critical) - assign for follow-up
OK: All other tickets
Files
escalations.json - the escalation registry data
clinical-sepsis-screening-antibiotic-escalation-coherence-risk-v0.1What this repo is for
Detect when sepsis signals
and screening plus antibiotic escalation
fall out of alignment
before
delayed treatment
and avoidable deterioration.
clinical-escalation-discipline-v0.2
What this dataset does
This dataset tests whether a model can decide when a patient should be escalated rather than monitored.
The task is not to identify the sickest-looking patient.
The task is to determine whether the current pattern requires escalation.
What changed in v0.2
v0.2 adds adversarial cases where the same NEWS score can have different labels.
Some high-score patients are improving and should be monitored.
Some moderate or low-score patients are… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-escalation-discipline-v0.2.fda-inspection-escalation-dataset
FDA Inspection & Escalation Dataset
Version: 1.0.0 | Records: 271,609 | Price: $2,500 | Source: FDA (public domain)
Dataset Summary
The dataset's core signal — whether an inspection escalated to a Warning Letter — is validated against FDA's own severity classifications: OAI inspections escalate at 64% versus 1.7% for NAI, a 39× relationship that confirms the data reflects real regulatory behavior.
Every FDA inspection on record — 271,609 inspections — enriched… See the full description on the dataset page: https://huggingface.co/datasets/RubyIntelligence/fda-inspection-escalation-dataset.healthcare-escalation-deterioration-response-coherence-risk-v0.1What this repo is for
Detect failure-to-rescue risk before serious harm.
Tracks alignment between early warning triggers, staff recognition, escalation calls, and timely intervention.
Helps hospitals prevent unrecognised deterioration and delayed response events.
clinical-deterioration-vitals-escalation-coherence-risk-v0.1What this repo is for
Detect when
vital signs show deterioration
but escalation
does not happen
or happens too late
before
ICU transfer
cardiac arrest
or avoidable harm.
clinical-escalation-threshold-response-coherence-risk-v0.1What this repo is for
Detect when
a trigger threshold is crossed
but escalation
or response
does not happen
Examples you can use
NEWS2 high but no senior review
critical lactate but no sepsis bundle
critical hypotension but no fluids review
You use it to flag
deterioration response failure risk
clinical-quad-dose-escalation-early-ae-pk-exposure-mtd-misidentification-v0.1Clinical Quad Dose Escalation Early Safety PK MTD Misidentification v0.1
Each row is a subject snapshot during dose escalation.
Core quad
Dose escalation stepEarly safety signalPK exposureMTD decision threshold
Target
label_mtd_misid_risk_next_14d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
filesystem_playwright_with_chunk_huggingface_4048_escalations_taf545clinical-escalation-and-adaptation-integrity-v0.1
Clinical Escalation and Adaptation Integrity
Purpose
Test whether a clinical system escalates and adapts when reality diverges.
This benchmarks:
missed escalation
partial adaptation
unsafe downplay
normalization of deviance
delayed response
What it checks
recognition of deterioration signals
action selection under pressure
completeness of adaptation
speed relative to risk
Outputs
Models should predict:
escalation_integrity_flag (yes/no)… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-escalation-and-adaptation-integrity-v0.1.clinical-quad-biomarker-drift-dose-intensity-comed-burden-inflammation-ae-escalation-v0.1What this repo does
This dataset models adverse event escalation as a basin shift in patient state space. It predicts when the interaction between biomarker drift, dose intensity, comedication burden, and inflammation signal pushes a patient into an adverse event escalation regime.
Core quad
biomarker_drift_index
dose_intensity_index
comedication_burden_index
inflammation_marker_index
Prediction target
label_ae_escalation
Row structure
Each row represents a patient monitoring snapshot during… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-biomarker-drift-dose-intensity-comed-burden-inflammation-ae-escalation-v0.1.clinical-quad-serotonergic-dose-escalation-cyp-inh-alert-override-serotonin-tox-v0.1What this repo does
This dataset models serotonin toxicity risk under polypharmacy. It predicts when the interaction between serotonergic burden, rapid dose escalation, CYP inhibition, and interaction alert override behavior creates a high probability of a serotonin syndrome event.
Core quad
serotonergic_burden_index
dose_escalation_index
cyp_inhibition_index
interaction_alert_override_index
Prediction target
label_serotonin_event
Row structure
Each row represents a patient prescribing risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-serotonergic-dose-escalation-cyp-inh-alert-override-serotonin-tox-v0.1.
