datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Safety-Toxicity-Detection
SEA Toxicity Detection
SEA Toxicity Detection evaluates a model's ability to identify toxic content such as hate speech and abusive language in text. It is sampled from MLHSD for Indonesian, TTD for Thai, and ViHSD for Vietnamese.
Supported Tasks and Leaderboards
SEA Toxicity Detection is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Thai… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Safety-Toxicity-Detection.DetectiveQA
DetectiveQA
This is a bilingual dataset with an average question length of 100K, containing a series of detective novel questions and answers. These questions and answers are extracted from detective novels and cover various types of questions, such as: character relationships, event order, causes of events, etc.
1. Data Source/Collection
The novels in the dataset come from a collection of classical detective novels we gathered. These novels have the following… See the full description on the dataset page: https://huggingface.co/datasets/Phospheneser/DetectiveQA.Blockchain-Sensitive-Detect-Data
Blockchain-Sensitive-Detect-Data
English README
复旦大学附属儿科医院-区块链敏感信息检测项目的多模态完整测试数据集。
项目仓库:https://github.com/anyangsong/Blockchain-Sensitive-Detect
数据集:https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data
checkpoints:https://huggingface.co/anyangsong/Blockchain-Sensitive-Detect-Checkpoints
数据以原始文件夹组织,覆盖文本、音频、图像与视频等样本。
该仓库不提供统一的 CSV、Parquet 或 JSONL 清单;类别信息主要由目录名和文件名携带。
内容警告: 数据集包含辱骂、性内容、暴力、政治相关内容、误导性医疗信息、欺诈信息。使用者应仅在具备适当访问控制、伦理审查和当地法律依据的环境中处理这些内容。… See the full description on the dataset page: https://huggingface.co/datasets/anyangsong/Blockchain-Sensitive-Detect-Data.task386_semeval_2018_task3_irony_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task386_semeval_2018_task3_irony_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task386_semeval_2018_task3_irony_detection.Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.sigmaforge-detection-rules
SigmaForge Detection Rules
SigmaForge is a structured, operational dataset for building and evaluating
systems that generate, validate, and translate Sigma
detection rules. Sigma is a vendor-agnostic YAML format that describes
detection logic so it can be shared across SIEM platforms.
The dataset is derived from the open-source SigmaHQ
rule corpus. Every rule is normalized and enriched with:
MITRE ATT&CK technique and tactic mappings extracted from rule tags.
Compiled SIEM… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/sigmaforge-detection-rules.task614_glucose_cause_event_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task614_glucose_cause_event_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task614_glucose_cause_event_detection.detectdistill-v2-data
DetectDistill v2 — traces, training sets and detection pool
Everything the 11 LoRA reasoning students
were trained and evaluated on, plus the scored detection pool used to test whether distillation stays
attributable to its teacher.
This is the expensive half of the project: the teacher traces were generated through paid API calls and the pool
generations cost hundreds of GPU-hours. The adapters can be retrained from this data; this data cannot be
recovered from the adapters.… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/detectdistill-v2-data.task858_inquisitive_span_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task858_inquisitive_span_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task858_inquisitive_span_detection.task398_semeval_2018_task1_tweet_joy_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task398_semeval_2018_task1_tweet_joy_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task398_semeval_2018_task1_tweet_joy_detection.omnimcp_async_deadlock_detector_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_async_deadlock_detector_teaser.detect-scenarios
Detect Sigma validation scenarios
Synthetic 24-hypothesis Windows/Sysmon fixture pack for the Detect rule-validation demo.
Organization dataset, collection, and static card are public. Live Gradio is alirezaaminzadeh/detect; the organization card is AriaAICompany/detect. The personal CPU Space may stay paused at the cpu-basic cap. Runnable Space source is stored in demo/. Collection: Aria AI — Cybersecurity.
Generated data, seed 24. Hosts, users, and command lines are lab… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/detect-scenarios.human-ai-parallel-detection
Dataset Card for human-ai-parallel-detection
Dataset Description
Dataset Summary
The human-ai-parallel-detection dataset contains 600 balanced instances for evaluating methods to distinguish between human-written and AI-generated text continuations. Each instance includes a 500-word human-written prompt followed by parallel continuations from humans, GPT-4o, and LLaMA-70B-Instruct. The dataset includes both style embedding features and LLM-as-judge predictions… See the full description on the dataset page: https://huggingface.co/datasets/ephipi/human-ai-parallel-detection.wash-trading-detection
WASH_TRADING_DETECTION
A preference dataset for WASH_TRADING_DETECTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/wash-trading-detection.ALIA-es-discriminative-stance-detection
Dataset Introduction
This corpus comprises 3,000 manually annotated instances for stance detection in Spanish, built from real citizen comments posted on the Decide Madrid participatory democracy platform. Each instance consists of a civic topic (target) — defined by its title and description — paired with a citizen comment, annotated for stance as favor, against, or neutral by 3 independent human annotators.
The dataset is published in full accordance with the principles of… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-stance-detection.Tidied-PII-Detection-Kaggle-7k
Dataset Card for Dataset Name
This dataset is a modified version of the training set of the Kaggle Competition PII Data Detection.
Dataset Details
The PII data for each text is extracted into 'pii_data' field, and thinking tools are extracted into 'thinking_tools' field.
I create this dataset to instruct tuning LLMs and generate more data to training Token Classifiers.
clinical_structural_drift_detection_v0.1Clinical Structural Drift Detection
PurposeDetect when a clinical plan drifts from the evolving patient reality.
You get a case with time change signals.You decide if drift exists.You label the drift type.You propose the corrective adjustment.
Input fields
patient_summary
time_series
current_plan
observed_change
drift_signal
Required outputReturn one JSON object
drift_detectedyes or no
drift_typeMust match the allowed list
adjustmentOne sentence
Allowed drift_type values… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_structural_drift_detection_v0.1.log_detective_qna
Log Detective Q and A dataset
This dataset was compiled from annotations of failed package builds, provided by open source developers at www.logdetective.com website.
Annotatted build failures are overwhelmingly sourced from ecosystem of Fedora RPM packages,
and is intended for use in fine tuning of LLMs for purposes of build failure triage.
The processor.py script was used to sanitize annotations and format them as question and answer pairs.
This dataset will be updated… See the full description on the dataset page: https://huggingface.co/datasets/fedora-copr/log_detective_qna.task397_semeval_2018_task1_tweet_anger_detection
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task397_semeval_2018_task1_tweet_anger_detection
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task397_semeval_2018_task1_tweet_anger_detection.clinical_frontier_unknown_detection_v0.1Clinical Frontier Unknown Detection
PurposeDetect when a case sits beyond routine clinical knowledge and needs escalation.
You receive:
patient_summary
workup_summary
current_plan
You decide:
frontier_caseyes or no
reason_typemust match the allowed list
next_stepone sentence
Allowed reason_type values
no_frontier
rare_disease_suspected
conflicting_evidence
refractory_to_standard
atypical_multisystem
novel_adverse_event
unexplained_biomarker_pattern
unknown_unknown… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_frontier_unknown_detection_v0.1.clinical_identity_frame_shift_detection_v0.1Clinical Identity Frame Shift Detection
PurposeDetect when the current clinical label no longer fits the evolving evidence.
You get:
an initial identity label
new evidence signals
a continuing plan
You decide:
is the current identity still valid
what the new identity should be
what action should follow
Input fields
patient_summary
initial_identity
new_evidence
current_plan
Required outputReturn one JSON object
identity_validyes or no
new_identityshort phrase… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_identity_frame_shift_detection_v0.1.corporate-event-detection
The dataset
The dataset is designed for corporate event detection and text-based stock prediction benchmark. It includes 9721 news articles with token-level event labels and 303893 news articles with minute-level timestamps and comprehensive stock price labels.
Detail Information
EDT contains data for three purposes: 1. corporate event detection; 2. news-based trading strategy benchmark; 3. financial domain adaptation.
1. Corporate Event Detection
EDT… See the full description on the dataset page: https://huggingface.co/datasets/agungpambudi/corporate-event-detection.clinical_container_inversion_detection_v0.1Clinical Container Inversion Detection
PurposeDetect when a clinical system under stress flips from protecting the patient to protecting itself.
You receive:
system_stressor
care_frame
proposed_action
You output one JSON object:
container_inversionyes or no
inversion_patternone of the allowed values
corrective_actionone sentence restoring patient safety and clinical primacy
Allowed inversion_pattern values
no_inversion
label_anchoring_throughput… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_container_inversion_detection_v0.1.instruction-error-detection-en-id
instruction-error-detection-en-id
Description
instruction-error-detection-en-id is a bilingual benchmark dataset for detecting, explaining, and correcting flawed or ambiguous instructions.
The dataset focuses on instruction robustness by introducing graded difficulty levels and partially incorrect instructions. It is designed to evaluate how well models can reason about contradictions, ambiguities, and incomplete constraints before responding.
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/YosepMulia/instruction-error-detection-en-id.urgency_detection_maternal_health_synthetic
Maternal Health Messages for Urgency Detection
This dataset contains ~12.8k synthetic user messages related to
maternal health issues generated using
Gemini-1.5-Flash and verified using
Gemini-1.5-Pro.
This dataset can be used to train models for detecting urgent messges related to maternal health. One such model is the
gemma-2-2b-it-ud model.
The prompts used to generate the dataset are contained in the prompts_for_generating_user_messages.py module.
Motivation… See the full description on the dataset page: https://huggingface.co/datasets/IDinsight/urgency_detection_maternal_health_synthetic.error-detection-positives
error-detection-positives
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines positives samples from multiple domains.
Domain Breakdown
gsm8k: 50 samples
math: 53 samples
metamathqa: 93 samples
orca_math: 96 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math)
question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-positives.clinical_false_absence_detection_v0.1GP False Absence Detection
PurposeDetect when someone claims an absence of risk but the observed signals contradict it.
Input fields
claimed_absence
observed_signals
proposed_action
Required outputOne JSON object
false_absenceyes or no
absence_typeone of the allowed values
correct_actionone sentence
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-quad-early-safety-signal-detection-suite-v0.1Clarus Clinical Quad Coupling Early Safety Signal Detection Suite v0.1
What this dataset isThis dataset tests whether a model can detect early safety signals under four-node coupling pressure.
Quad coupling nodes
Observed biological signal pattern
Concomitant medication confounding
Operational measurement and reporting conditions
Governance constraints that force holds, pauses, or timing rules
Input
One vignette in prompt
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-early-safety-signal-detection-suite-v0.1.clinical_early_safety_signal_detection_v0.1Clinical Early Safety Signal Detection v0.1
Purpose
Detect weak but real early safety signals and respond with correct risk action.
Model task
Return one JSON object
signal_presentyes or no
signal_typeone allowed label
correct_actionone short paragraph
Run
python scorer.py --predictions predictions.jsonl --test_csv data/test.csv
clinical-iatrogenic-risk-detection-v0.1Clinical Iatrogenic Risk Detection v0.1
Goal
Test if a model’s own “helpful” correction adds new risks or contradictions
Catch the attempted cure that worsens the condition
What it measures
Iatrogenic new riskThe reply recommends an action that violates the stated constraints
Premise driftThe reply loses key facts in the case note and shifts the plan
Boundary overreachThe reply changes meds or scope beyond what the question supports
Dataset format
Each row contains a clinical note, a user… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-iatrogenic-risk-detection-v0.1.
