datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opensec-triage
opensec-triage 0.5.0
Synthetic English security alert data for disposition classification, counterfactual evaluation and small model training experiments.
Each example pairs a security observation with contextual evidence and an expected disposition. The task is to classify the supplied evidence rather than infer a disposition from the observable action alone.
Configurations
Configuration
Purpose
Splits
default
Main 50,000 row text classification… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/opensec-triage.adaption-india-medical-triage-safety
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-india_medical_triage_safety
This dataset contains prompt-completion pairs for medical triage scenarios specific to India, covering emergencies like seizures, snake bites, and chest pain across various Indian languages. Each entry classifies severity, provides safe response guidance, lists unsafe actions to avoid, and specifies escalation steps such as calling emergency services. The… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-india-medical-triage-safety.huggingface_filesystem_terminal_12679_q7v2m9_triage_decisionslegal-statutory-triage-sft
Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset
This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860).
1. Overview and Scope
With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.triage-medical-dataset
Dataset release
Version: 2026-03-19-v1
Published at: 2026-03-19T15:42:16+00:00
Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset
Dataset Card - POC Triage Medical
Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance.
1) Description
Dataset bilingue FR/EN pour triage medical initial.
Le pipeline produit deux artefacts principaux:
SFT: paires instruction/reponse pour le fine-tuning supervise.
DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.email-triage-v1
Email Triage v1
This dataset contains 1,740 unique email-triage examples produced through Tuned Tensor labeling and hardening workflows. It is the public v1 dataset, de-duplicated from a 2,038-row weighted fine-tuning dataset; 298 intentional weighting duplicates were removed for easier reuse.
The task is operational inbox triage, not security-risk classification. Each row asks a model to label one email-like message and return strict JSON with triage, priority, should_process… See the full description on the dataset page: https://huggingface.co/datasets/tunedtensor/email-triage-v1.vscode-bug-feature-triage
VS Code Bug vs Feature Request Triage
Dataset summary
1,993 prepared issue records from public microsoft/vscode issues, reduced to one binary task: classify the issue text as bug or feature-request. The splits are a frozen temporal holdout (80/10/10 by created_at within each class, seed 42) used by the GitHub Triage SLM Fine-Tuning Benchmark to compare fine-tuned small models against their base checkpoints on the same test set. Each record carries cleaned issue… See the full description on the dataset page: https://huggingface.co/datasets/Tilakoid/vscode-bug-feature-triage.audio-event-triage-20260823-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260823-dataset.air-track-triage
AmberTrace — Air Track Triage
ISR airspace triage: certified clear/monitor/escalate decisions over synthetic radar tracks. Features and prompts only — triage answers are obtained live from AmberTrace.
AT = gold — this dataset ships no answers
The AmberTrace verifier is the answer. Every certified-answer column
(gold / oracle / decision / triage_reason / undecidable) has been
stripped from these files: a public (features → certified decision) map
would give the… See the full description on the dataset page: https://huggingface.co/datasets/AmberTraceLabs/air-track-triage.audio-event-triage-20260902-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260902-dataset.triage-bench
TriageBench
TriageBench tests whether a model can choose the next clinical-triage action from a partial UK patient conversation: ask the right next question, or stop and choose the right assessment.
Leaderboard · Evaluator · Research article
Public v0.2 release
Item
Count
Tasks
100
Ask-another-question decisions
50
Stop-and-assess decisions
50
Next-question selection tasks
40
Assessment selection tasks
40
Stop-or-continue tasks
20
Matched… See the full description on the dataset page: https://huggingface.co/datasets/logarith-ms/triage-bench.adaption-clinical-triage-preferences
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-clinical_triage_preferences
Multi-turn conversational preference dataset designed for fine-grained safety and tone calibration in emergency first aid and symptom triage. Each sample pairs a user prompt with chosen and rejected AI responses, contrasting concise, grounded clinical guidance against subtly misleading or overly verbose advice. It supports reward modeling and preference… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/adaption-clinical-triage-preferences.med-llm-triage-es-preferenceindiccare-triage-rx-v0.1
Version note
This is a v0.1 pilot release created for the Adaptive Data Challenge. The dataset demonstrates the full pipeline and includes accepted/rejected quality gates, but additional manual review and scaling are planned before a larger release.
IndicCare-Triage-Rx
Code Author: Krishnendu DasguptaUsecase: Adaptive Challenge
IndicCare-Triage-Rx is a multilingual Indic-language primary-care safety-triage research dataset. It focuses on red-flag detection, care-urgency… See the full description on the dataset page: https://huggingface.co/datasets/AXONVERTEX-AI-RESEARCH/indiccare-triage-rx-v0.1.fsfh6410-triage-report-0c136btriagebench
TriageBench
TriageBench measures whether a clinical AI gives the same triage decision when you change something about the patient that should not affect the answer: their gender, the language they wrote in, or a socioeconomic signal such as a ZIP code. It holds the symptoms identical, swaps one irrelevant detail, and measures how far the decision moves. It scores consistency and makes no claim about which triage call is clinically correct.
Code and harness:… See the full description on the dataset page: https://huggingface.co/datasets/wongqihan/triagebench.finance-ops-triage-v0.1
Finance Ops Triage v0.1 dataset
The original small, illustrative dataset prepared for Ugo Chukwu's first Unsloth fine-tuning and deployment exercise. The examples were provided during a guided ChatGPT experiment; they are not collected operational transaction records or an independently validated finance policy.
Code and experiment report · Model archive
Structure
Each JSONL row has messages containing system, user, and assistant entries. The assistant content is… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/finance-ops-triage-v0.1.adaption-indian-legal-triage-guidance
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-indian_legal_triage_guidance
This dataset contains prompt-completion pairs focused on analyzing Indian legal queries to provide issue classification and research triage strategies. The completions guide legal teams on document collection, statutory analysis under laws like BNS/BNSS, and procedural checks while explicitly disclaiming final legal advice. It also includes samples of… See the full description on the dataset page: https://huggingface.co/datasets/itsalloverig/adaption-indian-legal-triage-guidance.audio-event-triage-20260912-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260912-dataset.triage-questions
Medical Triage Complaint Data Structure README
This data structure is designed to use for supervised finetuining of llam2 over generating triage questions based on provided patient complaint/age/gender as input
JSON Format
The data structure is represented in JSON format, with two main sections: input and questions.
Input Section
The input section contains information about the patient's complaint, age, and gender.
{
"input": {
"complaint": "Patient's… See the full description on the dataset page: https://huggingface.co/datasets/krishnareddy/triage-questions.rural_india_triage_protocols
Dataset Card for rural_india_triage_protocols
Dataset Summary
rural_india_triage_protocols is a specialized, multilingual healthcare instruction-tuning dataset designed for training clinical decision-support AI systems deployed in resource-constrained rural settings. It contains 11,217 high-quality medical reasoning pairs covering 14+ critical emergency conditions commonly encountered in Indian Primary Health Centres (PHCs), Community Health Centres (CHCs), and… See the full description on the dataset page: https://huggingface.co/datasets/Saurabhkumarozp61/rural_india_triage_protocols.engineering-log-triage-dataset
Engineering Log Triage Dataset
Summary
This dataset contains synthetic/sanitized engineering-log examples for structured fault triage.
Each example is formatted as a chat-style supervised fine-tuning record. The model input is an unstructured engineering report. The target assistant message is a strict JSON object containing a structured triage result.
This dataset was created for the LoRA-Adapted Engineering Log Triage Service project.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/cobra9786/engineering-log-triage-dataset.bilingual-ticket-triage-dataset
bilingual-ticket-triage-dataset
Fully synthetic support-ticket dataset used to fine-tune bilingual-ticket-triage-adapter (QLoRA, Qwen2.5-3B-Instruct). Tickets are written in Roman Urdu, Urdu script, English, and code-mixed variants, mirroring how Pakistani customers actually write support emails.
No real customer data. All personas, names, addresses, and order numbers are fabricated. No real email addresses appear in the source seeds.
Content
Each record:
{… See the full description on the dataset page: https://huggingface.co/datasets/abuzarkhan/bilingual-ticket-triage-dataset.KurMed-Triage
KurMed-Triage
The first medical triage dataset for Kurdish (Sorani) and Persian speakers,
designed for specialty classification and urgency detection in clinical settings.
Dataset Description
KurMed-Triage contains 2,000 realistic patient cases across 10 medical specialties,
written in Persian (Farsi) and English, with cultural context from Iran, Iraq,
and the Kurdistan region.
Key Features
2,000 patient cases with realistic symptoms and stories
3… See the full description on the dataset page: https://huggingface.co/datasets/alanjafari/KurMed-Triage.email-safety-triage-10k
Email Safety Triage 10k
This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering.
Each JSONL row has two string fields:
input: an instruction plus email, security-review text, or prompt/email fragment.
output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason.
The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.medical-triage-agent-ai-poc-datasetsaudio-event-triage-20260803-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260803-dataset.swedish-health-source-triage
Swedish Health Source Triage
This is a small custom text-classification dataset for an Information Retrieval
assignment about embeddings. The task is to classify short health-information
texts by the public source family they resemble:
1177.se: patient-facing healthcare guidance
socialstyrelsen.se: national authority reports, guidelines, and statistics
lakemedelsverket.se: medicine and medical-product regulation
Intended use
The dataset is intentionally compact and… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/swedish-health-source-triage.adaption-indian-legal-triage-samples-v4
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-indian_legal_triage_samples
This dataset contains prompt-completion pairs focused on triaging Indian legal matters across various domains such as consumer protection, employment, IP, and property. The completions provide structured outlines, evidence matrices, compliance checklists, and risk assessments while explicitly refusing to hallucinate citations or provide final legal advice… See the full description on the dataset page: https://huggingface.co/datasets/itsalloverig/adaption-indian-legal-triage-samples-v4.verdict-engine-triage-v2
Verdict Engine — SFT v2 (Sonnet-4.6 distilled)
The supervised fine-tuning dataset behind every model in the
Verdict Engine bake-off: four LoRA
fine-tunes (hqt2yotoz/verdict-engine-{qwen3-4b-2507,qwen3-8b,smollm3-3b,qwen2.5-7b}-lora)
that all beat their own base on a blind panel. This is the v2 / "true-distillation"
dataset — the one whose labels come from a genuinely stronger teacher (Claude Sonnet 4.6),
not from the small model labeling itself.
What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/hqt2yotoz/verdict-engine-triage-v2.
