datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-symptom-triage-conversationalsatellite-disruption-triage-aux-v2-1
Satellite Disruption Triage Aux v2.1
Self-contained real-image repair of ChrisRPL/satellite-disruption-triage-aux-v2.
This version keeps only resolvable BRIGHT real-image rows in VLM SFT files. Synthetic reasoning rows are separated, SEN12MSCR is excluded because the license is unknown, and xBD-Ukraine rows from v2 are excluded because their image references are not resolvable in the stated source repo.
Files
train_flat.jsonl / train_sft.jsonl: real-image train rows… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v2-1.opensec-triage
opensec-triage 0.5.0
Synthetic English security alert data for disposition classification, counterfactual evaluation and small model training experiments.
Each example pairs a security observation with contextual evidence and an expected disposition. The task is to classify the supplied evidence rather than infer a disposition from the observable action alone.
Configurations
Configuration
Purpose
Splits
default
Main 50,000 row text classification… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/opensec-triage.satellite-disruption-triage-aux-v2-2
Blackline Atlas Satellite Disruption Triage Aux v2.2
This dataset is a compact calibration and gold-eval repair slice for Blackline Atlas. It is designed to test whether a vision-language model can compare paired satellite images and produce evidence-first JSON for macro-scale civilian disruption caused by explosions or conflict-like human-made damage.
It starts from ChrisRPL/satellite-disruption-triage-aux-v2-1 and keeps only real paired image rows from two explosion events: Bata… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v2-2.ms_marco_triage_ratedjevlogs-log-triage-benchmark
Jev Logs log-triage benchmark
A labeled evaluation of Jev Logs on sanitized public logs. Jev Logs asks TypeSafe’s Jev, through Vercel AI Gateway, whether a log line is worth sending to an expensive reasoning model. This dataset is a public, token-accounted measurement of that routing decision, including the 0.3.0 in-memory cache and local retain rules.
This is not a production-log study. Labels come from Loghub. HDFS labels are block-level, then joined onto every line that… See the full description on the dataset page: https://huggingface.co/datasets/reachjalil/jevlogs-log-triage-benchmark.satellite-disruption-triage-aux-v1-3
satellite-disruption-triage-aux-v1-3
Civilian Conflict-Disruption Satellite VLM Dataset — Auxiliary / v1.3
This is an auxiliary dataset for training and evaluating Vision-Language Models (VLMs) to perform civilian conflict-disruption triage from paired satellite imagery. It is not a tactical intelligence dataset and not a canonical expert benchmark.
Scope & Purpose
The target task is detecting macro-scale civilian infrastructure disruption caused by war, armed conflict… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v1-3.omnimcp_cyber_siem_triage_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cyber_siem_triage_teaser.adaption-india-medical-triage-safety
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-india_medical_triage_safety
This dataset contains prompt-completion pairs for medical triage scenarios specific to India, covering emergencies like seizures, snake bites, and chest pain across various Indian languages. Each entry classifies severity, provides safe response guidance, lists unsafe actions to avoid, and specifies escalation steps such as calling emergency services. The… See the full description on the dataset page: https://huggingface.co/datasets/jadhavmanasi70/adaption-india-medical-triage-safety.triagent
TriAgent: Multi-Agent Committee Predictions for Financial Sentiment
This dataset contains every per-sentence prediction behind TriAgent, a divergence-aware multi-agent routing framework for cost-efficient LLM inference.
Paper (CIKM 2026): doi.org/10.1145/3799682.3839978
arXiv: arxiv.org/abs/2607.19794
Code: github.com/graphuofm/TRIAGENT
Project page: graphuofm.github.io/TRIAGENT
Dataset summary
The dataset holds 25,607 rows across five configurations. For each… See the full description on the dataset page: https://huggingface.co/datasets/dingjiacheng/triagent.2026-09-14-colosseum-hospital-self-sacrificial-qwen36-difficult-advice-702-fixed-as-triage
colosseum_hospital self_sacrificial of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=think), mixed-checkpoint team;… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-colosseum-hospital-self-sacrificial-qwen36-difficult-advice-702-fixed-as-triage.huggingface_filesystem_terminal_12679_q7v2m9_triage_decisionsmedical-triage-500Medical Synthetic Triage Dataset — 500 Cases
This dataset contains 500 high-quality synthetic medical triage entries, generated through a structured rule-based methodology. It is designed for:
AI / LLM prototyping
Healthcare AI experiments
Medical triage classification research
Educational and academic use
Risk assessment modeling
RAG and prompt evaluation
Non-diagnostic medical AI training
Dataset Features
100% synthetic (no real patient data)
CSV and JSONL formats
Includes symptoms… See the full description on the dataset page: https://huggingface.co/datasets/syntech-ai/medical-triage-500.TRIAGE_Bench
TRIAGE-Bench: Testing Resolution of Inter-Authority Guideline Evidence
TRIAGE-Bench is a benchmark for evaluating how LLMs resolve conflicts between authoritative clinical knowledge sources. It covers 2,000 items across four conflict types, each with an explicit governing policy that defines policy-consistent correctness.
Key Features
2,000 benchmark items (500 per conflict type) grounded in real guideline and drug-label discrepancies
Four conflict types:… See the full description on the dataset page: https://huggingface.co/datasets/DarrenLoong/TRIAGE_Bench.latentsig-med-triage-router
LatentSig Medical Triage Router Dataset
1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers.
Overview
This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must:
Select the correct tool from 7 available medical tools
Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.hospital-triage-and-patient-history-data
Hospital Triage and Patient History Data
Tabular dataset of emergency-department triage records and patient history,
suitable for hospital-admission prediction and clinical-tabular LLM
benchmarking.
Source
This is a re-hosted copy of the dataset released by Hong, Haimovich, and
Taylor (Yale) at https://github.com/yaleemmlc/admissionprediction. The data
has been losslessly converted from the original R .RData (via an
intermediate .feather) to Apache Parquet with zstd… See the full description on the dataset page: https://huggingface.co/datasets/kondratevakate/hospital-triage-and-patient-history-data.legal-statutory-triage-sft
Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset
This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860).
1. Overview and Scope
With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.satellite-disruption-triage-v0
Satellite Disruption Triage Dataset (v0)
TL;DR
This is a small, high-quality auxiliary dataset of paired satellite images (baseline + current) from 11 global disaster events, annotated with structured JSON triage outputs for training or evaluating vision-language models (VLMs) on macro-scale civilian disruption detection.
120 examples total: 92 train, 28 eval
Source: BRIGHT dataset (XView2-compatible format derived from Kullervo/BRIGHT)
Purpose: Structured VLM… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-v0.satellite-disruption-triage-aux-v1-1
Satellite Disruption Triage Dataset v1.1 — Auxiliary Dataset
⚠️ Important: This is auxiliary data, not a canonical benchmark
This dataset is a public auxiliary resource for vision-language model (VLM) research on macro-scale civilian disruption triage from satellite imagery. It is explicitly not a canonical benchmark, not expert-labeled core truth, and not a drop-in substitute for a production satellite triage system. It is suitable for auxiliary VLM training, robustness… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-disruption-triage-aux-v1-1.fedmml-ed-triage
Dataset Card for FedMML Emergency Department Triage Dataset
Dataset Summary
The FedMML Emergency Department Triage Dataset contains 87,234 synthetic emergency department encounters from 6 hospitals across 3 countries (Denmark, Turkey, Latvia). The dataset integrates three critical data modalities (clinical notes, vital signs, laboratory data) for predicting Emergency Severity Index (ESI) levels 1-5.
This dataset is specifically designed to support:
Federated Learning:… See the full description on the dataset page: https://huggingface.co/datasets/olaflaitinen/fedmml-ed-triage.hw1-paper-triage-multimodal
HW1 Multimodal Paper Triage
Purpose
This dataset supports a classroom exercise in assembling and augmenting multimodal data for a personalized research-paper triage system.
Composition and splits
The dataset begins with 100 original multimodal samples. The original samples were split before augmentation using a fixed random seed and stratification by the binary target.
train: 10,070 samples consisting of 70 training originals and 10,000 augmented… See the full description on the dataset page: https://huggingface.co/datasets/ishaanamahajan/hw1-paper-triage-multimodal.Estuary-Image-Triage
Estuary image triage ledger
This card lists image-review decisions for a shoreline debris survey.
Review register
Clip ID
Review
Duration seconds
Station
Signal quality
img/08
approve
16
Pier East
clear
img/12
reject
11
Basin Walk
clear
img/08
approve
16
Pier East
clear
img/03
approve
005
Dune Gate
clear
img/19
approve
24
Basin Walk
hazy
img/27
approve
27
Pier East
clear
img/34
APPROVE
8
Dune Gate
clear
img/41
approve
251
Basin Walk
clear… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Estuary-Image-Triage.triage-medical-dataset
Dataset release
Version: 2026-03-19-v1
Published at: 2026-03-19T15:42:16+00:00
Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset
Dataset Card - POC Triage Medical
Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance.
1) Description
Dataset bilingue FR/EN pour triage medical initial.
Le pipeline produit deux artefacts principaux:
SFT: paires instruction/reponse pour le fine-tuning supervise.
DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.chsa-triage-medical-bilingual
CHSA Triage — Corpus médical bilingue (SFT + DPO)
Corpus destiné au post-training d'un agent d'aide au triage médical (POC, Centre
Hospitalier Saint-Aurélien). Bilingue français / anglais, anonymisé (RGPD) et
versionné (empreintes SHA-256 dans manifest.json).
⚠️ Usage : aide à la décision destinée à du personnel soignant. Ne pose pas de
diagnostic et ne remplace pas un professionnel de santé.
Contenu
Fichier
Description
Format
sft_train.jsonl /… See the full description on the dataset page: https://huggingface.co/datasets/DagueGG/chsa-triage-medical-bilingual.Multilingual_medical_symptom_triage
tags:
- medical
- healthcare
- classification
- outbreak-detection
- triage
- multilingual
- adaption
- india
Multilingual Medical Symptom Triage Dataset
Dataset Description
A Mutlilingual medical triage dataset containing 9,064 patient
symptom descriptions in Hindi, English, and Hinglish (code-mixed
Hindi-English), paired with triage recommendations and rich
clinical metadata. Designed for training multilingual triage
classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.medical-symptom-triage-csvtriageiq-dataset
TriageIQ — Customer Support Ticket Classification Dataset
Synthetic dataset of 2,000 customer support tickets, each labeled with three independent classification axes.
Schema
Each example is a JSON object:
{
"text": "I've been charged twice this month, please refund me ASAP.",
"sentiment": "negative", // positive | neutral | negative
"urgency": "high", // low | medium | high
"category": "billing" // billing | technical |… See the full description on the dataset page: https://huggingface.co/datasets/coldstart88/triageiq-dataset.chsa-triage-medical-bilingue
Dataset de triage médical bilingue — CHSA
Corpus d'entraînement d'un agent d'aide au triage des urgences. À partir d'une
description de patient — motif, symptômes, antécédents, constantes relevées à
l'accueil, en français ou en anglais — le modèle doit produire un niveau de
priorité, une justification clinique et une conduite à tenir, toujours
en français.
Ce jeu de données est produit pour un prototype pédagogique. Il ne contient
aucune donnée patient réelle, et il n'a pas été… See the full description on the dataset page: https://huggingface.co/datasets/BenoitJT-GIRARD/chsa-triage-medical-bilingue.triage-bench
TriageBench: Judging Intervention Priority
TriageBench provides LLM-council judgments of which steps to repair first in failed multi-agent executions. It extends Who&When with candidate rankings, individual judgments, short rationales, and ranking uncertainty, alongside the original decisive-error labels.
This is the dataset accompanying TRIAGE: Severity-Ranked Multi-Agent Failure Attribution, by Xizhi Wang, accepted to AACL 2026 (Main Conference). TRIAGE code is available on… See the full description on the dataset page: https://huggingface.co/datasets/jimmywang585/triage-bench.email-triage-v1
Email Triage v1
This dataset contains 1,740 unique email-triage examples produced through Tuned Tensor labeling and hardening workflows. It is the public v1 dataset, de-duplicated from a 2,038-row weighted fine-tuning dataset; 298 intentional weighting duplicates were removed for easier reuse.
The task is operational inbox triage, not security-risk classification. Each row asks a model to label one email-like message and return strict JSON with triage, priority, should_process… See the full description on the dataset page: https://huggingface.co/datasets/tunedtensor/email-triage-v1.
