datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triage-medical-dataset
Dataset release
Version: 2026-03-19-v1
Published at: 2026-03-19T15:42:16+00:00
Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset
Dataset Card - POC Triage Medical
Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance.
1) Description
Dataset bilingue FR/EN pour triage medical initial.
Le pipeline produit deux artefacts principaux:
SFT: paires instruction/reponse pour le fine-tuning supervise.
DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.legal-statutory-triage-sft
Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset
This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860).
1. Overview and Scope
With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.email-triage-v1
Email Triage v1
This dataset contains 1,740 unique email-triage examples produced through Tuned Tensor labeling and hardening workflows. It is the public v1 dataset, de-duplicated from a 2,038-row weighted fine-tuning dataset; 298 intentional weighting duplicates were removed for easier reuse.
The task is operational inbox triage, not security-risk classification. Each row asks a model to label one email-like message and return strict JSON with triage, priority, should_process… See the full description on the dataset page: https://huggingface.co/datasets/tunedtensor/email-triage-v1.chsa-triage-medical-qa-fr-en
CHSA Triage POC — corpus médical bilingue SFT et préférences DPO
Jeu de données d'un projet d'étude : Proof of Concept d'assistant de triage médical initial (Qwen3-1.7B-Base, SFT LoRA puis DPO). Code, rapport et preuves : github.com/ppluton/medical-triage-llm-poc.
Usage pédagogique uniquement. Ce jeu n'est ni validé cliniquement ni destiné à un dispositif médical, un diagnostic, une prescription ou une décision clinique. Il ne contient aucun label de priorité de triage.… See the full description on the dataset page: https://huggingface.co/datasets/Pedro1321/chsa-triage-medical-qa-fr-en.triagebench
TriageBench
TriageBench measures whether a clinical AI gives the same triage decision when you change something about the patient that should not affect the answer: their gender, the language they wrote in, or a socioeconomic signal such as a ZIP code. It holds the symptoms identical, swaps one irrelevant detail, and measures how far the decision moves. It scores consistency and makes no claim about which triage call is clinically correct.
Code and harness:… See the full description on the dataset page: https://huggingface.co/datasets/wongqihan/triagebench.finance-ops-triage-v0.1
Finance Ops Triage v0.1 dataset
The original small, illustrative dataset prepared for Ugo Chukwu's first Unsloth fine-tuning and deployment exercise. The examples were provided during a guided ChatGPT experiment; they are not collected operational transaction records or an independently validated finance policy.
Code and experiment report · Model archive
Structure
Each JSONL row has messages containing system, user, and assistant entries. The assistant content is… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/finance-ops-triage-v0.1.engineering-log-triage-dataset
Engineering Log Triage Dataset
Summary
This dataset contains synthetic/sanitized engineering-log examples for structured fault triage.
Each example is formatted as a chat-style supervised fine-tuning record. The model input is an unstructured engineering report. The target assistant message is a strict JSON object containing a structured triage result.
This dataset was created for the LoRA-Adapted Engineering Log Triage Service project.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/cobra9786/engineering-log-triage-dataset.email-safety-triage-10k
Email Safety Triage 10k
This dataset contains 10,000 supervised examples for classifying email and email-adjacent content for operational triage, phishing/spam risk, and prompt-attack filtering.
Each JSONL row has two string fields:
input: an instruction plus email, security-review text, or prompt/email fragment.
output: compact strict JSON with triage, priority, risk, should_process, confidence, and reason.
The dataset is intended for fine-tuning and evaluating classifiers… See the full description on the dataset page: https://huggingface.co/datasets/weijianzhg/email-safety-triage-10k.verdict-engine-triage-v2
Verdict Engine — SFT v2 (Sonnet-4.6 distilled)
The supervised fine-tuning dataset behind every model in the
Verdict Engine bake-off: four LoRA
fine-tunes (hqt2yotoz/verdict-engine-{qwen3-4b-2507,qwen3-8b,smollm3-3b,qwen2.5-7b}-lora)
that all beat their own base on a blind panel. This is the v2 / "true-distillation"
dataset — the one whose labels come from a genuinely stronger teacher (Claude Sonnet 4.6),
not from the small model labeling itself.
What it teaches… See the full description on the dataset page: https://huggingface.co/datasets/hqt2yotoz/verdict-engine-triage-v2.
