datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human_anatomy_qa_with_difficulty
Truth, Trust, and Trouble (TTT) – Medical Anatomy QA Benchmark
This repository hosts the dataset introduced in the EMNLP Industry Track 2025 paper “Truth, Trust, and Trouble: Medical AI on the Edge.”
The dataset contains 1,077 high-quality, clinically validated True/False anatomy questions, designed to evaluate medical LLMs along three critical axes:
Honesty (factual alignment)
Helpfulness (semantic relevance & completeness)
Harmlessness (safety under clinical constraints)
This… See the full description on the dataset page: https://huggingface.co/datasets/ekplatebiryani/human_anatomy_qa_with_difficulty.Human-Like-Gut-Health-DPO-QnA
Gut Health DPO Dataset
Overview
This dataset contains 200 carefully curated examples for Direct Preference Optimization (DPO) training in the domain of gut health and digestive wellness. Each example consists of a user prompt, a "chosen" response (preferred), and a "rejected" response (less preferred), designed to train AI models to provide high-quality, medically responsible advice on digestive health topics.
Dataset Structure
The dataset is provided in CSV… See the full description on the dataset page: https://huggingface.co/datasets/Sid3503/Human-Like-Gut-Health-DPO-QnA.human-confabulation-benchmark
Human-Confabulated Hallucination Benchmark
Read this before you report a number
Known confound: authorship. In this dataset every grounded_response was written by a language model from a source, and every fabricated_response was written by a human from memory. Authorship is therefore perfectly correlated with the label. A detector can score very highly here by recognising who wrote the text rather than whether it is grounded.
This is not a hypothetical. In… See the full description on the dataset page: https://huggingface.co/datasets/groundlens/human-confabulation-benchmark.
