datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mfaqWe present the first multilingual FAQ dataset publicly available. We collected around 6M FAQ pairs from the web, in 21 different languages.clinical-trials
Clinical Trials Dataset
A comprehensive dataset of clinical trials sourced from ClinicalTrials.gov, featuring structured metadata, detailed study information, and pre-computed semantic embeddings for machine learning applications in biomedical research.
Dataset Description
This dataset provides a rich collection of clinical trial information systematically collected from the official ClinicalTrials.gov database. Each record contains detailed study metadata, eligibility… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/clinical-trials.clinical-trials-qa
Clinical Trials QA Dataset
A multi-tier question-answering benchmark for evaluating Retrieval-Augmented Generation (RAG) systems on clinical trial protocols from ClinicalTrials.gov.
Dataset Summary
This dataset provides question-answer pairs across four difficulty tiers, designed to benchmark RAG systems on real-world clinical trial documentation. Questions span four reasoning categories and require retrieval from protocol PDFs.
Key Features:
4 difficulty tiers based on… See the full description on the dataset page: https://huggingface.co/datasets/Parexel/clinical-trials-qa.cli-commands-explained
Overview
This dataset is a collection of 16,098 command line instructions sourced from Commandlinefu and Cheatsheets. It includes an array of commands, each with an id, title, description, date, url to source, author, votes, and flag indicating if the description is AI generated. The descriptions are primarily authored by the original contributors, for entries where descriptions were absent, they have been generated using NeuralBeagle14-7B. Out of the total entries, 10,039… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/cli-commands-explained.pentabrid-reproducibility
Pentabrid 27B: reproducibility package
Everything required to recompute the results of a controlled evaluation of fine-tuning
configurations for medical question answering. Openly available with no access
restrictions.
Contents
Path
Description
per_item/medxpertqa_*.jsonl
Per-item predictions for all six checkpoints on 2,450 MedXpertQA-Text items. Fields: id, gold, extracted_answer, correct, explicit_marker_present, n_markers, response_chars… See the full description on the dataset page: https://huggingface.co/datasets/Clinical-Reasoning-Hub/pentabrid-reproducibility.epikg-clinicalbench
ClinicalBench: Assertion-Aware Clinical QA Benchmark
ClinicalBench is a benchmark for evaluating clinical question-answering systems on epistemic assertion reasoning over longitudinal patient records from MIMIC-IV.
It accompanies the paper:
ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Longitudinal Clinical QA
(Paper under review; preprint forthcoming)
Benchmark Overview
400 questions across 9 assertion categories
43 MIMIC-IV patients with… See the full description on the dataset page: https://huggingface.co/datasets/alexstinard/epikg-clinicalbench.clinical-trial-outcomes-predictions
Clinical Trial Outcomes Prediction Dataset
A dataset of 1,366 binary forecasting questions about clinical trial outcomes, automatically generated and labeled using Lightning Rod Labs' Future-as-Label methodology.
Dataset Description
This dataset contains questions about pharmaceutical clinical trials from 2023-2024, paired with verified outcomes (success/failure). Each question asks whether a specific trial will meet its endpoints, receive FDA approval, or complete by a… See the full description on the dataset page: https://huggingface.co/datasets/3rdSon/clinical-trial-outcomes-predictions.atlaspi-historical-geography
AtlasPI — Historical Geography Dataset
1,006 historical geopolitical entities · 643 events · 55 periods · 104 dynasty chains · 252 cities · 41 trade routes
The first open dataset specifically designed for AI agents working on
historical geography questions. Apache 2.0 licensed. Includes real GeoJSON
boundaries from academic sources, not placeholder polygons.
Temporal range: 4500 BCE → 2024 CE
Geographic coverage: all inhabited continents (Asia 31%, Africa 18%,
Americas 17%… See the full description on the dataset page: https://huggingface.co/datasets/clirim911/atlaspi-historical-geography.finject
FInject Dataset Card
FInject is a financial unanswerability benchmark built by transforming answerable financial reasoning problems into controlled unanswerable variants. Each row preserves the original question and pairs an answerable original context with a perturbed context that is no longer sufficient to support a unique answer.
Dataset Summary
Seed source: 78 answerable hard problems from FinanceReasoning.
Final release size: 426 unanswerable variants.… See the full description on the dataset page: https://huggingface.co/datasets/pnu-clink/finject.earth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.surgeon-tested-clinical-ai-benchmark
Surgeon-Tested Clinical AI Benchmark (TH-CAB v1.1)
An independent, reproducible evaluation of large language models (LLMs) on real, de-identified cancer cases — scored by a practicing surgeon item-by-item against current clinical guidelines.
Homepage & full leaderboard: https://tanhaosheng.asia
Methodology (citable authority, TH-CAB v1.1): https://tanhaosheng.asia/methodology/
Open data layer: https://tanhaosheng.asia/data/
This is a benchmark / research dataset, not clinical… See the full description on the dataset page: https://huggingface.co/datasets/tanhaosheng/surgeon-tested-clinical-ai-benchmark.Med-ART_Clinical_Agent_EHR_Dataset
ART — Action-based Reasoning Tasks (Subset)
120-task stratified sample from the ART benchmark introduced in:
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
Ananya Mantravadi, Shivali Dalmia, Abhishek Mukherji
arXiv:2601.08988
ART is a programmatically generated clinical decision benchmark built on real FHIR patient data. It targets three dominant error categories in medical AI reasoning — retrieval failures, aggregation errors, and conditional logic… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Med-ART_Clinical_Agent_EHR_Dataset.
