datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.mediflow
MediFlow
A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents.
t-SNE 2D Plot of MediFlow Embeddings by Task Types
Dataset Splits
mediflow: 2.5M instruction data for SFT alignment.
mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment.
Main Columns
instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.whiteglove-medical-medlineplus-2025
WhiteGlove Medical Knowledge Corpus
MedlinePlus 2025 — Spectral Curation Pipeline
Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government)
Dataset Summary
A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.russian-nmo-medical-mcq
Russian NMO Medical MCQ
Choose language / Выберите язык: Русский | English
Русский
Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа.
В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными
вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для
тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA.
Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/IMDb-Media.triage-medical-dataset
Dataset release
Version: 2026-03-19-v1
Published at: 2026-03-19T15:42:16+00:00
Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset
Dataset Card - POC Triage Medical
Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance.
1) Description
Dataset bilingue FR/EN pour triage medical initial.
Le pipeline produit deux artefacts principaux:
SFT: paires instruction/reponse pour le fine-tuning supervise.
DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.medium-web-pentesting
Medium Web Pentesting Articles
Dataset Description
A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body.
This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.cc-mediation
CC-Mediation
A cross-cultural conflict-mediation benchmark grounded in the Developmental
Model of Intercultural Sensitivity (DMIS). Each scenario is a culturally
grounded conflict dialogue with a mediation intervention and its
post-intervention trajectory, organised as a preference pair
(positive vs negative continuation) so downstream effects are measurable.
📄 Paper: CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation (Lee, Zhang, Yow, Deng;… See the full description on the dataset page: https://huggingface.co/datasets/Suhyunlee/cc-mediation.Olympiads_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 13284
Filtered size: 13240
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.bulgarian-medical-cpt-100m
Bulgarian text for MOSS continued pretraining
Exactly 100 million training tokens: 10M medical and 90M general Bulgarian.
An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers.
No model training has been performed as part of this dataset build.
Medical data
Exactly 10,000,000 training tokens and 50,000 additional validation tokens,
including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.Numina_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 37133
Filtered size: 37133
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.medical_data_for_slm
🏥 Medical SLM Pretraining Dataset Card
This dataset is a high-quality, cleaned collection of medical text designed for pretraining small language models (SLMs). It aggregates data from three primary authoritative sources, focusing on general medicine and clinical guidelines.
📊 Dataset Summary
Total Documents: ~44,400
Estimated Tokens: ~44.7 Million
Primary Language: English
Configurations:
documents: Raw cleaned text records.
chunks: Tokenized and packed 1024-token… See the full description on the dataset page: https://huggingface.co/datasets/Saminx22/medical_data_for_slm.bulgarian-medical-cpt-10m
Bulgarian text for MOSS continued pretraining
Exactly 10 million training tokens: 3M medical and 7M general Bulgarian.
An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers.
No model training has been performed as part of this dataset build.
Medical data
Exactly 3,000,000 training tokens and 50,000 additional validation tokens,
including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-10m.Nemotron-Math-v2-Medium-10k
Nemotron-Math-v2-Medium-10k
A lightweight 10,500-problem subset of
nvidia/Nemotron-Math-v2
for long-horizon Python-TIR reinforcement learning. It contains 1,500 problems
from each metadata.reason_high_with_tool.pass bucket 1 through 7. A
deterministic seed-42 shuffle assigns 500 examples to validation and 10,000
to train.
This Hugging Face release intentionally contains no teacher traces. The
full messages/tools aggregation is retained as a separate local artifact.… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Nemotron-Math-v2-Medium-10k.test-medicina
Medschool-Test, or "Test di Medicina"
Is your LLM able to pass a National Entrance Exam for the Italian Medical School?
This the GitHub repo for our Hugging Face dataset designed for evaluating Large Language Models (LLMs) on a broad range of questions from the national entrance exams for the Italian medical school (ORIGINAL WEBSITE).
The dataset includes multiple-choice questions from various subjects such as biology, chemistry, physics, mathematics, world… See the full description on the dataset page: https://huggingface.co/datasets/room-b007/test-medicina.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.khmer-medical-qa
Khmer Medical Q&A Dataset
Update Notice
✅ Dataset Updated (Aug 15, 2024): We identified and fixed an alignment issue. Everything is working properly now!
Dataset Description
This dataset contains 18,756 high-quality medical question-answer pairs translated from English to Khmer, designed for training medical AI assistants in the Khmer language.
Features
index: Sequential row index (0-18755)
question_en: Medical question in English
response_en:… See the full description on the dataset page: https://huggingface.co/datasets/khopilot/khmer-medical-qa.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.MediFlowThinks
MediFlow
A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents.
t-SNE 2D Plot of MediFlow Embeddings by Task Types
Dataset Splits
mediflow: 2.5M instruction data for SFT alignment.
mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment.
Main Columns
instruction:… See the full description on the dataset page: https://huggingface.co/datasets/IAMRonHIT/MediFlowThinks.star-mediumSynthetic data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions.
Project page: https://dhruveshp.com/projects/ilm
IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of… See the full description on the dataset page: https://huggingface.co/datasets/RyanHalliwell/IMDb-Media.drbodebench_medicamentos
Medication-Focused Clinical Benchmark from DrBodeBench
Dataset Details
To evaluate retrieval capabilities in higher-level reasoning scenarios, we created a second benchmark derived from the Portuguese medical benchmark DrBodeBench. This benchmark aggregates questions from Brazilian medical examinations, including the Revalida and the FUVEST direct-access residency exam. From DrBodeBench, we curated a specific subset of questions that exclusively pertains to… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/drbodebench_medicamentos.ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513
GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps
This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps.
Summary
Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513
Source rows: 500 K2-Think Codeforces traces
Source manifest-selected Qwen3.5 labels: 10,000 steps
GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.real-world-medical-mistakes-dataset
Real-World Medical Mistakes Dataset
A curated dataset of 100 de-identified clinical reports from Internal Medicine and Emergency Departments, each containing a physician-inserted realistic medical error. Designed for training and evaluating AI systems that detect critical patient safety errors in clinical documentation.
Dataset Description
Overview
This dataset was created as part of the Clinipal project — an AI-powered clinical error detection system. Three… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/real-world-medical-mistakes-dataset.spai-ss6-corpus-medical-o1-verifiable
SPAI SS6 Medical O1 Verifiable Thai Index
Index repo for the imported Thai medical verifiable-problem dataset config.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: medical_o1_verifiable_problem_thai
Rows in canonical config: 40,906
Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-o1-verifiable.medicine
medicine
Thai public medical and health web corpus collected for research and LLM dataset experimentation.
Dataset Contents
Split: train
Records: 3035 deduplicated articles
Format: Parquet
Latest collection profile: free_1000
Latest generated at: 2026-06-06T16:25:20.801225+00:00
Source And Method
URLs are collected from public sitemap XML files on configured Thai public
sources, then crawled with robots.txt checks, rate limiting, Thai text… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/medicine.spai-ss6-corpus-medical-health-web
SPAI SS6 Thai Medical Health Web Corpus
Thai public medical and health web articles collected by the local scraping pipeline.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: default
Rows in canonical config: 3,660
Parquet size in canonical config: 0.01 GB
Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.Medical-R1-Distill-Data-m1k
Medical-R1-Distill-Data
