datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical_medqa_dialogue_en
Description
The dataset is from medalpaca/medical_meadow_mediqa, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_medqa_dialogue_en.MedQAOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
En split
Just edited columns. Contents are same.
Ko split
Train
The train dataset is translated by "solar-1-mini-translate-enko".
Test
The test dataset is translated by DeepL Pro.
reference-free COMET score: 0.7989 (Unbabel/wmt23-cometkiwi-da-xxl)
Citation information:
@article{jin2020disease… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/MedQA.MedQAData
MedQAData
A multilingual medical question-answering dataset covering 31 medical specialties
with 35,481 clinical Q&A pairs, translated into English, French, and Moroccan Darija.
This dataset is part of the BRAIN HEALTH project, intended for research and educational
purposes around multilingual clinical NLP.
Languages
English (en) — primary source language
French (fr) — full / partial translations
Moroccan Darija (ar / darija) — translations into Moroccan Arabic dialect… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQAData.medqa_corpus_enMedQA Textbook (English) with emphasis on domain of Clinical Medicine and other subsets.MedQAData-v2
MedQAData-v2
A fully multilingual medical question-answering dataset covering 31 medical specialties
with 35,481 clinical Q&A pairs, each provided in English, French, and Moroccan Darija.
This is the v2 release, with every field filled — no more gaps.
Part of the BRAIN HEALTH project, for research and educational use around multilingual clinical NLP.
What's new in v2
Compared to v1:
Improvement
v1
v2
Columns
12
14 (added context_question_en… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQAData-v2.MedQA-Evol-Korean
MedQA-Evol
Original Data: TsinghuaC3I/UltraMedical.
Translated into Korean by "solar-1-mini-translate-enko".
001_MedQA_rawOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
MedQADataEnglishSaad
Health QA English — Medical Question Answering Dataset
Dataset Description
A curated dataset of 13,812 medical question-answer pairs sourced from real patient-doctor consultations. Each entry contains a patient's clinical scenario, a focused medical question, and a doctor's professional response, enriched with named medical entities (symptoms, diseases, medications, tests).
Key Features
13,812 high-quality entries across 15 medical specialties
Structured… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedQADataEnglishSaad.Open-MedQA-Nexus
Open Nexus MedQA
This dataset combines various publicly available medical datasets like ChatDoctor, icliniq, etc., into a unified format for training and evaluating medical question-answering models.
Dataset Details
Open Nexus MedQA is a comprehensive dataset designed to facilitate the development of advanced medical question answering systems. It integrates diverse medical data sources, meticulously processed to provide a uniform format. The format includes:… See the full description on the dataset page: https://huggingface.co/datasets/exafluence/Open-MedQA-Nexus.MedQA-Pretrain-Corpus
MedQA Pretrain Corpus
A medical pretraining corpus combining multiple high-quality medical
text sources for training my mini small language model (mini-SLM) on
medical domain knowledge just for learning, fun and research purpose.
Dataset Description
This corpus was built for causal language model pretraining
on the medical domain. It is NOT a finetuning dataset.
Total Size
~1.5 GB of clean medical text
Sources Combined
Medical… See the full description on the dataset page: https://huggingface.co/datasets/salisai/MedQA-Pretrain-Corpus.MedQADataDarijaSaad
Health QA Darija — Medical QA in Moroccan Arabic (الدارجة المغربية)
Dataset Description
A curated dataset of 8,129 medical question-answer pairs in Moroccan Darija (الدارجة المغربية). Each entry contains a patient scenario, a focused medical question, and a doctor's response — all in authentic Darija. Enriched with named medical entities (symptoms, diseases, medications, tests).
🇲🇦 First large-scale medical QA dataset in Moroccan Darija — addressing the critical gap in… See the full description on the dataset page: https://huggingface.co/datasets/BrainHealthAI/MedQADataDarijaSaad.
