datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA-USMLE-4-optionsOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
med_qa
Dataset Card for MedQA
In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA,
collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and
traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. Together
with the question data, we also collect and release a large-scale corpus from medical textbooks from which the… See the full description on the dataset page: https://huggingface.co/datasets/bigbio/med_qa.MedQA-USMLE-4-options-hfOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Citation information:
@article{jin2020disease,
title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={arXiv preprint arXiv:2009.13081},
year={2020}
}
medqamedical_meadow_medqa
Dataset Card for MedQA
Dataset Summary
This is the data and baseline source code for the paper: Jin, Di, et al. "What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams."
From https://github.com/jind11/MedQA:
The data that contains both the QAs and textbooks can be downloaded from this google drive folder. A bit of details of data are explained as below:
For QAs, we have three sources: US, Mainland of China, and… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medqa.medqa-enUploaded version of the MedQA english subset: https://github.com/jind11/MedQA
Citation
@article{jin2021disease,
title={What disease does this patient have? a large-scale open domain question answering dataset from medical exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={Applied Sciences},
volume={11},
number={14},
pages={6421},
year={2021},
publisher={MDPI}
}
MedQA_US_testMedQA-Diag
🔭 Overview
R2MED: First Reasoning-Driven Medical Retrieval Benchmark
R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems.
Dataset
#Q
#D
Avg. Pos
Q-Len
D-Len
Biology
103
57359
3.6
115.2
83.6
Bioinformatics77
47473
2.9
273.8
150.5
Medical Sciences
88
34810
2.8
107.1
122.7
MedXpertQA-Exam
97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedQA-Diag.medqaMedQA-Mixtral-CoT
Dataset Card for medqa-cot
Synthetically enhanced responses to the medqa dataset using mixtral.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a
hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.MedQAData-EN
MedQAData-EN
English-focused medical Q&A dataset derived from MedQAData-v2.
Field
Description
context_question
Full patient narrative
question
Short direct question reformulated from context (1 line, ends with ?)
answer
Concise doctor answer (filler removed)
language
English
urgency
low / medium / high / critical
speciality
Medical specialty
article_title
Reference article title
entities
Dict with keys age, medicament, sympt, medical_field, disease, Test… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQAData-EN.MedQA-Darija-MultiLingual
MedQA-Darija-MultiLingual
The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija.
A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region.
Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.MedQA-USMLE-4-options-hfmedqa
MedQA (USMLE 4-option, US subset + English textbook corpus)
Dataset Summary
This dataset is a re-upload of the English USMLE 4-option question subset and the English textbook corpus from the original jind11/MedQA release introduced by Jin et al. in What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams.
The original MedQA release contains question sets in English, Simplified Chinese, and Traditional Chinese, and also… See the full description on the dataset page: https://huggingface.co/datasets/awinml/medqa.medqa_usmle
Dataset Card for "medqa_usmle"
More Information needed
Medprompt-MedQA-CoT
Medprompt-MedQA-CoT
Dataset Summary
Medprompt-MedQA-CoT is a retrieval-augmented database created to enhance contextual reasoning in multiple-choice medical question answering (MCQA). The dataset follows a Chain-of-Thought (CoT) reasoning format, providing step-by-step justifications for each question before identifying the correct answer.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Medprompt-MedQA-CoT.MedQAWant to fine-tune this dataset on LLaMA-Factory? Check this repository for preprocessing: llm-merging datasets
I automatically converted the dataset into the default format that can be previewed on huggingface.
Dataset Card for MedQA
In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA,
collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and
traditional Chinese, and… See the full description on the dataset page: https://huggingface.co/datasets/fzkuji/MedQA.med-qa-orpo-dpo
MED QA ORPO-DPO Dataset
This dataset is restructured from several existing datasource on medical literature and research, hosted here on hugging face. The dataset is shaped in question, choosen
and rejected pairs to match the ORPO-DPO trainset requirements.
Features
The dataset consists of the following features:
question: MCQ or yes/no/maybe based questions on medical questions
direct-answer: correct answer to the above question
chosen: the correct answer along with… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/med-qa-orpo-dpo.medqa-usmlemedqa
Dataset Card for MedQA
Homepage: https://github.com/jind11/MedQA
This is an unofficial curation of the MedQA dataset, uploaded here with minimal (i.e., no content-modifying) processing.
Paper: What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams (MDPI)
Languages: English (en), Taiwanese (tw), and Chinese (zh).
Dataset Subsets
This dataset contains multiple configs:
QA with four possible answers (as reported… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/medqa.medical_medqa_dialogue_en
Description
The dataset is from medalpaca/medical_meadow_mediqa, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_medqa_dialogue_en.med_qaIn this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA,
collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and
traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. Together
with the question data, we also collect and release a large-scale corpus from medical textbooks from which the reading
comprehension models can obtain necessary knowledge for answering the questions.MedQAOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
En split
Just edited columns. Contents are same.
Ko split
Train
The train dataset is translated by "solar-1-mini-translate-enko".
Test
The test dataset is translated by DeepL Pro.
reference-free COMET score: 0.7989 (Unbabel/wmt23-cometkiwi-da-xxl)
Citation information:
@article{jin2020disease… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/MedQA.nejm-medqa-diagnostic-reasoning-datasetDownloaded from Supplemental Information of the article "Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine
" [link]
Savage, T., Nayak, A., Gallo, R. et al. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digit. Med. 7, 20 (2024). https://doi.org/10.1038/s41746-024-01010-1
agentclinic_medqa
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Release
[05/18/2024] 🤗 We added support for HuggingFace models!
[05/17/2024] We release new results and support for GPT-4o!
[05/13/2024] 🔥 We release AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environment. We propose a multimodal benchmark based on language agents which simulate the clinical environment. Checkout the paper and the… See the full description on the dataset page: https://huggingface.co/datasets/katielink/agentclinic_medqa.MedQAData
MedQAData
A multilingual medical question-answering dataset covering 31 medical specialties
with 35,481 clinical Q&A pairs, translated into English, French, and Moroccan Darija.
This dataset is part of the BRAIN HEALTH project, intended for research and educational
purposes around multilingual clinical NLP.
Languages
English (en) — primary source language
French (fr) — full / partial translations
Moroccan Darija (ar / darija) — translations into Moroccan Arabic dialect… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQAData.MedQA-USMLE
MedQA-USMLE
HuggingFace upload of the MedQA-USMLE dataset with deduping. If used, please cite the original authors using the citation below.
A small number of exact-duplicate questions were identified within train and us_qbank. The question text was identical, but the options were formatted slightly differently or had a different distractor. The main difference was the listed correct letter, so the incorrect duplicates were removed. Each split was then reindexed to keep indices… See the full description on the dataset page: https://huggingface.co/datasets/mkieffer/MedQA-USMLE.MedQA_train
Dataset Card for "MedQA_train"
More Information needed
gbaker_medqa_usmle_4_options_hf_generic_to_brandmedqa-cot-llama31
medqa-cot-llama31
Synthetically enhanced responses to the MedQa dataset. Used to train Aloe-Beta model.
Dataset Details
Dataset Description
To increase the quality of answers from the training splits of the MedQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medqa-cot-llama31.
