datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA_DutchTranslation of the English version of MedQA,
to Dutch using the GPT 4.1 mini LLM by OpenAI.
Attribution
If you use this dataset please use the following to credit the creators of MedQA:
@article{jin2021disease,
title={What disease does this patient have? a large-scale open domain question answering dataset from medical exams},
author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter},
journal={Applied Sciences}… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/MedQA_Dutch.BioASQ_11B_DutchBio ASQ challenge 11b
This can be used to finetune a decoder model for Q/A interaction,
alternatively it can be used to create (question, positive, negative)
triplets to train a sentence encoder using SBERT.
Reference:
@inbook{Nentidis_2023,
title={Overview of BioASQ 2023: The Eleventh BioASQ Challenge on Large-Scale Biomedical Semantic Indexing and Question Answering},
ISBN={9783031424489},
ISSN={1611-3349},
url={http://dx.doi.org/10.1007/978-3-031-42448-9_19}… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/BioASQ_11B_Dutch.MedQA_SymptomDisease_small_DutchA Dutch translation of this huggingface dataset using GPT4.1 mini, with the courtesy of Prognosis.
