CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GBaker /MedQA-USMLE-4-optionsOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams Citation information: @article{jin2020disease, title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={arXiv preprint arXiv:2009.13081}, year={2020} } text10K<n<100K100 likes90k downloads4y agoHugging Face02bigbio /med_qa Dataset Card for MedQA In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA, collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. Together with the question data, we also collect and release a large-scale corpus from medical textbooks from which the… See the full description on the dataset page: https://huggingface.co/datasets/bigbio/med_qa.text100K<n<1M158 likes13k downloads22d agoHugging Face03GBaker /MedQA-USMLE-4-options-hfOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams Citation information: @article{jin2020disease, title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={arXiv preprint arXiv:2009.13081}, year={2020} } text10K<n<100K24 likes11k downloads4y agoHugging Face04openlifescienceai /medqatext10K<n<100K6 likes9.5k downloads2y agoHugging Face05medalpaca /medical_meadow_medqa Dataset Card for MedQA Dataset Summary This is the data and baseline source code for the paper: Jin, Di, et al. "What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams." From https://github.com/jind11/MedQA: The data that contains both the QAs and textbooks can be downloaded from this google drive folder. A bit of details of data are explained as below: For QAs, we have three sources: US, Mainland of China, and… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medqa.textquestion-answering10K<n<100K117 likes8.3k downloads3y agoHugging Face06davidheineman /medqa-enUploaded version of the MedQA english subset: https://github.com/jind11/MedQA Citation @article{jin2021disease, title={What disease does this patient have? a large-scale open domain question answering dataset from medical exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={Applied Sciences}, volume={11}, number={14}, pages={6421}, year={2021}, publisher={MDPI} } text10K<n<100K0 likes3.7k downloads1y agoHugging Face07xuxuxuxuxu /MedQA_US_testtext1K<n<10K0 likes1.4k downloads1y agoHugging Face08R2MED /MedQA-Diag 🔭 Overview R2MED: First Reasoning-Driven Medical Retrieval Benchmark R2MED is a high-quality, high-resolution synthetic information retrieval (IR) dataset designed for medical scenarios. It contains 876 queries with three retrieval tasks, five medical scenarios, and twelve body systems. Dataset #Q #D Avg. Pos Q-Len D-Len Biology 103 57359 3.6 115.2 83.6 Bioinformatics77 47473 2.9 273.8 150.5 Medical Sciences 88 34810 2.8 107.1 122.7 MedXpertQA-Exam 97… See the full description on the dataset page: https://huggingface.co/datasets/R2MED/MedQA-Diag.texttext-retrieval10K<n<100K0 likes1.3k downloads1y agoHugging Face09truehealth /medqatext10K<n<100K6 likes906 downloads3y agoHugging Face10HPAI-BSC /MedQA-Mixtral-CoT Dataset Card for medqa-cot Synthetically enhanced responses to the medqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.textmultiple-choice10K<n<100K9 likes764 downloads2y agoHugging Face11Williamsanderson /MedQAData-EN MedQAData-EN English-focused medical Q&A dataset derived from MedQAData-v2. Field Description context_question Full patient narrative question Short direct question reformulated from context (1 line, ends with ?) answer Concise doctor answer (filler removed) language English urgency low / medium / high / critical speciality Medical specialty article_title Reference article title entities Dict with keys age, medicament, sympt, medical_field, disease, Test… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQAData-EN.textquestion-answering10K<n<100K0 likes684 downloads5mo agoHugging Face12Williamsanderson /MedQA-Darija-MultiLingual MedQA-Darija-MultiLingual The largest open trilingual medical Q&A dataset with directly-playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low-resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase. After programmatic validation (Stage 1 LOF outlier… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQA-Darija-MultiLingual.audioquestion-answering100K<n<1M4 likes514 downloads5mo agoHugging Face13openlifescienceai /MedQA-USMLE-4-options-hftext10K<n<100K1 likes494 downloads2y agoHugging Face14awinml /medqa MedQA (USMLE 4-option, US subset + English textbook corpus) Dataset Summary This dataset is a re-upload of the English USMLE 4-option question subset and the English textbook corpus from the original jind11/MedQA release introduced by Jin et al. in What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. The original MedQA release contains question sets in English, Simplified Chinese, and Traditional Chinese, and also… See the full description on the dataset page: https://huggingface.co/datasets/awinml/medqa.textquestion-answering10K<n<100K1 likes432 downloads5mo agoHugging Face15augtoma /medqa_usmle Dataset Card for "medqa_usmle" More Information needed text10K<n<100K1 likes345 downloads3y agoHugging Face16HPAI-BSC /Medprompt-MedQA-CoT Medprompt-MedQA-CoT Dataset Summary Medprompt-MedQA-CoT is a retrieval-augmented database created to enhance contextual reasoning in multiple-choice medical question answering (MCQA). The dataset follows a Chain-of-Thought (CoT) reasoning format, providing step-by-step justifications for each question before identifying the correct answer. Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Medprompt-MedQA-CoT.question-answering10K<n<100K1 likes334 downloads1y agoHugging Face17fzkuji /MedQAWant to fine-tune this dataset on LLaMA-Factory? Check this repository for preprocessing: llm-merging datasets I automatically converted the dataset into the default format that can be previewed on huggingface. Dataset Card for MedQA In this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA, collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and traditional Chinese, and… See the full description on the dataset page: https://huggingface.co/datasets/fzkuji/MedQA.text100K<n<1M3 likes329 downloads2y agoHugging Face18empirischtech /med-qa-orpo-dpo MED QA ORPO-DPO Dataset This dataset is restructured from several existing datasource on medical literature and research, hosted here on hugging face. The dataset is shaped in question, choosen and rejected pairs to match the ORPO-DPO trainset requirements. Features The dataset consists of the following features: question: MCQ or yes/no/maybe based questions on medical questions direct-answer: correct answer to the above question chosen: the correct answer along with… See the full description on the dataset page: https://huggingface.co/datasets/empirischtech/med-qa-orpo-dpo.textquestion-answering100K<n<1M7 likes318 downloads2y agoHugging Face19nnilayy /medqa-usmletext10K<n<100K1 likes305 downloads2y agoHugging Face20mathewhe /medqa Dataset Card for MedQA Homepage: https://github.com/jind11/MedQA This is an unofficial curation of the MedQA dataset, uploaded here with minimal (i.e., no content-modifying) processing. Paper: What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams (MDPI) Languages: English (en), Taiwanese (tw), and Chinese (zh). Dataset Subsets This dataset contains multiple configs: QA with four possible answers (as reported… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/medqa.textquestion-answering100K<n<1M0 likes299 downloads11mo agoHugging Face21lamhieu /medical_medqa_dialogue_en Description The dataset is from medalpaca/medical_meadow_mediqa, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it. Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth. Structure View online through viewer. Note We advise you to reconsider before use, thank you. If you find it useful, please like and follow… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/medical_medqa_dialogue_en.texttext-generation10K<n<100K2 likes249 downloads2y agoHugging Face22katielink /med_qaIn this work, we present the first free-form multiple-choice OpenQA dataset for solving medical problems, MedQA, collected from the professional medical board exams. It covers three languages: English, simplified Chinese, and traditional Chinese, and contains 12,723, 34,251, and 14,123 questions for the three languages, respectively. Together with the question data, we also collect and release a large-scale corpus from medical textbooks from which the reading comprehension models can obtain necessary knowledge for answering the questions.1 likes231 downloads3y agoHugging Face23ChuGyouk /MedQAOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams En split Just edited columns. Contents are same. Ko split Train The train dataset is translated by "solar-1-mini-translate-enko". Test The test dataset is translated by DeepL Pro. reference-free COMET score: 0.7989 (Unbabel/wmt23-cometkiwi-da-xxl) Citation information: @article{jin2020disease… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/MedQA.texttext-generation10K<n<100K3 likes209 downloads2y agoHugging Face24katielink /nejm-medqa-diagnostic-reasoning-datasetDownloaded from Supplemental Information of the article "Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine " [link] Savage, T., Nayak, A., Gallo, R. et al. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digit. Med. 7, 20 (2024). https://doi.org/10.1038/s41746-024-01010-1 tabularn<1K8 likes198 downloads3y agoHugging Face25katielink /agentclinic_medqa AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments Release [05/18/2024] 🤗 We added support for HuggingFace models! [05/17/2024] We release new results and support for GPT-4o! [05/13/2024] 🔥 We release AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environment. We propose a multimodal benchmark based on language agents which simulate the clinical environment. Checkout the paper and the… See the full description on the dataset page: https://huggingface.co/datasets/katielink/agentclinic_medqa.imagen<1K2 likes168 downloads2y agoHugging Face26Williamsanderson /MedQAData MedQAData A multilingual medical question-answering dataset covering 31 medical specialties with 35,481 clinical Q&A pairs, translated into English, French, and Moroccan Darija. This dataset is part of the BRAIN HEALTH project, intended for research and educational purposes around multilingual clinical NLP. Languages English (en) — primary source language French (fr) — full / partial translations Moroccan Darija (ar / darija) — translations into Moroccan Arabic dialect… See the full description on the dataset page: https://huggingface.co/datasets/Williamsanderson/MedQAData.textquestion-answering10K<n<100K0 likes162 downloads5mo agoHugging Face27mkieffer /MedQA-USMLE MedQA-USMLE HuggingFace upload of the MedQA-USMLE dataset with deduping. If used, please cite the original authors using the citation below. A small number of exact-duplicate questions were identified within train and us_qbank. The question text was identical, but the options were formatted slightly differently or had a different distractor. The main difference was the listed correct letter, so the incorrect duplicates were removed. Each split was then reindexed to keep indices… See the full description on the dataset page: https://huggingface.co/datasets/mkieffer/MedQA-USMLE.tabularquestion-answering10K<n<100K0 likes160 downloads8mo agoHugging Face28clinicalnlplab /MedQA_train Dataset Card for "MedQA_train" More Information needed text10K<n<100K3 likes144 downloads3y agoHugging Face29AIM-Harvard /gbaker_medqa_usmle_4_options_hf_generic_to_brandtabular1K<n<10K0 likes144 downloads2y agoHugging Face30HPAI-BSC /medqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedQa dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medqa-cot-llama31.textmultiple-choice10K<n<100K3 likes141 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.