CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GBaker /MedQA-USMLE-4-optionsOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams Citation information: @article{jin2020disease, title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={arXiv preprint arXiv:2009.13081}, year={2020} } text10K<n<100K100 likes91k downloads4y agoHugging Face02GBaker /MedQA-USMLE-4-options-hfOriginal dataset introduced by Jin et al. in What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams Citation information: @article{jin2020disease, title={What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams}, author={Jin, Di and Pan, Eileen and Oufattole, Nassim and Weng, Wei-Hung and Fang, Hanyi and Szolovits, Peter}, journal={arXiv preprint arXiv:2009.13081}, year={2020} } text10K<n<100K24 likes12k downloads4y agoHugging Face03openlifescienceai /MedQA-USMLE-4-options-hftext10K<n<100K1 likes507 downloads2y agoHugging Face04augtoma /medqa_usmle Dataset Card for "medqa_usmle" More Information needed text10K<n<100K1 likes384 downloads3y agoHugging Face05nnilayy /medqa-usmletext10K<n<100K1 likes308 downloads2y agoHugging Face06augtoma /usmle_step_1 Dataset Card for "usmle_self_eval_step1" More Information needed textn<1K1 likes290 downloads3y agoHugging Face07augtoma /usmle_step_2 Dataset Card for "usmle_self_eval_step2" More Information needed textn<1K1 likes224 downloads3y agoHugging Face08LeoZotos /usmle_textbooksTextbooks: Epidemiology https://vtechworks.lib.vt.edu/items/3f46b25c-c8e7-46bc-a7ef-64e0db8556be (ISBN 13: 9781957213651) Neuroscience for Pre-Clinical Students https://pressbooks.lib.vt.edu/neuroscience/ (ISBN 978-1-949373-80-6) Microbiology, Pharmacology, and Immunology for Pre-Clinical Students https://pressbooks.lib.vt.edu/micropharmimmuno/ (ISBN 978-1-962841-04-7) text10K<n<100K0 likes199 downloads13d agoHugging Face09augtoma /usmle_step_3 Dataset Card for "usmle_self_eval_step3" More Information needed textn<1K0 likes182 downloads3y agoHugging Face10awacke1 /USMLE-Test-Traintext10K<n<100K4 likes163 downloads3y agoHugging Face11mkieffer /MedQA-USMLE MedQA-USMLE HuggingFace upload of the MedQA-USMLE dataset with deduping. If used, please cite the original authors using the citation below. A small number of exact-duplicate questions were identified within train and us_qbank. The question text was identical, but the options were formatted slightly differently or had a different distractor. The main difference was the listed correct letter, so the incorrect duplicates were removed. Each split was then reindexed to keep indices… See the full description on the dataset page: https://huggingface.co/datasets/mkieffer/MedQA-USMLE.tabularquestion-answering10K<n<100K0 likes157 downloads8mo agoHugging Face12AIM-Harvard /gbaker_medqa_usmle_4_options_hf_generic_to_brandtabular1K<n<10K0 likes148 downloads2y agoHugging Face13kernelvectortech /usmle-crackers-question-bank USMLE Crackers Question Bank 198,379 medical multiple-choice questions, every one assigned a topic and a chapter from a closed taxonomy of 20 topics and 228 chapters. This is a re-annotation of two existing open datasets, not new questions. What it adds is complete, consistent categorization: Upstream MedQA has no topic labels at all. Upstream MedMCQA has 21 coarse subjects, one of which is literally Unknown, and a topic_name field that is null on 53% of rows and spread over 2… See the full description on the dataset page: https://huggingface.co/datasets/kernelvectortech/usmle-crackers-question-bank.textquestion-answering100K<n<1M0 likes118 downloads24d agoHugging Face14HuggingSara /usmle_self_assessment Dataset Card for "usmle_self_assesment" More Information needed textn<1K0 likes91 downloads3y agoHugging Face15AIM-Harvard /gbaker_medqa_usmle_4_options_hf_originaltabular1K<n<10K0 likes89 downloads2y agoHugging Face16agentN0 /usmle-step1-form31 USMLE Step 1 — Form 31, Form 30 & NBME Form 27 Structured multiple-choice questions extracted from USMLE Step 1 and NBME practice forms. Files data/questions.jsonl — 183 USMLE Step 1 Form 31 questions (35 SOTA-cropped images) questions_nbme27.jsonl — 198 NBME Form 27 questions (41 SOTA-cropped images) questions_form30.jsonl — 200 NBME Form 30 questions (45 SOTA-cropped images) Structure Each record contains: id — unique identifier (e.g. form31_page-0… See the full description on the dataset page: https://huggingface.co/datasets/agentN0/usmle-step1-form31.textn<1K0 likes79 downloads4mo agoHugging Face17AIM-Harvard /gbaker_medqa_usmle_4_options_hf_brand_to_generictabular1K<n<10K0 likes66 downloads2y agoHugging Face18ssswwwxxx /medqa-usmle-4-options MedQA USMLE Four Options (Parquet) This repository provides the English USMLE four-option subset of bigbio/med_qa as one Parquet file: medqa_usmle_4_options.parquet. It contains 12,723 multiple-choice questions. The original data partitions are preserved in the split column: Split Rows train 10,178 dev 1,272 test 1,273 Schema question: exam question text options: mapping of answer labels A–D to option text answer_idx: correct answer label… See the full description on the dataset page: https://huggingface.co/datasets/ssswwwxxx/medqa-usmle-4-options.text10K<n<100K0 likes64 downloads2mo agoHugging Face19dynamoai-ml /MedQA-USMLE-4-MultiTurnRobust MedQA Multi-Turn Robustness Benchmark Paper: Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMsCode: https://github.com/bmanczak/medqa_deep_robustnessVenue: NeurIPS 2025 Workshop - The Second Workshop on GenAI for Health: Potential, Trust, and Policy Compliance 1,050 USMLE questions with adversarial follow-up contexts that test whether medical LLMs maintain correct answers across conversation turns. Why This Dataset Medical LLMs achieve… See the full description on the dataset page: https://huggingface.co/datasets/dynamoai-ml/MedQA-USMLE-4-MultiTurnRobust.textquestion-answering1K<n<10K1 likes57 downloads11mo agoHugging Face20Aurigene-AI /MedQA-USMLE-4-options Mirrored by Aurigene AI Discovery stage: Evidence and literature US Medical Licensing Exam style questions in four-option multiple choice form. Rows: 11,451 (phrases_no_exclude_test.jsonl 1,273, phrases_no_exclude_train.jsonl 10,178) Pairs with Aurigene-AI/BioMistral-7B from our model catalogue. Upstream: GBaker/MedQA-USMLE-4-options - all credit to the original authors and to the researchers who produced the underlying data; the dataset card and licence below are theirs.… See the full description on the dataset page: https://huggingface.co/datasets/Aurigene-AI/MedQA-USMLE-4-options.text10K<n<100K0 likes57 downloads15d agoHugging Face21GBaker /MedQA-USMLE-4-options-hf-MPNet-IR Dataset Card for "MedQA-USMLE-4-options-hf-MPNet-IR" More Information needed text10K<n<100K5 likes56 downloads4y agoHugging Face22LeoZotos /fineweb-edu-usmle FineWeb-Edu USMLE This is a paragraph-level subset of LeoZotos/fineweb-edu-topics ranked by usmle_similarity. The 2.5B configuration is the highest-ranked core. The 5B configuration contains that same core plus the extension; the shared core files are stored only once. Token budgets use allenai/OLMo-2-0425-1B at revision stage1-step1907359-tokens4001B and include one EOS document boundary per paragraph. The paragraph crossing each target is retained, so the actual token count is… See the full description on the dataset page: https://huggingface.co/datasets/LeoZotos/fineweb-edu-usmle.tabulartext-generation10M<n<100M0 likes56 downloads3d agoHugging Face23AIDx /USMLEtext10K<n<100K0 likes49 downloads3y agoHugging Face24maximegmd /MedQA-USMLE-4-options-clean MedQA-USMLE-4-options-clean Dataset Overview MedQA-USMLE-4-options-clean is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. Physicians analyzed the 1373 questions in the original dataset and moved 52 questions that were either malformed or incomplete to another split incomplete. Key Features Relabeled malformed/incorrect questions Dataset Details Size: 1373 Language: English Data Source… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MedQA-USMLE-4-options-clean.textquestion-answering1K<n<10K0 likes43 downloads2y agoHugging Face25Detsutut /MedQA-USMLE-combined-synonym-firsttabular10K<n<100K0 likes43 downloads2y agoHugging Face26agentN0 /usmle-step1-qbank-v3tabular1K<n<10K0 likes43 downloads3mo agoHugging Face27lemon-mint /GBaker-MedQA-USMLE-4-options-Koreantext10K<n<100K0 likes42 downloads2y agoHugging Face28shuyuej /MedQA-USMLE-Benchmark 💻 Dataset Usage Run the following command to load the testing set (1,273 examples): from datasets import load_dataset dataset = load_dataset("shuyuej/MedQA-USMLE-Benchmark", split="test") print(dataset) text1K<n<10K1 likes40 downloads2y agoHugging Face29agentN0 /usmle-qbanktabular1K<n<10K0 likes37 downloads3mo agoHugging Face30tansutt /MedQA-USMLE-4-options-hftextquestion-answering10K<n<100K0 likes35 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.