datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA-MM
MedQA-MM Identifier Release
Paper repository ·
Hugging Face dataset
MedQA-MM is a 1,000-item shortcut-mitigated medical multimodal multiple-choice benchmark constructed from MedThinkVQA, MedXpertQA-MM, and the Health and Medicine portion of MMMU. This public release is intentionally identifier-only.
It does not contain source questions, answer choices, gold answers, images, clinical text, or repaired payloads. It provides stable source locators, pinned source revisions, and a… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedQA-MM.MedQA-CS-ExamBenchmarking LLMs Clinical Skills for Patient-Centered Diagnostics and Documentation
Project github: https://github.com/bio-nlp/MedQA-CS
MedQA-CS-Student dataset: https://huggingface.co/datasets/bio-nlp-umass/MedQA-CS-Student
combined_bionlp_task_dataset_model_cardseval-gliner2-ner-bionlp2004-boundary-smoothing-validationeval-gliner2-ner-bionlp2004-boundary-smoothing-testMedQA-CS-StudentTask_models_BioNLP_results_GPTrevise_bigbio_task_models_w_BioNLPrevise_bigbio_dataset_models_w_BioNLPeval-gliner2-ner-bionlp2004-affine-validationeval-gliner2-ner-bionlp2004-affine-boundary-smoothing-testeval-gliner2-ner-bionlp2004-affine-boundary-smoothing-validationeval-gliner2-ner-bionlp2004-affine-testrevise_merged_models_w_BioNLPmerged_task_dataset_models_BioNLP_resultsbigbio_dataset_models_BioNLP_results
