datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQA-USMLE
MedQA-USMLE
HuggingFace upload of the MedQA-USMLE dataset with deduping. If used, please cite the original authors using the citation below.
A small number of exact-duplicate questions were identified within train and us_qbank. The question text was identical, but the options were formatted slightly differently or had a different distractor. The main difference was the listed correct letter, so the incorrect duplicates were removed. Each split was then reindexed to keep indices… See the full description on the dataset page: https://huggingface.co/datasets/mkieffer/MedQA-USMLE.gbaker_medqa_usmle_4_options_hf_generic_to_brandgbaker_medqa_usmle_4_options_hf_originalgbaker_medqa_usmle_4_options_hf_brand_to_genericfineweb-edu-usmle
FineWeb-Edu USMLE
This is a paragraph-level subset of LeoZotos/fineweb-edu-topics ranked by
usmle_similarity. The 2.5B configuration is the highest-ranked
core. The 5B configuration contains that same core plus the extension; the
shared core files are stored only once.
Token budgets use allenai/OLMo-2-0425-1B at revision
stage1-step1907359-tokens4001B and include one EOS document boundary per
paragraph. The paragraph crossing each target is retained, so the actual token
count is… See the full description on the dataset page: https://huggingface.co/datasets/LeoZotos/fineweb-edu-usmle.MedQA-USMLE-combined-synonym-firstusmle-step1-qbank-v3usmle-qbankMedQA-USMLE-synonym-replacementMedQA dataset perturbed using knowledge-based synonym replacement technique with BAT
MedQA-USMLE-back-translatedMedQA dataset perturbed using back-translation technique with BAT
MedQA-USMLE-combinedgbaker_medqa_usmle_4_optionsMedQA-USMLE
MedQA-USMLE
HuggingFace upload of the MedQA-USMLE dataset with deduping. If used, please cite the original authors using the citation below.
A small number of exact-duplicate questions were identified within train and us_qbank. The question text was identical, but the options were formatted slightly differently or had a different distractor. The main difference was the listed correct letter, so the incorrect duplicates were removed. Each split was then reindexed to keep indices… See the full description on the dataset page: https://huggingface.co/datasets/Muraleedharan/MedQA-USMLE.usmle-step1-qbankusmle-step1-qbank-v2medqa-usmle-fhir-produsmle_extendedLLMusmle_diffcultyclass
Dataset Card for "usmle_diffcultyclass"
More Information needed
