datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grading_group1OnePO-Medical-20K
OnePO-Medical-20K
📄 Paper |
💻 GitHub
⚡ Introduction
OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages.
One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage.
Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics.
Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.tmp_2nd_50_samplesgraspbench_masked_criteriatmp_crg_samplesgrading_group2task_labelsgraspbench_oh3_common_841_tasksoneportfit_sweeps_december2025
