onepo
Datasets
All datasets matching “onepo”standard_group1jun_27_1000_g5_200grading_group1skill_use_eval_group2skill_using_eval_dataset
rubrics/<skill_name>/: three LLM-as-judge prompts + one .json detail
run_env/<skill_name>/: .claude/ (skill) + all the environment files
user_query/<skill_name>/: an user prompt.
OnePO-Medical-20K
OnePO-Medical-20K
📄 Paper |
💻 GitHub
⚡ Introduction
OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages.
One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage.
Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics.
Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.
