datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
standard_group1jun_27_1000_g5_200grading_group1skill_use_eval_group2skill_using_eval_dataset
rubrics/<skill_name>/: three LLM-as-judge prompts + one .json detail
run_env/<skill_name>/: .claude/ (skill) + all the environment files
user_query/<skill_name>/: an user prompt.
OnePO-Medical-20K
OnePO-Medical-20K
📄 Paper |
💻 GitHub
⚡ Introduction
OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages.
One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage.
Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics.
Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.opus-4_8-openhands-thoughtful-scored-v1humanities_runs_1kimi-2_6-openhands-rollout-0803opus-4_8-claude_code-rollout-xhg_0805opus-4_8-claude_code-rollout-xhg_0804glm-5.2-openhands-rollout-0806g5_200_tsk
jun_27_1000_g5_200
Lossless, restorable archives for the jun_27_1000_g5_200 batch.
Contents
jun_27_1000_g5_200__pure_harbor_tasks.tar.zst: 183 rollout-ready Harbor tasks and their complete raw skill inputs. Pipeline intermediates, rollout/judge outputs, and .setup_synthesis directories are excluded. dhh-coder is excluded because its setup requires human classification.
jun_27_1000_g5_200__rubrics.tar.zst: 330 validated final rubric JSON files (166 completion and… See the full description on the dataset page: https://huggingface.co/datasets/OnepointfiveHz/g5_200_tsk.skill_use_eval_hardsonnet-openhands-0815glm-5.2-openhands-rollout-0805g8_100_tsk
jun_27_1000_g8_100
Lossless, restorable payload archives for the selected task data in the jun_27_1000_g8_100 batch.
Contents
jun_27_1000_g8_100__pure_harbor_tasks.tar.zst: 92 rollout-ready Harbor tasks and their complete raw skill inputs. Pipeline intermediates, rollout/judge outputs, .setup_synthesis, nested .git metadata, and macOS Finder/AppleDouble files are excluded.
jun_27_1000_g8_100__rubrics.tar.zst: 176 validated final rubric JSON files (88 completion… See the full description on the dataset page: https://huggingface.co/datasets/OnepointfiveHz/g8_100_tsk.all-model-failed-rollout-trajectoryqwen-3_5-OH-rollout-1_7-28opus-4_8-openhands-rollout-xhg_0807minimax-openhands-rollout-0818gemini-3.6-flash-openhands-rollout-0820tmp_2nd_50_samplesgpt-oss-rolloutgpt-oss-rollout-3_7-28qwen-3_5-rollout-3_7-28glm-5.2-openhands-rollout-0801_07opus-4_8-openhands-rollout-xhg_0806haiku-openhands-rollout-0816qwen_3_5_rollout
