Siye01/LLM-Fusion-Train
Multi-Domain RLVR Training Data Per-domain reinforcement-learning-with-verifiable-rewards (RLVR) training sets for five domains, plus the mixed-domain blend used as a joint-training baseline. Every subset uses the verl RLHF parquet schema: data_source, prompt, ability, reward_model, extra_info. Subsets Subset Rows Size Math 38,131 11.6 MiB Science 50,000 45.6 MiB Code 19,169 1,432.5 MiB IF 16,575 9.3 MiB Agent 10,229 0.2 MiB Mix 87,699 1,505.1… See the full description on the dataset page: https://huggingface.co/datasets/Siye01/LLM-Fusion-Train.
0146
Update README.md
Upload folder using huggingface_hub
initial commit
