Siye01/LLM-Fusion-Train
Multi-Domain RLVR Training Data Per-domain reinforcement-learning-with-verifiable-rewards (RLVR) training sets for five domains, plus the mixed-domain blend used as a joint-training baseline. Every subset uses the verl RLHF parquet schema: data_source, prompt, ability, reward_model, extra_info. Subsets Subset Rows Size Math 38,131 11.6 MiB Science 50,000 45.6 MiB Code 19,169 1,432.5 MiB IF 16,575 9.3 MiB Agent 10,229 0.2 MiB Mix 87,699 1,505.1… See the full description on the dataset page: https://huggingface.co/datasets/Siye01/LLM-Fusion-Train.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face