CoolFace
Datasetpublic

Siye01/LLM-Fusion-Train

Multi-Domain RLVR Training Data Per-domain reinforcement-learning-with-verifiable-rewards (RLVR) training sets for five domains, plus the mixed-domain blend used as a joint-training baseline. Every subset uses the verl RLHF parquet schema: data_source, prompt, ability, reward_model, extra_info. Subsets Subset Rows Size Math 38,131 11.6 MiB Science 50,000 45.6 MiB Code 19,169 1,432.5 MiB IF 16,575 9.3 MiB Agent 10,229 0.2 MiB Mix 87,699 1,505.1… See the full description on the dataset page: https://huggingface.co/datasets/Siye01/LLM-Fusion-Train.

sourceHugging Faceapache-2.0updated 27d agoView on Hugging Face
0likes146downloads
3 commits on main
7250d3227d ago

Update README.md

Siye01
0a4879827d ago

Upload folder using huggingface_hub

Siye01
60fbf5827d ago

initial commit

Siye01