CoolFace
Datasetpublic

wflying/instruction-following-rl-66k

Instruction Following RL 66K Dataset overview instruction-following-rl-66k is an English training dataset for instruction-following reinforcement learning (RL/RLVR), containing 66,418 examples. It is derived primarily from AllenAI's IF_multi_constraints_upto5, whose instructions contain up to five verifiable constraints drawn from IFEval and IFBench-Train. Each record is first validated for its JSON, prompt, and metadata structure. A predefined… See the full description on the dataset page: https://huggingface.co/datasets/wflying/instruction-following-rl-66k.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
0likes104downloads

wflying/instruction-following-rl-66k · main · files are served by the source, never re-hosted here