CoolFace
Datasetpublic

wflying/instruction-following-rl-66k

Instruction Following RL 66K Dataset overview instruction-following-rl-66k is an English training dataset for instruction-following reinforcement learning (RL/RLVR), containing 66,418 examples. It is derived primarily from AllenAI's IF_multi_constraints_upto5, whose instructions contain up to five verifiable constraints drawn from IFEval and IFBench-Train. Each record is first validated for its JSON, prompt, and metadata structure. A predefined… See the full description on the dataset page: https://huggingface.co/datasets/wflying/instruction-following-rl-66k.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
0likes102downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
wflying/instruction-following-rl-66k · CoolFace