CoolFace
Datasetpublic

wflying/instruction-following-rl-content-constrained-30k

Instruction-Following RL Content-Constrained 30K Dataset summary Instruction-Following RL Content-Constrained 30K is a training dataset for precise instruction following and reinforcement learning from verifiable rewards (RLVR). The current cleaned revision contains 29,520 heterogeneous, single-turn user prompts. Each prompt combines a substantive task with one to five explicit output constraints, such as keyword inclusion or exclusion, response length… See the full description on the dataset page: https://huggingface.co/datasets/wflying/instruction-following-rl-content-constrained-30k.

sourceHugging Faceodc-byupdated 2mo agoView on Hugging Face
0likes60downloads
7 commits on main
d7f24432mo ago

Update data_type and simplify dataset card

wflying
d37c0e22mo ago

Rename data_type to if_content_constrained

wflying
7a7a2802mo ago

Update card after filtering unsatisfiable records

wflying
86ca28c2mo ago

Remove 480 unsatisfiable paragraph-constraint records

wflying
53f639a2mo ago

Add detailed dataset card

wflying
37f79f72mo ago

Upload content-constrained instruction-following RL dataset

wflying
a8c18eb2mo ago

initial commit

wflying