CoolFace
Datasetpublic

dusersad12/verl-deepscaler-curated

verl DeepScaleR Curated A cleaned, de-duplicated and evaluation-safe training split derived from the DeepScaleR-Preview-Dataset, reformatted for rule-based-reward RL post-training with verl. Total examples: 38,783 (from 41,705 raw records read across three source batches). Row format Each row follows the verl dataset_row template: field value data_source "DeepScaleR" prompt [{"role": "user", "content": <problem text>}] ability "math" reward_model… See the full description on the dataset page: https://huggingface.co/datasets/dusersad12/verl-deepscaler-curated.

sourceHugging Facemitupdated 8d agoView on Hugging Face
0likes59downloads

dusersad12/verl-deepscaler-curated · main · files are served by the source, never re-hosted here