CoolFace
Datasetpublic

pre-to-post-olmo/rl-math-skyeasy25k-omi2

rl-math (deepscaler-easy + OpenMathInstruct-2) Easy-biased math RL dataset (verl format; prompts [system, user] with the OpenThoughts thinking system prompt; rule-based math_verify reward). train (train_deepscaler10k_diff5_omi2_sysprompt.parquet, 24,362): 10K from Skywork-OR1-RL-Data deepscaler subset filtered to 1.5B difficulty < 5, merged with 14,862 unique non-augmented gsm8k+math problems from nvidia/OpenMathInstruct-2. (breakdown: deepscaler 9,802 / OMI2 14,560) test… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/rl-math-skyeasy25k-omi2.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes44downloads
Dataset Card

rl-math (deepscaler-easy + OpenMathInstruct-2)

Easy-biased math RL dataset (verl format; prompts [system, user] with the OpenThoughts thinking system prompt; rule-based math_verify reward).

  • —train (train_deepscaler10k_diff5_omi2_sysprompt.parquet, 24,362): 10K from Skywork-OR1-RL-Data deepscaler subset filtered to 1.5B difficulty < 5, merged with 14,862 unique non-augmented gsm8k+math problems from nvidia/OpenMathInstruct-2. (breakdown: deepscaler 9,802 / OMI2 14,560)
  • —test (test_indist_500_deepscaler_sysprompt.parquet, 500): in-distribution held-out split, disjoint from train.