datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rl-math-skyeasy25k-omi2
rl-math (deepscaler-easy + OpenMathInstruct-2)
Easy-biased math RL dataset (verl format; prompts [system, user] with the
OpenThoughts thinking system prompt; rule-based math_verify reward).
train (train_deepscaler10k_diff5_omi2_sysprompt.parquet, 24,362):
10K from Skywork-OR1-RL-Data deepscaler subset filtered to 1.5B difficulty < 5,
merged with 14,862 unique non-augmented gsm8k+math problems from
nvidia/OpenMathInstruct-2. (breakdown: deepscaler 9,802 / OMI2 14,560)
test… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/rl-math-skyeasy25k-omi2.rlm-trajectories-seed
HotCopy RLM Trajectories (Seed)
A 12-row seed corpus of synthetic Recursive Language Model trajectories
emitted by the HotCopy two-tier agentic CLI (the orchestrator root, sub-call workers).
Why this dataset exists
The Recursive Language Model paper (Zhang, Kraska, Khattab — MIT CSAIL, 2026,
arxiv.org/abs/2512.24601) reports that
"Fine-tuning Qwen3-8B on 1,000 RLM trajectories improved performance 28.3%"
— a strong signal that the shape of RLM execution can be taught from… See the full description on the dataset page: https://huggingface.co/datasets/HotCopyAI/rlm-trajectories-seed.
