CoolFace
Datasetpublic

fxevangelinenyu/rl-forgetting-math-benchmarks

RL-Forgetting math benchmarks Train and test/benchmark sets used in the RL-Forgetting-Exp study of replay-buffer freshness. All parquets share the verl RL schema (data_source, prompt, ability, reward_model, extra_info). Layout polaris_full/ train.parquet # 52,309 prompts (Polaris-full training set) test.parquet # 800 prompts (held-out test, 100/difficulty) deepscaler/ train.parquet # 8,192 prompts (skywork_deepscaler_easy_8192, fixed… See the full description on the dataset page: https://huggingface.co/datasets/fxevangelinenyu/rl-forgetting-math-benchmarks.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes53downloads
2 commits on main
c844b6e1mo ago

Add polaris_full + deepscaler train/test benchmark parquets

fxevangelinenyu
9d3f08a1mo ago

initial commit

fxevangelinenyu