CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RL-Forgetting-Experiments-3 /mbpp-code-rl MBPP for code RL (deduplicated against MBPP+) MBPP prepared for RLVR training in verl, with two independent hold-outs so both MBPP+ and MBPP's own canonical test split stay reportable after training on this data. split rows contents train 320 MBPP canonical train + validation + prompt, minus everything in MBPP+ test 378 exactly the problems in evalplus/mbppplus heldout_mbpp_test 276 MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.texttext-generationn<1K0 likes121 downloads14d agoHugging Face02fxevangelinenyu /rl-forgetting-math-benchmarks RL-Forgetting math benchmarks Train and test/benchmark sets used in the RL-Forgetting-Exp study of replay-buffer freshness. All parquets share the verl RL schema (data_source, prompt, ability, reward_model, extra_info). Layout polaris_full/ train.parquet # 52,309 prompts (Polaris-full training set) test.parquet # 800 prompts (held-out test, 100/difficulty) deepscaler/ train.parquet # 8,192 prompts (skywork_deepscaler_easy_8192, fixed… See the full description on the dataset page: https://huggingface.co/datasets/fxevangelinenyu/rl-forgetting-math-benchmarks.text10K<n<100K0 likes53 downloads1mo agoHugging Face03RL-Forgetting-Exp-2 /polaris_math_rltext10K<n<100K0 likes38 downloads15d agoHugging Face04leonli66 /rl-forgetting-polaris-8gpu-results Polaris 8-GPU experiment results Four completed training runs and all 122 unique evaluation points (976 shards). Qwen3-8B-Base runs finish at step 2000; Llama-3.2-3B-Instruct runs at step 4000. Each model has a no-replay arm and a hard-cooldown replay arm with lambda 0.1. Evaluation Each point covers all 800 prompts with 160 samples per prompt. The shared base model is step 0; trained checkpoints are evaluated every 100 steps. Sampling uses temperature 0.6, top-p… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/rl-forgetting-polaris-8gpu-results.0 likes33 downloads6d agoHugging Face05RL-Forgetting-Exp /mbpp-code-rl0 likes29 downloads21d agoHugging Face06RL-Forgetting-Exp-2 /qwen3b_code_sft_data_s3000 likes19 downloads2d agoHugging Face07RL-Forgetting-Exp-2 /llama3b_math_sft_data_s20000 likes18 downloads2d agoHugging Face08RL-Forgetting-Exp-2 /qwen3_1p7b_s500_code_sft_data0 likes18 downloads2d agoHugging Face09RL-Forgetting-Experiments-3 /qwen2.5-3b-math-kk-sft-artifacts Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts Training metrics, per-rank training manifests, exact SFT configs, and full persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms. Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF model is complete for inference/evaluation, but its later optimizer/prev-params serialization failed, so no resumable training-state claim is made. See delivery_manifest.json for source lineage and… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts.text-generation0 likes17 downloads2d agoHugging Face10RL-Forgetting-Exp-2 /llama3b_code_sft_data0 likes17 downloads2d agoHugging Face11RL-Forgetting-Exp-2 /qwen3_1p7b_code_sft_data0 likes14 downloads2d agoHugging Face12RL-Forgetting-Exp-2 /polaris_math_eval_600tabularn<1K0 likes12 downloads17d agoHugging Face13RL-Forgetting-Exp-2 /qwen3b_code_sft_data0 likes12 downloads2d agoHugging Face14RL-Forgetting-Exp-2 /qwen3b_math_sft_data_s40000 likes6 downloads10d agoHugging Face15RL-Forgetting-Exp-2 /llama3b_math_sft_data_s26640 likes1 downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.