datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbpp-code-rl
MBPP for code RL (deduplicated against MBPP+)
MBPP prepared for RLVR training in verl,
with two independent hold-outs so both MBPP+ and MBPP's own canonical test
split stay reportable after training on this data.
split
rows
contents
train
320
MBPP canonical train + validation + prompt, minus everything in MBPP+
test
378
exactly the problems in evalplus/mbppplus
heldout_mbpp_test
276
MBPP's canonical test split (task_id 11-510) that is not in MBPP+… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/mbpp-code-rl.qwen2.5-3b-math-kk-sft-artifacts
Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts
Training metrics, per-rank training manifests, exact SFT configs, and full
persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms.
Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF
model is complete for inference/evaluation, but its later optimizer/prev-params
serialization failed, so no resumable training-state claim is made. See
delivery_manifest.json for source lineage and… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts.
