RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts
Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts Training metrics, per-rank training manifests, exact SFT configs, and full persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms. Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF model is complete for inference/evaluation, but its later optimizer/prev-params serialization failed, so no resumable training-state claim is made. See delivery_manifest.json for source lineage and… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts.
Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts
Training metrics, per-rank training manifests, exact SFT configs, and full persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms. Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF model is complete for inference/evaluation, but its later optimizer/prev-params serialization failed, so no resumable training-state claim is made. See delivery_manifest.json for source lineage and exact coverage.
Math SFT uses 274,786 successful no-buffer RL rollouts from steps 1-4000; KK uses 812,631 from steps 1-1000. Saturated prompt groups were dropped. Ordered arms preserve RL-step/line order; shuffled arms use global seed 42. Evaluations use Polaris600 at temperature 0.6 for Math and the fixed 700-prompt KK test set at temperature 0.8, with n=160 and k=1..128.
dataset_inputs/ contains the exact prepared SFT JSONL and evaluation parquet files, deduplicated once per domain. training_history/wandb/ contains the four full W&B history exports and their validation/merge manifest. For KK ordered, combine W&B steps 1-3174 with the genuine local metrics row at step 3175.
