CoolFace
Datasetpublic

RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts

Qwen2.5-3B Math and Knights-and-Knaves SFT artifacts Training metrics, per-rank training manifests, exact SFT configs, and full persisted evaluation outputs for the ordered/shuffled Math and KK SFT arms. Each arm has ten checkpoint evaluations at n=160. The KK ordered step-3175 HF model is complete for inference/evaluation, but its later optimizer/prev-params serialization failed, so no resumable training-state claim is made. See delivery_manifest.json for source lineage and… See the full description on the dataset page: https://huggingface.co/datasets/RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts.

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes43downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
RL-Forgetting-Experiments-3/qwen2.5-3b-math-kk-sft-artifacts · CoolFace