CoolFace
Datasetpublic

Xiaofeng77/answer-only-gp-l-only-10k

Debunk the Myth of SFT Generalization Dataset This dataset is associated with the paper "Debunk the Myth of SFT Generalization". The paper challenges the prevailing view that supervised fine-tuning (SFT) primarily memorizes training data and fails to generalize, in contrast to reinforcement learning (RL). It demonstrates that SFT can generalize as well as—or better than—RL when trained with appropriate data, achieved through prompt diversity and Chain-of-Thought (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/answer-only-gp-l-only-10k.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes52downloads
settings

This repository belongs to Xiaofeng77 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameanswer-only-gp-l-only-10k
visibilitypublic
licencenot set
gatedno
ownerXiaofeng77
Account settings
Xiaofeng77/answer-only-gp-l-only-10k · CoolFace