Richard-Sieg-TH-Koln/anlp-tinystories-dpo-gpt2
TinyStoriesInstruct DPO preference pairs Preference pairs for the Advanced NLP block course at TH Koeln, block 07 (DPO demo). Built by sampling two completions per held-out prompt from the SFT'd model (Richard-Sieg-TH-Koln/anlp-sft-sanity-checkpoint) at temperature 1.0, scoring each by how many of the prompt's required words (the Words: field in TinyStoriesInstruct) it actually contains, and keeping the higher-scoring completion as chosen and the other as rejected. Ties are… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Sieg-TH-Koln/anlp-tinystories-dpo-gpt2.
This repository belongs to Richard-Sieg-TH-Koln on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
