CoolFace
Datasetpublic

Richard-Sieg-TH-Koln/anlp-tinystories-dpo-gpt2

TinyStoriesInstruct DPO preference pairs Preference pairs for the Advanced NLP block course at TH Koeln, block 07 (DPO demo). Built by sampling two completions per held-out prompt from the SFT'd model (Richard-Sieg-TH-Koln/anlp-sft-sanity-checkpoint) at temperature 1.0, scoring each by how many of the prompt's required words (the Words: field in TinyStoriesInstruct) it actually contains, and keeping the higher-scoring completion as chosen and the other as rejected. Ties are… See the full description on the dataset page: https://huggingface.co/datasets/Richard-Sieg-TH-Koln/anlp-tinystories-dpo-gpt2.

sourceHugging Facecdla-sharing-1.0updated 2d agoView on Hugging Face
0likes64downloads
settings

This repository belongs to Richard-Sieg-TH-Koln on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameanlp-tinystories-dpo-gpt2
visibilitypublic
licencecdla-sharing-1.0
gatedno
ownerRichard-Sieg-TH-Koln
Account settings
Richard-Sieg-TH-Koln/anlp-tinystories-dpo-gpt2 · CoolFace