CoolFace
Datasetpublic

salisai/hh-rlhf-helpful-dpo-10k

HH-RLHF Helpful DPO Preference Pairs · 10k 10,000 real human preference pairs for teaching a tiny language model (≤50M params) what a good assistant sounds like — more helpful, more natural, less evasive. Why this dataset exists This is the preference-tuning stage of an end-to-end tiny-model training pipeline: Pretraining ──► SFT ──► DPO (this dataset) ──► Tiny Edge Assistant After SFT teaches the model how to speak, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/salisai/hh-rlhf-helpful-dpo-10k.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes46downloads
6 commits on main
6678ff91mo ago

Update README.md

salisai
93495281mo ago

Upload logo.svg with huggingface_hub

salisai
80911911mo ago

Upload README.md with huggingface_hub

salisai
14a43a91mo ago

Upload train.parquet with huggingface_hub

salisai
2c319461mo ago

Upload train.parquet with huggingface_hub

salisai
ee9f1db1mo ago

initial commit

salisai