salisai/hh-rlhf-helpful-dpo-10k
HH-RLHF Helpful DPO Preference Pairs · 10k 10,000 real human preference pairs for teaching a tiny language model (≤50M params) what a good assistant sounds like — more helpful, more natural, less evasive. Why this dataset exists This is the preference-tuning stage of an end-to-end tiny-model training pipeline: Pretraining ──► SFT ──► DPO (this dataset) ──► Tiny Edge Assistant After SFT teaches the model how to speak, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/salisai/hh-rlhf-helpful-dpo-10k.
Update README.md
Upload logo.svg with huggingface_hub
Upload README.md with huggingface_hub
Upload train.parquet with huggingface_hub
Upload train.parquet with huggingface_hub
initial commit
