CoolFace
Datasetpublic

Kyleyee/train_data_Helpful_drdpo_7b

HH-RLHF-Helpful-Base Dataset Summary The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_Helpful_drdpo_7b.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes40downloads
settings

This repository belongs to Kyleyee on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nametrain_data_Helpful_drdpo_7b
visibilitypublic
licencenot set
gatedno
ownerKyleyee
Account settings
Kyleyee/train_data_Helpful_drdpo_7b · CoolFace