preference-alignment
triple-preference-ultrafeedback-40K
Dataset Card for llama3-ultrafeedback-armorm
This dataset was used to train tpo-alignment/Llama-3-8B-TPO-L-40k, tpo-alignment/Llama-3-8B-TPO-40k, and tpo-alignment/Mistral-7B-TPO-40k.
Dataset Creation
This dataset is built based on the UltraFeedback. We reconstruct UltraFeedback to select three preferences per prompt. First, we rank the responses based on the scores provided in the base dataset. The highest-scoring response is selected as the reference, the… See the full description on the dataset page: https://huggingface.co/datasets/tpo-alignment/triple-preference-ultrafeedback-40K.BeaverTails-single-dimension-preferencepreference_alignment_ultra_cutq-alignment-dynamic-preference-datapreference_alignment_totalq-alignment-preference-data-v5
