datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval-imdb-drpo-134567-dpo-5000eval-tldr-dpo-drpo-0.9tmp-sft-1000train_data_tldr_for_drpo
TL;DR Dataset for DRPO
Summary
The TL;DR dataset is a processed version of Reddit posts
Data Structure
Columns:
"prompt": The unabridged Reddit post.
"a1": A summary of the post.
"a2": An alternative summary of the post.
"rank": The rank of the summary, where 1 indicates the first summary is preferred and 0 indicates the second summary is preferred.
This structure enables models to learn the relationship between detailed content and its abbreviated form… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_tldr_for_drpo.narrative-arc-tp-drpotrain_data_hh_for_drpo
HH-RLHF-Helpful-Base Dataset
Summary
The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_hh_for_drpo.eval-tldr-dpo-ppo-drpo-dm-sft-1000-cut2eval-tldr-dpo-ppo-drpo-dm-sft-1000kp_cfr_drpo_1200_v2semiautomatic-aestheticsdrpo_hh_qwen2.5_1.5b_with_ref_btprefkp_cfr_drpo_12000eval-imdb-drpo-3-drpo-4-1000tldr_test_tiny_data_drpo
TL;DR Dataset
Summary
The TL;DR dataset is a processed version of Reddit posts, specifically curated to train models using the TRL library for summarization tasks. It leverages the common practice on Reddit where users append "TL;DR" (Too Long; Didn't Read) summaries to lengthy posts, providing a rich source of paired text data for training summarization models.
Data Structure
Format: Conversational
Type: Preference
Columns:
"prompt": The user query.… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/tldr_test_tiny_data_drpo.drpo_ultrafeedback_qwen2.5-1.5b_first_iter_20kdrpo_hh_qwen2.5_1.5b_with_ref_prob_vllm_conveval-imdb-drpo-1-3-4-dpo-1000eval-tldr-dpo-drpo-0.75tmp-sft-ppo-1000eval-tldr-dpo-ppo-drpo-dm-sft-1000-cutdrpo_hh_qwen2.5_1.5b_with_ref_prob_sampledkp_cfr_drpo_12000_nonadversarialdrpo_hh_qwen2.5_1.5bbankingDRPO_data_from_ultrafeed_new_templatehh_helpfulness_drpo_from_sftDRPO_data_from_ultrafeeddrpo_ultrafeedback_qwen2.5-1.5b-1eval-imdb-drpo-sft-dpo-5000DRPO_first_iterdrpo_ultrafeedback_qwen2.5-1.5b-2drpo_hh_qwen2.5_1.5b_with_ref_prob_vllm
