CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /helpful-instructions Dataset Card for Helpful Instructions Dataset Summary Helpful Instructions is a dataset of (instruction, demonstration) pairs that are derived from public datasets. As the name suggests, it focuses on instructions that are "helpful", i.e. the kind of questions or tasks a human user might instruct an AI assistant to perform. You can load the dataset as follows: from datasets import load_dataset # Load all subsets helpful_instructions =… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/helpful-instructions.text100K<n<1M24 likes3.4k downloads4y agoHugging Face02RLHFlow /HH-RLHF-Helpful-standardWe process the helpful subset of Anthropic-HH into the standard format. The filtering script is as follows. def filter_example(example): if len(example['chosen']) != len(example['rejected']): return False if len(example['chosen']) % 2 != 0: return False n_rounds = len(example['chosen']) for i in range(len(example['chosen'])): if example['chosen'][i]['role'] != ['user', 'assistant'][i % 2]: return False if… See the full description on the dataset page: https://huggingface.co/datasets/RLHFlow/HH-RLHF-Helpful-standard.text100K<n<1M4 likes560 downloads2y agoHugging Face03trl-internal-testing /hh-rlhf-helpful-base-trl-style TRL's Anthropic HH Dataset We preprocess the dataset using our standard prompt, chosen, rejected format. Reproduce this dataset Download the anthropic_hh.py from the https://huggingface.co/datasets/trl-internal-testing/hh-rlhf-helpful-base-trl-style/tree/0.1.0. Run python examples/datasets/anthropic_hh.py --push_to_hub --hf_entity trl-internal-testing text10K<n<100K14 likes533 downloads2y agoHugging Face04trl-lib /hh-rlhf-helpful-base HH-RLHF-Helpful-Base Dataset Summary The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/hh-rlhf-helpful-base.text10K<n<100K3 likes239 downloads2y agoHugging Face05HuggingFaceH4 /helpful_instructions_splitsThis splits the original helpful_instructions dataset into train and test splits. text10K<n<100K3 likes140 downloads4y agoHugging Face06bcui19 /chat-v2-anthropic-helpfulnesstext100K<n<1M1 likes107 downloads3y agoHugging Face07HuggingFaceH4 /helpful-anthropic-raw Dataset Card for "helpful-raw-anthropic" This is a dataset derived from Anthropic's HH-RLHF data of instructions and model-generated demonstrations. We combined training splits from the following two subsets: helpful-base helpful-online To convert the multi-turn dialogues into (instruction, demonstration) pairs, just the first response from the Assistant was included. This heuristic captures the most obvious answers, but overlooks more complex questions where multiple turns were… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/helpful-anthropic-raw.text10K<n<100K7 likes94 downloads4y agoHugging Face08Kyleyee /train_data_SFT_Helpful HH-RLHF-Helpful-Base Dataset Summary The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_SFT_Helpful.text10K<n<100K0 likes91 downloads2y agoHugging Face09pvduy /rm_hh_helpful_only Dataset Card for "rm_hh_helpful_only" More Information needed text100K<n<1M0 likes79 downloads3y agoHugging Face10Baidicoot /anthropic-helpful-harmless-rlhftext100K<n<1M0 likes79 downloads2y agoHugging Face11qgallouedec /hh-rlhf-helpful-base-trl-style TRL's Anthropic HH Dataset We preprocess the dataset using our standard prompt, chosen, rejected format. Reproduce this dataset Download the anthropic_hh.py from the https://huggingface.co/datasets/qgallouedec/hh-rlhf-helpful-base-trl-style/tree/0.1.0. Run python examples/datasets/anthropic_hh.py --push_to_hub --hf_entity qgallouedec text10K<n<100K0 likes79 downloads2y agoHugging Face12Dahoas /rm_instruct_helpful_preferences Dataset Card for "rm_instruct_helpful_preferences" More Information needed text10K<n<100K5 likes71 downloads4y agoHugging Face13HuggingFaceH4 /h4-anthropic-hh-rlhf-helpful-base-gentext10K<n<100K5 likes70 downloads2y agoHugging Face14cybershiptrooper /backdoored_helpful_only_completions_probe_type_linear_threshold_0_4text10K<n<100K0 likes67 downloads1y agoHugging Face15cybershiptrooper /backdoored_helpful_only_completions_probe_type_linear_threshold_0_65text10K<n<100K0 likes63 downloads1y agoHugging Face16MWilinski /hh-rlhf-helpful-base-rollouts-gpt-oss-20b-diverse-openroutertextn<1K0 likes63 downloads6mo agoHugging Face17trl-lib /ultrafeedback-gpt-3.5-turbo-helpfulness UltraFeedback GPT-3.5-Turbo Helpfulness Dataset Summary The UltraFeedback GPT-3.5-Turbo Helpfulness dataset contains processed user-assistant interactions filtered for helpfulness, derived from the openbmb/UltraFeedback dataset. It is designed for fine-tuning and evaluating models in alignment tasks. Data Structure Format: Conversational Type: Unpaired preference Column: "pompt": The input question or instruction provided to the model. "completion": The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness.text10K<n<100K4 likes62 downloads2y agoHugging Face18Ray2333 /RiC_harmless_helpfulThe hhrlhf dataset for RiC (https://huggingface.co/papers/2402.10207) training with harmless (R1) and helpful (R2) rewards. The 'input_ids' are obtained from Llama2 tokenizer. If you want to use other base models, replace it using other tokenizers. Note: the rewards are already normalized accroding to their corresponding mean and std. The mean and std data for R1 and R2 are saved into all_reward_stat_harmhelp_Rlarge.npy. The mean and std for R1 and R2 is (-0.94732502, 1.92034349)… See the full description on the dataset page: https://huggingface.co/datasets/Ray2333/RiC_harmless_helpful.tabular100K<n<1M0 likes61 downloads2y agoHugging Face19Kyleyee /train_data_Helpful_implicit_prompt HH-RLHF-Helpful-Base Dataset Summary The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_Helpful_implicit_prompt.text10K<n<100K0 likes57 downloads2y agoHugging Face20cybershiptrooper /backdoored_helpful_only_completions_probe_type_linear_threshold_0_45text10K<n<100K0 likes54 downloads1y agoHugging Face21thobauma /Anthropic-helpful-basetext10K<n<100K0 likes53 downloads2y agoHugging Face22Kyleyee /train_data_Helpful_drdpo_7b_sft_1e HH-RLHF-Helpful-Base Dataset Summary The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_Helpful_drdpo_7b_sft_1e.text10K<n<100K0 likes52 downloads1y agoHugging Face23lewtun /helpful-anthropic-raw Dataset Card for "helpful-anthropic-raw" More Information needed text10K<n<100K0 likes45 downloads4y agoHugging Face24cybershiptrooper /backdoored_helpful_only_completions_probe_type_linear_threshold_0_7text10K<n<100K0 likes45 downloads1y agoHugging Face25salisai /hh-rlhf-helpful-dpo-10k HH-RLHF Helpful DPO Preference Pairs · 10k 10,000 real human preference pairs for teaching a tiny language model (≤50M params) what a good assistant sounds like — more helpful, more natural, less evasive. Why this dataset exists This is the preference-tuning stage of an end-to-end tiny-model training pipeline: Pretraining ──► SFT ──► DPO (this dataset) ──► Tiny Edge Assistant After SFT teaches the model how to speak, this dataset… See the full description on the dataset page: https://huggingface.co/datasets/salisai/hh-rlhf-helpful-dpo-10k.texttext-generation10K<n<100K0 likes44 downloads1mo agoHugging Face26HuggingFaceH4 /helpful-self-instruct-raw Dataset Card for "helpful-self-instruct-raw" This dataset is derived from the finetuning subset of Self-Instruct, with some light formatting to remove trailing spaces and <|endoftext|> tokens. text10K<n<100K2 likes42 downloads4y agoHugging Face27tzwilliam0 /Safe_dpo_helpfultext10K<n<100K0 likes41 downloads1y agoHugging Face28Kyleyee /train_data_Helpful_drdpo_7b HH-RLHF-Helpful-Base Dataset Summary The HH-RLHF-Helpful-Base dataset is a processed version of Anthropic's HH-RLHF dataset, specifically curated to train models using the TRL library for preference learning and alignment tasks. It contains pairs of text samples, each labeled as either "chosen" or "rejected," based on human preferences regarding the helpfulness of the responses. This dataset enables models to learn human preferences in generating helpful responses… See the full description on the dataset page: https://huggingface.co/datasets/Kyleyee/train_data_Helpful_drdpo_7b.text10K<n<100K0 likes40 downloads1y agoHugging Face29august66 /ultrafeedback_helpful_basetext10K<n<100K0 likes34 downloads7mo agoHugging Face30AlignmentResearch /Helpfultext10K<n<100K0 likes33 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.