datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reward-modeling-short-tokenized
Dataset Card for "reward-modeling-short-tokenized"
More Information needed
reward-modeling-long-tokenized
Dataset Card for "reward-modeling-long-tokenized"
More Information needed
reward_modeling_dataset
Dataset Card for "reward_modeling_dataset"
More Information needed
pedagogical-rewardmodel-datareward_model_embeddingsreward_model_biases_attack_promptsreward_model_anthropic_88
Dataset Card for "reward_model_anthropic_88"
More Information needed
Scigraph2_BT_RewardModelingDataset_removing15kfiltered_overlap_devreward_model_anthropic
Dataset Card for "reward_model_anthropic"
More Information needed
reward_model_anthropic_8
Dataset Card for "reward_model_anthropic_8"
More Information needed
RewardModel_8RewardModel_9RewardModel_3RewardModel_training_dataRewardModel_4reward-model-data-falcon
Guidelines
These guidelines are based on the paper Training Language Models to Follow Instructions with Human Feedback
You are given a text-based description of a task, submitted by a user.
This task description may be in the form of an explicit instruction (e.g. "Write a story about a wise frog."). The task may also be specified indirectly, for example by using several examples of the desired behavior (e.g. given a sequence of movie reviews followed by their sentiment, followed by… See the full description on the dataset page: https://huggingface.co/datasets/argilla/reward-model-data-falcon.RewardModel_2RewardModel-DR-HH-Seed1reward-modeling-eval-tokenized
Dataset Card for "reward-modeling-eval-tokenized"
More Information needed
reward-models-biases-docsRewardModel_7reward_model_biasesreward_model_datareward-model-no-topic-predictions
Dataset Card for "reward-model-no-topic-predictions"
More Information needed
RewardModel_1RewardModel_5reward-modelRewardModel_6reward_model_ranking_dataset_RLHF
Dataset Card for "reward_model_ranking_dataset_RLHF"
More Information needed
reward_model_dataset
Reward Model Dataset
We only store the rewards here, rather than have the full traces. The rewards list per sample is mapped to 1 of three classes:
-1 -> bad score on the question (False or C/D/E on subjective)
0 -> NA, meaning the question was not relevant
1 -> good score on the question (True, or A/B on subjective)
Below is a table of the unique questions encountered, with their corresponding indices, dimensions, and principles.
question_idx
dimension
dimension_idx
principle… See the full description on the dataset page: https://huggingface.co/datasets/jdineen/reward_model_dataset.
