CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01andersonbcdefg /reward-modeling-short-tokenized Dataset Card for "reward-modeling-short-tokenized" More Information needed 100K<n<1M2 likes337 downloads3y agoHugging Face02andersonbcdefg /reward-modeling-long-tokenized Dataset Card for "reward-modeling-long-tokenized" More Information needed 100K<n<1M1 likes147 downloads3y agoHugging Face03HumanDynamics /reward_modeling_dataset Dataset Card for "reward_modeling_dataset" More Information needed text10K<n<100K2 likes77 downloads3y agoHugging Face04dmacjam /pedagogical-rewardmodel-datatext10K<n<100K1 likes77 downloads10mo agoHugging Face05ayganyavuz /reward_model_embeddingstext10K<n<100K0 likes60 downloads4mo agoHugging Face06abhayesian /reward_model_biases_attack_promptstabular1K<n<10K0 likes51 downloads1y agoHugging Face07Deojoandco /reward_model_anthropic_88 Dataset Card for "reward_model_anthropic_88" More Information needed tabular1K<n<10K1 likes50 downloads4y agoHugging Face08Vidushee /Scigraph2_BT_RewardModelingDataset_removing15kfiltered_overlap_devtext100K<n<1M0 likes47 downloads7mo agoHugging Face09Deojoandco /reward_model_anthropic Dataset Card for "reward_model_anthropic" More Information needed tabular100K<n<1M0 likes44 downloads4y agoHugging Face10Deojoandco /reward_model_anthropic_8 Dataset Card for "reward_model_anthropic_8" More Information needed tabular1K<n<10K1 likes43 downloads4y agoHugging Face11tranthanhnguyenai1 /RewardModel_8text100K<n<1M0 likes32 downloads1y agoHugging Face12tranthanhnguyenai1 /RewardModel_9text100K<n<1M0 likes28 downloads1y agoHugging Face13tranthanhnguyenai1 /RewardModel_3text100K<n<1M0 likes27 downloads1y agoHugging Face14HFXM /RewardModel_training_datatext10K<n<100K0 likes26 downloads2y agoHugging Face15tranthanhnguyenai1 /RewardModel_4text100K<n<1M0 likes25 downloads1y agoHugging Face16argilla /reward-model-data-falcon Guidelines These guidelines are based on the paper Training Language Models to Follow Instructions with Human Feedback You are given a text-based description of a task, submitted by a user. This task description may be in the form of an explicit instruction (e.g. "Write a story about a wise frog."). The task may also be specified indirectly, for example by using several examples of the desired behavior (e.g. given a sequence of movie reviews followed by their sentiment, followed by… See the full description on the dataset page: https://huggingface.co/datasets/argilla/reward-model-data-falcon.text1K<n<10K1 likes22 downloads3y agoHugging Face17tranthanhnguyenai1 /RewardModel_2text100K<n<1M0 likes19 downloads1y agoHugging Face18mamba413 /RewardModel-DR-HH-Seed1tabularn<1K0 likes18 downloads2y agoHugging Face19andersonbcdefg /reward-modeling-eval-tokenized Dataset Card for "reward-modeling-eval-tokenized" More Information needed 10K<n<100K2 likes16 downloads3y agoHugging Face20abhayesian /reward-models-biases-docstext10K<n<100K0 likes15 downloads1y agoHugging Face21tranthanhnguyenai1 /RewardModel_7text100K<n<1M0 likes14 downloads1y agoHugging Face22abhayesian /reward_model_biasestabular10K<n<100K0 likes13 downloads1y agoHugging Face23davide221 /reward_model_datatabularn<1K0 likes11 downloads3y agoHugging Face24AlekseyKorshuk /reward-model-no-topic-predictions Dataset Card for "reward-model-no-topic-predictions" More Information needed tabular1K<n<10K0 likes10 downloads3y agoHugging Face25tranthanhnguyenai1 /RewardModel_1text100K<n<1M0 likes10 downloads1y agoHugging Face26tranthanhnguyenai1 /RewardModel_5text100K<n<1M0 likes10 downloads1y agoHugging Face27elmanz /reward-modeltext10K<n<100K0 likes9 downloads2y agoHugging Face28tranthanhnguyenai1 /RewardModel_6text100K<n<1M0 likes9 downloads1y agoHugging Face29sayan1101 /reward_model_ranking_dataset_RLHF Dataset Card for "reward_model_ranking_dataset_RLHF" More Information needed text1K<n<10K0 likes8 downloads3y agoHugging Face30jdineen /reward_model_dataset Reward Model Dataset We only store the rewards here, rather than have the full traces. The rewards list per sample is mapped to 1 of three classes: -1 -> bad score on the question (False or C/D/E on subjective) 0 -> NA, meaning the question was not relevant 1 -> good score on the question (True, or A/B on subjective) Below is a table of the unique questions encountered, with their corresponding indices, dimensions, and principles. question_idx dimension dimension_idx principle… See the full description on the dataset page: https://huggingface.co/datasets/jdineen/reward_model_dataset.text10K<n<100K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.