CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01andersonbcdefg /red_teaming_reward_modeling_pairwise Dataset Card for "red_teaming_reward_modeling_pairwise" More Information needed text10K<n<100K7 likes183 downloads3y agoHugging Face02andersonbcdefg /red_teaming_reward_modeling_pairwise_no_as_an_ai Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai" More Information needed text10K<n<100K6 likes131 downloads3y agoHugging Face03andersonbcdefg /sharegpt_reward_modeling_pairwise_no_as_an_ai Dataset Card for "sharegpt_reward_modeling_pairwise_no_as_an_ai" More Information needed text10K<n<100K3 likes97 downloads3y agoHugging Face04leffff /Diffusion-Reward-Modeling-for-Text-Rendering-Dataset 🖼️ Text-to-Image Rendering Dataset A dataset of 14k text prompts for image generation with text rendering evaluation 📚 Dataset Overview This dataset contains 14,000 text prompts specifically designed for: Image generation with text rendering Evaluating text preservation in generated images Training diffusion models for better text rendering Each prompt comes with: Pre-extracted target text for rendering 5 Stable Diffusion 3 generated latents (70k total) Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.tabulartext-to-image10K<n<100K7 likes80 downloads1y agoHugging Face05HumanDynamics /reward_modeling_dataset Dataset Card for "reward_modeling_dataset" More Information needed text10K<n<100K2 likes77 downloads3y agoHugging Face06dmacjam /pedagogical-rewardmodel-datatext10K<n<100K1 likes77 downloads10mo agoHugging Face07andersonbcdefg /gpteacher_reward_modeling_pairwise Dataset Card for "gpteacher_reward_modeling_pairwise" More Information needed text1K<n<10K2 likes66 downloads3y agoHugging Face08ayganyavuz /reward_model_embeddingstext10K<n<100K0 likes60 downloads4mo agoHugging Face09andersonbcdefg /sharegpt_reward_modeling_pairwise Dataset Card for "sharegpt_reward_modeling_pairwise" More Information needed text10K<n<100K1 likes53 downloads3y agoHugging Face10abhayesian /reward_model_biases_attack_promptstabular1K<n<10K0 likes51 downloads1y agoHugging Face11Deojoandco /reward_model_anthropic_88 Dataset Card for "reward_model_anthropic_88" More Information needed tabular1K<n<10K1 likes50 downloads3y agoHugging Face12athirorg /USS-reward-model-qwen-FT-v2tabularn<1K1 likes49 downloads16d agoHugging Face13Vidushee /Scigraph2_BT_RewardModelingDataset_removing15kfiltered_overlap_devtext100K<n<1M0 likes47 downloads7mo agoHugging Face14LangAGI-Lab /Mind2Web-cleaned-lite-reward-model Dataset Card for "Mind2Web-cleaned-lite-reward-model" More Information needed text1K<n<10K0 likes46 downloads2y agoHugging Face15Deojoandco /reward_model_anthropic Dataset Card for "reward_model_anthropic" More Information needed tabular100K<n<1M0 likes44 downloads4y agoHugging Face16Deojoandco /reward_model_anthropic_8 Dataset Card for "reward_model_anthropic_8" More Information needed tabular1K<n<10K1 likes43 downloads4y agoHugging Face17athirorg /USS-reward-model-qwen-FT-utt-v2-classtabular1K<n<10K0 likes41 downloads17d agoHugging Face18LangAGI-Lab /Mind2Web-cleaned-lite-reward-model-w-cot Dataset Card for "Mind2Web-cleaned-lite-reward-model-w-cot" More Information needed text1K<n<10K0 likes40 downloads2y agoHugging Face19matCercola18 /quotient-margins-reward-models Quotient Margins for Reward Models — data release Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for Reward Models. The short version of the paper. Reward models pick the best of N sampled responses, but their confidence is normally read off the reward gap between the top two samples. When several candidates express the same underlying behaviour, that gap is a within-class spacing and its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.texttext-generation0 likes38 downloads11h agoHugging Face20tranthanhnguyenai1 /RewardModel_8text100K<n<1M0 likes32 downloads1y agoHugging Face21HanningZhang /RAG-Reward-Modeling-v2 Dataset Card for "RAG-Reward-Modeling-v2" More Information needed text10K<n<100K0 likes32 downloads1y agoHugging Face22LangAGI-Lab /Mind2Web-cleaned-lite-reward-model-w-cot-v2 Dataset Card for "Mind2Web-cleaned-lite-reward-model-w-cot-v2" More Information needed text1K<n<10K0 likes31 downloads2y agoHugging Face23athirorg /USS-reward-model-qwen-FT-utterancetabular1K<n<10K0 likes30 downloads3mo agoHugging Face24tranthanhnguyenai1 /RewardModel_9text100K<n<1M0 likes28 downloads1y agoHugging Face25tranthanhnguyenai1 /RewardModel_3text100K<n<1M0 likes27 downloads1y agoHugging Face26HFXM /RewardModel_training_datatext10K<n<100K0 likes26 downloads2y agoHugging Face27tranthanhnguyenai1 /RewardModel_4text100K<n<1M0 likes25 downloads1y agoHugging Face28athirorg /USS-reward-model-qwen-binarytabularn<1K0 likes25 downloads3mo agoHugging Face29heitorefer /repro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces Agent traces Agent sessions published from a Trackio Logbook. textn<1K0 likes25 downloads2mo agoHugging Face30argilla /reward-model-data-falcon Guidelines These guidelines are based on the paper Training Language Models to Follow Instructions with Human Feedback You are given a text-based description of a task, submitted by a user. This task description may be in the form of an explicit instruction (e.g. "Write a story about a wise frog."). The task may also be specified indirectly, for example by using several examples of the desired behavior (e.g. given a sequence of movie reviews followed by their sentiment, followed by… See the full description on the dataset page: https://huggingface.co/datasets/argilla/reward-model-data-falcon.text1K<n<10K1 likes22 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.