datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
red_teaming_reward_modeling_pairwise
Dataset Card for "red_teaming_reward_modeling_pairwise"
More Information needed
red_teaming_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "red_teaming_reward_modeling_pairwise_no_as_an_ai"
More Information needed
sharegpt_reward_modeling_pairwise_no_as_an_ai
Dataset Card for "sharegpt_reward_modeling_pairwise_no_as_an_ai"
More Information needed
Diffusion-Reward-Modeling-for-Text-Rendering-Dataset
🖼️ Text-to-Image Rendering Dataset
A dataset of 14k text prompts for image generation with text rendering evaluation
📚 Dataset Overview
This dataset contains 14,000 text prompts specifically designed for:
Image generation with text rendering
Evaluating text preservation in generated images
Training diffusion models for better text rendering
Each prompt comes with:
Pre-extracted target text for rendering
5 Stable Diffusion 3 generated latents (70k total)
Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.reward_modeling_dataset
Dataset Card for "reward_modeling_dataset"
More Information needed
pedagogical-rewardmodel-datagpteacher_reward_modeling_pairwise
Dataset Card for "gpteacher_reward_modeling_pairwise"
More Information needed
reward_model_embeddingssharegpt_reward_modeling_pairwise
Dataset Card for "sharegpt_reward_modeling_pairwise"
More Information needed
reward_model_biases_attack_promptsreward_model_anthropic_88
Dataset Card for "reward_model_anthropic_88"
More Information needed
USS-reward-model-qwen-FT-v2Scigraph2_BT_RewardModelingDataset_removing15kfiltered_overlap_devMind2Web-cleaned-lite-reward-model
Dataset Card for "Mind2Web-cleaned-lite-reward-model"
More Information needed
reward_model_anthropic
Dataset Card for "reward_model_anthropic"
More Information needed
reward_model_anthropic_8
Dataset Card for "reward_model_anthropic_8"
More Information needed
USS-reward-model-qwen-FT-utt-v2-classMind2Web-cleaned-lite-reward-model-w-cot
Dataset Card for "Mind2Web-cleaned-lite-reward-model-w-cot"
More Information needed
quotient-margins-reward-models
Quotient Margins for Reward Models — data release
Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for
Reward Models.
The short version of the paper. Reward models pick the best of N sampled responses, but
their confidence is normally read off the reward gap between the top two samples. When
several candidates express the same underlying behaviour, that gap is a within-class spacing and
its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.RewardModel_8RAG-Reward-Modeling-v2
Dataset Card for "RAG-Reward-Modeling-v2"
More Information needed
Mind2Web-cleaned-lite-reward-model-w-cot-v2
Dataset Card for "Mind2Web-cleaned-lite-reward-model-w-cot-v2"
More Information needed
USS-reward-model-qwen-FT-utteranceRewardModel_9RewardModel_3RewardModel_training_dataRewardModel_4USS-reward-model-qwen-binaryrepro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces
Agent traces
Agent sessions published from a Trackio Logbook.
reward-model-data-falcon
Guidelines
These guidelines are based on the paper Training Language Models to Follow Instructions with Human Feedback
You are given a text-based description of a task, submitted by a user.
This task description may be in the form of an explicit instruction (e.g. "Write a story about a wise frog."). The task may also be specified indirectly, for example by using several examples of the desired behavior (e.g. given a sequence of movie reviews followed by their sentiment, followed by… See the full description on the dataset page: https://huggingface.co/datasets/argilla/reward-model-data-falcon.
