datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Unified-FeedbackCollections of pairwise feedback datasets.
openai/summarize_from_feedback
openai/webgpt_comparisons
Dahoas/instruct-synthetic-prompt-responses
Anthropic/hh-rlhf
lmsys/chatbot_arena_conversations
openbmb/UltraFeedback
argilla/ultrafeedback-binarized-preferences-cleaned
berkeley-nest/Nectar
Codes to reproduce the dataset: jdf-prog/UnifiedFeedback
Dataset formats
{
"id": "...",
"conv_A": [
{
"role": "user",
"content": "...",
},
{
"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.Code-Feedback OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
Introduction
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and related… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Code-Feedback.summarize_from_feedbackSummarize from Feedback contains the human feedback data released by the "Learning to summarize from human feedback" paper.vision-feedback-mix-binarized
Dataset Card for Vision-Feedback-Mix-Binarized
Introduction
This dataset aims to provide large-scale vision feedback data.
It is a combination of the following high-quality vision feedback datasets:
zhiqings/LLaVA-Human-Preference-10K: 9,422 samples
MMInstruction/VLFeedback: 80,258 samples
YiyangAiLab/POVID_preference_data_for_VLLMs: 17,184 samples
openbmb/RLHF-V-Dataset: 5,733 samples
openbmb/RLAIF-V-Dataset: 83,132 samples
We also offer a cleaned version in… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vision-feedback-mix-binarized.text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.summarize_from_feedback_smallfeedback_data_training
Repair replay update — September 15, 2026
The split still contains 161,030 weighted rows, with the same category counts:
Category
Rows
Share
Distinct examples before → after
One-shot
79,970
49.66%
35,197 → 35,197
Regular repairs
60,931
37.84%
40,530 → 48,726
Rollout-derived deep repairs
20,129
12.50%
436 → 1,825
This adds 9,585 distinct checked repair examples while preserving every legacy distinct row and every one-shot row's multiplicity. The new examples… See the full description on the dataset page: https://huggingface.co/datasets/formalmathatepfl/feedback_data_training.Code-Feedback
Dataset Card for CodeFeedback
This is a formatted version of m-a-p/Code-Feedback to store the conversations in the same format as the OpenAI SDK.
Feedback-Collection
Dataset Card
Dataset Summary
The Feedback Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.\
Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\
In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Feedback-Collection.cosmo-1B-claude_3h_feedbackSWE-bench__style-3__fs-oracle_large_tokenlength
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
vietnamese_students_feedbackStudents’ feedback is a vital resource for the interdisciplinary research involving the combining of two different
research fields between sentiment analysis and education.
Vietnamese Students’ Feedback Corpus (UIT-VSFC) is the resource consists of over 16,000 sentences which are
human-annotated with two different tasks: sentiment-based and topic-based classifications.
To assess the quality of our corpus, we measure the annotator agreements and classification evaluation on the
UIT-VSFC corpus. As a result, we obtained the inter-annotator agreement of sentiments and topics with more than over
91% and 71% respectively. In addition, we built the baseline model with the Maximum Entropy classifier and achieved
approximately 88% of the sentiment F1-score and over 84% of the topic F1-score.summarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
SWE-bench__style-3__fs-oracle
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
llm-human-feedback-collector-chat-interface-dpowb-feedbacks
Dataset Card for Wildberries products
Dataset Summary
The dataset contains product reviews from the Russian marketplace Wildberries, collected by mining about The dataset was collected by bruteforcing possible product identifiers (about 230 million) and querying all available feedbacks for them. The data are stored in zstd-archives containing jsonl-files. The 'nmId' in the dataset usually corresponds to the valid product article on the site, but sometimes reviews are… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-feedbacks.text-2-image-Rich-Human-Feedback-32k
Building upon Google's research Rich Human Feedback for Text-to-Image Generation, and the
smaller, previous version of this dataset, we have collected over 3.7 million responses from 307'415 individual humans for the open-image-preference-v1 dataset using Rapidata via the Python API. Collection took less than 2 weeks.
If you get value from this dataset and would like to see more in the future, please consider liking it ♥️
Overview
We asked humans to evaluate AI-generated… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback-32k.reflection_eval_prompt1fast-feedback-benchmark
fast-feedback: offline benchmark data for cheap scaling-recipe ranking
Curated snapshot (2026-08-11) of public training-curve datasets used to build an
offline ground-truth benchmark for evaluating cheap feedback mechanisms that rank
LLM scaling recipes / architectures without training full ladder families.
A feedback mechanism is a budgeted procedure ("train these small rungs to N tokens,
inspect the curves, output a ranking"). This data lets such procedures be evaluated
by… See the full description on the dataset page: https://huggingface.co/datasets/jhyuckkim/fast-feedback-benchmark.VayuChat_Feedbackruncam-feed-camera-egocentric-rgb-imu
RunCam Feed Camera - Egocentric RGB + IMU Sample Dataset
A small sample dataset captured with the RunCam Feed Camera for egocentric video and synchronized motion-sensor workflows.
Capture Device
Video: H.265 MP4, 1920x1080, 60 fps for V01-V06
Nominal video bitrate: 18 Mbps
Horizontal field of view: 126 degrees
Device weight: approximately 26 g
IMU: ICM-42607, 6-axis
IMU sampling rate: 800 Hz for the recordings in this sample
Firmware reported in the GCSV files:… See the full description on the dataset page: https://huggingface.co/datasets/RunCam/runcam-feed-camera-egocentric-rgb-imu.openhands-feedback
OpenHands Feedback Dataset 🙌
Dataset Description
What is OpenHands Feedback?
The OpenHands Feedback Dataset is a collection of user interactions and feedback with the OpenHands AI coding assistant. This dataset contains real-world examples of how users interact with AI coding assistants, including both successful and unsuccessful interactions, along with user feedback on the quality and helpfulness of the responses.
The dataset currently contains 275 examples… See the full description on the dataset page: https://huggingface.co/datasets/OpenHands/openhands-feedback.ultra_feedback_dutch_cleaned
Ultra Feedback Dutch Cleaned
This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used.
After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.saf_communication_networks_english
Dataset Card for "saf_communication_networks_english"
Dataset Summary
Short Answer Feedback (SAF) dataset is a short answer dataset introduced in Your Answer is Incorrect... Would you like to know why? Introducing a Bilingual Short Answer Feedback Dataset (Filighera et al., ACL 2022) as a way to remedy the lack of content-focused feedback datasets. This version of the dataset contains 31 English questions covering a range of college-level communication networks topics -… See the full description on the dataset page: https://huggingface.co/datasets/Short-Answer-Feedback/saf_communication_networks_english.reflection_eval_prompt2test_reflection_eval_promptCode-feedback-sharegpt-renamed2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tokultra-feedback-pairedSeeTRUE-Feedback
Dataset Card for SeeTRUE-Feedback
Dataset Description
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Fields
Data Splits
Dataset Creation
Licensing Information
Citation Information
Dataset Description
The SeeTRUE-Feedback dataset is a diverse benchmark for the meta-evaluation of image-text matching/alignment feedback. It aims to overcome limitations in current benchmarks, which primarily focus on predicting a matching score between 0-1.… See the full description on the dataset page: https://huggingface.co/datasets/mismatch-quest/SeeTRUE-Feedback.
