datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Unified-FeedbackCollections of pairwise feedback datasets.
openai/summarize_from_feedback
openai/webgpt_comparisons
Dahoas/instruct-synthetic-prompt-responses
Anthropic/hh-rlhf
lmsys/chatbot_arena_conversations
openbmb/UltraFeedback
argilla/ultrafeedback-binarized-preferences-cleaned
berkeley-nest/Nectar
Codes to reproduce the dataset: jdf-prog/UnifiedFeedback
Dataset formats
{
"id": "...",
"conv_A": [
{
"role": "user",
"content": "...",
},
{
"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.Code-Feedback OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
[🏠Homepage]
|
[🛠️Code]
Introduction
OpenCodeInterpreter is a family of open-source code generation systems designed to bridge the gap between large language models and advanced proprietary systems like the GPT-4 Code Interpreter. It significantly advances code generation capabilities by integrating execution and iterative refinement functionalities.
For further information and related… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Code-Feedback.summarize_from_feedbackSummarize from Feedback contains the human feedback data released by the "Learning to summarize from human feedback" paper.vision-feedback-mix-binarized
Dataset Card for Vision-Feedback-Mix-Binarized
Introduction
This dataset aims to provide large-scale vision feedback data.
It is a combination of the following high-quality vision feedback datasets:
zhiqings/LLaVA-Human-Preference-10K: 9,422 samples
MMInstruction/VLFeedback: 80,258 samples
YiyangAiLab/POVID_preference_data_for_VLLMs: 17,184 samples
openbmb/RLHF-V-Dataset: 5,733 samples
openbmb/RLAIF-V-Dataset: 83,132 samples
We also offer a cleaned version in… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vision-feedback-mix-binarized.text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.summarize_from_feedback_smallCode-Feedback
Dataset Card for CodeFeedback
This is a formatted version of m-a-p/Code-Feedback to store the conversations in the same format as the OpenAI SDK.
feedback_data_training
Repair replay update — September 15, 2026
The split still contains 161,030 weighted rows, with the same category counts:
Category
Rows
Share
Distinct examples before → after
One-shot
79,970
49.66%
35,197 → 35,197
Regular repairs
60,931
37.84%
40,530 → 48,726
Rollout-derived deep repairs
20,129
12.50%
436 → 1,825
This adds 9,585 distinct checked repair examples while preserving every legacy distinct row and every one-shot row's multiplicity. The new examples… See the full description on the dataset page: https://huggingface.co/datasets/formalmathatepfl/feedback_data_training.Feedback-Collection
Dataset Card
Dataset Summary
The Feedback Collection is a dataset designed to induce fine-grained evaluation capabilities into language models.\
Recently, proprietary LLMs (e.g., GPT-4) have been used to evaluate long-form responses. In our experiments, we found that open-source LMs are not capable of evaluating long-form responses, showing low correlation with both human evaluators and GPT-4.\
In our paper, we found that by (1) fine-tuning feedback generated by GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/prometheus-eval/Feedback-Collection.summarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
SWE-bench__style-3__fs-oracle_large_tokenlength
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
vietnamese_students_feedbackStudents’ feedback is a vital resource for the interdisciplinary research involving the combining of two different
research fields between sentiment analysis and education.
Vietnamese Students’ Feedback Corpus (UIT-VSFC) is the resource consists of over 16,000 sentences which are
human-annotated with two different tasks: sentiment-based and topic-based classifications.
To assess the quality of our corpus, we measure the annotator agreements and classification evaluation on the
UIT-VSFC corpus. As a result, we obtained the inter-annotator agreement of sentiments and topics with more than over
91% and 71% respectively. In addition, we built the baseline model with the Maximum Entropy classifier and achieved
approximately 88% of the sentiment F1-score and over 84% of the topic F1-score.text-2-image-Rich-Human-Feedback-32k
Building upon Google's research Rich Human Feedback for Text-to-Image Generation, and the
smaller, previous version of this dataset, we have collected over 3.7 million responses from 307'415 individual humans for the open-image-preference-v1 dataset using Rapidata via the Python API. Collection took less than 2 weeks.
If you get value from this dataset and would like to see more in the future, please consider liking it ♥️
Overview
We asked humans to evaluate AI-generated images… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback-32k.wb-feedbacks
Dataset Card for Wildberries products
Dataset Summary
The dataset contains product reviews from the Russian marketplace Wildberries, collected by mining about The dataset was collected by bruteforcing possible product identifiers (about 230 million) and querying all available feedbacks for them. The data are stored in zstd-archives containing jsonl-files. The 'nmId' in the dataset usually corresponds to the valid product article on the site, but sometimes reviews are… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-feedbacks.SWE-bench__style-3__fs-oracle
Dataset Card for "SWE-bench__style-3__fs-oracle"
More Information needed
llm-human-feedback-collector-chat-interface-dpocosmo-1B-claude_3h_feedback2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tokultra_feedback_dutch_cleaned
Ultra Feedback Dutch Cleaned
This is a cleaned version of BramVanroy/ultra_feedback_dutch, based on the cleaning done by Argilla on the original Ultra Feedback dataset. Another difference is that we only include GEITje 7B Ultra and GPT-4-Turbo. GEITje chat, which was used in the original dataset, is not used.
After cleaning I also generated replies for other models (like TowerInstruct, Mistral), but the results were too poor (in Dutch) to include so we only kept the GEITje Ultra and… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultra_feedback_dutch_cleaned.fast-feedback-benchmark
fast-feedback: offline benchmark data for cheap scaling-recipe ranking
Curated snapshot (2026-08-11) of public training-curve datasets used to build an
offline ground-truth benchmark for evaluating cheap feedback mechanisms that rank
LLM scaling recipes / architectures without training full ladder families.
A feedback mechanism is a budgeted procedure ("train these small rungs to N tokens,
inspect the curves, output a ranking"). This data lets such procedures be evaluated
by… See the full description on the dataset page: https://huggingface.co/datasets/jhyuckkim/fast-feedback-benchmark.omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_rightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 20,
"total_frames": 10481,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_vla_probe_rock_real_left_probe_right.VayuChat_Feedbackmano_feedbackreflection_eval_prompt1smar-feedback
SMAR AQI 3B -- Sovereign Feedback Dataset
User-curated quality signals from the SMAR AQI 3B chat workspace (smar.neugi.com/chat).
What Is This?
Every time a user clicks thumbs up/down on a chatbot response, that feedback flows here automatically. This dataset powers the continuous improvement loop for SMAR AQI 3B -- a sovereign AI model trained from scratch on Indonesian, English, and Arabic.
Lifecycle (A-Z)
1. Collection (Automatic)… See the full description on the dataset page: https://huggingface.co/datasets/atharia179/smar-feedback.openhands-feedback
OpenHands Feedback Dataset 🙌
Dataset Description
What is OpenHands Feedback?
The OpenHands Feedback Dataset is a collection of user interactions and feedback with the OpenHands AI coding assistant. This dataset contains real-world examples of how users interact with AI coding assistants, including both successful and unsuccessful interactions, along with user feedback on the quality and helpfulness of the responses.
The dataset currently contains 275 examples… See the full description on the dataset page: https://huggingface.co/datasets/OpenHands/openhands-feedback.omy_f3m_motor_feedback_test_baseball_aThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1399,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_baseball_a.summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.Irish_Prompt_Response_Human_Feedbackomy_f3m_motor_feedback_test_papercup_aThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy_f3m_motor_feedback",
"total_episodes": 3,
"total_frames": 1800,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jyLee0111/omy_f3m_motor_feedback_test_papercup_a.
