datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vision-feedback-mix-binarized
Dataset Card for Vision-Feedback-Mix-Binarized
Introduction
This dataset aims to provide large-scale vision feedback data.
It is a combination of the following high-quality vision feedback datasets:
zhiqings/LLaVA-Human-Preference-10K: 9,422 samples
MMInstruction/VLFeedback: 80,258 samples
YiyangAiLab/POVID_preference_data_for_VLLMs: 17,184 samples
openbmb/RLHF-V-Dataset: 5,733 samples
openbmb/RLAIF-V-Dataset: 83,132 samples
We also offer a cleaned version in… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/vision-feedback-mix-binarized.text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.text-2-image-Rich-Human-Feedback-32k
Building upon Google's research Rich Human Feedback for Text-to-Image Generation, and the
smaller, previous version of this dataset, we have collected over 3.7 million responses from 307'415 individual humans for the open-image-preference-v1 dataset using Rapidata via the Python API. Collection took less than 2 weeks.
If you get value from this dataset and would like to see more in the future, please consider liking it ♥️
Overview
We asked humans to evaluate AI-generated images… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback-32k.VayuChat_FeedbackSeeTRUE-Feedback
Dataset Card for SeeTRUE-Feedback
Dataset Description
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Fields
Data Splits
Dataset Creation
Licensing Information
Citation Information
Dataset Description
The SeeTRUE-Feedback dataset is a diverse benchmark for the meta-evaluation of image-text matching/alignment feedback. It aims to overcome limitations in current benchmarks, which primarily focus on predicting a matching score between 0-1.… See the full description on the dataset page: https://huggingface.co/datasets/mismatch-quest/SeeTRUE-Feedback.vision-feedback-mix-binarized
Dataset Card for Vision-Feedback-Mix-Binarized
Introduction
This dataset aims to provide large-scale vision feedback data.
It is a combination of the following high-quality vision feedback datasets:
zhiqings/LLaVA-Human-Preference-10K: 9,422 samples
MMInstruction/VLFeedback: 80,258 samples
YiyangAiLab/POVID_preference_data_for_VLLMs: 17,184 samples
openbmb/RLHF-V-Dataset: 5,733 samples
openbmb/RLAIF-V-Dataset: 83,132 samples
We also offer a cleaned version in… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/vision-feedback-mix-binarized.bb-trade-idp-feedback
🏦 Bangladesh Bank Trade Finance IDP — Multi-User Collaborative Fine-Tuning Dataset
This dataset contains human-reviewed, verified, and corrected document extractions for the 8 official Bangladesh Bank regulatory trade-finance document types.
It is completely self-contained and structured for immediate Vision-Language Model (VLM) fine-tuning anytime from any environment (Colab, Kaggle, GPU cluster, or local), with built-in multi-annotator merge support and incremental delta… See the full description on the dataset page: https://huggingface.co/datasets/jihadv4/bb-trade-idp-feedback.vision-feedback-mix-binarized-cleaned
Dataset Card for Vision-Feedback-Mix-Binarized-Cleaned
Introduction
This dataset represents a cleaned version on wangclnlp/vision-feedback-mix-binarized.
Descriptions of the base datasets, including the data format and the procedure for mixing data, can be found in this link.
Our Methods for Cleaning Vision Feedback Data
Our goal is to select vision feedback samples where the preferred outputs are significantly differentiated from the dispreferred ones, and the… See the full description on the dataset page: https://huggingface.co/datasets/NiuTrans/vision-feedback-mix-binarized-cleaned.BlindLoop-Difficulty-Feedback
BlindLoop Difficulty Feedback
This is the public, hash-bound release of BlindLoop Section 3. Coding agents
generated executable visual-question tasks; each task's inverse program checked
the answer from rendered pixels. For complete feedback transactions, the exact
same five images were evaluated by three frontier VLMs and the resulting
difficulty signal was returned to the next generation episode.
Contents
Config
Unit
Rows
tasks
generated task
266… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/BlindLoop-Difficulty-Feedback.summarize_from_feedback_oai_preprocessing_1711138793
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138793"
More Information needed
summarize_from_feedback_oai_preprocessing_1711138084
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138084"
More Information needed
bb-trade-idp-feedback
🏦 Bangladesh Bank Trade Finance IDP — Multi-User Collaborative Fine-Tuning Dataset
This dataset contains human-reviewed, verified, and corrected document extractions for the 8 official Bangladesh Bank regulatory trade-finance document types.
It is completely self-contained and structured for immediate Vision-Language Model (VLM) fine-tuning anytime from any environment (Colab, Kaggle, GPU cluster, or local), with built-in multi-annotator merge support and incremental delta… See the full description on the dataset page: https://huggingface.co/datasets/bisalsaha/bb-trade-idp-feedback.text-2-video-Rich-Human-Feedback
Rapidata Video Generation Rich Human Feedback Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~4 hours total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~22'000 human annotations were collected to evaluate AI-generated videos (using Sora) in 5 different categories.
Prompt - Video… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-Rich-Human-Feedback.bb-trade-idp-feedback_V1
🏦 Bangladesh Bank Trade Finance IDP — Multi-User Collaborative Fine-Tuning Dataset
This dataset contains human-reviewed, verified, and corrected document extractions for the 8 official Bangladesh Bank regulatory trade-finance document types.
It is completely self-contained and structured for immediate Vision-Language Model (VLM) fine-tuning anytime from any environment (Colab, Kaggle, GPU cluster, or local), with built-in multi-annotator merge support and incremental delta… See the full description on the dataset page: https://huggingface.co/datasets/jihadv4/bb-trade-idp-feedback_V1.multi-feedbackmulti-feedback-corrvibe-feedbacksmulti-feedback-1DAPO-Math-17k-gemma-feedbackOmni-MATH-gemma-feedbackDualLoop-Difficulty-Feedback
DualLoop Difficulty Feedback
This release contains the DualLoop model-feedback discovery data. Coding agents
generated executable visual-question tasks; each task's inverse program checked
the answer from rendered pixels. For complete feedback transactions, the exact
same five images were evaluated by three frontier VLMs and the resulting
difficulty signal was returned to the next generation episode.
Contents
Config
Unit
Rows
tasks
generated task
266… See the full description on the dataset page: https://huggingface.co/datasets/DualLoop/DualLoop-Difficulty-Feedback.feedbacklorafeedbacklora2feedback_saes
