datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-2-image-Rich-Human-Feedback
Building upon Google's research Rich Human Feedback for Text-to-Image Generation we have collected over 1.5 million responses from 152'684 individual humans using Rapidata via the Python API. Collection took roughly 5 days.
If you get value from this dataset and would like to see more in the future, please consider liking it.
Overview
We asked humans to evaluate AI-generated images in style, coherence and prompt alignment. For images that contained flaws, participants were… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback.text-2-image-Rich-Human-Feedback-32k
Building upon Google's research Rich Human Feedback for Text-to-Image Generation, and the
smaller, previous version of this dataset, we have collected over 3.7 million responses from 307'415 individual humans for the open-image-preference-v1 dataset using Rapidata via the Python API. Collection took less than 2 weeks.
If you get value from this dataset and would like to see more in the future, please consider liking it ♥️
Overview
We asked humans to evaluate AI-generated… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-image-Rich-Human-Feedback-32k.text-2-video-Rich-Human-Feedback
Rapidata Video Generation Rich Human Feedback Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~4 hours total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~22'000 human annotations were collected to evaluate AI-generated videos (using Sora) in 5 different categories.
Prompt - Video… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-Rich-Human-Feedback.Dojo-HumanFeedback-DPO
Dataset Description:
Dojo-HumanFeedback-DPO is a preference dataset designed to improve interface generation capabilities in large language models (LLMs). The dataset contains 12500 high-quality, synthetic chosen-rejected preference pairs, in the specific domain of generating frontend interfaces using HTML, CSS, and JavaScript.
The dataset format is optimized for Direct Preference Optimization (DPO), but can potentially be used in other machine learning contexts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tensorplex-labs/Dojo-HumanFeedback-DPO.lave-human-feedback
LAVE human judgments
This repository contains the human judgment data for Improving Automatic VQA Evaluation Using Large Language Models. Details about the data collection process and crowdworker population can be found in our paper, specifically in section 5.2 and appendix A.1.
Fields:
dataset: VQA dataset of origin for this example (vqav2, vgqa, okvqa).
model: VQA model that generated the predicted answer (blip2, promptcap, blip_vqa, blip_vg).
qid: question ID coming from the… See the full description on the dataset page: https://huggingface.co/datasets/mair-lab/lave-human-feedback.human_feedbackagile-cymru-synthetic-dataset-with-human-feedback
