datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.Hidream_t2i_human_preference
Rapidata Hidream I1 full Preference
This T2I dataset contains over 195k human responses from over 38k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Hidream I1 full across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Hidream_t2i_human_preference.VIP-200K
Dataset Documentation
Video Files: If you need the pre-packaged video clips (.tar archives), you can download them directly from HiDream-ai/VIP-200K-Video.
Dataset Overview
This dataset containing annotated face and contextual segment information from video content. Each entry represents a person with detected faces and corresponding video segments.
JSON Structure
[
{
"faces": [
{ /* Face Object */ }
],
"segments": [
{ /* Segment… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/VIP-200K.
