datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ReCo-Data
ReCo-Data Dataset Card
Introduction
ReCo-Data is a large-scale, high-quality video editing dataset comprising 500K+ instruction-video pairs. This card provides its statistics, collection pipeline, and dataset format.
1. Dataset Statistics
Statistics
Figure Caption:
(a) Overview of scale
(b) Task distribution showing balanced quantities: Replace (156.6K), Style (130.6K), Remove (121.6K), and Add (115.6K). Human evaluation on 200 randomly… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Data.ReactID-Data
ReactID-Data
✨ Summary
ReactID-Data is a large-scale, high-quality dataset for subject-driven video generation (Subject-to-Video). It contains 4.1M subject–text–video triples with instance detection/segmentation, face detection, multi-dimensional quality scores, structured entity labels, and timeline annotations with temporally segmented events. The dataset also supports generation tasks beyond Subject-to-Video.
📁 Data Structure
ReactID-Data/
├──… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReactID-Data.Hidream_o1-RoboLab-resultsHiDream-I1-ArtistsHidream_t2i_human_preference
Rapidata Hidream I1 full Preference
This T2I dataset contains over 195k human responses from over 38k individual annotators, collected in just ~1 Day using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Hidream I1 full across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Hidream_t2i_human_preference.ReCo-Bench
ReCo-Bench
Project Page | Paper | GitHub | ReCo-Data
This is the official ReCo-Bench dataset introduced in the paper "Region-Constraint In-Context Generation for Instructional Video Editing". ReCo-Bench is a VLLM-based evaluation benchmark designed to comprehensively and effectively assess video editing quality.
Usage
After downloading the repository, you can start the evaluation directly by running the following script:
bash run_eval_via_gemini.sh
VLLM-based… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/ReCo-Bench.VIP-200K-Video
VIP-200K-Video
Overview
This repository contains only the video files (.tar archives) from the VIP-200K dataset.
For the complete dataset (including JSON annotations, face frames, and segment metadata), please download HiDream-ai/VIP-200K.
File Structure
video/
├── vip200k_train_0001_of_0100.tar
├── vip200k_train_0002_of_0100.tar
├── ...
└── vip200k_train_0100_of_0100.tar
Each .tar archive contains video clips organized by YouTube video ID:
{video_id}/… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/VIP-200K-Video.VIP-200K
Dataset Documentation
Video Files: If you need the pre-packaged video clips (.tar archives), you can download them directly from HiDream-ai/VIP-200K-Video.
Dataset Overview
This dataset containing annotated face and contextual segment information from video content. Each entry represents a person with detected faces and corresponding video segments.
JSON Structure
[
{
"faces": [
{ /* Face Object */ }
],
"segments": [
{ /* Segment… See the full description on the dataset page: https://huggingface.co/datasets/HiDream-ai/VIP-200K.HiDream
HidDream Text To Image Installation + Worklfow:
https://www.stablediffusiontutorials.com/2025/04/hidream-model.html
FLUX VS HiDream Image Comparision test:
https://www.stablediffusiontutorials.com/2025/05/flux-vs-hidream.html
HidDream E1.1 EDIT(Image Editor) Installation + Worklfow:
https://www.stablediffusiontutorials.com/2025/07/hidream-e11-image-editing-on-low-vrams.html
passport-id-face-injection-qwen-hidream-v1
Passport/ID Face-Injection Detection Benchmark v1
A demographically-stratified, paired real-vs-synthetic benchmark of identity-preserving
face-injection attacks framed for the passport / ID / remote-KYC threat model. This is an
evaluation benchmark, not training data.
Version: v1 (pristine cohort, pre-perturbation). Generated 2026-07-15/16.
Total images: 2,994 = 998 bona-fide reals + 1,996 synthetic attacks (998 per generator).
Structure: paired. Each real reference has… See the full description on the dataset page: https://huggingface.co/datasets/danb21/passport-id-face-injection-qwen-hidream-v1.kyc-passport-deepfake-pad-hidream-o1-qwen-image-edit
KYC Passport Deepfake / Presentation-Attack Benchmark
HiDream-O1 · Qwen-Image-Edit · 23 document-reproduction conditions
A presentation-attack detection (PAD) evaluation set for identity-document face
imagery: real reference faces and generator-attributed synthetic faces, each carried
through 23 acquisition conditions that emulate how a passport photo actually reaches a
KYC system: print and scan, photocopy, fax, phone recapture, JPEG chains, and
background… See the full description on the dataset page: https://huggingface.co/datasets/danb21/kyc-passport-deepfake-pad-hidream-o1-qwen-image-edit.hidreamlorasdusk-Hidream-E1-dataset22hidream_l1_fullbc8-hidream-datasetHidream_challegedHiDream_Outputs_20260512_143450_925SCHiDream_Outputs_20260512_143850_J6CJFHiDream_Outputs_20260512_144418_C73QJ
