datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VideoEspresso_train_multi_image
VideoEspresso
This dataset is the multi-image version.
Leaderboard
Model
Params
Frames
Overall
Narrative Analysis
Event Dynamic
Preparation Steps
Causal Analysis
Theme Analysis
Contextual Analysis
Influence Analysis
Role Analysis
Interaction Analysis
Behavior Analysis
Emotion Analysis
Cooking Process
Traffic Analysis
Situation Analysis
LLaVA-Video
72B
64
66.3%
68.4%
66.2%
74.5%
62.7%
62.3%
71.6%
62.5%
63.5%
67.7%
63.2%
60.0%
75.5%
76.7%
74.0%
LLaVA-OneVision… See the full description on the dataset page: https://huggingface.co/datasets/hshjerry0315/VideoEspresso_train_multi_image.multi-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.Qwen2.5_Image_Results
