datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MVU-Eval-Team/MVU-Eval-Data.CameraClone-Dataset
CamCloneMaster: Enabling reference-based camera control for video generation
Paper:https://arxiv.org/abs/2506.03140
Project Page:https://camclonemaster.github.io/
Dataset:https://huggingface.co/datasets/KwaiVGI/CameraClone-Dataset
Training & Inference Code:https://github.com/KwaiVGI/CamCloneMaster
Camera Clone Dataset
1. Dataset Introduction
TL;DR: The Camera Clone Dataset, introduced in CamCloneMaster, is a large-scale synthetic dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/CameraClone-Dataset.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/: directories… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Talker-T2AV-Data.StatePlay-Dataset
StatePlay SF3 Dataset
This repository contains the curated Street Fighter III
training data used by StatePlay.
Paper: https://huggingface.co/papers/2607.26754
Code: https://github.com/Jimntu/StatePlay
Dataset contents
The SF3 subset contains 10,000 five-second gameplay clips. Every row in
SF3/metadata_state_polish.csv references:
video: a gameplay clip at SF3/clips/*/video.mp4;
action: an aligned action and game-state table at
SF3/clips/*/actions.parquet;
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/onepiece1999/StatePlay-Dataset.MVU-Eval-Data
MVU-Eval Dataset
Paper | Code | Project Page
Dataset Description
The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/stephenbasd/MVU-Eval-Data.egomonth-dataset
EgoMonth Dataset
Overview
EgoMonth is a month-level egocentric video question-answering benchmark for evaluating long-term spatiotemporal memory in multimodal large language models. The dataset focuses on daily-life first-person videos and QA tasks that require temporal indexing, spatial grounding, multi-video reasoning, and long-horizon memory.
This repository provides QA metadata, structured annotations, representative anonymized sample videos, and baseline… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-egomonth/egomonth-dataset.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.video-data
TalorData Video Data
Rich, up-to-date video metadata + ready-to-transcribe MP4s from YouTube.
TalorData Video Data is a large, constantly refreshed dataset of YouTube videos. Each record bundles the video's pre-cut MP4 alongside an auto-generated, timestamped transcript — ready for fine-tuning, RAG, search, summarization, and multimodal inference.
Sample Files (Preview)
This repository hosts a public preview/sample of the record schema. Open sample-metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/TalorDataHQ/video-data.Tiktok-Videos
TikTok Video Analytics Dataset
Sample TikTok video dataset with comprehensive engagement metrics and metadata. Each row represents a single TikTok video with content and detailed analytics.
This is a sample dataset. To access the full version or request any custom dataset tailored to your needs, contact DataHive at contact@datahive.ai.
Files Included
train.csv – TikTok video analytics data
What's included
Video URLs and identifiers
Comprehensive engagement… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/Tiktok-Videos.ReactiveGWM-Datasets
ReactiveGWM-Datasets: Strategy-Aligned Rollouts for Reactive Game World Models
📚 Datasets-Introduction
ReactiveGWM-Datasets is the strategy-aligned training corpus that powers
ReactiveGWM, a game world
model that decouples player control from NPC autonomy. To learn that
decoupling, the model needs supervision that pairs each gameplay clip with
both a per-frame action stream (what the player did) and a high-level
NPC description (what the NPC tried to do, and under… See the full description on the dataset page: https://huggingface.co/datasets/INV-WZQ/ReactiveGWM-Datasets.pilates-dataset
Project Overview
THE PILATES SESSION is an end-to-end data science project that transforms the classical Pilates mat repertoire into an intelligent, AI-powered lesson-planning engine.
The project follows a structured data science workflow, beginning with the creation of a synthetic dataset - generated through careful prompt engineering on top of a clinical knowledge base - to capture the classical repertoire in a rich, structured form. This is followed by a… See the full description on the dataset page: https://huggingface.co/datasets/lia-prop13/pilates-dataset.bambu-timelapse-dataset
Bambu Timelapse Dataset
Dataset Summary
Disclaimer: This dataset is an independent community project and is not affiliated with or endorsed by Bambu Lab in any official capacity. We simply curate footage captured on its consumer printers to enable open research.
The Bambu Timelapse Dataset is an open, community‑driven collection of time‑lapse videos captured on Bambu Lab 3‑D printers (P1 series, X1 series and variants).
Its goal is to provide a high‑quality video corpus… See the full description on the dataset page: https://huggingface.co/datasets/v2thegreat/bambu-timelapse-dataset.Object_Grasping_DatasetThis dataset contains ground-truth annotations for videos depicting interactions with everyday objects. Specifically, each video captures a user reaching toward and grasping an object. The annotations are intended to support the evaluation of hand-object interaction models and related computer vision tasks.
The dataset consists of videos selected from the Something-Something V2 dataset, complemented by supplementary videos recorded specifically for this work. The supplementary videos feature… See the full description on the dataset page: https://huggingface.co/datasets/IoannisKap/Object_Grasping_Dataset.Stroke_Prediction_Dataset
Stroke Prediction Dataset Analysis
Presentation Video
Project Overview
The goal of this project is to predict the likelihood of a patient suffering a stroke based on demographic, health, and lifestyle parameters. Stroke is a leading cause of death and long-term disability worldwide, and early identification of high-risk individuals can significantly improve prevention strategies and clinical outcomes.
Dataset Summary
Source: Kaggle – Stroke… See the full description on the dataset page: https://huggingface.co/datasets/nadiCR7/Stroke_Prediction_Dataset.EgoAVU_data
[CVPR2026] EgoAVU
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding
See our github for the code and setup instructions.
Check out our homepage and paper for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual understanding. EgoAVU enriches existing egocentric narrations by integrating human actions with environmental context, explicitly linking visible objects and the sounds produced during interactions… See the full description on the dataset page: https://huggingface.co/datasets/jun111111/EgoAVU_data.Talker-T2AV-Data_trainer
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/prakhar-adaf/Talker-T2AV-Data_trainer.VidData
Technical Documentation for the Text-to-Video Dataset “VidData”
1. Introduction
This dataset contains 1006 annotated videos of everyday scenes, used for training and evaluating AI models in video generation and recognition. It is structured to meet the needs of Text-to-Video models and motion analysis.
2. Dataset Specifications
2.1. Generation Criteria
Maximum video duration: 10 seconds maximum
Video themes:
Walking
Exercising
Writing
Shopping… See the full description on the dataset page: https://huggingface.co/datasets/Databoost/VidData.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/Prakhar-kumar/Talker-T2AV-Data.shipping_databrazilian-cultural-video-dataset
Bamboo Data Brazilian Cultural Video Dataset (Sample)
⚠️ License Notice: Evaluation Only
This is a sample of the Bamboo Data brazilian cultural video dataset, provided for internal evaluation purposes ONLY. The use of this data is strictly limited by the license defined below.
Any use for training, fine-tuning, or inference of AI/ML models, or any commercial activity, is strictly prohibited with this sample.
Dataset Description
The Bamboo Data… See the full description on the dataset page: https://huggingface.co/datasets/bamboodata/brazilian-cultural-video-dataset.Talker-T2AV-Data
Talker-T2AV-Data
Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
Paper (arXiv 2604.23586) ·
Code (GitHub) ·
Model ·
Samples
Clean training data package for Talker-T2AV. Paths in metadata/train.csv are relative to the dataset root after extracting the shard archives.
Contents
metadata/train.csv: training index used by Talker-T2AV.
shards/*.tar: clean archive shards grouped by modality and dataset source.
audio/, motion/, video/:… See the full description on the dataset page: https://huggingface.co/datasets/dsenflam/Talker-T2AV-Data.Urdu-Multimodal-Emotion-Datasetolympic-data-analysis
The analysis behind the glory: 120 Years of Data
Project Overview
In this project, I performed a comprehensive Exploratory Data Analysis (EDA) on a dataset covering 120 years of Olympic history. My main goal was to transform a messy, historical dataset into a clean, analyzed resource to uncover the key physical, demographic, and geopolitical factors that determine an athlete's success.
Dataset Description
The dataset provides a wide view of Olympic athletes… See the full description on the dataset page: https://huggingface.co/datasets/grasimus/olympic-data-analysis.Face_Emotion_Insight_Recognition_Dataset
Multimodal Emotion & Physiological Analysis
Project Overview
This project explores the relationship between physiological signals and facial micro expressions to improve emotion recognition in therapeutic settings. This research assists in determining which facial and physiological signals should be prioritized to detect clinical "misalignment" or hidden distress during therapy sessions.
Dataset Selection & Description
Source: The dataset is sourced from Kaggle (Face Emotion & Physiological… See the full description on the dataset page: https://huggingface.co/datasets/ofekponzo/Face_Emotion_Insight_Recognition_Dataset.OmniShow_example_dataset
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
Donghao Zhou1,*, Guisheng Liu2,*, Hao Yang2, Jiatong Li2,†, Jingyu Lin3, Xiaohu Huang4,
Yichen Liu2, Xin Gao2, Cunjian Chen3, Shilei Wen2,§, Chi-Wing Fu1, Pheng-Ann Heng1,§
1The Chinese University of Hong Kong, 2ByteDance, 3Monash University, 4The University of Hong Kong
*Equal contribution, †Project lead, §Corresponding author
🌍 Useful Links
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/donghao-zhou/OmniShow_example_dataset.AVSR-Vietnamese-Datasetforensics-grpo-data
forensics-grpo-data
Generated-video dataset + annotations used to train
sdzt/forensics-grpo.
📂 Repository layout
forensics-grpo-data/
├── video/ # 5,388 .mp4 clips, packed as one .tar per generator
│ ├── 01_vidu.tar # 9.7 GB — Vidu
│ ├── 02_wan.tar # 28 GB — Wan
│ ├── 03_fcvg.tar # 27 GB — FCVG
│ ├── 04_scifi.tar # 34 GB — SciFi
│ ├── 05_ltx.tar # 6.3 GB — LTX
│… See the full description on the dataset page: https://huggingface.co/datasets/sdzt/forensics-grpo-data.Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction
Title
An Annotated Dataset for Hearing-Impaired Speech-to-Text Correction
Abstract
This dataset is specifically designed for the task of correcting speech to text errors in hearing-impaired individuals.
It includes one hour of real speech files of hearing-impaired individuals, automatic speech recognition (ASR) output text, and manually corrected standard text.
We searched for an hour of audio from hearing-impaired individuals to ensure that the voice was authentic and… See the full description on the dataset page: https://huggingface.co/datasets/llllliuuy/Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction.PAD3-dataset-video-captiondataset
