datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
videos-testVideo-MMELLaVA-Video-178K
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.GoPro-Raw-Videos
Raw GoPro Videos for Four Robotic Manipulation Tasks
[Project Page]
[Paper]
[Code]
[Models]
[Processed Dataset]
This repository contains raw GoPro videos of robotic manipulation tasks collected in-the-wild using UMI, as described in the paper "Data Scaling Laws in Imitation Learning for Robotic Manipulation". The dataset covers four tasks:
Pour Water
Arrange Mouse
Fold Towel
Unplug Charger
Dataset Folders:
arrange_mouse and pour_water: Each folder contains data… See the full description on the dataset page: https://huggingface.co/datasets/Fanqi-Lin/GoPro-Raw-Videos.code-world-model-project-page-videos
Code World Model Project Page Videos
Public research-demo video assets used by the Code World Model project page.
The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
VideoChat3-LV116k
VideoChat3-LV116K
VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments.
The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.qvhighlights-videos
QVHighlights Videos
All 25124 videos from the QVHighlights benchmark (train + val + test splits).
Total size: 135.3 GB.
Layout
Files are sharded into subdirectories by the first character of the filename
(HuggingFace caps each directory at 10,000 files):
<first-char>/<youtube-id>_<start>_<end>.mp4
Source
Lei et al., "QVHighlights: Detecting Moments and Highlights in Videos via Natural
Language Queries" (NeurIPS 2021).
Original archive:… See the full description on the dataset page: https://huggingface.co/datasets/ayushsdev/qvhighlights-videos.test-videosVBench-2.0_sampled_videos
Sample Videos of VBench-2.0
This dataset is used in the paper:👉 arXiv:2503.21755
Vera-Layered-Video-Dataset
Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing
Hongkai Zheng¹²* ·
Ta-Ying Cheng² ·
Benjamin Klein² ·
Yisong Yue¹ ·
Zhuoning Yuan²†
¹California Institute of Technology ²Netflix, Inc.
*Work done during an internship at Netflix †Project Lead
TL;DR: A layered diffusion framework for video editing. Vera jointly generates an edit layer, an alpha… See the full description on the dataset page: https://huggingface.co/datasets/netflix/Vera-Layered-Video-Dataset.videoshort_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.hoigen-filtered-videos
HOIGen Filtered Videos Dataset
This dataset contains 28562 filtered videos from the HOIGen-1M dataset based on the allowlist.
Dataset Structure
The videos are organized in the same structure as the original HOIGen dataset:
filtered_videos/
├── videos_part_1/
├── videos_part_2/
├── ...
└── videos_part_100/
Usage
from huggingface_hub import hf_hub_download
# Download a specific video
video_path = hf_hub_download(… See the full description on the dataset page: https://huggingface.co/datasets/charlychan123/hoigen-filtered-videos.Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.vitra-ego4d-videovivid-video-instructllava-video-178k-siglip-tokens-ftov-new
LLaVA-Video-178K SigLIP Token Cache (LLaVA-OV fine-tuned vision tower)
Derived data (vision-encoder features of video frames), not a
redistribution of the source videos. Source:
lmms-lab/LLaVA-Video-178K -- its card
restricts use to academic research and education, and its annotations come
from GPT-4-class models (see the OpenAI usage policy).
Complete: 85000 clips.
Subset
Folders: 0_30_s_academic_v0_1, 0_30_s_youtube_v0_1, 30_60_s_academic_v0_1… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Nasri/llava-video-178k-siglip-tokens-ftov-new.video-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.VideoOdysseyarxiv.org/abs/2605.22907
vlabench_composite_ft_lerobot_videoveo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.KABR-raw-videos
Dataset Card for KABR Raw Videos: Unprocessed Drone Footage for Kenyan Animal Behavior Analysis
Dataset Summary
This dataset contains the raw, unprocessed drone video footage collected during the creation of the KABR (Kenyan Animal Behavior Recognition) dataset.
Unlike the processed KABR mini-scene dataset which contains extracted video clips with behavioral annotations,
this collection provides the original full-frame drone videos captured at the Mpala Research… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR-raw-videos.3D-dungeon-crawler-video-v2
3D Dungeon Crawler Video v2
32,000 deterministic 28-second observational Unity episodes.
The canonical split contains 16,000 pretrain, 14,000 training,
1,000 test, and 1,000 evaluation episodes.
Unity renders at 512x288 for supersampling. Videos are stored at
256x144, 30 fps, H.264. Training samples every third frame,
yielding 280 frames and an 18x32 visual-token grid per episode.
manifest.jsonl is authoritative for asset paths. Each record points to one MP4 and one
NPZ… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/3D-dungeon-crawler-video-v2.fixtures_videosDemo_videovlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks Dataset - LeRobot v3.0
Dataset Description
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially.
Compared with the v2.0 version and the RLDS version of the dataset, this release stores visual observations in a video-compressed format rather than as individual image files. This design provides significant advantages in both storage efficiency and data loading performance.… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_ft_lerobot_video.Videos4CameraBenchKABR-mini-scene-raw-videos
Dataset Card for Kenyan Animal Behavior Recognition (KABR) Mini-Scene Raw Videos
Dataset Summary
This dataset is comprised of a collection of 10+ hours of drone videos focused on Kenyan wildlife that contains behaviors of giraffes, plains zebras, and Grevy's zebras.
Animals can be identified with bounding box coordinates provided, and behavior annotations can be recovered by linking the labels back to these bounding boxes from the mini-scene annotations provided in our… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR-mini-scene-raw-videos.train_video_and_instruction
ShareGPTVideo Training Data
All dataset and models can be found at ShareGPTVideo.
Contents:
Train 300k video frames: contains video frames used for SFT and DPO model, which is a subset of total 900k.
ActivityNet 50k + vidal 150k + webvid 100k.
Train 600k video frames: contains the rest 600k frames, the total 900k frames are used for pre-training stage. If you just do finetuning using our video QA, you can just download the 300k above.
900k composition is 400k WebVid +… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPTVideo/train_video_and_instruction.keystroke-typing-videos
Keystroke Typing Videos of Reuters
Recordings of typing randomly sampled sentences (<= 150 characters) from nltk Reuters dataset. Keystroke data is provided too.
