datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
video-vec2wav2-tokenizer
video-vec2wav2-tokenizer
Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command
video2dataset) that turns a folder of videos into clean AI training datasets
for speech recognition (ASR) and text-to-speech (TTS).
videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json
Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg
audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.video-vec2wav2-tokenizer-2
video-vec2wav2-tokenizer-2
Version 2 - continuation shard of the video-to-AI-dataset tokenizer project.
Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command
video2dataset) that turns a folder of videos into clean AI training datasets
for speech recognition (ASR) and text-to-speech (TTS).
videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json
Video processing —… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-2.video-vec2wav2-tokenizer-3
video-vec2wav2-tokenizer-3
Version 3 - continuation shard of the video-to-AI-dataset tokenizer project.
Version 2 - continuation shard of the video-to-AI-dataset tokenizer project.
Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command
video2dataset) that turns a folder of videos into clean AI training datasets
for speech recognition (ASR) and text-to-speech (TTS).
videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──►… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-3.videos-testDL3DV-ALL-video
DL3DV-Dataset
This repo has all the original videos of DL3DV-10K Dataset. We are working hard to review all the dataset to avoid sensitive information. Thank you for your patience.
Download
If you have enough space, you can use git to download a dataset from huggingface. See this link.
If you do not have enough space, we further provide a download script here to download a subset. The usage:
usage: download.py [-h] --odir ODIR --subset {1K,2K,3K,4K,5K,6K,7K,8K,9K,10K}… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-ALL-video.Video-MMELLaVA-Video-178K
Dataset Card for LLaVA-Video-178K
Uses
This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy.
Data Sources
For the training of LLaVA-Video, we utilized video-language data from five primary sources:
LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.videommeVideoChat-Flash-Training-Data
🦜 VideoChat-Flash-Training-Data
This repos contains all annotaions and most videos for training VideoChat-Flash.
📕 How to use the LongVid data?
For video_dir like longvid_subset/coin_grounding_10k_zip, you need to concat this dir to a zip file as follows:
cat ego4dhcap_eventunderstanding_2k_zip/* > ego4dhcap_eventunderstanding_2k.zip
✏️ Citation
@article{li2024videochatflash,
title={VideoChat-Flash: Hierarchical Compression for Long-Context… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/VideoChat-Flash-Training-Data.GoPro-Raw-Videos
Raw GoPro Videos for Four Robotic Manipulation Tasks
[Project Page]
[Paper]
[Code]
[Models]
[Processed Dataset]
This repository contains raw GoPro videos of robotic manipulation tasks collected in-the-wild using UMI, as described in the paper "Data Scaling Laws in Imitation Learning for Robotic Manipulation". The dataset covers four tasks:
Pour Water
Arrange Mouse
Fold Towel
Unplug Charger
Dataset Folders:
arrange_mouse and pour_water: Each folder contains data… See the full description on the dataset page: https://huggingface.co/datasets/Fanqi-Lin/GoPro-Raw-Videos.code-world-model-project-page-videos
Code World Model Project Page Videos
Public research-demo video assets used by the Code World Model project page.
The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
VideoChat3-LV116k
VideoChat3-LV116K
VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments.
The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.ResearchData_P1qvhighlights-videos
QVHighlights Videos
All 25124 videos from the QVHighlights benchmark (train + val + test splits).
Total size: 135.3 GB.
Layout
Files are sharded into subdirectories by the first character of the filename
(HuggingFace caps each directory at 10,000 files):
<first-char>/<youtube-id>_<start>_<end>.mp4
Source
Lei et al., "QVHighlights: Detecting Moments and Highlights in Videos via Natural
Language Queries" (NeurIPS 2021).
Original archive:… See the full description on the dataset page: https://huggingface.co/datasets/ayushsdev/qvhighlights-videos.video-to-data-robot-dexterity-task-library-and-dataset
Video to Data: Robot Dexterity Task Library and Dataset
Dataset Description
This dataset contains samples of human demonstrations on manipulation tasks retargeted to bimanual Sharpa robot hands and episodes of robot executions that mimic the original human demonstrations. The former allows a Video to Data user to easily experiment with the Video to Data grounding pipeline, and the latter is an example of the grounded robot data that can be generated with the Video… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-to-data-robot-dexterity-task-library-and-dataset.Egoverse_videosVBench_sampled_video
VBench Sampled Video
test-videosVBench-2.0_sampled_videos
Sample Videos of VBench-2.0
This dataset is used in the paper:👉 arXiv:2503.21755
hoigen-filtered-videos
HOIGen Filtered Videos Dataset
This dataset contains 28562 filtered videos from the HOIGen-1M dataset based on the allowlist.
Dataset Structure
The videos are organized in the same structure as the original HOIGen dataset:
filtered_videos/
├── videos_part_1/
├── videos_part_2/
├── ...
└── videos_part_100/
Usage
from huggingface_hub import hf_hub_download
# Download a specific video
video_path = hf_hub_download(… See the full description on the dataset page: https://huggingface.co/datasets/charlychan123/hoigen-filtered-videos.omni-refiner-videovideoTotal_Editing_Synthetic_Video_Albedo_Fullbdd100k_videos
BDD100K videos
This repository stores the original bdd100k_videos.zip as byte-for-byte split files. The archive is not recompressed. Refer to the BDD100K license and terms before using or redistributing the data.
Restore
Download all bdd100k_videos.zip.part-* files from bdd100k_videos/, then run:
cat bdd100k_videos.zip.part-* > bdd100k_videos.zip
md5sum -c bdd100k_videos.zip.md5
Source MD5: 253d9a2f9d89d2b09d8d93f397aecdd7. There are 19 parts of up to 100.00 GiB… See the full description on the dataset page: https://huggingface.co/datasets/linxxx3/bdd100k_videos.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.Vera-Layered-Video-Dataset
Dataset for Vera: A Layered Diffusion Model for Content-Preserving Video Editing
Hongkai Zheng¹²* ·
Ta-Ying Cheng² ·
Benjamin Klein² ·
Yisong Yue¹ ·
Zhuoning Yuan²†
¹California Institute of Technology ²Netflix, Inc.
*Work done during an internship at Netflix †Project Lead
TL;DR: A layered diffusion framework for video editing. Vera jointly generates an edit layer, an alpha… See the full description on the dataset page: https://huggingface.co/datasets/netflix/Vera-Layered-Video-Dataset.gs-videos-v2Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.Grounded-VideoLLMvideo-data-1
