datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Egocentric-100K
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here.
Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
100,405
Total Frames
10.8 billion
Video Clips
2,010,759
Median Clip Length
180.0 seconds
Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.robotwin2.0-fastwam
RobotWin 2.0 (Preprocessed LeRobot v2.1 Release)
This repository releases our preprocessed RoboTwin / RobotWin 2.0 dataset in LeRobot v2.1 format for the open-source release of Fast-WAM: Do World Action Models Need Test-time Future Imagination?
This is not the official upstream RoboTwin release. It is our paper-specific processed version prepared to support training, evaluation, and reproducibility for our project.
This Hugging Face repository distributes the dataset as split… See the full description on the dataset page: https://huggingface.co/datasets/yuanty/robotwin2.0-fastwam.OpenVid-1M-wds
OpenVid-1M — WebDataset repackaging
This repository is a sequential-read-optimized WebDataset repackaging of nkp37/OpenVid-1M by Nan et al. (ICLR 2025). The video content is identical to the original — only the on-disk layout is changed so it can be streamed efficiently from a single HTTP/NFS connection.
What differs from the original
Aspect
Original nkp37/OpenVid-1M
This repository
Format
Per-video mp4 files zipped
WebDataset .tar shards (~2 GB each)… See the full description on the dataset page: https://huggingface.co/datasets/Dev-Jahn/OpenVid-1M-wds.describe-anything-dataset
Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for Describe Anything Datasets
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.Senorita
Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
If you use Señorita-2M in your research, please cite our work as follows:@article{zi2025senorita,
title={Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists},
author={Bojia Zi and Penghui Ruan and Marco Chen and Xianbiao Qi and Shaozhe Hao and Shihao Zhao and Youze Huang and Bin Liang and Rong Xiao and Kam-Fai Wong}… See the full description on the dataset page: https://huggingface.co/datasets/SENORITADATASET/Senorita.sdoml-lite
SDOML-lite
SDOML-lite is a lightweight alternative to the SDOML dataset specifically designed for machine learning applications in solar physics, providing continuous full-disk images of the Sun with magnetic field and extreme ultraviolet data in several wavelengths. The data source is the Solar Dynamics Observatory (SDO) space telescope, a NASA mission that has been in operation since 2010.
NASA’s SDO mission has generated over 20 petabytes of high-resolution solar imagery… See the full description on the dataset page: https://huggingface.co/datasets/oxai4science/sdoml-lite.sign-bibles
bible-nlp/sign-bibles
This dataset is still being generated and currently includes only test files
This dataset contains sign language videos from the Digital Bible Library (DBL), processed for machine learning applications. The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
Dataset Structure
Each sample contains:
["mp4"] the original video
["json"] Metadata, including bible reference, copyright information… See the full description on the dataset page: https://huggingface.co/datasets/bible-nlp/sign-bibles.SCARED-C
Dataset Card for SCARED-C
SCARED-C is a corrected version of the SCARED endoscopic depth estimation dataset. By replacing the original kinematics-based camera poses with poses re-estimated through Structure-from-Motion (COLMAP) followed by a metric scale recovery step, SCARED-C expands the number of reliable RGB-D pairs in SCARED from 35 keyframes to 17,135 frames — a roughly 490× increase in reliably labeled real-tissue surgical data.
Left: original SCARED depth maps misaligned… See the full description on the dataset page: https://huggingface.co/datasets/juseonghan/SCARED-C.DL3DV-Evaluation
DL3DV Testing Split Download Instructions
This repo contains all 55 scenes for evaluation. Note: it is an independent dataset, and none of its scenes overlap with those in DL3DV-10K. Have a galance on the preview page: https://dl3dv-10k.github.io/DL3DV-Testing-Split-Preview/.
Download
As the whole benchmark dataset is ~500G, a python script to download and untar files.
Environment Setup
The download script relies on huggingface hub, tqdm. You can download by… See the full description on the dataset page: https://huggingface.co/datasets/DL3DV/DL3DV-Evaluation.SportsAction
Dataset Card for MultiSports
Dataset Summary
Spatio-temporal action localization is an important and challenging problem in video understanding. Previous action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions. MultiSports is a multi-person dataset of spatio-temporal localized sports actions. Please refer to this paper for more details. Please refer to this repository for evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/SportsAction.asdftrain_raw_video
ShareGPTVideo Raw ActivityNet Videos for Train data
All dataset and models can be found at ShareGPTVideo.
Contents:
Due to our scene split, we provide our processed activityNet videos corresponding to test frames in
train video frames
the processing script is process_activitynet.py
fastwam-robotwin2-tarballs
RobotWin 2.0 (Preprocessed LeRobot v2.1 Release)
This repository releases our preprocessed RoboTwin / RobotWin 2.0 dataset in LeRobot v2.1 format for the open-source release of Fast-WAM: Do World Action Models Need Test-time Future Imagination?
This is not the official upstream RoboTwin release. It is our paper-specific processed version prepared to support training, evaluation, and reproducibility for our project.
This Hugging Face repository distributes the dataset as split… See the full description on the dataset page: https://huggingface.co/datasets/huiliu123/fastwam-robotwin2-tarballs.sign-dictionary-isl
Dataset Card for Sign Dictionary Dataset
This dataset contains Indian sign language videos with one gloss per video. There are 3077 seperate lex items or glosses included.
The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
Dataset Details
There is a total of 2.5 hours of sign videos.
How to use
import webdataset as wds
import numpy as np
import json
import tempfile
import os
import cv2
def… See the full description on the dataset page: https://huggingface.co/datasets/bridgeconn/sign-dictionary-isl.navidromeopen-sora-pexels-subset
Open-Sora Pexels Dataset (Captioned Only)
A curated subset of the Pexels videos from LanguageBind/Open-Sora-Plan-v1.1.0, converted to layered WebDataset format. Every video has at least one caption.
Dataset Summary
Statistic
Value
Total Videos
9,750
Total Caption Entries
31,910
Captions from 513f source
4,452
Captions from 65f source
27,458
Video Shards
~120
Total Size
~120 GB
Caption Sources
Captions are merged from two Open-Sora… See the full description on the dataset page: https://huggingface.co/datasets/zengxianyu/open-sora-pexels-subset.wikivideo
Paper and Code
Associated with the paper: WikiVideo (https://arxiv.org/abs/2504.00939)
Associated with the github repository (https://github.com/alexmartin1722/wikivideo)
Download instructions
The dataset can be found on huggingface. However, you can't use the datasets library to access the videos because everything is tarred. Instead you need to locally download the dataset and then untar the videos (and audios if you use those).
Step 1: Install git-lfs
The… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/wikivideo.test_raw_video_data
ShareGPTVideo Raw Videos for Testing data
All dataset and models can be found at ShareGPTVideo.
Contents:
In case of need, this contains raw videos corresponding to test frames in
Test video frames
MCoRec
CHiME-9 Task 1: Multi-Modal Context-aware Recognition (MCoRec)
The MCoRec dataset contains video recording sessions. A single recording session typically features multiple conversations, each involving two or more participants. A 360° camera, placed in the middle of the room, captures the central view which contains all participants. The audio for the sessions is also captured by the microphone integrated into the 360° camera. Each session can involve a maximum of 8 active speakers… See the full description on the dataset page: https://huggingface.co/datasets/MCoRecChallenge/MCoRec.IllumiCraft
IllumiCraft Dataset
This repository contains the dataset released with:
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark, Ming-Hsuan Yang
🔗 Links
📄 Paper: https://arxiv.org/abs/2506.03150
🌐 Project Page: https://yuanze-lin.me/IllumiCraft_page/
💻 GitHub: https://github.com/yuanze-lin/IllumiCraft
🎥 YouTube: https://youtu.be/qAV58sADEzo
🤗 Checkpoints:… See the full description on the dataset page: https://huggingface.co/datasets/YuanzeLin/IllumiCraft.StorageFisheye-example
Dataset Summary
Underwater video frames of river herring with bounding-box annotations.
I2V-pairs-m1xm3
I2V Pairs (m1xm3)
Paired image-to-video generations for preference / DPO research.
Each entry contains two videos generated from the same source image (first frame
of source_video) and same caption, but with two different sampling configs
(method m1: CFG=7, steps=15 vs method m3: CFG=3, steps=50) on Wan2.2-TI2V-5B @ 720p.
Stats
1284 clean pairs (after filtering for unreadable / static / first-frame-drift)
Source videos: OpenVid (parts 112-113), 5s clips at 720p… See the full description on the dataset page: https://huggingface.co/datasets/qgfvadfuvads/I2V-pairs-m1xm3.mario_data
📘 Mario Gameplay Dataset (522k Frames)
This dataset contains 522,000 gameplay frames from Super Mario Bros Level 1-1, designed specifically for generative modeling, diffusion models, and action-conditioned video generation.
The dataset provides raw frames, per-frame action labels, and alive/death markers, enabling both vision-only and action-conditioned modeling.
Project code: https://github.com/zhoufeiyn/ltpclean
Game environment: https://github.com/zhoufeiyn/super-mario1_1env… See the full description on the dataset page: https://huggingface.co/datasets/FeiyanZhou/mario_data.Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos-Transfer1-7B-Sample-AV-Data-Example
Cosmos | Code | Paper | Paper Website
Dataset Description:
This dataset contains 10 sample data points intended to help users better utilize our Cosmos-Transfer1-7B-Sample-AV model. It includes HD Map annotations and LiDAR data, with no personally identifiable information such as faces or license plates. This dataset is intended for research and development only.
Dataset Owner(s):
NVIDIA
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Transfer1-7B-Sample-AV-Data-Example.sign-dictionary-isl
Dataset Card for Sign Dictionary Dataset
This dataset contains Indian sign language videos with one gloss per video. There are 3077 seperate lex items or glosses included.
The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
Dataset Details
There is a total of 2.5 hours of sign videos.
How to use
import webdataset as wds
import numpy as np
import json
import tempfile
import os
import cv2
def… See the full description on the dataset page: https://huggingface.co/datasets/Bharath1330/sign-dictionary-isl.k600_test_ds
K600 Frozen Test Clips (for reproducible evaluation)
This repository provides a frozen set of extracted clips used for evaluation in the paper:
Paper
https://huggingface.co/papers/2511.20928
https://arxiv.org/abs/2511.20928
Why this dataset exists
Kinetics videos are hosted on YouTube, and availability can change over time. This dataset exists to provide a stable evaluation set aligned with the paper’s experiments.
What’s included… See the full description on the dataset page: https://huggingface.co/datasets/DrGil/k600_test_ds.connan_30k
Connan 30K Video Caption Dataset
Private Hugging Face backup of 30,617 short anime video clips and English visual captions.
Layout
metadata/manifest.parquet # sample index, captions, original paths, shard/member mapping
wds/*.tar # WebDataset shards, about 10GB each
dataset_info.json # summary metadata
Each WebDataset sample uses a stable key such as 001_shot_00003 and contains:
001_shot_00003.mp4
001_shot_00003.json
The JSON sidecar… See the full description on the dataset page: https://huggingface.co/datasets/hdhacker/connan_30k.sign-dictionary-isl
Dataset Card for Sign Dictionary Dataset
This dataset contains Indian sign language videos with one gloss per video. There are 3077 seperate lex items or glosses included.
The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0).
Dataset Details
There is a total of 2.5 hours of sign videos.
How to use
import webdataset as wds
import numpy as np
import json
import tempfile
import os
import cv2
def… See the full description on the dataset page: https://huggingface.co/datasets/lakshmi0110/sign-dictionary-isl.triangle
