datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Egocentric-100K
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here.
Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
100,405
Total Frames
10.8 billion
Video Clips
2,010,759
Median Clip Length
180.0 seconds
Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.EgoDemo
EgoDemo
A 50-hour sample from EgoSuite-Open100K, covering every annotated subset plus two raw-video variants.
Collection ·
EgoStandard ·
EgoPro ·
Project page
Explore EgoSuite-Open100K ↗
EgoSuite-Open100K Overview
Collection:
EgoSuite-Open100K
SKU
Sub-SKU
Format
Planned Duration
EgoStandard
EgoStand
Hand Pose
80,000 h… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/EgoDemo.EgoLifeData cleaning, stay tuned! Please refer to https://egolife-ai.github.io/ first for general info.
Checkout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information.
Code: https://github.com/egolife-ai/EgoLife
ABot-World-Explorer-500h
ABot World Explorer 500h
ABot World Explorer 500h contains 30,969 action-conditioned video episodes
associated with the data infrastructure described in
ABot-World-0. Each episode preserves an MP4,
dataset-native keyboard actions, captions, and one COLMAP text sparse model.
Dataset facts
Item
Value
Episodes
30,969
Source objects
185,814
Semantic splits
None
License
Apache-2.0
The repository name is an identifier, not an audited… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.Video-MMEEgoPro
EgoPro
The 10,000-hour head-and-wrist line of EgoSuite-Open100K.
Data Bucket ·
Collection ·
EgoDemo ·
EgoStandard ·
Project page
Explore EgoSuite-Open100K ↗
Data location: EgoPro is distributed through the LightwheelAI/EgoPro Bucket. This Git repository is the dataset card and access point; download the data from the Bucket.
Overview
EgoPro pairs synchronized head- and wrist-view video with 3D hand pose. Its body subset adds full-body pose. LeRobot and… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/EgoPro.tracker-pov
Eidon Tracker POV
1,274 hours of egocentric video paired with 7-point IMU arm tracking, recorded during ordinary household work.
Contributors wore a head-mounted camera and a seven-sensor IMU harness while doing real chores in their own homes: laundry, cleaning, dishes, cooking. Each recording pairs first-person video with 24 Hz orientation data for both hands, both forearms, both upper arms, and the chest.
This is a complete, final release. Eidon AI (Solidic Labs Inc) has wound… See the full description on the dataset page: https://huggingface.co/datasets/eidon-ai/tracker-pov.X-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.EgoBrain
[ICLR 2026] EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
東京大学 The Univerisity of Tokyo X 微軟亞洲研究院 Microsoft Research Asia
Nie Lin
·
Yansen Wang
·
Dongqi Han
·
Weibang Jiang
·
Jingyuan Li
·
Ryosuke Furuta
·
Yoichi Sato*
·
Dongsheng Li*
·
*(Co-corresponding authors)*
This is the official dataset repository of our ICLR 2026 paper "EgoBrain: Synergizing Minds and Eyes For Human Action… See the full description on the dataset page: https://huggingface.co/datasets/ut-vision/EgoBrain.EBench-Dataseteval-resultsEPIC-KITCHENShumanoid-everyday
Humanoid Everyday
A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
Overview
Humanoid Everyday is a large-scale, diverse humanoid manipulation dataset designed for open-world robotic learning and embodied intelligence.
It contains over 260 tasks across 7 major categories, covering dexterous manipulation, human–humanoid interaction, and locomotion-integrated activities.All data were collected through a human-supervised teleoperation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/USC-PSI-Lab/humanoid-everyday.epic_kitchens_100
Motivation
The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here.
Source
You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts
Citation
@INPROCEEDINGS{Damen2018EPICKITCHENS,
title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset},
author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/awsaf49/epic_kitchens_100.robocasa365-pretrain-mg
Pretraining (MimicGen) — atomic
MimicGen-generated rollouts across 60 atomic tasks (~10,000 demos/task). 1,615 hours total, generated by scripted augmentation from human demonstrations.
Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable.
Stats
Episodes: 536,030
Frames: 116,246,439 (20 fps → 1615 h)
Tasks: 720 (natural-language phrasings; underlying RoboCasa task classes: 60)
Cameras: 3 × 256×256 h264… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-mg.A12d12s12short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.g1-moves
G1 Moves
Dataset (you are here) · Showcase (interactive gallery) · Code (scripts & docs)
60 motion capture clips for the Unitree G1 humanoid robot (edition EDU, 29 DOF), captured from real performers in Austin, TX using MOVIN TRACIN markerless motion capture and video2robot monocular video extraction. Each clip is provided at multiple pipeline stages: raw mocap (BVH/FBX), retargeted robot joint trajectories (PKL), processed RL training data (NPZ), and trained ONNX… See the full description on the dataset page: https://huggingface.co/datasets/exptech/g1-moves.AVUTBenchmark
Audio-centric Video Understanding Benchmark (AVUT)
This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut.
Code Repository: https://github.com/lark-png/AVUT
Paper: https://arxiv.org/pdf/2503.19951
Introduction
The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/AVUTBenchmark.egoscaler-v2
EgoScalerV2 Dataset
This dataset accompanies our work on Developing Vision-Language-Action Model from Egocentric Videos. It provides 6DoF object trajectories paired with egocentric visual observations and natural-language action descriptions, formatted in the LeRobot v2.0 schema so it can be consumed directly by LeRobot-compatible pipelines.
🌐 Project page: https://biscue5.github.io/egovla-project-page/
📄 Paper: Developing Vision-Language-Action Model from Egocentric Videos… See the full description on the dataset page: https://huggingface.co/datasets/Biscue5/egoscaler-v2.epic_kitchens_100
EPIC-KITCHENS-100 (Mirror)
This repository provides a mirror of the EPIC-KITCHENS-100 dataset videos for easier access and high-speed downloading via the Hugging Face Hub.
Important Note
This mirror is uploaded for personal convenience and may not contain the entire dataset. If you need the full, official, and most up-to-date version of the dataset (including all annotations and subsets), please visit the official website.
Official Links
Official… See the full description on the dataset page: https://huggingface.co/datasets/a1raman/epic_kitchens_100.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.EgoSteer-RealWorld
EgoSteer Real-World Bimanual Teleoperation Dataset
EgoSteer-RealWorld is the real-robot dataset collected and used in
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos.
It contains 54,454 teleoperated episodes (192 hours, 20.75 M frames) of bimanual dexterous manipulation across 193 tasks,
recorded on a RealMan dual-arm robot with two Ruiyan dexterous hands and two RGB-D cameras (head and chest), with… See the full description on the dataset page: https://huggingface.co/datasets/EgoSteer/EgoSteer-RealWorld.ExpVid
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
We present ExpVid, a benchmark to evaluate MLLMs on scientific experiment videos. ExpVid comprises 10 tasks across 3 levels, curated from a collection of 390 lab experiment videos spanning 13 disciplines.
How to Use
from datasets import load_dataset
dataset = load_dataset("OpenGVLab/ExpVid")
All task annotation .jsonl files are stored under annotations/level_*.
Each annotation includes the field:… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/ExpVid.EgoLife_EyeTracking_EyeGazevitra-ego4d-video2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/elonelonelon/2025-challenge-demos.PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card
Dataset Description
PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.egodex_custom
