datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Egocentric-100K
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here.
Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
100,405
Total Frames
10.8 billion
Video Clips
2,010,759
Median Clip Length
180.0 seconds
Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.AgiBotWorld-Beta
Key Features 🔑
1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours.
100+ real-world scenarios across 5 target domains.
Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots
200+ types of tasks:
Contact-rich manipulation
Long-horizon planning
Multi-robot collaboration
87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc.
Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.Open-Sora-Plan-v1.1.0
Annotation
We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973
Pexels
Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
🎉NEWS
[2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance!
[2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026!
[2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens.
[2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.hot3d
HOT3D-Clips
This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset.
Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here.
See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.).
More details can be found in the HOT3D paper and BOP 2024 report.
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse
Vchitect Team1
1Shanghai Artificial Intelligence Laboratory
Paper |
Project Page |
Data Overview
The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.
It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.wds_objectnetEmilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.voxbox
VoxBox
This dataset is a curated collection of bilingual speech corpora annotated clean transcriptions and rich metadata incluing age, gender, and emotion.
Dataset Structure
.
├── audios/
│ └── aishell-3/ # Audio files (organised by sub-corpus)
│ └── ...
└── metadata/
├── aishell-3.jsonl
├── casia.jsonl
├── commonvoice_cn.jsonl
├── ...
└── wenetspeech4tts.jsonl # JSONL metadata files
Each JSONL file corresponds to a… See the full description on the dataset page: https://huggingface.co/datasets/SparkAudio/voxbox.cc12m-wds
Dataset Card for Conceptual Captions 12M (CC12M)
Dataset Summary
Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training.
Its data collection pipeline is a relaxed version of the one used in Conceptual Captions 3M (CC3M).
Usage
This instance of Conceptual Captions is in webdataset .tar format. It can be used with webdataset library or upcoming releases of Hugging Face datasets.… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc12m-wds.RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.WildGUI
WildGUI
This repository hosts a personally reprocessed annotation release for WildGUI, the dataset introduced by Video2GUI.
The original Video2GUI project builds WildGUI from large-scale Internet tutorial videos for GUI agent pretraining. This repository focuses on the open annotation artifacts: the records were regenerated and cleaned following the full annotation workflow, then reformatted to make the data easier to inspect, reuse, and reproduce. It also ships the screenshot… See the full description on the dataset page: https://huggingface.co/datasets/xwm/WildGUI.yodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.AgiBotWorld-Alpha
⚠️Important Notice !!!
Dear Users,
The Alpha Dataset has been updated as follows:
Frame Loss Data Removal: Several episodes with frame loss issues have been removed. For the complete list of removed episode IDs, please refer to this document.
Changes in Episode Count: The updated Alpha Dataset retains the original 36 tasks. The new version has been enriched with additional interactive objects, extending the total duration from 474.12… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Alpha.PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.SekaiTimeLens-100K
TimeLens-100K
📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data
✨ Dataset Description
TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro.
📊 Dataset Statistics
Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.fine-t2i
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv]
by Xu Ma, Yitian Zhang,
Qihua Dong, Yun Fu
Northeastern Univeristy
Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient).
🆕 What's New
[2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️
[2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.MINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.BEDLAM-depth
Dataset Mirror of BEDLAM Dataset (Depth Data Subset)
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM
Dataset Information
Depth maps (EXR, 32-bit, 3.8TB)
Camera ground truth information is not included but can be found in the BEDLAM dataset mirror
Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.wds_imagenet_sketchLAION-Audio-300MMINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.soundscapesDuplexConv
DuplexConv
DuplexConv is a large-scale Chinese multi-channel conversational speech dataset with LLM-assisted annotations, developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family.
Companion dataset: SmoothConv on HuggingFace (100 hours, expert human annotation). DuplexConv and SmoothConv share the same conversational domains and a unified data design. SmoothConv focuses on high-quality human annotations for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/qualialabsAI/DuplexConv.imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
cc3m-wds
Dataset Card for Conceptual Captions (CC3M)
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc3m-wds.ImageNetV23d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.OVO-Bench
OVO-Bench: How Far is Your Video-LLMs from Real-World Online VideO Understanding?
🔥🔥OVO-Bench is accepted by CVPR 2025!🔥🔥
Important Note: Current codebase is modified compared to our initial arXiv paper. We strongly recommend that any use of OVO-Bench should be based on current edition.
Introduction
🌟 Three distinct problem-solving modes
Backward Tracing: trace back to past events to answer the question.Real-Time Visual Perception:… See the full description on the dataset page: https://huggingface.co/datasets/JoeLeelyf/OVO-Bench.
