datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
🎉NEWS
[2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance!
[2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026!
[2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens.
[2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.wds_imagenet_sketchBEDLAM-depth
Dataset Mirror of BEDLAM Dataset (Depth Data Subset)
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM
Dataset Information
Depth maps (EXR, 32-bit, 3.8TB)
Camera ground truth information is not included but can be found in the BEDLAM dataset mirror
Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
ImageNetV2T2I-CoReBench-Images
T2I-CoReBench-Images
📖 Overview
T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.wds_imagenet-rwds_imagenet-ags-images-v2wds_imagenetv2GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imagenet-22k-wds
Dataset Summary
This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds)
This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.wds_imagenetcwds_imagenet1kTreeOfLife-10M
Dataset Card for TreeOfLife-10M
Dataset Summary
With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.BEDLAM2-depth
Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset)
Project site: https://bedlam2.is.tuebingen.mpg.de/
Please register at project site for additional information and data in its Download section.
Related Hugging Face dataset mirror: BEDLAM2
Dataset Information
Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB)
Multilayer EXR details
16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R)
Color image without motion blur
Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.imagenet-12k-wds
Dataset Summary
This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm.
The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k
NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.idl-wds
Dataset Card for Industry Documents Library (IDL)
Dataset Summary
Industry Documents Library (IDL) is a document dataset filtered from UCSF documents library with 19 million pages kept as valid samples.
Each document exists as a collection of a pdf, a tiff image with the same contents rendered, a json file containing extensive Textract OCR annotations from the idl_data project, and a .ocr file with the original, older OCR annotation. In each pdf, there may be from 1 to up… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/idl-wds.Xeno-Canto-6s-16khz
Xeno-Canto Bird Sound Dataset
This repository provides access to the Xeno-Canto bird sound dataset (checkpoint from 2022-07-18) used in the benchmark BIRB, specifically pre-processed to facilitate training deep learning models.
The dataset has been processed using CNN14 from PANNs, a model pre-trained on AudioSet, to select 6-second windows with the highest bird sound activation. All audio has been downsampled to 16kHz and converted into Pytorch format (.pt), optimizing it for… See the full description on the dataset page: https://huggingface.co/datasets/ilyassmoummad/Xeno-Canto-6s-16khz.ET-Instruct-164K
E.T. Instruct 164K
arXiv | Project Page | GitHub
E.T. Instruct 164K is a large-scale instruction-tuning dataset tailored for fine-grained event-level and time-sensitive video understanding. It contains 101K meticulously collected videos under diverse domains and 9 event-level understanding tasks with well-designed instruction-response pairs. The average video length is around 146 seconds.
📦 Download Dataset
You may download the dataset using the following command.… See the full description on the dataset page: https://huggingface.co/datasets/PolyU-ChenLab/ET-Instruct-164K.ilias
ILIAS is a large-scale test dataset for evaluation on Instance-Level Image retrieval At Scale. It is designed to support future research in image-to-image and text-to-image retrieval for particular objects and serves as a benchmark for evaluating representations of foundation or customized vision and vision-language models, as well as specialized retrieval techniques.
website | download | arxiv | github
Composition
The dataset includes 1,000 object instances across… See the full description on the dataset page: https://huggingface.co/datasets/vrg-prague/ilias.open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.IRRISIGHT
IRRISIGHT
IRRISIGHT is a large-scale multimodal dataset to address water availability problems in agriculture. It is designed to support supervised and semi-supervised learning tasks related to agricultural water use monitoring.
Due to the space constraints, we uploaded the files across multiple repositories as follows:
To download Pennsylvania and Maryland, use the current repository (OBH30/IRRISIGHT).
To download Arizona, Arkansas, Florida, Georgia, New Jersey, North Carolina… See the full description on the dataset page: https://huggingface.co/datasets/OBH30/IRRISIGHT.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.biggest-ru-bookA bigger version of its5Q/bigger-ru-book, the smaller set being a subset of this one. Almost 1000 hours of high-quality audio.
T2I-ImageNet-Normal3D_Visual_Illusion_Depth_Estimation
3D Visual Illusion Depth Estimation Dataset
Dataset Summary
The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks.
Dataset Structure
Each sample in the dataset includes:
left: Left-view RGB image
right: Right-view RGB image
depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.
