datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LAION-Audio-300Mcc3m-wds
Dataset Card for Conceptual Captions (CC3M)
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc3m-wds.3d_optical_flow_droid
3D Optical Flow DROID Dataset
Processed DROID robotics dataset with optical flow and scene flow annotations.
Dataset Structure
Organized by lab, each trajectory in separate tar.gz archive:
IPRL/IPRL+2023-06-19+Mon_Jun_19_23:27:48_2023.tar.gz
CLVR/CLVR+2023-...tar.gz
... (15 labs, ~33K trajectories)
Each trajectory contains:
metadata.json - Trajectory metadata
trajectory.h5 - Robot state and actions
camera_left/, camera_right/ - Camera data
rgb/ - RGB images
depth/ -… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/3d_optical_flow_droid.gs-videos-v3laion-3m
WebDataset shards
Each sample contains only:
{image_hash}.png
{image_hash}.json
RevealLayer-100K
RevealLayer Open Dataset
RevealLayer Open is the open-source dataset accompanying RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition.
Paper: https://arxiv.org/html/2605.11818v1 Accepted by ICML 2026
RevealLayer studies box-guided layered image decomposition for natural images. Given an RGB image and instance bounding boxes, the task is to decompose the scene into a clean background and object-level foreground layers, where each… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/RevealLayer-100K.marianne_pdf_3emo_webds_23DCoMPaT200
3DCoMPaT200 Dataset
The 3DCoMPaT200 dataset is a comprehensive collection of 3D objects with compositional part annotations. This repository contains various formats and versions of the dataset organized for different use cases.
📁 Directory Structure
2D Folder
Contains train, validation, and test data in tar format for 10 compositions:
Training set
Validation set
Test set
Each file contains 2D representations of the objects with their corresponding… See the full description on the dataset page: https://huggingface.co/datasets/CoMPaT/3DCoMPaT200.IRRISIGHT
IRRISIGHT
IRRISIGHT is a large-scale multimodal dataset to address water availability problems in agriculture. It is designed to support supervised and semi-supervised learning tasks related to agricultural water use monitoring.
Due to the space constraints, we uploaded the files across multiple repositories as follows:
To download Pennsylvania and Maryland, use the current repository (OBH30/IRRISIGHT).
To download Arizona, Arkansas, Florida, Georgia, New Jersey, North Carolina… See the full description on the dataset page: https://huggingface.co/datasets/OBH30/IRRISIGHT.ophnet_3demo_parleremo_webdscc3mwds_sun397CounterStrike-1K-360-wds
CounterStrike-1K — 360p WebDataset shards
This repo contains the 360p shards of CounterStrike-1K. Use the main repo to browse the manifest, schema, and subsets.
360p is the recommended resolution for most training pipelines — the actions/state/events/metadata sidecars are identical to the 720p shards, so you can swap resolutions without touching downstream code.
Quickstart
Start a fresh uv project and add the loader:
mkdir cs1k-demo && cd cs1k-demo
uv init
uv add… See the full description on the dataset page: https://huggingface.co/datasets/ArnieRamesh/CounterStrike-1K-360-wds.3D_Visual_Illusion_Depth_Estimation
3D Visual Illusion Depth Estimation Dataset
Dataset Summary
The 3D Visual Illusion Depth Estimation Dataset is designed for research on stereo and monocular depth estimation in 3D visual illusion scenes.It contains left and right stereo images, depth maps estimated from DepthAnything V2, and illusion-region masks.
Dataset Structure
Each sample in the dataset includes:
left: Left-view RGB image
right: Right-view RGB image
depth: Monocularly estimated depth… See the full description on the dataset page: https://huggingface.co/datasets/AdamYao/3D_Visual_Illusion_Depth_Estimation.qwen3-tts-multilingual-emotional-speechhg38_cactus447wayMemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
ipapack_plus_train_3ConsistCompose3M
ConsistCompose3M: A 3M-Scale Dataset for Unified Multimodal Layout Control in Image Composition
Overview
ConsistCompose3M is a large-scale dataset (~3M samples) dedicated to layout-controllable multi-instance image composition, with significant improvements in scale, diversity, quality and adaptability. It provides millions of diverse multi-instance scenes, identity-preserving samples filtered by CLIP/DINO similarity, and structured spatial-semantic supervision… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/ConsistCompose3M.emo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
ms1mv3-wds
MS-Celeb-1M (v3)
This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper.
There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.UAV123parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.VideoTemp-o3 VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
Illustration of the agentic pipeline in VideoTemp-o3. Given a video QA pair, the model performs on-demand temporal grounding to locate the most relevant segment, then refines it iteratively. Finally, it produces a reliable answer grounded in the pertinent visual evidence.
Data Source
The question and answer pairs used for training VideoTemp-o3 are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Keye/VideoTemp-o3.CALVIN-3D_PCD-ABC_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABC_D.Emilia-with-Emotion-Annotations3CALVIN-3D_PCD-ABCD_D
| FALCON | From Spatial to Actions:Grounding Vision-Language-Action Model in Spatial Foundation Priors (ICLR 2026)
Zhengshen Zhang
Hao Li
Yalun Dai
Zhengbang Zhu
Lei Zhou
Chenchen Liu
Dong Wang
Francis E. H. Tay
Sijin Chen
Ziwei Liu
Yuxiao Liu*†
Xinghang Li*
Pan Zhou*
*Corresponding Author
†Project Lead… See the full description on the dataset page: https://huggingface.co/datasets/FALCON-VLA/CALVIN-3D_PCD-ABCD_D.
