datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.mesh4d_datasetgpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.ai4g-flood-dataset
Flood Detection Dataset
Introduction
This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide.
Key features:
Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.RGB-Event-ISP-DatasetCRA5-Dataset
Climate science data can be compressed efficiently by dual-stage extreme compression with a variational auto-encoder transformer
Introduction and get started
CRA5 dataset now is available at OneDrive
Paper Summary
We introduce VAEformer, a variational autoencoder transformer designed for the extreme compression of climate data. Addressing the storage challenges of massive datasets like ERA5, VAEformer utilizes a… See the full description on the dataset page: https://huggingface.co/datasets/taohan10200/CRA5-Dataset.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.nerf-gs-datasetsI keep a collection compiled of existing datasets from various sources for training NeRFs or Splats. This dataset is most of that collection. All of the individual scenes also have a trained Gaussian Splat.
https://rishit-dagli.github.io/2025/03/28/nerf-gs-datasets.html
cocoffhq-datasetFlickr-Faces-HQ Dataset (FFHQ) dataset: https://github.com/NVlabs/ffhq-dataset
The dataset consists of 70,000 high-quality PNG images at 1024×1024 resolution and contains considerable variation in terms of age, ethnicity and image background.
It also has good coverage of accessories such as eyeglasses, sunglasses, hats, etc. The images were crawled from Flickr, thus inheriting all the biases of that website, and automatically aligned and cropped using dlib.
Only images under permissive… See the full description on the dataset page: https://huggingface.co/datasets/marcosv/ffhq-dataset.MolmoAct-DatasetThis dataset was created using LeRobot.
Dataset Description
This dataset contains MolmoAct Dataset in lerobot format. All contents in this dataset were collected in-house by Ai2.
Quick links:
📂 All Models
📂 All Data
📃 Paper
🎥 Blog Post
🎥 Video
Code
License and Use
This dataset is licensed under CC BY-4.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{molmoact2025… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Dataset.nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.mv2-dataset
MV2 Dataset
MV2 is a multi-view and multi-vehicle urban driving dataset designed for research in novel view synthesis, neural rendering, 3D reconstruction, cross-view scene understanding, and autonomous-driving perception. The dataset contains synchronized image sequences captured from multiple viewpoints, including ground vehicles and aerial views, along with camera parameters required for geometry-aware learning and rendering.
The dataset is released for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/sanjay810/mv2-dataset.fashion_mnist
Dataset Card for FashionMNIST
Dataset Summary
Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. We intend Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for benchmarking machine learning algorithms. It shares the same image size and structure of training and testing… See the full description on the dataset page: https://huggingface.co/datasets/zalando-datasets/fashion_mnist.KAIST-Multispectral-Pedestrian-Detection-Datasetc2p-dataset
Concept-to-Pixel Dataset
Dataset Summary
This repository hosts the prepared dataset package used by Concept-to-Pixel (C2P), a prompt-free universal medical image segmentation framework associated with the ECCV 2026 paper:
Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation
The repository is intended as a ready-to-use training package for the official C2P codebase rather than a single cleaned benchmark export. In addition to images and masks, it… See the full description on the dataset page: https://huggingface.co/datasets/yunyundi/c2p-dataset.conceptual_captions
Dataset Card for Conceptual Captions
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.CASTLE2024
What is CASTLE?
The CASTLE dataset is a large-scale, multimodal dataset designed for advancing research in lifelogging, human activity recognition, and multimodal retrieval. It provides a rich collection of time-aligned sensor and video data for analysis and benchmarking. See the Paper (or its arXiv pre-print) for more details.
You can check our website for more details.
Characteristics
Captured over four days in a controlled environment
10 participants engaged… See the full description on the dataset page: https://huggingface.co/datasets/CASTLE-Dataset/CASTLE2024.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.sample-datasetsimg2cad-dataset
Img2CAD Annotated CAD Dataset
This dataset contains annotated CAD models with semantic part information for training CAD reverse engineering models.
Dataset Description
The Img2CAD Annotated CAD Dataset is a comprehensive collection of 3D CAD models with:
Raw annotated CAD data in HDF5 format
Rendered images from multiple viewpoints
Semantic part labels and annotations
Train/test splits for three furniture categories
Categories
Chair: ~2,000+ annotated chair… See the full description on the dataset page: https://huggingface.co/datasets/qq456cvb/img2cad-dataset.multimodal_data_annotator_datasetMaterials dataset consisting of spatial and time resolved versions of the same object. Specially curated for the annotator such that for each object, time resolved signal may be viewed alongside the RGB and for different graphs/forms
jojokanbao-dataset
Marxism Dataset
中文书籍与报刊数字馆藏。
书籍镜像已与 B2 当前发布内容对齐:58 个书目、127 个 Item,使用可读 Dataset ID。《毛泽东年谱》统一为一个书目。
书籍目录
程序目录
同步与来源记录
人民日报数据
其他报刊数据
书籍正文、注释与媒体来自经过哈希校验的 B2 Delivery;未补造 Delivery 中不存在的原始导入元数据。报刊与人民日报原始 PDF 在本次书籍同步中保持不变。
v1.0
MuLAn: : A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
MuLAn is a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multilayer, instance-wise RGBA decompositions, and over 100K instance images. It is composed of MuLAn-COCO and MuLAn-LAION sub-datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance… See the full description on the dataset page: https://huggingface.co/datasets/mulan-dataset/v1.0.ViFailback-Dataset
ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance
A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols.
ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/ViFailback-Dataset.Obstacle-Detection-Dataset-YOLO
ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision
24,326-image, 25-class YOLO dataset for obstacle detection
This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.regression-dataset-for-docling-parse
Regression Dataset for docling-parse
This repository contains the reference dataset used as a regression test corpus for
docling-parse.
Its purpose is to make parser and renderer changes safe: when behavior changes in
docling-parse, the test
suite can compare the current output against the expected artifacts stored in this
dataset.
Correct workflow to add new files
cp /path/to/new.pdf regression/new.pdf
git add regression/new.pdf
git lfs status
git commit -s -m… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/regression-dataset-for-docling-parse.goldsrc-models-datasetemit-test-dataset
Dataset Card for EMIT-MSeg Dataset
If you use this dataset, please cite our article:
@misc{herec2026fastmethanedetectionpipeline,
title={A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet},
author={Jonáš Herec and Vít Růžička and Rado Pitoňák and Jan Sedmidubsky},
year={2026},
eprint={2606.03675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.03675},
}… See the full description on the dataset page: https://huggingface.co/datasets/onboard-coop/emit-test-dataset.
