datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.witLLaVA-OneVision-2-Data
LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training.
At a Glance
The dataset is split across two Hugging Face repositories because of its size:
Repository
What it contains
Part 1 (this repository)
~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.datacomp200m
Datacomp200m
This is a smaller version of the datacomp_1b dataset.
Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows.
The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling.
Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.mesh4d_datasetdatacomp_pools
DataComp Pools
This repository contains metadata files for DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace (https://huggingface.co/terms-of-service), which covers… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_pools.gpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.LLaVA-OneVision-1.5-Instruct-Data
LLaVA-OneVision-1.5 Instruction Data
Paper | Code
📌 Introduction
This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.Honey-Data-15M
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.ai4g-flood-dataset
Flood Detection Dataset
Introduction
This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide.
Key features:
Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.dataGenEvolve-Data-Bench
GenEvolve Data and Bench
This repository contains the open-source data release for GenEvolve:
Config
Directory
Records
Images
Purpose
sft
GenEvolve-Data-SFT/
9,000 trajectories
50,291 reference images
supervised cold-start trajectories
rl
GenEvolve-Data-RL/
3,175 prompts
3,175 GT images
self-evolution / RL training prompts
bench
GenEvolve-Bench/
594 prompts
594 GT images
held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.datacomp_xlarge
DataComp XLarge Pool
This repository contains metadata files for the xlarge pool of DataComp. For details on how to use the metadata, please visit our website and our github repository.
We distribute the image url-text samples and metadata under a standard Creative Common CC-BY-4.0 license. The individual images are under their own copyrights.
Terms and Conditions
We have terms of service that are similar to those adopted by HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/datacomp_xlarge.DataCompDR-1B
Dataset Card for DataCompDR-1B
This dataset contains synthetic captions, embeddings, and metadata for DataCompDR-1B.
The metadata has been generated using pretrained image-text models on DataComp-1B.
For details on how to use the metadata, please visit our github repository.
Dataset Details
Dataset Description
DataCompDR is an image-text dataset and an enhancement to the DataComp dataset.
We reinforce the DataComp dataset using our multi-modal… See the full description on the dataset page: https://huggingface.co/datasets/apple/DataCompDR-1B.PPTAgent-parsed_dataRGB-Event-ISP-DatasetCRA5-Dataset
Climate science data can be compressed efficiently by dual-stage extreme compression with a variational auto-encoder transformer
Introduction and get started
CRA5 dataset now is available at OneDrive
Paper Summary
We introduce VAEformer, a variational autoencoder transformer designed for the extreme compression of climate data. Addressing the storage challenges of massive datasets like ERA5, VAEformer utilizes a… See the full description on the dataset page: https://huggingface.co/datasets/taohan10200/CRA5-Dataset.EO-Data1.5M
🤖 EO-Data-1.5M
A Large-Scale Interleaved Vision-Text-Action Dataset for Embodied AI
The first large-scale interleaved embodied dataset emphasizing temporal dynamics and causal dependencies among vision, language, and action modalities.
📊 Dataset Overview
EO-Data-1.5M is a massive, high-quality multimodal embodied reasoning dataset designed for training generalist robot… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/EO-Data1.5M.LLaVA-OneVision-Data
Dataset Card for LLaVA-OneVision
[2024-09-01]: Uploaded VisualWebInstruct(filtered), it's used in OneVision Stage
almost all subsets are uploaded with HF's required format and you can use the recommended interface to download them and follow our code below to convert them.
the subset of ureader_kg and ureader_qa are uploaded with the processed jsons and tar.gz of image folders.
You may directly download them from the following url.… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data.nerf-gs-datasetsI keep a collection compiled of existing datasets from various sources for training NeRFs or Splats. This dataset is most of that collection. All of the individual scenes also have a trained Gaussian Splat.
https://rishit-dagli.github.io/2025/03/28/nerf-gs-datasets.html
cocoMolmoAct-DatasetThis dataset was created using LeRobot.
Dataset Description
This dataset contains MolmoAct Dataset in lerobot format. All contents in this dataset were collected in-house by Ai2.
Quick links:
📂 All Models
📂 All Data
📃 Paper
🎥 Blog Post
🎥 Video
Code
License and Use
This dataset is licensed under CC BY-4.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Citation
@misc{molmoact2025… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Dataset.ffhq-datasetFlickr-Faces-HQ Dataset (FFHQ) dataset: https://github.com/NVlabs/ffhq-dataset
The dataset consists of 70,000 high-quality PNG images at 1024×1024 resolution and contains considerable variation in terms of age, ethnicity and image background.
It also has good coverage of accessories such as eyeglasses, sunglasses, hats, etc. The images were crawled from Flickr, thus inheriting all the biases of that website, and automatically aligned and cropped using dlib.
Only images under permissive… See the full description on the dataset page: https://huggingface.co/datasets/marcosv/ffhq-dataset.nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/CADS-dataset.mv2-dataset
MV2 Dataset
MV2 is a multi-view and multi-vehicle urban driving dataset designed for research in novel view synthesis, neural rendering, 3D reconstruction, cross-view scene understanding, and autonomous-driving perception. The dataset contains synchronized image sequences captured from multiple viewpoints, including ground vehicles and aerial views, along with camera parameters required for geometry-aware learning and rendering.
The dataset is released for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/sanjay810/mv2-dataset.fashion_mnist
Dataset Card for FashionMNIST
Dataset Summary
Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. We intend Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for benchmarking machine learning algorithms. It shares the same image size and structure of training and testing… See the full description on the dataset page: https://huggingface.co/datasets/zalando-datasets/fashion_mnist.PDE_Inverse_Problem_Benchmarking
PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems
This is the official dataset for the paper PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems.
Code: GitHub - ASK-Berkeley/PDEInvBench
Sample Usage
You can use the provided script from the codebase to batch download the data:
pip install huggingface_hub
python3 huggingface_pdeinv_download.py --dataset… See the full description on the dataset page: https://huggingface.co/datasets/DabbyOWL/PDE_Inverse_Problem_Benchmarking.LongE2V-data
LongE2V Dataset
This repository contains the preprocessed dataset (including BS-ERGB, ECD, MVSEC, and HQF) for LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models (SIGGRAPH 2026).
Project Page | Paper | GitHub Repository
Dataset Description
LongE2V is a unified video diffusion framework that reconstructs high-quality, stable, and temporally coherent videos from sparse event streams. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/fansam39/LongE2V-data.
