datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenMind
The OpenMind Dataset: A large-scale Head-And-Neck 3D MRI Dataset for self-supervised learning
Description
The OpenMind Dataset is a large-scale 3D MRI dataset of the head and neck region featuring 114k MRI Images. Its purpose is to provide access of large amounts of 3D medical imaging data to accelerate the development of self-supervised learning methods for 3D medical imaging. This data was pooled from exactly 800 datasets from the OpenNeuro platform and… See the full description on the dataset page: https://huggingface.co/datasets/MIC-DKFZ/OpenMind.gpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.Honey-Data-15M
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.OmniDocBench
OmniDocBench
English | 简体中文
OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics:
Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more.
Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.OpenPathNet
OpenPathNet Dataset
This README describes the OpenPathNet dataset (the release referred to as Link 1 in the OpenPathNet project documentation). The dataset is generated by the OpenPathNet toolchain from real-world Miami and Boston urban areas based on OpenStreetMap (OSM), and then simulated with NVIDIA Sionna ray tracing for RF multipath propagation / channel modeling research and AI tasks.
The dataset is also carefully cleaned to ensure good building coverage in every scene.… See the full description on the dataset page: https://huggingface.co/datasets/liu-lz/OpenPathNet.OpenFake
Dataset Card for OpenFake
Known issues
Prompt–image misalignment in the synthetic split (reported November 2025, fix pending)
For five of the eighty generators, the prompt field attached to synthetic
images does not correspond to the prompt actually used to generate that image.
Affected generators:
flux-realism
sd-3.5
sdxl-realvis-v5
sd-1.5-dreamshaper
sd-1.5-epicdream
This affects approximately 19.77% of synthetic images. It was first reported in
discussion… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.MVBench
MVBench
Important Update
[18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference.
We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MVBench.Open-Qwen2VL-Data
Introduction
This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources.
Project page: https://victorwz.github.io/Open-Qwen2VL
Code: https://github.com/Victorwz/Open-Qwen2VL
Dataset
ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1
datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.open-schematics
Open Schematics Dataset
The largest dataset of electronic schematics and PCB layouts on the internet, built as an engineering reference for schematic and PCB layout work. It's a self-growing, autonomous dataset that continuously scans the web for new engineering designs and updates itself accordingly.
Dataset Description
Each record corresponds to one schematic file and includes the raw source, rendered images, structured metadata, and all associated PCB files… See the full description on the dataset page: https://huggingface.co/datasets/bshada/open-schematics.CoW-Bench
CoW-Bench
Dataset | Evaluation Code
Authors: OpenRaiser
CoW-Bench is a comprehensive benchmark for evaluating video and image generation models' understanding of Composition of World (CoW), focusing on spatial relationships, object interactions, and temporal dynamics in multi-modal content generation.
Associated Paper
This dataset is associated with the following paper:
The Trinity of Consistency as a Defining Principle for General World Models
arXiv:… See the full description on the dataset page: https://huggingface.co/datasets/OpenRaiser/CoW-Bench.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.messy-prompt-datasets
🎨 Messy Prompt Dataset
🎨 A mixed collection of AI image prompts (500+). A bit of everything — raw and uncurated. Truly open source: No login, no ads, no redirection. Just pure data for AI creators.
This project is a growing collection of diverse image generation prompts gathered from social platforms like Twitter/X. The entire dataset contains 500+ images, all structured into a comprehensive dataset.
Due to GitHub's limitations with large file storage, the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/messy-prompt-datasets.TextEdit
TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models
Danni Yang,
Sitao Chen,
Changyao Tian
If you find our work helpful, please give us a ⭐ or cite our paper. See the InternVL-U technical report appendix for more details.
🎉 News
[2026/03/06] TextEdit benchmark released.
[2026/03/06] Evaluation code and initial baselines released.
[2026/03/06] Leaderboard updated with latest models.
📖 Introduction… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.open-image-preferences-v1
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.OpenMath-Vision-CoT-10kGUI-Odyssey
Dataset Card for GUI Odyssey
News⭐️
A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉
👉 Please use the latest version and refer to the updated README for the most up-to-date information.
We highly recommend using the new version for all training and evaluation!
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Latest Version of Dataset: hflqf88888/GUIOdyssey
Paper: https://arxiv.org/pdf/2406.08451
Introduction
GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.attributiontag-vl-22k-openrouter-imagesGameQA-140K
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale, significantly outperforming its Qwen3-VL-8B-Instruct base.
[2026/07] 🔥Peking University and WeChat AI use our Game-RL data… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.dalle-3-dataset
Dataset Card for LAION DALL·E 3 Discord Dataset
Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration.
Source Code: The code used to generate this data can be found here.
Contributors
Zach Nagengast
Eduardo Pach
Seva Maltsev
Ben Egan
The LAION community
Data Attributes
caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.Uni-GUI-OpenMobile
Uni-GUI-OpenMobile
A mobile GUI agent trajectory dataset collected on open-source Android applications via AndroidWorld, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,640
Total Steps
~25.9K
Platform
Android Mobile (1080x2400)
Applications
19 open-source apps
Coordinate System
Normalized to [0, 1000]… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-OpenMobile.VisRAG-Ret-Train-Synthetic-data
Dataset Description
This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries.
Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset.
Name
Source
Description
# Pages
Textbooks
https://openstax.org/
College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World.
Code: https://github.com/iamwangyabin/OpenSDI
OpenING_Generators_OutputsEgoNormia
EgoNormia: Benchmarking Physical-Social Norm Understanding
MohammadHossein Rezaei*,
Yicheng Fu*,
Phil Cuvin*,
Caleb Ziems,
Yanzhe Zhang,
Hao Zhu,
Diyi Yang,
🌎Website |
🤗 Dataset |
📄 arXiv |
📄 HF Paper
EgoNormia
EgoNormia is a challenging QA benchmark that tests VLMs' ability to reason over norms in context.
The datset consists of 1,853 physically grounded egocentric
interaction clips from Ego4D… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/EgoNormia.openimages
FineOpenimages — Open Images V7 boxed subset in the unified detection format
Source: official Open Images bbox CSVs + CVDF-hosted image tars (open-images-dataset S3 bucket).
Converted by the finedet project into a unified, AutoTrain-compatible layout:
image / width / height / objects{bbox, category} with COCO-format
[x, y, w, h] boxes in absolute pixels. Boxes are clipped to the image and
empty boxes dropped; category ids are densified per the category tables
below.… See the full description on the dataset page: https://huggingface.co/datasets/finedet/openimages.Open4DHOI
Open-Source Release Manifest
Generated from upload_records.json for records with annotation_progress == 4.
Required release items
video.mp4: source video to publish.
obj_init.obj: object mesh to publish.
mask_dir/: object masks.
human_mask_dir/: human masks.
motion/result.pt: reconstructed human motion.
motion/hand_pose.npz: SMPL-X hand pose parameters.
kp_record_new.json: point annotations.
Generated files
release_manifest.json: full per-record… See the full description on the dataset page: https://huggingface.co/datasets/acane2/Open4DHOI.OpenPath
