datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Honey-Data-15M
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.OmniDocBench
OmniDocBench
English | 简体中文
OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics:
Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more.
Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.Open-Qwen2VL-Data
Introduction
This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources.
Project page: https://victorwz.github.io/Open-Qwen2VL
Code: https://github.com/Victorwz/Open-Qwen2VL
Dataset
ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1
datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.open-image-preferences-v1
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.dalle-3-dataset
Dataset Card for LAION DALL·E 3 Discord Dataset
Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration.
Source Code: The code used to generate this data can be found here.
Contributors
Zach Nagengast
Eduardo Pach
Seva Maltsev
Ben Egan
The LAION community
Data Attributes
caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.ChartVerse-SFT-600KChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page.
This dataset contains non-trivial samples filtered by failure rate (r > 0), ensuring that every sample provides meaningful learning signal. Samples that are too easy (r = 0, where the model always answers correctly) are… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-600K.MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-586K
The Hardest 33% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1).
Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning.
🎯 Key Highlights
586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.MMFineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason
Closing the Multimodal Reasoning Gap via Open Data-Centric Methods
Average score across mathematical reasoning and multimodal understanding benchmarks.
📖 Overview
MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking.
🎯 Key Highlights
1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.ChartVerse-SFT-1.8MChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page.
This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-1.8M.MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking
MMFineReason-SFT-123K
The Hardest 7% — Less Data, More Reasoning
📖 Overview
MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0).
🎯 Key Highlights
123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.open-image-preferences-v1-binarized
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1-binarized.Honey-Data-1M
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-1M.US-Real-time-gun-detection-in-CCTV-An-open-problem-datasetMolRecBench-Wild
MolRecBench-Wild (2026-08-19)
MolRecBench-Wild is a real-world benchmark for optical chemical structure
recognition. This release contains 5,024 molecular structure images and
their CARBON molecular-graph annotations from 818 source articles.
This is the repository's authoritative 2026-08-19 release. It differs from the
5,029-sample snapshot described in the first arXiv version of the paper.
Dataset structure
The dataset has one test split. Images are embedded in… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/MolRecBench-Wild.open-gis-raster-cog-datasetsChartVerse-RL-40KChartVerse-RL-40K is a curated dataset of the most challenging chart reasoning samples for Reinforcement Learning, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page.
This dataset contains samples with the highest failure rates — the most difficult samples that strong VLMs struggle with but can still solve occasionally. These samples provide the strongest learning signal for RL training.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-RL-40K.ROVR-Open-Dataset
ROVR Open Dataset
Introduction
Welcome to the ROVR Open Dataset repository! This dataset is designed to empower autonomous driving and robotics research by providing rich, real-world data captured from ADAS cameras and LiDAR sensors. The dataset spans 50+ countries with over 20 million kilometers of driving data, making it ideal for training and developing advanced AI algorithms for depth estimation, object detection, and semantic segmentation.… See the full description on the dataset page: https://huggingface.co/datasets/ROVR-Network/ROVR-Open-Dataset.google-open-images-hair-style-dataset
Google Open Images — Hair Style Dataset
🇺🇸 English | 🇰🇷 한국어
Overview
This dataset is a curated custom subset of the Google Open Images V7 dataset, specifically filtered to include images of humans with various hair styles.It is intended for use in computer vision research and applications such as hair style classification, person detection, and instance segmentation.
Split
Purpose
train
Model training
validation
Model evaluation / hyperparameter tuning… See the full description on the dataset page: https://huggingface.co/datasets/hwany79/google-open-images-hair-style-dataset.Bee-Training-Data-Stage1
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage1.OpenDatasetFlowersITI
🌸 OpenDatasetFlowers – ImageToImage
A large-scale paired image dataset of AI-generated flower photographs and their corresponding line-art sketches, designed for image-to-image translation tasks such as sketch-to-photo synthesis, pix2pix training, and contour-guided generation.
Dataset Summary
The dataset provided as OpenDatasetFlowersITI.zip contains approximately 80k images (40k pairs) across 129 flower categories. Each pair consists of a line-art sketch… See the full description on the dataset page: https://huggingface.co/datasets/erikpro007/OpenDatasetFlowersITI.gaissian_atc_opendata_bridge_crack_x1.5
Dataset Card for "gaissian_atc_opendata_bridge_crack_x1.5"
More Information needed
digit3a_hardware_door_open_rgb_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "digit_third_arm",
"total_episodes": 33,
"total_frames": 13730,
"total_tasks": 1,
"total_videos": 33,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:33"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kelvinzhaozg/digit3a_hardware_door_open_rgb_dataset.open-omniparser-datasetBenchmark_Hallucinations_Data__run_Baseer__morph_openopen_genmoji_data
GenMoji Dataset
This repository hosts the GenMoji Dataset, a collection of Apple emojis sourced from Emojigraph, along with their respective captions.
Dataset Overview
Total Examples: 3,770
Features:
image: An emoji image file.
caption: grinning face emoji.
Example Usage
To load the dataset, use the Hugging Face datasets library:
from datasets import load_dataset
dataset = load_dataset("emon-j/open_genmoji_data")
print(dataset)
example =… See the full description on the dataset page: https://huggingface.co/datasets/emon-j/open_genmoji_data.SA-RxnDiagram-15k
U-RxnDiagram-15k Dataset (Sci-Align)
🌌 The Sciverse Data Foundation
Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research.
Sciverse consists of three core… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SA-RxnDiagram-15k.
