CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Open-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes58k downloads7mo agoHugging Face02opendatalab /OmniDocBench OmniDocBench English | 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.image1K<n<10K106 likes26k downloads3mo agoHugging Face03weizhiwang /Open-Qwen2VL-Data Introduction This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources. Project page: https://victorwz.github.io/Open-Qwen2VL Code: https://github.com/Victorwz/Open-Qwen2VL Dataset ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1 datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.imageimage-text-to-text10M<n<100M25 likes15k downloads1y agoHugging Face04Goku-OpenLab /open-models-prompt-datasets 🖼️ Open Models Prompt Dataset 🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.image1K<n<10K1 likes6.4k downloads2mo agoHugging Face05OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.5k downloads8mo agoHugging Face06data-is-better-together /open-image-preferences-v1 Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.imagetext-to-image1K<n<10K31 likes4.9k downloads2y agoHugging Face07OpenDatasets /dalle-3-dataset Dataset Card for LAION DALL·E 3 Discord Dataset Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration. Source Code: The code used to generate this data can be found here. Contributors Zach Nagengast Eduardo Pach Seva Maltsev Ben Egan The LAION community Data Attributes caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.image10K<n<100K27 likes3.7k downloads2y agoHugging Face08turing-motors /Japan-Open-Driving-Dataset-Sample Japan Open Driving Dataset Sample Overview This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan. The data is stored in nuScenes format and can be loaded with the nuscenes-devkit. In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.image10K<n<100K5 likes2k downloads6mo agoHugging Face09opendatalab /ChartVerse-SFT-600KChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains non-trivial samples filtered by failure rate (r > 0), ensuring that every sample provides meaningful learning signal. Samples that are too easy (r = 0, where the model always answers correctly) are… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-600K.imagevisual-question-answering100K<n<1M11 likes1.9k downloads8mo agoHugging Face10OpenDataArena /MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking MMFineReason-SFT-586K The Hardest 33% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-586K is a difficulty-filtered subset of MMFineReason-1.8M, containing the hardest 33% of samples where Qwen3-VL-4B-Thinking do not consistently succeed. (pass rate ≠ 1). Specifically, this subset removes all easy samples (pass rate = 1) under Qwen3-VL-4B-Thinking, retaining only instances that require non-trivial multimodal reasoning. 🎯 Key Highlights 586K… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-586K-Qwen3-VL-235B-Thinking.image100K<n<1M6 likes1.8k downloads8mo agoHugging Face11OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.6k downloads7mo agoHugging Face12Open-Bee /Bee-Training-Data-Stage2 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.imageimage-to-text10M<n<100M6 likes1.6k downloads7mo agoHugging Face13opendatalab /ChartVerse-SFT-1.8MChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains all verified correct samples without failure rate filtering. Unlike SFT-600K which excludes easy samples (r=0), SFT-1800K includes the complete set of truth-anchored QA pairs for maximum coverage and scale.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-SFT-1.8M.imagevisual-question-answering1M<n<10M139 likes1.3k downloads7mo agoHugging Face14OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes611 downloads8mo agoHugging Face15data-is-better-together /open-image-preferences-v1-binarized Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1-binarized.image1K<n<10K59 likes316 downloads2y agoHugging Face16Open-Bee /Honey-Data-1M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-1M.imageimage-text-to-text1M<n<10M21 likes304 downloads7mo agoHugging Face17jsalazar /US-Real-time-gun-detection-in-CCTV-An-open-problem-datasetimage10K<n<100K3 likes240 downloads2y agoHugging Face18opendatalab /MolRecBench-Wild MolRecBench-Wild (2026-08-19) MolRecBench-Wild is a real-world benchmark for optical chemical structure recognition. This release contains 5,024 molecular structure images and their CARBON molecular-graph annotations from 818 source articles. This is the repository's authoritative 2026-08-19 release. It differs from the 5,029-sample snapshot described in the first arXiv version of the paper. Dataset structure The dataset has one test split. Images are embedded in… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/MolRecBench-Wild.imageimage-to-text1K<n<10K3 likes208 downloads1mo agoHugging Face19rafiqns /open-gis-raster-cog-datasetsimagen<1K0 likes200 downloads2mo agoHugging Face20opendatalab /ChartVerse-RL-40KChartVerse-RL-40K is a curated dataset of the most challenging chart reasoning samples for Reinforcement Learning, developed as part of the opendatalab/ChartVerse project. For more details about our method, datasets, and full model series, please visit our Project Page. This dataset contains samples with the highest failure rates — the most difficult samples that strong VLMs struggle with but can still solve occasionally. These samples provide the strongest learning signal for RL training.… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/ChartVerse-RL-40K.imagevisual-question-answering10K<n<100K12 likes169 downloads8mo agoHugging Face21ROVR-Network /ROVR-Open-Dataset ROVR Open Dataset Introduction Welcome to the ROVR Open Dataset repository! This dataset is designed to empower autonomous driving and robotics research by providing rich, real-world data captured from ADAS cameras and LiDAR sensors. The dataset spans 50+ countries with over 20 million kilometers of driving data, making it ideal for training and developing advanced AI algorithms for depth estimation, object detection, and semantic segmentation.… See the full description on the dataset page: https://huggingface.co/datasets/ROVR-Network/ROVR-Open-Dataset.imageimage-to-3d10K<n<100K6 likes144 downloads10mo agoHugging Face22hwany79 /google-open-images-hair-style-dataset Google Open Images — Hair Style Dataset 🇺🇸 English | 🇰🇷 한국어 Overview This dataset is a curated custom subset of the Google Open Images V7 dataset, specifically filtered to include images of humans with various hair styles.It is intended for use in computer vision research and applications such as hair style classification, person detection, and instance segmentation. Split Purpose train Model training validation Model evaluation / hyperparameter tuning… See the full description on the dataset page: https://huggingface.co/datasets/hwany79/google-open-images-hair-style-dataset.imageimage-classification1K<n<10K0 likes56 downloads7mo agoHugging Face23Open-Bee /Bee-Training-Data-Stage1 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage1.imageimage-to-text100K<n<1M4 likes51 downloads7mo agoHugging Face24erikpro007 /OpenDatasetFlowersITI 🌸 OpenDatasetFlowers – ImageToImage A large-scale paired image dataset of AI-generated flower photographs and their corresponding line-art sketches, designed for image-to-image translation tasks such as sketch-to-photo synthesis, pix2pix training, and contour-guided generation. Dataset Summary The dataset provided as OpenDatasetFlowersITI.zip contains approximately 80k images (40k pairs) across 129 flower categories. Each pair consists of a line-art sketch… See the full description on the dataset page: https://huggingface.co/datasets/erikpro007/OpenDatasetFlowersITI.image10K<n<100K0 likes49 downloads5mo agoHugging Face25alphaca /gaissian_atc_opendata_bridge_crack_x1.5 Dataset Card for "gaissian_atc_opendata_bridge_crack_x1.5" More Information needed imagen<1K0 likes34 downloads3y agoHugging Face26kelvinzhaozg /digit3a_hardware_door_open_rgb_datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "digit_third_arm", "total_episodes": 33, "total_frames": 13730, "total_tasks": 1, "total_videos": 33, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:33"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kelvinzhaozg/digit3a_hardware_door_open_rgb_dataset.imagerobotics10K<n<100K0 likes32 downloads10mo agoHugging Face27Francesco /open-omniparser-datasetimage1K<n<10K1 likes30 downloads1y agoHugging Face28muhammad0-0hreden /Benchmark_Hallucinations_Data__run_Baseer__morph_openimagen<1K0 likes28 downloads22d agoHugging Face29emon-j /open_genmoji_data GenMoji Dataset This repository hosts the GenMoji Dataset, a collection of Apple emojis sourced from Emojigraph, along with their respective captions. Dataset Overview Total Examples: 3,770 Features: image: An emoji image file. caption: grinning face emoji. Example Usage To load the dataset, use the Hugging Face datasets library: from datasets import load_dataset dataset = load_dataset("emon-j/open_genmoji_data") print(dataset) example =… See the full description on the dataset page: https://huggingface.co/datasets/emon-j/open_genmoji_data.imageimage-to-image1K<n<10K1 likes23 downloads2y agoHugging Face30opendatalab /SA-RxnDiagram-15k U-RxnDiagram-15k Dataset (Sci-Align) 🌌 The Sciverse Data Foundation Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research. Sciverse consists of three core… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SA-RxnDiagram-15k.image10K<n<100K3 likes19 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.