CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Fsoft-AIC /RobotDesign1M RobotDesign1M: A Large-scale Dataset for Robot Design Understanding RobotDesign1M is a large-scale, multimodal dataset for robot design understanding, built from image–text data curated from scientific literature across a wide range of robotics domains. It is designed to support research on design-aware foundation models, including design image generation, visual question answering about designs, and design image retrieval. 📄 Paper: RobotDesign1M: A Large-scale Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Fsoft-AIC/RobotDesign1M.imageimage-text-to-text1M<n<10M8 likes70k downloads2mo agoHugging Face02markov-ai /cad-environments CAD Environments CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling. Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.imagen<1K17 likes57k downloads2mo agoHugging Face03ai-for-good-lab /ai4g-flood-dataset Flood Detection Dataset Introduction This dataset accompanies the paper Mapping global floods with 10 years of satellite radar data (Nature Communications, 2025) and contains global flood detections derived from Sentinel-1 Synthetic Aperture Radar (SAR) imagery using a deep learning change detection model. The dataset spans October 2014 – September 2024, offering a longitudinal view of flood-prone areas worldwide. Key features: Cloud-penetrating SAR data for consistent… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/ai4g-flood-dataset.imagen<1K16 likes51k downloads11mo agoHugging Face04MeiGen-AI /GenEvolve-Data-Bench GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose sft GenEvolve-Data-SFT/ 9,000 trajectories 50,291 reference images supervised cold-start trajectories rl GenEvolve-Data-RL/ 3,175 prompts 3,175 GT images self-evolution / RL training prompts bench GenEvolve-Bench/ 594 prompts 594 GT images held-out evaluation benchmarkAll metadata is provided in both JSONL and Parquet. The Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/MeiGen-AI/GenEvolve-Data-Bench.imagetext-to-image10K<n<100K2 likes42k downloads4mo agoHugging Face05AI4Math /MathVista Dataset Card for MathVista Dataset Description Paper Information Dataset Examples Leaderboard Dataset Usage Data Downloading Data Format Data Visualization Data Source Automatic Evaluation License Citation Dataset Description MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts. It consists of three newly created datasets, IQTest, FunctionQA, and PaperQA, which address the missing visual domains and are tailored to evaluate logical… See the full description on the dataset page: https://huggingface.co/datasets/AI4Math/MathVista.imagemultiple-choice1K<n<10K225 likes24k downloads3y agoHugging Face06lmms-lab-encoder /ai2d@misc{kembhavi2016diagram, title={A Diagram Is Worth A Dozen Images}, author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon Seo and Hannaneh Hajishirzi and Ali Farhadi}, year={2016}, eprint={1603.07396}, archivePrefix={arXiv}, primaryClass={cs.CV} } image1K<n<10K24 likes16k downloads2y agoHugging Face07ai-habitat /OVMM_objects3d7 likes16k downloads2y agoHugging Face08nunchaku-ai /cdn Nunchaku CDN imagen<1K6 likes13k downloads2mo agoHugging Face09Metaverse-AI-Lab /M3DLayout M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation We are continuously scaling up our layout collection and will release more results as soon as they are ready. Please stay tuned and follow our work for updates! In text-driven 3D scene generation, object layout serves as a crucial intermediate representation that bridges high-level language instructions with detailed geometric output. It not only provides a structural blueprint for… See the full description on the dataset page: https://huggingface.co/datasets/Metaverse-AI-Lab/M3DLayout.3d100K<n<1M7 likes13k downloads5mo agoHugging Face10ai4ce /NYC-CDimage1 likes12k downloads8mo agoHugging Face11ait4x /polyu-storyworld-charactersimagen<1K0 likes11k downloads5mo agoHugging Face12iisc-aim /BMD-45 BMD-45: Bengaluru Mobility Dataset A large-scale CCTV vehicle detection benchmark for Indian urban traffic Dataset Summary BMD-45 is a large-scale, India-specific vehicle detection dataset captured from 3,679 operational CCTV cameras across Bengaluru — one of the world's most traffic-congested megacities. Statistic Value Total images 45,986 (1920×1080 RGB) Total annotations ≈ 481,947 bounding boxes Vehicle classes 14 fine-grained categories Camera… See the full description on the dataset page: https://huggingface.co/datasets/iisc-aim/BMD-45.imageobject-detection1K<n<10K4 likes9.7k downloads6mo agoHugging Face13math-ai /BlueMO BlueMO 🚀 BlueMO: A Comprehensive Collection of Challenging Mathematical Olympiad Problems from the Little Blue Book Series   BlueMO is a comprehensive and challenging dataset comprising mathematical olympiad problems paired with detailed solutions, meticulously curated from the esteemed "Little Blue Book" (小蓝书) series (Second Edition)—a vital resource for Chinese students training for national and international olympiad math competitions.Designed to advance and… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/BlueMO.imagequestion-answering1K<n<10K3 likes9.2k downloads8mo agoHugging Face14alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes7.7k downloads2mo agoHugging Face15KMK040412 /aitw-processed-labeled-full AiTW Processed Full with App Labels This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset. Why This Exists AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.imageimage-text-to-text1M<n<10M0 likes7.6k downloads4mo agoHugging Face16sii-rhos-ai /ViFailback-Dataset ViFailback Dataset: Real-World Robotic Manipulation Failure Dataset with Visual Symbol Guidance A real-world dataset for diagnosing, correcting, and learning from robotic manipulation failures via visual symbols. ViFailback is a large-scale, real-world robotic manipulation failure dataset introduced in the CVPR 2026 paper "Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols". It introduces visual… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/ViFailback-Dataset.imagevisual-question-answeringn<1K7 likes7.3k downloads25d agoHugging Face17mila-ai4h /mid-space MID-Space: Aligning Diverse Communities’ Needs to Inclusive Public Spaces A new version of the dataset will be released soon, incorporating user identity markers and expanded annotations. LIVS PAPER Click below to see more: Overview The MID-Space dataset is designed to align AI-generated visualizations of urban public spaces with the preferences of diverse and marginalized communities in Montreal. It includes textual prompts, Stable Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/mila-ai4h/mid-space.imagetext-to-image1K<n<10K1 likes7.3k downloads2y agoHugging Face18AI-Lab-Makerere /beans Dataset Card for Beans Dataset Summary Beans leaf dataset with images of diseased and health leaves. Supported Tasks and Leaderboards image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any. Languages English Dataset Structure Data Instances A sample from the training set is provided below: { 'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.imageimage-classification1K<n<10K47 likes7.1k downloads3y agoHugging Face19weikaih /ai2thor-vsi-bench-1k-v2imagen<1K0 likes5.9k downloads1y agoHugging Face20AISHELL /RealMANimage3 likes5.7k downloads2y agoHugging Face21lmarena-ai /VisionArena-Chat VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes 200k single and multi-turn chats between users and VLM's collected on Chatbot Arena. WARNING: Images may contain inappropriate content. Dataset Details 200K conversations 45 VLM's 138 languages ~43k unique images Question Category Tags (Captioning, OCR, Entity Recognition, Coding, Homework, Diagram, Humor, Creative Writing, Refusal) Dataset Description 200,000… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/VisionArena-Chat.imagevisual-question-answering100K<n<1M15 likes5.6k downloads2y agoHugging Face22blanchon /AID Aerial Image Dataset (AID) Description The Aerial Image Dataset (AID) is a scene classification dataset consisting of 10,000 RGB images, each with a resolution of 600x600 pixels. These images have been extracted using Google Earth and cover various scenes from regions and countries around the world. AID comprises 30 different scene categories, with several hundred images per class. The new dataset is made up of the following 30 aerial scene types: airport, bare… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/AID.imageimage-classification10K<n<100K9 likes4.9k downloads3y agoHugging Face23math-ai /olympiadbenchimagen<1K7 likes4.8k downloads1y agoHugging Face24aimagelab /RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors. imageimage-classification2 likes4.5k downloads1y agoHugging Face25xiaoleezuishuai /airbot-fold-cloth-mcapimage1K<n<10K0 likes4.4k downloads2mo agoHugging Face261thesudden /AIOU26_checkpointsimage0 likes4.2k downloads21d agoHugging Face27amir-kazemi /aidovecl-vehicle-detection-classification-localization AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models. Citation Notice Please ensure that all publications and presentations using this data reference the following paper: Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.imageobject-detection1K<n<10K0 likes4k downloads5mo agoHugging Face28DurYi /AirGoal-10k AirGoal-10k AirGoal-10k is an aerial image-goal navigation dataset released with UA-NWM: Uncertainty-Aware World Model for Aerial Image-Goal Navigation. Project page: https://duryi.github.io/UA-NWM-Project-Page/Code: https://github.com/DurYi/UA-NWMPaper: https://arxiv.org/abs/2608.05597 Dataset Summary AirGoal-10k contains 11,000 aerial navigation trajectories for image-goal navigation. Each trajectory contains 12 RGB observations and trajectory metadata. The test… See the full description on the dataset page: https://huggingface.co/datasets/DurYi/AirGoal-10k.imagerobotics100K<n<1M1 likes4k downloads28d agoHugging Face29aialliance /GEOBench-VLM GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks Summary While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, which is critical for applications such as environmental monitoring, urban planning, and disaster management. Some of the unique… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/GEOBench-VLM.image1K<n<10K16 likes3.9k downloads1y agoHugging Face30zhaojiao /ai-images-100kimagen<1K0 likes3.8k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.