CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WanyueZhang /MulSeT MulSeT: A Benchmark for Multi-view Spatial Understanding Tasks Paper: Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Code: https://github.com/WanyueZhang-ai/spatial-understanding A high-level overview of the MulSeT benchmark. The dataset challenges models to integrate information from two distinct viewpoints of a 3D scene to answer spatial reasoning questions. 📝 Dataset Summary MulSeT is a comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/WanyueZhang/MulSeT.image6 likes89k downloads11mo agoHugging Face02racineai /VDR_MEGA_MultiDomain_DocRetrieval Visual Document Retrieval Dataset Overview This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks. Dataset Structure The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.imagevisual-document-retrieval1M<n<10M24 likes68k downloads6mo agoHugging Face03kohsei /MultiBanana-Benchmark🍌 MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation 🍌 CVPR 2026 (Main) This repository provides the datasets for “MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation” by Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo and Hiroki Furuta Paper Link https://arxiv.org/abs/2511.22989 Github Repository For the usage of this benchmark, please see Github… See the full description on the dataset page: https://huggingface.co/datasets/kohsei/MultiBanana-Benchmark.imagetext-to-image1K<n<10K5 likes21k downloads3mo agoHugging Face04richidubey /KAIST-Multispectral-Pedestrian-Detection-Datasetimage10K<n<100K4 likes14k downloads2y agoHugging Face050001AMA /multimodal_data_annotator_datasetMaterials dataset consisting of spatial and time resolved versions of the same object. Specially curated for the annotator such that for each object, time resolved signal may be viewed alongside the RGB and for different graphs/forms image10K<n<100K0 likes10k downloads8mo agoHugging Face06osunlp /Multimodal-Mind2Web Dataset Summary Multimodal-Mind2Web is the multimodal version of Mind2Web, a dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. In this dataset, we align each HTML document in the dataset with its corresponding webpage screenshot image from the Mind2Web raw dump. This multimodal version addresses the inconvenience of loading images from the ~300GB Mind2Web Raw Dump. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/Multimodal-Mind2Web.image10K<n<100K99 likes10k downloads2y agoHugging Face07multimodal-reasoning-lab /Zebra-CoT Zebra‑CoT A diverse large-scale dataset for interleaved vision‑language reasoning traces. Dataset Description Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games. Dataset Structure Each example in Zebra‑CoT consists of: Problem statement:… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT.imageany-to-any100K<n<1M78 likes9.2k downloads8mo agoHugging Face08mulan-dataset /v1.0 MuLAn: : A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation MuLAn is a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multilayer, instance-wise RGBA decompositions, and over 100K instance images. It is composed of MuLAn-COCO and MuLAn-LAION sub-datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance… See the full description on the dataset page: https://huggingface.co/datasets/mulan-dataset/v1.0.imagetext-to-imagen<1K27 likes8.5k downloads2y agoHugging Face09trl-internal-testing /zen-multi-imageimagen<1K1 likes8.3k downloads3mo agoHugging Face10alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes6.8k downloads2mo agoHugging Face11zwcolin /clevr-multichange CLEVR-Multi-Change (30–40 objects) Two-image change-captioning data used in "Stateful Visual Encoders for Vision-Language Models" (the Multi-object Visual Differencing task). Each example is a before/after pair of a CLEVR scene with 30–40 objects and 4 simultaneous changes (add / delete / move / replace), rendered at 768×768 with a wide camera angle. Built with the CLEVR-Multi-Change engine (Johnson et al. 2017; Qiu et al. 2021). Code & paper:… See the full description on the dataset page: https://huggingface.co/datasets/zwcolin/clevr-multichange.imageimage-to-text100K<n<1M0 likes6.1k downloads4mo agoHugging Face12limingcv /MultiGen-20M_train Dataset Card for "MultiGen-20M_train" This dataset is constructed from UniControl, and used for evaluation of the paper ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback ControlNet++ Github repository: https://github.com/liming-ai/ControlNet_Plus_Plus image1M<n<10M6 likes5.8k downloads2y agoHugging Face13multimodalart /facesyntheticsspigacaptioned Dataset Card for "face_synthetics_spiga_captioned" This is a copy of the Microsoft FaceSynthetics dataset with SPIGA-calculated landmark annotations, and additional BLIP-generated captions. For a copy of the original FaceSynthetics dataset with no extra annotations, please refer to pcuenq/face_synthetics. Here is the code for parsing the dataset and generating the BLIP captions: from transformers import pipeline dataset_name = "pcuenq/face_synthetics_spiga" faces =… See the full description on the dataset page: https://huggingface.co/datasets/multimodalart/facesyntheticsspigacaptioned.image100K<n<1M35 likes5.6k downloads4y agoHugging Face14multimodalart /lora-fusing-preferencesimage1K<n<10K12 likes5.2k downloads2y agoHugging Face15limingcv /MultiGen-20M_depth Dataset Card for "MultiGen-20M_depth" More Information needed image1M<n<10M6 likes4.7k downloads3y agoHugging Face16WindyLab /MultiAgent-Content MultiAgent UE5 Content MultiAgent-Unreal 项目的配套 UE5 资产库。 快速克隆 #### 安装 Git LFS #### ## Mac 系统 brew install git-lfs git lfs install ## Linux 系统 sudo apt-get install git-lfs git lfs install #### 克隆到项目目录 #### cd unreal_project ## 有🪜 git clone https://huggingface.co/datasets/WindyLab/MultiAgent-Content Content ## 中国用户 git clone https://hf-mirror.com/datasets/WindyLab/MultiAgent-Content Content 📂 目录结构 Content/ ├── Agent/ # 智能体配置 ├──… See the full description on the dataset page: https://huggingface.co/datasets/WindyLab/MultiAgent-Content.3dn<1K0 likes4.5k downloads2mo agoHugging Face17Voxel51 /kitscenes-multimodal KITScenes Multimodal — FiftyOne Dataset A FiftyOne build of KITScenes Multimodal (KIT-MRT), a high-fidelity European urban autonomous-driving dataset. Each frame is a synchronized capture from a full robotaxi sensor suite — nine global-shutter cameras giving 360° coverage, seven long-range lidars, and three 4D imaging radars — paired with production-grade Lanelet2 HD-map labels, projected lidar depth, the future ego path, and image instance predictions. This build packages… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/kitscenes-multimodal.imageobject-detection10K<n<100K11 likes3.8k downloads3mo agoHugging Face18lfsm /multimodal_wikiimage1M<n<10M1 likes3.6k downloads3y agoHugging Face19mulsi /fruit-vegetable-conceptsimage1K<n<10K5 likes3.3k downloads2y agoHugging Face20electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.8k downloads2mo agoHugging Face21Wenyan0110 /Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks. The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks. If you find our research helpful, please cite our paper: @article{xu2025finmultitime, title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Wenyan0110/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.imagen<1K12 likes2.8k downloads1y agoHugging Face22artefactory /ledger-long-context-multi-kpi the LEDGER Long-Context Multi-KPI extraction datasets and benchmarks. OCR'd annual reports with ground-truth KPI values for financial information extraction benchmarking. Dataset Description This dataset pairs OCR-extracted annual report text (from DeepSeek OCR) with structured KPI ground-truth values. It is designed for evaluating LLM-based financial information extraction, retrieval, and needle-in-a-haystack tasks. Configs Config Reports… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/ledger-long-context-multi-kpi.imagetable-question-answering1K<n<10K15 likes2.7k downloads2mo agoHugging Face23thelfer /multimodal_supernovaeimage1K<n<10K1 likes2.7k downloads2y agoHugging Face24Multimodal-Fatima /VQAv2_train Dataset Card for "VQAv2_train" More Information needed image100K<n<1M6 likes2.3k downloads3y agoHugging Face25LianeMarilin /CADBench-Extended-Multimodal-Dataset Dataset Card Dataset Description CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics. Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.3dimage-to-textn<1K2 likes2k downloads24d agoHugging Face26suijinru /MultiViewBench MultiView-Bench Evaluation data for MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs (arXiv:2607.08970). This repository currently contains the synthetic image-and-text portion of the benchmark. The real-world subset is not included; see the scope note below. Contents 2,400 evaluation samples across 24 benchmark variants 7,900 PNG views (2.94 GiB logical image data) One, three, or six images per sample English question text… See the full description on the dataset page: https://huggingface.co/datasets/suijinru/MultiViewBench.imagevisual-question-answering1K<n<10K1 likes1.9k downloads2mo agoHugging Face27hanhuark /BoilingBench-Multimodal BoilingBench-Multimodal (NED3-017) BoilingBench-Multimodal is a family of research datasets from the NED³ laboratory for machine learning, computer vision, acoustic sensing, and multimodal heat-transfer analysis. The family contains four multimodal pool-boiling datasets, one human-annotated image dataset, one hydrophone-only pool-boiling dataset, and one infrared immersion-cooling dataset. This folder is a data distribution, not a Python package. The original acquisition files… See the full description on the dataset page: https://huggingface.co/datasets/hanhuark/BoilingBench-Multimodal.imagetabular-regression1K<n<10K1 likes1.8k downloads1mo agoHugging Face28haolunwang /aloha_multiviewimage100K<n<1M0 likes1.8k downloads13d agoHugging Face29UniqueData /multiple-sclerosis-dataset Multiple Sclerosis Dataset, Brain MRI Object Detection & Segmentation Dataset The dataset consists of .dcm files containing MRI scans of the brain of the person with a multiple sclerosis. The images are labeled by the doctors and accompanied by report in PDF-format. The dataset includes 13 studies, made from the different angles which provide a comprehensive understanding of a multiple sclerosis as a condition. MRI study angles in the dataset 💴 For… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/multiple-sclerosis-dataset.imageimage-to-imagen<1K3 likes1.8k downloads1y agoHugging Face30Multimodal-Fatima /FGVC_Aircraft_train Dataset Card for "FGVC_Aircraft_train" More Information needed image1K<n<10K3 likes1.5k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.