CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alibaba-multimodal-industrial-ai /IndustryBench-MIPU IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images. Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.imageimage-to-text10K<n<100K7 likes7.6k downloads2mo agoHugging Face02electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.7k downloads1mo agoHugging Face03LianeMarilin /CADBench-Extended-Multimodal-Dataset Dataset Card Dataset Description CADBench Extended Multimodal Dataset is an independently produced public extension for multimodal CAD reconstruction research. It contains 100 CAD samples with clean and perturbed meshes, STEP/STL/OBJ/GLB representations, single-view and four-view renders, PBR images, bilingual descriptions, prompt variants, QA, geometry metadata, grading signals, and manually reviewed visual semantics. Tasks: image-to-text, text-to-image… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/CADBench-Extended-Multimodal-Dataset.3dimage-to-textn<1K1 likes1.7k downloads21d agoHugging Face04FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes346 downloads2y agoHugging Face05huanngzh /DeepFashion-MultiModal-Parts2Whole DeepFashion MultiModal Parts2Whole Dataset Details Dataset Description This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/huanngzh/DeepFashion-MultiModal-Parts2Whole.imagetext-to-image10K<n<100K9 likes239 downloads2y agoHugging Face06fluid-concepts /multimodal-peer-collaboration-samplesgated Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges. ▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.audion<1K1 likes131 downloads5d agoHugging Face07aitf-komdigi /KomdigiITS-DFK3-Multimodalimage10K<n<100K0 likes47 downloads3mo agoHugging Face08LIAGM /DeepFashion-MultiModal-Parts2Whole DeepFashion MultiModal Parts2Whole Dataset Details Dataset Description This human image dataset comprising about 41,500 reference-target pairs. Each pair in this dataset includes multiple reference images, which encompass human pose images (e.g., OpenPose, Human Parsing, DensePose), various aspects of human appearance (e.g., hair, face, clothes, shoes) with their short textual labels, and a target image featuring the same individual (ID) in the same outfit… See the full description on the dataset page: https://huggingface.co/datasets/LIAGM/DeepFashion-MultiModal-Parts2Whole.imagetext-to-image10K<n<100K2 likes42 downloads2y agoHugging Face09f13rnd /multimodal-example Multimodal Example Dataset Small example dataset for testing multimodal (vision-language) fine-tuning with ms-swift. Structure ├── train.jsonl # 10 training samples ├── test.jsonl # 2 validation samples ├── images/ # All referenced images (400x300 JPEG) │ ├── dog_portrait.jpg │ ├── forest_river.jpg │ ├── laptop_desk.jpg │ ├── mountain_lake.jpg │ ├── ocean_rocks.jpg │ ├── coffee_cup.jpg │ ├── bookshelf.jpg │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/f13rnd/multimodal-example.imagen<1K0 likes40 downloads6mo agoHugging Face10Sigoso12 /multimodal-ir-zh-tw Multimodal IR Traditional Chinese Dataset Latest exported JSON files for the multimodal information-retrieval project. Files File Records Description SHA-256 multimodal_documents.g4.captioned.s2tw.json 769,245 Latest document corpus with image captions and Simplified-to-Traditional Chinese conversion e4e3f7609fde3297d89e436a11751dc5a394109f7aba2ae487fe9323d198802d multimodal_pretrain_pairs.json 351,979 Latest large pretraining/query pairs with text… See the full description on the dataset page: https://huggingface.co/datasets/Sigoso12/multimodal-ir-zh-tw.imagetext-retrieval1K<n<10K0 likes37 downloads1mo agoHugging Face11obaydata /svg-multimodal-rubrics SVG Multimodal Rubrics A multimodal dataset of SVG code generation samples with natural language descriptions and evaluation rubrics. Each sample pairs a detailed prompt (Markdown) with its corresponding SVG source code, covering animations, 3D scenes, games, and visual effects. Designed for training and evaluating models on visual code generation — generating complex, interactive SVG artwork from natural language descriptions. Overview Item Details Samples… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/svg-multimodal-rubrics.imagetext-generationn<1K0 likes35 downloads6mo agoHugging Face12hsuya /data_multimodalgatedimageimage-to-textn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.