datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.IndustryBench-MIPU
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
Multi-Image Industrial Product Understanding Benchmark — evaluating MLLMs on structured attribute extraction from real-world industrial product images.
Industrial product specifications are scattered across multiple heterogeneous images — specification tables, nameplates, technical drawings. IndustryBench-MIPU tests whether MLLMs can reliably recover them through four… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench-MIPU.aimotive-multimodal
Dataset Card for aiMotive Multimodal Dataset
The aiMotive Multimodal Dataset is a 176-scene autonomous driving dataset
with synchronized and calibrated LiDAR, camera, and radar sensors providing
360-degree field-of-view coverage with sensor redundancy. Scenes were
captured in highway, urban, and suburban environments across three countries
during daytime, night, and rain. The dataset contains 26,583 annotated
frames with 3D bounding boxes for 14 object classes (425k+ instances)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/aimotive-multimodal.FGVC_Aircraft_train
Dataset Card for "FGVC_Aircraft_train"
More Information needed
FGVC_Aircraft_test
Dataset Card for "FGVC_Aircraft_test"
More Information needed
africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
industry_multimodal_v1household_multimodal_v1Kairos-Multimodal-Reasoning
A dataset for training models in multimodal reasoning tasks
Usage
from datasets import load_dataset
ds = load_dataset("Aquiles-ai/Kairos-Multimodal-Reasoning")
print(ds.features)
print(ds["train"]["source"])
Preview of dataset examples
We've built a playground so you can see some of the examples included in the dataset.
Link: https://kairos-example.vercel.app/
Dataset used in the blog post: Kairos: Building a Multimodal Model with LFM2.5 and… See the full description on the dataset page: https://huggingface.co/datasets/Aquiles-ai/Kairos-Multimodal-Reasoning.rand-1m-multimodalmvai-doctag-r3-v1
MVAI DocTag R3 v1
Basic Information
Field
Value
Dataset ID
multimodal-vision-ai/mvai-doctag-r3-v1
Version
v1
Owner
Dizzar, Hohai University / multimodal-vision-ai
Dataset type
Image + DocTags text annotations
Intended use
OCR and document layout research, especially image-to-DocTags SFT / GRPO
This dataset is the formal closure version of the existing HohaiR3 DocTags data
asset. It contains the HohaiR3 gold, human-labeled, HTML-rendered… See the full description on the dataset page: https://huggingface.co/datasets/multimodal-vision-ai/mvai-doctag-r3-v1.IndustryBench
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper
IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension
Details
Total questions
2,049
Languages
Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.car-driving-multimodal-v1
Car Driving Multimodal Dataset v1
Overview
The Car Driving Multimodal Dataset v1 is a publication-quality egocentric dataset designed to advance research in human-object interactions (HOI), egocentric vision, and autonomous cabin activities. It captures a driver's hand interactions, hand poses, activity recognition labels, and synthetic inertial/depth annotations from a first-person perspective.
This dataset was generated using the Hand Egocentric Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/deepannotate-ai/car-driving-multimodal-v1.cattle-farming-multimodal-v1
Cattle Farming Multimodal Dataset v1
Overview
The Cattle Farming Multimodal Dataset v1 is a publication-quality egocentric agricultural dataset designed to advance research in human-object interactions (HOI), egocentric vision, and embodied AI in farming/veterinary contexts. It captures fine-grained hand interactions, hand poses, activity recognition labels, and synthetic inertial/depth annotations from a first-person perspective during animal interaction and… See the full description on the dataset page: https://huggingface.co/datasets/deepannotate-ai/cattle-farming-multimodal-v1.ai-residency-multimodal-captioning-datavegetable-harvesting-multimodal-v1
Vegetable Harvesting Multimodal Dataset v1
Overview
The Vegetable Harvesting Multimodal Dataset v1 is a publication-quality first-person (egocentric) dataset capturing hands, objects, poses, and activity annotations for agricultural manual tasks. It contains two major subset activities:
Vegetable Harvesting: Manual crop collection and tool usage sequence.
Vegetable Plucking: Fine-grained picking and placement of products.
This dataset was generated using the… See the full description on the dataset page: https://huggingface.co/datasets/deepannotate-ai/vegetable-harvesting-multimodal-v1.multimodal-fusion-remote-sensing-dataKomdigiITS-DFK3-Multimodallumos_multimodal_ground_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_ground_iterative.multimodal_hallucination_resultsmultimodal-ai-taxonomy
Multimodal AI Taxonomy
A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities.
Dataset Description
This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to:
Navigate the complex landscape of multimodal AI models
Filter models by specific input/output modality combinations
Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.lumos_multimodal_plan_iterative
🪄 Agent Lumos: Unified and Modular Training for Open-Source Language Agents
🌐[Website]
📝[Paper]
🤗[Data]
🤗[Model]
🤗[Demo]
We introduce 🪄Lumos, Language Agents with Unified Formats, Modular Design, and Open-Source LLMs. Lumos unifies a suite of complex interactive tasks and achieves competitive performance with GPT-4/3.5-based and larger open-source agents.
Lumos has following features:
🧩 Modular Architecture:
🧩 Lumos consists of planning, grounding… See the full description on the dataset page: https://huggingface.co/datasets/ai2lumos/lumos_multimodal_plan_iterative.FGVC_Aircraft_test_facebook_opt_350m_Attributes_Caption_ns_3333
Dataset Card for "FGVC_Aircraft_test_facebook_opt_350m_Attributes_Caption_ns_3333"
More Information needed
ai-code-multimodal-en
Dataset: AI Code Generation, Multimodal AI & Small Language Models (EN)
Description
English dataset covering AI coding assistants, multimodal AI, Small Language Models (SLMs), and GraphRAG.
This dataset is designed for research, training, and application development in the field of AI applied to software development and cybersecurity.
Articles Covered
AI Code Generation: Copilot, Cursor, Claude Code - Comparison of 12 leading AI coding assistants
Computer… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-en.ai-code-multimodal-fr
Dataset : IA Generation de Code, IA Multimodale & Small Language Models (FR)
Description
Dataset francophone couvrant les outils d'assistance au codage par IA, l'IA multimodale, les Small Language Models (SLM) et GraphRAG.
Ce dataset est concu pour la recherche, la formation et le developpement d'applications dans le domaine de l'IA appliquee au developpement logiciel et a la cybersecurite.
Articles couverts
IA pour la Generation de Code : Copilot, Cursor… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-code-multimodal-fr.Viet-multimodal-open-r1-8k-verifiedFGVC_Aircraft_train_embeddings
Dataset Card for "FGVC_Aircraft_train_embeddings"
More Information needed
AI4Math_MathVista-multimodal-rollout8SONDER-Mini-Portfolio-0001-Multimodal-AI-Training-Dataset================================================================
SONDER MINI PORTFOLIO 0001
================================================================
Dataset: SONDER Mini Portfolio 0001
Creator: Chanda Mandisa Lowrance, PhD
Released: 20260421
Version: 1.0
Price: $500.00
Purchase: https://chandamandisa.gumroad.com/l/sonder-portfolio-mini
================================================================
ABOUT… See the full description on the dataset page: https://huggingface.co/datasets/chandapanda/SONDER-Mini-Portfolio-0001-Multimodal-AI-Training-Dataset.
