datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
propagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.Multi-modal-Self-instruct
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Evaluation
Citation
You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip.
Dataset Description
Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.Kairos-Multimodal-Reasoning
A dataset for training models in multimodal reasoning tasks
Usage
from datasets import load_dataset
ds = load_dataset("Aquiles-ai/Kairos-Multimodal-Reasoning")
print(ds.features)
print(ds["train"]["source"])
Preview of dataset examples
We've built a playground so you can see some of the examples included in the dataset.
Link: https://kairos-example.vercel.app/
Dataset used in the blog post: Kairos: Building a Multimodal Model with LFM2.5 and… See the full description on the dataset page: https://huggingface.co/datasets/Aquiles-ai/Kairos-Multimodal-Reasoning.Multimodal-Robustness-BenchmarkMedical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.GeoperceptionEuclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
Dataset Card for Geoperception
A Benchmark for Low-level Geometric Perception
Dataset Details
Dataset Description
Geoperception is a benchmark focused specifically on accessing model's low-level visual perception ability in 2D geometry.
It is sourced from the Geometry-3K corpus, which offers precise logical forms for geometric diagrams, compiled from popular high-school… See the full description on the dataset page: https://huggingface.co/datasets/euclid-multimodal/Geoperception.Multimodal-STEM-HLE-plus-plus
multimodal-STEM-HLE++
A high-value multimodal STEM dataset designed and empirically proven to push state-of-the-art LLMs beyond their current limits.
Explore the full multimodal-STEM-HLE++ dataset: https://go.turing.com/mm-stem-hle
Why This Dataset
Post-training with RL is now the primary driver of frontier model improvement. The bottleneck is finding data at the right difficulty for current SOTA models. MMLU is saturated (>90%). HLE, once considered unsolvable, is now… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Multimodal-STEM-HLE-plus-plus.Iranian_olympiad_of_informatics_multimodal_questionssolarhive-community-solar-multimodal
SolarHive Community Solar Dataset
Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded).
A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.ck12-tqa-multimodal
CK-12 TQA Multimodal: Textbook Question Answering with Images
Dataset Description
Dataset Summary
CK-12 TQA Multimodal is a comprehensive multimodal dataset for science education, containing 26,260 questions paired with 6,206 images from middle school science textbooks. This dataset is sourced from CK-12 Foundation's open educational resources and includes both text-only questions and diagram-based visual reasoning questions.
This is the complete multimodal… See the full description on the dataset page: https://huggingface.co/datasets/notefill/ck12-tqa-multimodal.dutch-central-exam-mcq-multimodal-subsetMultimodal Multiple Choice Questions of the Dutch Central Exam 1999-2024
What?
This dataset contains only multimodal multiple choice questions from the Dutch Central Exam (High School level). From Wikipedia:
The Eindexamen (Dutch pronunciation: [ˈɛi̯ntɛksamən]) or centraal examen (CE) is the matriculation exam in the Netherlands, which takes place in a student's final year of high school education (voortgezet onderwijs; "continued education"). The exam is regulated by the Dutch Secondary… See the full description on the dataset page: https://huggingface.co/datasets/jjzha/dutch-central-exam-mcq-multimodal-subset.multimodal-privacy
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
Recent advances in multi-modal Large Language Models (M-LLMs) have demonstrated a powerful ability to synthesize implicit information from disparate sources, including images and text. These resourceful data from social media also introduce a significant and underexplored privacy risk: the inference of sensitive personal attributes from seemingly daily media content. However, the lack of benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/xaddh/multimodal-privacy.Driving_License_Nepali_MultimodalThis dataset is curated as part of the Cohere4AI project called "Multimodal-Multilingual Exam Collection".
GATE_2022_Multimodal
GATE 2022 MULTIMODAL
This dataset has been curated as part of Cohere For AI's multimodal examination benchmark creation.
multimodal_mcq_greek_physics
