datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/stanford_kuka_multimodal_dataset.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.sigma_aldrich_safety_data-multimodalmultimodal_qa_dataset_v1Procedural-City-Multimodal-Dataset-V2
Synthetic Urban Multimodal Dataset - Volume 2 (v0.2)
This is the second major release of the Synthetic Urban Multimodal Dataset, part of an ongoing observation project on AI-generated data ecosystems.
Volume 2 introduces a completely new collection of urban environments to the series, providing 40,000 high-precision files (8,000 completely new unique scenes × 5 modalities) generated entirely via the "Constructive Furnace" (our custom Blender-Python pipeline).
What's New… See the full description on the dataset page: https://huggingface.co/datasets/jp-cypress/Procedural-City-Multimodal-Dataset-V2.multimodal_qa_dataset_v4_trainmultimodal_qa_dataset_v3_trainstanford_kuka_multimodal_datasetThis dataset was created using LeRobot.
meta/info.json
{
"codebase_version": "v2.0",
"data_path": "data/chunk-{episode_chunk:03d}/train-{episode_index:05d}-of-{total_episodes:05d}.parquet",
"robot_type": "unknown",
"total_episodes": 3000,
"total_frames": 149985,
"total_tasks": 1,
"total_videos": 3000,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:3000"
},
"keys": [
"observation.state"… See the full description on the dataset page: https://huggingface.co/datasets/aliberts/stanford_kuka_multimodal_dataset.Multimodal-RL-Data
Multimodal-Cold-Start Dataset
This dataset, named Multimodal-Cold-Start, is a crucial component of the two-stage training approach presented in the paper Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start. It is specifically used for the first stage: supervised fine-tuning (SFT) as a "cold start."
Curated by distilling data from Qwen2.5-VL-32B using rejection sampling, this dataset is designed to instill structured chain-of-thought reasoning patterns in… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/Multimodal-RL-Data.multimodal_qa_dataset_v2_image_focusMultiModalDataset
Dataset Card for MultiModal Dataset
Dataset Description
Dataset Summary
MultiModal Dataset is a curated collection of 85,000 samples spanning three modalities: text, images, and audio. It combines high-quality web content, image-caption pairs from COCO 2017, and audio samples from AudioSet to enable comprehensive multimodal model training and evaluation.
The dataset is organized into three subsets:
fineweb: 37,500 high-quality web text samples (>8… See the full description on the dataset page: https://huggingface.co/datasets/lv12/MultiModalDataset.dehset-vahset-super-dataset-v6-multimodalgemma4-multimodal-recipe-dataset
🍳 Gemma 4 Multimodal Recipe & Food Dataset
A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation.
🔗 Upstream & Source Datasets
This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources:
Dataset
Modality
Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.multimodal-rag-pharma-datamultimodal-dataset
Text and Image Retrieval Dataset
This dataset is designed for text and image retrieval tasks. It consists of parsed documents (corpus), generated queries, and relevance judgments (qrels).
Dataset Structure
The dataset contains three configurations: corpus, queries, and qrels.
1. Corpus (corpus)
Contains the document pages with their text and image content. The images are stored directly within the Parquet files.
corpus_id (string): Unique identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/eagerworks/multimodal-dataset.enriched-movie-dataset-with-multimodal-embeddings
Enriched Movie Dataset with Multimodal Embeddings
Dataset Description
This dataset provides rich metadata for over 44,000 movies, with a primary focus on providing a pre-computed, high-quality multimodal content embedding for each film.
It was created by fusing two popular Kaggle datasets: "The Movies Dataset" and the "IMDB Multimodal Vision & NLP Genre Classification" dataset. It has been further enriched with parsed text features and a unique 512-dimensional vector… See the full description on the dataset page: https://huggingface.co/datasets/ujwal-jibhkate/enriched-movie-dataset-with-multimodal-embeddings.Multimodal_urban_livability_evaluation_dataset
Multimodal Urban Livability Evaluation Dataset
📝 Dataset Description
This multimodal dataset is designed for multi-task urban livability evaluation, combining remote sensing imagery (RS), digital surface models (DSM), night light remote sensing (NLRS) imagery, and point-of-interest (POI) text data to predict multiple aspects of urban livability.
It covers various urban areas in the Netherlands and provides a rich set of ground-truth labels for 6 dimensions of livability.… See the full description on the dataset page: https://huggingface.co/datasets/Vinjou/Multimodal_urban_livability_evaluation_dataset.Multimodal-test-dataset-technicalindicators
KRX Investment Warning Prediction Dataset (OHLCV + Technical Indicators + Korean News)
Dataset Summary
This dataset is a test dataset for predicting Investment Warning (투자주의종목) designations in the Korean stock market (KRX).
It contains raw daily OHLCV price data, 13 technical indicators, and Korean news text (title + body), designed for multimodal anomaly detection / binary classification.
Important: No normalization/scaling is applied. All values are raw.
Date range:… See the full description on the dataset page: https://huggingface.co/datasets/k-datasoft/Multimodal-test-dataset-technicalindicators.multimodal_nli_dataset
Dataset Card for Multimodal and Multilingual NLI Dataset
Dataset Details
Paper: Beyond Similarity Scoring: Detecting Entailment and Contradiction in Multilingual and Multimodal Contexts, Interspeech 2025
Dataset Description
The Multimodal and Multilingual NLI Dataset supports multilingual and multimodal Natural Language Inference (NLI). It enables classification of entailment, contradiction, and neutrality across four modality combinations:
Text-Text (T-T)… See the full description on the dataset page: https://huggingface.co/datasets/oist/multimodal_nli_dataset.multimodal_doc_datamultimodal-modality-conflict-datasetMultimodal-test-dataset
KRX Investment Warning Prediction Dataset (OHLCV + Korean News)
Dataset Summary
This dataset is a test dataset for predicting Investment Warning (투자주의종목) designations in the Korean stock market (KRX).It contains raw daily OHLCV price data and Korean news text (title + body), designed for multimodal anomaly detection / binary classification.
Important: No technical indicators are included, and no normalization/scaling is applied.
Date range: 2025-07-01 ~ 2025-09-30… See the full description on the dataset page: https://huggingface.co/datasets/k-datasoft/Multimodal-test-dataset.multimodal-oncology-atlas
LH2 Data — Multimodal Oncology Dataset
A large-scale, multimodal oncology dataset built around a principle rare in the field: placing non-Caucasian patient populations at the centre, not the periphery.
Dataset Summary
The vast majority of oncology datasets used to train diagnostic, prognostic, and treatment AI models are drawn overwhelmingly from Caucasian, Western populations — a well-documented limitation that undermines model generalisability and equity in real-world… See the full description on the dataset page: https://huggingface.co/datasets/LH2-data-labs/multimodal-oncology-atlas.Multimodal-test-dataset
TIPS Multimodal Test Dataset
Dataset Description
This dataset contains test data for multimodal stock prediction using time series and text data.
Dataset Structure
labels: Binary labels for stock price prediction (0: down/neutral, 1: up)
time_series: Time series features for stock data
texts: Korean news text data related to stocks
Data Splits
This is the test split of the TIPS multimodal dataset.
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/makedeltakmlee/Multimodal-test-dataset.orbura-autoscientist-dataviz-multimodal-pilotmultimodal-spectroscopic-seed-datasets-fullmultimodal_sample_datahumor-multimodal-dataset-fixedspace-multimodal-dataset
Space Vision Dataset (Multimodal)
Overview
The Space Vision Dataset is a multimodal dataset consisting of images paired with descriptive captions. It includes a diverse collection of space-related imagery such as planetary views, telescopes, galaxies, and Mars rover scenes.
This dataset is designed for tasks like:
Image Captioning
Vision-Language Modeling (VLM)
Multimodal Retrieval
Contrastive Learning
Dataset Structure
Each sample contains:
image_id —… See the full description on the dataset page: https://huggingface.co/datasets/AIOmarRehan/space-multimodal-dataset.MultiModal_TISER_train-dataset
