datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChartNet
ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding
🌐 Homepage | 📖 arXiv
📝 Changelog
June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability)
May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI.
April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.Landslide4sense
Landslide4Sense
Dataset Description
This dataset is originally introduced in GitHub repo Landslide4Sense-2022.
The Landslide4Sense dataset has three splits, training/validation/test, consisting of 3799, 245, and 800 image patches, respectively. Each image patch is a composite of 14 bands that include:
Multispectral data from Sentinel-2: B1, B2, B3, B4, B5, B6, B7, B8, B9, B10, B11, B12.
Slope data from ALOS PALSAR: B13.
Digital elevation model (DEM) from ALOS… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/Landslide4sense.VAREX
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.hls_burn_scarsThis dataset contains Harmonized Landsat and Sentinel-2 imagery of burn scars and the associated masks for the years 2018-2021 over the contiguous United States. There are 804 512x512 scenes. Its primary purpose is for training geospatial machine learning models.cif-dataset
Cracks in the Foundation
A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories:
Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling
Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one.
Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples.
Splits
Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.VDR_ibm-research_REAL-MM-RAG
VDR_ibm-research_REAL-MM-RAG - Overview
Dataset Summary
VDR_ibm-research_REAL-MM-RAG is a multimodal dataset that combines text and image data, and support tasks such as DSE retrieval (RAG).
Dataset Creation
This dataset is a merge and shuffle of the following datasets in the VDR format:
ibm-research/REAL-MM-RAG_TechSlides
ibm-research/REAL-MM-RAG_TechReport
ibm-research/REAL-MM-RAG_FinTabTrainSet
ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_ibm-research_REAL-MM-RAG.Sombench-Ice-Prospectivity-Regression
SomBench Benchmark: Polar Ice Prospectivity Regression
Science theme: Polar volatiles
Task: Regression
Dataset Summary
A polar, multi-layer benchmark for predicting near-surface water-ice
prospectivity within ~10° latitude of each pole at 240 m/pixel. Following
the ice-prospectivity workflow of Coyan et al. (2025), the dataset includes a
group of physically motivated evidential layers (thermophysical,
illumination, and terrain) alongside a continuous prospectivity… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-Ice-Prospectivity-Regression.hurricane
Data Format Description for Hurricane Evaluation on Prithvi WxC
Overview
To evaluate the performance of Prithvi WxC on hurricanes, the surface and pressure data from the MERRA-2 dataset, comprising 160 variables used in training, is required. The complete evaluation dataset includes 75 different initial conditions for hurricanes that formed in the Atlantic Ocean between 2017 and 2023.
The scientific objective is to assess the zero-shot performance of Prithvi WxC in… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/hurricane.hls_merra2_gppFlux
Dataset Summary:
This dataset consists of Harmonized Landsat and Sentinel-2 multispectral reflectance imagery and MERRA-2 observations centered around eddy covariance flux towers and the corresponding Gross Primary Productivity (GPP) data at the towers. Its purpose is to serve as a finetuning dataset for geospatial foundation models for the task of regressing GPP flux observations from HLS and MERRA-2 data.
Dataset Structure:
The dataset consists of:
(1) HLS 6-band Tiff… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/hls_merra2_gppFlux.REAL-MM-RAG_FinReport
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport.Sombench-WAC-Crater-Detection
SomBench Benchmark: Robbins Crater Detection, WAC
Science theme: Impact processes
Task: Object detection
Dataset Summary
An impact-crater object-detection benchmark built from the
Robbins (2019) global lunar crater
catalog, a manually compiled, near-complete census of
lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are
converted to bounding boxes and packaged over LROC WAC visible tiles drawn
from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.REAL-MM-RAG_TechReport
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport.REAL-MM-RAG_TechSlides
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides.REAL-MM-RAG_TechSlides_BEIR
BEIR Version of REAL-MM-RAG_TechSlides
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_TechSlides
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_TechSlides
Content: 62 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides_BEIR.REAL-MM-RAG_FinReport_BEIR
BEIR Version of REAL-MM-RAG_FinReport
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_FinReport
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_FinReport
Content: 19 financial… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport_BEIR.REAL-MM-RAG_FinSlides_BEIR
BEIR Version of REAL-MM-RAG_FinSlides
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_FinSlides
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_FinSlides
Content: 65 quarterly… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides_BEIR.REAL-MM-RAG_TechReport_BEIR
BEIR Version of REAL-MM-RAG_TechReport
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_TechReport
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_TechReport
Content: 17 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport_BEIR.REAL-MM-RAG_FinTabTrainSet
REAL-MM-RAG_FinTabTrainSet
We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM) and filtered by a large language model (LLM). We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents.
For more information, see the project page:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet.AIX360-CEM-MAF-data
Data for uscase with AIX360 CEM-MAF
All data here corresponds to the notebook CEM-MAF example. The notebook automatically downloads the data. For each image, there are 3 files, a .png image file, a .npy file use to view image file, and a .npy file containg a latent features that can be use to generate the image with appropriate GAN. See notebook for more details. Images are based on a GAN trained on the CelebA [1] dataset of celebrity faces.
[1] Ziwei Liu, Ping Luo, Xiaogang… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AIX360-CEM-MAF-data.REAL-MM-RAG_FinSlides
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides.Sombench-IMP-Segmentation
SomBench Benchmark: Irregular Mare Patch (IMP) Segmentation
Science theme: Volcanic history
Task: Binary semantic segmentation
Dataset Summary
A binary semantic-segmentation benchmark for irregular mare patches
(IMPs): rare, morphologically subtle features interpreted as unusually young
volcanic landforms. Each sample is an LROC NAC image tile paired with a
binary IMP mask (IMP vs. background). The set is derived from published IMP
polygon annotations, framed as a… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-IMP-Segmentation.REAL-MM-RAG_FinTabTrainSet_rephrased
REAL-MM-RAG_FinTabTrainSet_rephrased
We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM. We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents. This is the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased.IBMR
IBMR Dataset: Insulator Burn Mark RGB-Point Cloud Dataset
Directory Structure
IBMR/
│
├── Sample 1/
│ ├── sample-1.png
│ ├── sample-1.pcd
│ └── GT/
│ ├── insulator-1.txt
│ └── burn_mark-1.txt
├── Sample 2/
│ ├── sample-2.png
│ ├── sample-2.pcd
│ └── GT/
│ ├── insulator-2.txt
│ └── burn_mark-2.txt
└── ...
Citation
If you use this dataset, please cite our paper:
@article{tang2025dimensional,
title={Dimensional Compensation… See the full description on the dataset page: https://huggingface.co/datasets/Junqiu-Tang/IBMR.burn_intensity
Dataset Summary
This dataset contains burn scar intensity data and Harmonized Landsat and Sentinel-2 (HLS) images for burn scar analysis across various time frames: pre-burn, during-burn, and post-burn.
Each file provides spatial information on burn scar intensity and top-of-atmosphere (TOA) reflectance values.
The dataset includes:
BS_files_raw.csv: The complete set of burn scar intensity data without filtering.
BS_files_with_less_than_25_percent_zeros.csv: Filtered dataset with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-nasa-geospatial/burn_intensity.ibm-hls-burn-originalSynthetic_Executable_Companion_IBM_Common_Stock_April_17_2026
Synthetic Executable Companion: IBM Common Stock Intraday Price Simulation from a Google Finance Snapshot
This repository contains a DBbun-generated synthetic simulation bundle built from a Google Finance snapshot of IBM Common Stock (NYSE: IBM). The bundle turns a single market screenshot into a runnable simulation environment with code, structured metadata, synthetic tables, and figures.
"AI Turned an IBM Stock Chart Into 500 Market Simulations" Video:… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/Synthetic_Executable_Companion_IBM_Common_Stock_April_17_2026.Examples
Data Examples
This repository incudes samples for TerraMind demos at https://github.com/IBM/terramind.
ibm-hls-burn-vectorizedibm-handwriting-campaign-word
IBM Handwriting Campaign Word Dataset
Source dataset: docling-project/ibm-handwriting-campaign
This directory contains a Hugging Face dataset export generated from the project source handwriting data.
Overview
Splits: train, validation, test
Total samples: 841 (train: 671, validation: 83, test: 87)
Format: Parquet files with images stored as bytes
Each record corresponds to a single word extracted from scanned forms and includes the word image, annotation JSON… See the full description on the dataset page: https://huggingface.co/datasets/Felix92/ibm-handwriting-campaign-word.puext690refExp_simp
Dataset Card for "puext690refExp_simp"
More Information needed
