datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TerraMesh
TerraMesh
A planetary‑scale, multimodal analysis‑ready dataset for Earth‑Observation foundation models: TerraMesh merges data from Sentinel‑1 SAR, Sentinel‑2 optical, Copernicus DEM, NDVI, and land‑cover sources into more than 9 million co‑registered patches ready for large‑scale representation learning.
You find more information about the data sampling and preprocessing in our paper: TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data.
Samples from the TerraMesh… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh.ChartNet
ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding
🌐 Homepage | 📖 arXiv
📝 Changelog
June 3, 2026 — Release of grounded_qa subset and completed reasoning subset (both subject to Notice Regarding Data Availability)
May 15, 2026 — Added link to 30K real-world charts and detailed captions dataset released by our collaborators Abaka AI/2077AI.
April 29, 2026 — Release of an additional 2.5 million row subset core_permissive (subject to… See the full description on the dataset page: https://huggingface.co/datasets/ibm-granite/ChartNet.duorc
Dataset Card for duorc
Dataset Summary
The DuoRC dataset is an English language dataset of questions and answers gathered from crowdsourced AMT workers on Wikipedia and IMDb movie plots. The workers were given freedom to pick answer from the plots or synthesize their own answers. It contains two sub-datasets - SelfRC and ParaphraseRC. SelfRC dataset is built on Wikipedia movie plots solely. ParaphraseRC has questions written from Wikipedia movie plots and the answers are… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/duorc.acp_bench
ACP Bench
🏠 Homepage •
📄 Paper •
📄 Paper
ACPBench is a benchmark dataset designed to evaluate the reasoning capabilities of large language models (LLMs) in the context of Action, Change, and Planning. It spans 13 diverse domains:
Blocksworld
Logistics
Grippers
Grid
Ferry
FloorTile
Rovers
VisitAll
Depot
Goldminer
Satellite
Swap
Alfworld
Task Types in ACPBench
ACPBench includes the following 8 reasoning tasks:
Action Applicability (app)… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/acp_bench.MedMentions-ZSAuto-BenchmarkCard
Dataset Card for Auto-BenchmarkCard
A catalog of validated AI evaluation benchmark descriptions, generated as an LLM-assisted, human-reviewed summary of the capabilities, attributes, and risks of the AI benchmarks. Each card provides structured metadata about benchmark purpose, methodology, data sources, risks, and limitations.
Dataset Details
Dataset Description
This dataset is a previous version. The current, maintained Auto-BenchmarkCard dataset… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Auto-BenchmarkCard.VAREX
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.mgsm
Dataset Card for MGSM
Dataset Summary
Copy and merge of this MGSM Dataset, Catalan version, Basque version, Galician version, but in training samples we removed the prompt formatting, e.g. removed Question: ... in question field or Answer: ... in answer field.
Multilingual Grade School Math Benchmark (MGSM) is a benchmark of grade-school math problems, proposed in the paper Language models are multilingual chain-of-thought reasoners.
The same 250 problems from GSM8K are… See the full description on the dataset page: https://huggingface.co/datasets/jbross-ibm-research/mgsm.cif-dataset
Cracks in the Foundation
A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories:
Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling
Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one.
Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples.
Splits
Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.TerraMesh-Masks
TerraMesh-Masks
TerraMesh-Masks is a dataset for open-vocabulary segmentation of satellite imagery. This dataset provides binary segmentation masks with captions that extend the samples from TerraMesh.
We also provide an human-verfied evaluation benchmark, called TerraMesh-Masks-Eval.
Examples from the training subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt. For development, you can… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks.Marco-Bench-MIF-languagesstruct-text
StructText — SEC_WikiDB & SEC_WikiDB_subset
Dataset card for the VLDB 2025 TaDA-workshop submission “StructText: A
Synthetic Table-to-Text Approach for Benchmark Generation with
Multi-Dimensional Evaluation” (under review).
from datasets import load_dataset
# default = SEC_WikiDB_unfiltered_all
ds = load_dataset(
"ibm-research/struct-text",
trust_remote_code=True)
# a specific configuration
subset = load_dataset(
"ibm-research/struct-text"… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/struct-text.Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.VDR_ibm-research_REAL-MM-RAG
VDR_ibm-research_REAL-MM-RAG - Overview
Dataset Summary
VDR_ibm-research_REAL-MM-RAG is a multimodal dataset that combines text and image data, and support tasks such as DSE retrieval (RAG).
Dataset Creation
This dataset is a merge and shuffle of the following datasets in the VDR format:
ibm-research/REAL-MM-RAG_TechSlides
ibm-research/REAL-MM-RAG_TechReport
ibm-research/REAL-MM-RAG_FinTabTrainSet
ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_ibm-research_REAL-MM-RAG.Sombench-WAC-Crater-Detection
SomBench Benchmark: Robbins Crater Detection, WAC
Science theme: Impact processes
Task: Object detection
Dataset Summary
An impact-crater object-detection benchmark built from the
Robbins (2019) global lunar crater
catalog, a manually compiled, near-complete census of
lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are
converted to bounding boxes and packaged over LROC WAC visible tiles drawn
from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.nestful
NESTFUL: Nested Function-Calling Dataset
NESTFUL is a benchmark to evaluate LLMs on nested sequences of API calls, i.e., sequences where the output of one API call is passed as input to
a subsequent call.
The NESTFUL dataset includes over 1800 nested sequences from two main areas: mathematical reasoning and coding tools. The mathematical reasoning portion is generated from
the MathQA dataset, while the coding portion is generated from the
StarCoder2-Instruct dataset.
All… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/nestful.REAL-MM-RAG_FinReport
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport.REAL-MM-RAG_TechReport
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport.REAL-MM-RAG_TechSlides
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides.bookcorpus-wikitext-ccnews-tinystories-chunkedbookcorpus-wikitext-ccnews-sometinystories-chunkedwatsonxDocsQA
Data Description
Corpus Dataset
The corpus dataset contains the following fields:
Field
Description
doc_id
Unique identifier for the document
title
Document title as it appears on the HTML page
document
Textual representation of the content
md_document
Markdown representation of the content
url
Origin URL of the document
Question-Answers Dataset
The QA dataset includes these fields:
Field
Description
question_id
Unique… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/watsonxDocsQA.REAL-MM-RAG_TechSlides_BEIR
BEIR Version of REAL-MM-RAG_TechSlides
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_TechSlides
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_TechSlides
Content: 62 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides_BEIR.REAL-MM-RAG_FinSlides_BEIR
BEIR Version of REAL-MM-RAG_FinSlides
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_FinSlides
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_FinSlides
Content: 65 quarterly… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides_BEIR.REAL-MM-RAG_FinReport_BEIR
BEIR Version of REAL-MM-RAG_FinReport
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_FinReport
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_FinReport
Content: 19 financial… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport_BEIR.REAL-MM-RAG_TechReport_BEIR
BEIR Version of REAL-MM-RAG_TechReport
Summary
This dataset is the BEIR-compatible version of the following Hugging Face dataset:
ibm-research/REAL-MM-RAG_TechReport
It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits.
REAL-MM-RAG_TechReport
Content: 17 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport_BEIR.REAL-MM-RAG_FinSlides
REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark
We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides.REAL-MM-RAG_FinTabTrainSet
REAL-MM-RAG_FinTabTrainSet
We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM) and filtered by a large language model (LLM). We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents.
For more information, see the project page:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet.TerraMesh-Masks-Eval
TerraMesh-Masks-Eval
TerraMesh-Masks-Eval is a human-verified benchmark dataset to evaluate open-vocabulary segmentation models on satellite imagery. This dataset provides binary segmentation masks with captions togehther with input samples from TerraMesh.
We also provide a training dataset, called TerraMesh-Masks.
Examples from the evaluation subset:
Usage
Download the data loading code from GitHub and install requirements with pip install -r requirements.txt.… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh-Masks-Eval.900K-Judgements
900K Judgements: A Large-Scale LLM-as-a-Judge Evaluation Dataset
Dataset Description
This dataset contains approximately 900,000 pairwise comparison judgements from multiple LLM judges evaluating model responses. The data was collected as part of the paper [`Mediocrity is the key for LLM as a Judge Anchor Selection'](https://arxiv.org/abs/2603.16848), investigating the impact of anchor selection in LLM-as-a-judge pairwise evaluation.
Dataset Summary
Total… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/900K-Judgements.
