datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlideVQA
SlideVQA
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
📖 arXiv 🌐 github
We introduce a new document VQA dataset, SlideVQA, for tasks wherein given a slide deck composed of multiple slide images and a corresponding question, a system selects a set of evidence images and answers the question.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{SlideVQA2023,
author = {Ryota Tanaka and… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/SlideVQA.InSight-doc-SFT-18k
InSight-doc-SFT-18k
Agentic Visual Perception for Long-Document Understanding
📄 Paper |
💻 Code |
🤗 Model |
🎯 RL Data |
🎬 Replay Demo |
🚀 Live Demo
Understand the big picture. Focus on the right details. Answer from the evidence.
InSight-doc-SFT-18k is the supervised fine-tuning corpus used to train the
InSight-doc long-document understanding agent. Each example is a complete
multimodal trajectory: the agent starts from low-resolution document… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/InSight-doc-SFT-18k.VDocRetriever-Pretrain-DocStructEarthReason
EarthReason📂:
The first large-scale benchmark dataset for geospatial pixel reasoning
📥Dwonload Dataset:
git lfs install
git clone https://huggingface.co/datasets/earth-insights/EarthReason
📦Additional Resources:
paper: ArXiv
project: earth-insights/SegEarth-R1
🚀Citation:
@article{li2025segearth,
title={SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model},
author={Li, Kaiyu and Xin, Zepeng and Pang, Li and Pang, Chao and… See the full description on the dataset page: https://huggingface.co/datasets/earth-insights/EarthReason.Slide_Insight_Images
About this Dataset
This Dataset contains data from several Presentation Slides. For each Slide the following information is available:
key: recordID_pdfNumber_slideNumber
image: each presentation slide as an PIL image
Zenodo Records Information
This repository contains data from Zenodo records.
Records
Zenodo Record 10008464Authors: Moore, JoshLicense: cc-by-4.0
Zenodo Record 10008465Authors: Moore, JoshLicense: cc-by-4.0
Zenodo Record… See the full description on the dataset page: https://huggingface.co/datasets/ScaDSAI/Slide_Insight_Images.Slide_Insight_Images_v2
About this Dataset
This Dataset contains several Presentation Slides as part of the NFDI4BIOIMAGE project SlideInsight to gain insights into presentation slides through multimodal AI models.
For each Slide the following information is available:
key: recordID_pdfNumber_slideNumber
image: each presentation slide as an PIL image
The corresponding embeddings and metadata can be found in this Huggingface Dataset.
Zenodo Records Information
This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/ScaDSAI/Slide_Insight_Images_v2.VisualMRC
VisualMRC
VisualMRC: Machine Reading Comprehension on Document Images
📖 arXiv 🌐 github
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{VisualMRC2021,
author = {Ryota Tanaka and
Kyosuke Nishida and
Sen Yoshida},
title =… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/VisualMRC.OpenDocVQA-Corpus
Dataset Card for OpenDocVQA
This is a training and evaluation corpus data file for VDocRAG, a new RAG framework that can directly understand diverse real-world documents purely from visual features.
Dataset Description
OpenDocVQA is the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats.
Supported Tasks and Leaderboards
Given a large collection of document images and a question… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/OpenDocVQA-Corpus.CMDR-Bench
Dataset Card for CMDR-Bench
CMDR-Bench is the first benchmark that jointly evaluates multi-page reasoning and multimodal document retrieval. It comprises 800 high-quality, human-annotated queries spanning four query categories and 255 documents across six domains, with an average document length of 183.5 pages.
Dataset Description
Task Definition
Given a query and a multi-page document, the task is to retrieve the top-k target pages that contain… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/CMDR-Bench.INSIGHTfixposV3_Keyframe
📊 Extraction Results & Issues
This section summarizes the frame extraction process, specifically detailing episodes where object-centric cropping was skipped due to the absence of Guide (Semantic Mask) information.
1. Summary Statistics
Total Episodes Processed: 3,917
Episodes without Guide (Crop Skipped): 125
Successfully Cropped: 3,792
2. Detailed Skip List by Task
The following episodes were processed without the Object-centric cropping step because no… See the full description on the dataset page: https://huggingface.co/datasets/Whalswp/INSIGHTfixposV3_Keyframe.airbnb-reviewsInsightEditinsight9-calibration-20260923
Looper Insight 9 Calibration Capture
Two calibration-oriented recordings captured with a Looper Insight 9 on 2026-09-23. The release contains native-orientation stereo grayscale images, RGB JPEG images, device timestamps, raw IMU measurements, auxiliary device-estimated poses, capture metadata, quality reports, review videos, and validated ROS1 bags.
This is a raw sensor-data release, not a calibrated benchmark or ground-truth trajectory dataset. No intrinsics, extrinsics, or… See the full description on the dataset page: https://huggingface.co/datasets/nerako/insight9-calibration-20260923.DE-DatasetDE-BenchmarkCMDR-Synth
Dataset Card for CMDR-Synth
This dataset is the training set for CMDR-Embed, comprising 39,796 query-image pairs over 4,754 documents.
Dataset Structure
Data Fields
Queries
{
'reasoning_type': String of reasoning type,
'query': String of query,
'pos_images': List of retrieval target page image paths,
'context_images': List of context page image paths that provide the necessary information to identify the 'pos_images'… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/CMDR-Synth.INSIGHTfixposV3_Keyframe_right_ver
📊 Extraction Results & Issues
This section summarizes the frame extraction process, specifically detailing episodes where object-centric cropping was skipped due to the absence of Guide (Semantic Mask) information.
1. Summary Statistics
Total Episodes Processed: 3,917
Episodes without Guide (Crop Skipped): 129
Successfully Cropped: 3,788
2. Detailed Skip List by Task
The following episodes were processed without the Object-centric cropping step because no… See the full description on the dataset page: https://huggingface.co/datasets/Whalswp/INSIGHTfixposV3_Keyframe_right_ver.
