datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlideVQA
SlideVQA
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
📖 arXiv 🌐 github
We introduce a new document VQA dataset, SlideVQA, for tasks wherein given a slide deck composed of multiple slide images and a corresponding question, a system selects a set of evidence images and answers the question.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{SlideVQA2023,
author = {Ryota Tanaka and… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/SlideVQA.InSight-doc-SFT-18k
InSight-doc-SFT-18k
Agentic Visual Perception for Long-Document Understanding
📄 Paper |
💻 Code |
🤗 Model |
🎯 RL Data |
🎬 Replay Demo |
🚀 Live Demo
Understand the big picture. Focus on the right details. Answer from the evidence.
InSight-doc-SFT-18k is the supervised fine-tuning corpus used to train the
InSight-doc long-document understanding agent. Each example is a complete
multimodal trajectory: the agent starts from low-resolution document… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/InSight-doc-SFT-18k.InSight-doc-RL-19k
InSight-doc-RL-19k
Agentic Visual Perception for Long-Document Understanding
📄 Paper |
💻 Code |
🤗 Model |
🧩 SFT Data |
🎬 Replay Demo |
🚀 Live Demo
Understand the big picture. Focus on the right details. Answer from the evidence.
InSight-doc-RL-19k is the reinforcement-learning corpus used to
train the InSight-doc long-document understanding agent after SFT. It contains
challenging document VQA prompts, unanswerable negatives, multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/InSight-doc-RL-19k.OpenDocVQA
Dataset Card for OpenDocVQA
This is a training and evaluation QA data file for VDocRAG, a new RAG framework that can directly understand diverse real-world documents purely from visual features.
Dataset Description
OpenDocVQA is the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats.
Supported Tasks and Leaderboards
Given a large collection of document images and a question, the… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/OpenDocVQA.VisualMRC
VisualMRC
VisualMRC: Machine Reading Comprehension on Document Images
📖 arXiv 🌐 github
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{VisualMRC2021,
author = {Ryota Tanaka and
Kyosuke Nishida and
Sen Yoshida},
title =… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/VisualMRC.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.OpenDocVQA-Corpus
Dataset Card for OpenDocVQA
This is a training and evaluation corpus data file for VDocRAG, a new RAG framework that can directly understand diverse real-world documents purely from visual features.
Dataset Description
OpenDocVQA is the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats.
Supported Tasks and Leaderboards
Given a large collection of document images and a question… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/OpenDocVQA-Corpus.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.expert-insights
Expert Insights
Expert profiles for Beau, Tate, and Wendy Thompson with specializations.
Details
Records: 3
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Thompson Mortgage Group
Publisher: Thompson Mortgage Group
Thompson Alpha Logic
Deep expert entity profiles with NMLS credentials, specialization routing, branded insight labels (Wendy's Wisdom, Beau's Brief, Tate's Take), and citation formats. Designed for AI entity disambiguation… See the full description on the dataset page: https://huggingface.co/datasets/wendymthompson/expert-insights.Problem-Solving-Insights-Based-on-Kazakh-Traditions
🇰🇿 Problem-Solving Insights Based on Kazakh Traditions
📖 Overview
Problem-Solving Insights Based on Kazakh Traditions is a instruction-tuning dataset designed to bridge the gap between ancient Kazakh wisdom and modern societal challenges.
📊 Dataset Statistics
General Metrics
Metric
Count
Total Samples
8,005
Total Words (approx.)
4,030,925
Avg. Words per Sample
503
Word Count Distribution (Per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Problem-Solving-Insights-Based-on-Kazakh-Traditions.NLP_Insights_2023_2024
Key Insignts from NLP papers (2023 -2024)
This dataset is processed and compiled by @hu_yifei as part of open-source effort from the Open Research Assistant Project.
It includes key insights extracted from top tier NLP conference papers:
Year
Venue
Paper Count
2024
eacl
225
2024
naacl
564
2023
acl
1077
2023
conll
41
2023
eacl
281
2023
emnlp
1048
2023
semeval
319
2023
wmt
101
Dataset Stats
Total number of papers: 3,640
Total rows (key… See the full description on the dataset page: https://huggingface.co/datasets/yifeihu/NLP_Insights_2023_2024.medical-insights-en-zh
Bilingual Medical Insights (EN-ZH) · 医学中英双语知识卡片样本
This dataset contains 50 bilingual English-Chinese medical insight cards, curated for AI model training, education, and research.
本数据集包含50条中英文医学知识卡片样本,适用于人工智能训练、医学教学与结构化语义分析任务。
✅ Fields Included | 字段结构
title / title_zh — Medical topic / 医学主题
narrative / narrative_zh — Context or background / 医学背景介绍
arguments / arguments_zh — Key points or findings / 论点要点
primary_theme — Major medical discipline (e.g. Psychiatry… See the full description on the dataset page: https://huggingface.co/datasets/jundai2003/medical-insights-en-zh.
