datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InsightVQA
InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark
Overview
InsightVQA is a large-scale dataset designed for hierarchical visual question answering that bridges emotion understanding and cognitive reasoning. While existing benchmarks predominantly focus on surface-level emotion recognition , InsightVQA introduces a structured paradigm to evaluate a model's ability to interpret emotional causes, ground evidence, and reason about… See the full description on the dataset page: https://huggingface.co/datasets/ziyul707/InsightVQA.SlideVQA
SlideVQA
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
📖 arXiv 🌐 github
We introduce a new document VQA dataset, SlideVQA, for tasks wherein given a slide deck composed of multiple slide images and a corresponding question, a system selects a set of evidence images and answers the question.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{SlideVQA2023,
author = {Ryota Tanaka and… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/SlideVQA.InSight-doc-SFT-18k
InSight-doc-SFT-18k
Agentic Visual Perception for Long-Document Understanding
📄 Paper |
💻 Code |
🤗 Model |
🎯 RL Data |
🎬 Replay Demo |
🚀 Live Demo
Understand the big picture. Focus on the right details. Answer from the evidence.
InSight-doc-SFT-18k is the supervised fine-tuning corpus used to train the
InSight-doc long-document understanding agent. Each example is a complete
multimodal trajectory: the agent starts from low-resolution document… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/InSight-doc-SFT-18k.locomoInSight
InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations
InSight evaluates multimodal agents on claim verification over interactive visualizations. Given an interactive HTML visualization and a natural-language proposition, an agent must explore the visualization over multiple turns (clicking, hovering, navigating) and classify the proposition as True, False, or NotEnoughInfo.
Paper: arXiv:2609.01383
Code / benchmark harness:… See the full description on the dataset page: https://huggingface.co/datasets/maevehutch/InSight.VDocRetriever-Pretrain-DocStructLight-INSIGHT-Bench
INSIGHT-Bench v1
A human-curated object-goal navigation benchmark: 1,097 episodes over 210 scenes, each
episode a short natural-language instruction, a start pose, a goal position and a success radius,
defined on Z-up, metre-scaled USD conversions of four scene sources -- HM3D, Matterport3D,
InteriorGS and Habitat-GS (3D Gaussian Splatting). It is evaluated in NVIDIA Isaac Sim by
the INSIGHT-Bench evaluation SDK, which publishes exactly one coordinate over these bytes:… See the full description on the dataset page: https://huggingface.co/datasets/LightOriginsHQ/Light-INSIGHT-Bench.LaSeRSOpenDocVQA
Dataset Card for OpenDocVQA
This is a training and evaluation QA data file for VDocRAG, a new RAG framework that can directly understand diverse real-world documents purely from visual features.
Dataset Description
OpenDocVQA is the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats.
Supported Tasks and Leaderboards
Given a large collection of document images and a question, the… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/OpenDocVQA.veri-bilimci-insight-diyalog-tr-16.2k
🇹🇷 Veri Bilimci Insight Diyalog Veri Seti (TR, 16.2K) — %100 Türkçe Metin
Gerçek dünya blogları, uzman soru-cevap içerikleri ve akademik makale metinlerinden üretilmiş; veri madenciliği ve uygulamalı veri bilimi karar diline odaklanan, %100 Türkçe çok turlu diyalog veri seti.
🧠 Bu Veri Seti Ne Amaçla Üretildi?
Amaç, modeli teorik tanım ezberinden çıkarıp bağlama göre karar veren veri bilimci davranışına yaklaştırmaktır.
Her örnekte yöntem seçimi, alternatif kıyası… See the full description on the dataset page: https://huggingface.co/datasets/zero9tech/veri-bilimci-insight-diyalog-tr-16.2k.dfs-glossary
DFS Glossary — Amharic & Afaan Oromoo
Expert-verified glossaries of Digital Financial Services (DFS) terminology in
Amharic (am) and Afaan Oromoo (om), published as structured,
machine-readable, openly-licensed data.
Open language infrastructure for two low-resource Ethiopian languages — for
developers, researchers, translators, and the financial-inclusion community.
Languages
Amharic (am, Ge'ez script) · Afaan Oromoo (om, Latin script)
Entries
87 Amharic + 86… See the full description on the dataset page: https://huggingface.co/datasets/shega-insight/dfs-glossary.Slide_Insight_Images
About this Dataset
This Dataset contains data from several Presentation Slides. For each Slide the following information is available:
key: recordID_pdfNumber_slideNumber
image: each presentation slide as an PIL image
Zenodo Records Information
This repository contains data from Zenodo records.
Records
Zenodo Record 10008464Authors: Moore, JoshLicense: cc-by-4.0
Zenodo Record 10008465Authors: Moore, JoshLicense: cc-by-4.0
Zenodo Record… See the full description on the dataset page: https://huggingface.co/datasets/ScaDSAI/Slide_Insight_Images.Slide_Insight_Images_v2
About this Dataset
This Dataset contains several Presentation Slides as part of the NFDI4BIOIMAGE project SlideInsight to gain insights into presentation slides through multimodal AI models.
For each Slide the following information is available:
key: recordID_pdfNumber_slideNumber
image: each presentation slide as an PIL image
The corresponding embeddings and metadata can be found in this Huggingface Dataset.
Zenodo Records Information
This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/ScaDSAI/Slide_Insight_Images_v2.VisualMRC
VisualMRC
VisualMRC: Machine Reading Comprehension on Document Images
📖 arXiv 🌐 github
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
Citation and contact
If you use this dataset, please cite our work:
@inproceedings{VisualMRC2021,
author = {Ryota Tanaka and
Kyosuke Nishida and
Sen Yoshida},
title =… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/VisualMRC.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.youtube-comment-insights-chatml
YouTube Comment Insights - ChatML
Overview
This dataset contains instruction-tuning samples for structured YouTube comment analysis.
The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models.
Each sample contains:
sentiment
tone
pros
cons
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.OpenDocVQA-Corpus
Dataset Card for OpenDocVQA
This is a training and evaluation corpus data file for VDocRAG, a new RAG framework that can directly understand diverse real-world documents purely from visual features.
Dataset Description
OpenDocVQA is the first unified collection of open-domain document visual question answering datasets, encompassing diverse document types and formats.
Supported Tasks and Leaderboards
Given a large collection of document images and a question… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/OpenDocVQA-Corpus.data-scientist-insight-dialog-en-16.5k
🇬🇧 Data Scientist Insight Dialogue Dataset (EN, 16.5K) — 100% English Text
A decision-focused, multi-turn dataset built from real-world blogs, expert Q&A content, and academic papers, tuned for data mining and applied data science reasoning in full English text purity.
🧠 Why This Dataset Exists
The goal is to train decision behavior, not only definition recall.
Each sample is structured to expose method choice, trade-off reasoning, risk signals, and validation… See the full description on the dataset page: https://huggingface.co/datasets/zero9tech/data-scientist-insight-dialog-en-16.5k.ORD-RXN-InsightCMDR-Bench
Dataset Card for CMDR-Bench
CMDR-Bench is the first benchmark that jointly evaluates multi-page reasoning and multimodal document retrieval. It comprises 800 high-quality, human-annotated queries spanning four query categories and 255 documents across six domains, with an average document length of 183.5 pages.
Dataset Description
Task Definition
Given a query and a multi-page document, the task is to retrieve the top-k target pages that contain… See the full description on the dataset page: https://huggingface.co/datasets/NTT-hil-insight/CMDR-Bench.Engineering_Jobs_Insight_Dataset
Engineering Job Market Analysis Dataset
Link to Data Sourcing Code
The code used for data sourcing is publicly available on GitHub:
GitHub Repository Link
Executive Summary
Motivation and Potential Applications
The rapid advancement of technology has significantly increased the demand for engineering professionals, particularly in computer science and related fields. Understanding current job market trends is crucial for:
Job Seekers: Aligning… See the full description on the dataset page: https://huggingface.co/datasets/yiqing111/Engineering_Jobs_Insight_Dataset.ai_data_defi_protocols
💎 Freemium Feed Notice: This public Hugging Face feed provides verified daily samples.Need complete raw historical archives, custom company contact points, or real-time REST webhook feeds?🌐 Upgrade to Institutional Full Feeds at alphaville.space or email lead strategist Sloane Valentine at alphaville-insights@agentmail.to.
📊 DeFi Protocol Intelligence & Alpha
Curated & Maintained Autonomously by Sloane Valentine @ Alphaville InsightsContact: alphaville-insights@agentmail.to… See the full description on the dataset page: https://huggingface.co/datasets/Alphaville-Insights/ai_data_defi_protocols.Pashto-Social-Insight-Reasoning-Dataset
Pashto Social Insight & Reasoning Dataset (PSIR)
Overview
The Pashto Social Insight & Reasoning (PSIR) dataset is a specialized collection designed to evaluate and enhance the sociological reasoning, cultural dynamics understanding, and analytical capabilities of AI models in the Pashto language. Born from an incremental "snowball effect" curation process, it captures deep contextual insights into social structures and community reasoning.
Structure… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Social-Insight-Reasoning-Dataset.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
ai_data_dev_jobs
💎 Freemium Feed Notice: This public Hugging Face feed provides verified daily samples.Need complete raw historical archives, custom company contact points, or real-time REST webhook feeds?🌐 Upgrade to Institutional Full Feeds at alphaville.space or email lead strategist Sloane Valentine at alphaville-insights@agentmail.to.
🚀 Developer Job Market Intelligence
Curated & Maintained Autonomously by Sloane Valentine @ Alphaville InsightsContact: alphaville-insights@agentmail.to… See the full description on the dataset page: https://huggingface.co/datasets/Alphaville-Insights/ai_data_dev_jobs.ai_data_hf_models
💎 Freemium Feed Notice: This public Hugging Face feed provides verified daily samples.Need complete raw historical archives, custom company contact points, or real-time REST webhook feeds?🌐 Upgrade to Institutional Full Feeds at alphaville.space or email lead strategist Sloane Valentine at alphaville-insights@agentmail.to.
🤖 AI/ML Model Intelligence & Due Diligence
Curated & Maintained Autonomously by Sloane Valentine @ Alphaville InsightsContact:… See the full description on the dataset page: https://huggingface.co/datasets/Alphaville-Insights/ai_data_hf_models.circle-packing-insight-loop
Circle-Packing Insight-Exploration Loop
Artifacts from an iterative GPT solver <-> proposer insight-exploration loop on the
21-circles-in-a-perimeter-4-rectangle packing problem (AlphaEvolve SOTA sum-of-radii
= 2.3658321334167627). Each round, 16 solvers propose a program + written explanation;
every program is scored; a proposer then mines all 16 attempts into an evolving insight
document that conditions the next round. Run: 16 solvers x 8 rounds.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/ars22/circle-packing-insight-loop.youtube-comment-insights-clean
YouTube Comment Insights - Clean Dataset
Overview
This dataset contains structured YouTube comment analytics data designed for visualization, analytics, and machine learning workflows.
Each sample contains:
comment
sentiment
tone
pros
cons
The dataset is intended for easy readability and downstream analytics tasks.
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON format
Files… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-clean.africa-somalia-explosive-weapons-monitor-data-by-insecurity-insight-87665212
Explosive Weapons Monitor Data by Insecurity Insight | Africa (Insecurity Insight)
987 rows - 1 Africa country/area - 2018-2025 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 987 rows from Insecurity Insight, covering Explosive Weapons Monitor Data by Insecurity Insight. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-explosive-weapons-monitor-data-by-insecurity-insight-87665212.1027_insight_prediction_with_cot
