datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CharXiv
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
NeurIPS 2024
🏠Home (🚧Still in construction) | 🤗Data | 🥇Leaderboard | 🖥️Code | 📄Paper
This repo contains the full dataset for our paper CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs, which is a diverse and challenging chart understanding benchmark fully curated by human experts. It includes 2,323 high-resolution charts manually sourced from arXiv preprints. Each chart is… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/CharXiv.MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.OmniGAIA
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is designed to evaluate long-horizon, multi-hop, open-form problem solving in realistic settings rather than short perception-only QA.
Benchmark Construction
The OmniGAIA construction… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/OmniGAIA.M3SciQA
🧑🔬 M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark For Evaluating Foundatio Models
EMNLP 2024 Findings
🖥️ Code
Introduction
In the realm of foundation models for scientific research, current benchmarks predominantly focus on single-document, text-only tasks and fail to adequately represent the complex workflow of such research. These benchmarks lack the $\textit{multi-modal}$, $\textit{multi-document}$ nature of scientific research, where comprehension… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/M3SciQA.ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.CHASE-QA
CHASE: Challenging AI with Synthetic Evaluations
The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.vlm_circuitVLM Circuit Datasets Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs Yaniv Nikankin ⋅ Dana Arad ⋅ Yossi Gandelsman ⋅ Yonatan Belinkov https://neurips.cc/virtual/2025/loc/san-diego/poster/119472
MathGames
🧮 MathGames
MathGames is a novel benchmark of 2,183 high-quality mathematical problems—both text-only and multimodal—designed to evaluate Large Language Models (LLMs) on open-ended mathematical and logical reasoning tasks.
It accompanies our EMNLP 2025 Main Track paper:
📄 Can Large Language Models Win the International Mathematical Games?
🌍 Overview
MathGames consists of carefully curated problems sourced from the International Mathematical and Logical Games… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/MathGames.
