datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lab-bench
LAB-Bench
The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.XLRS-Bench-lite
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite.ATM-Bench
ATM-Bench: Long-Term Personalized Referential Memory QA
ATM-Bench is the first benchmark for multimodal, multi-source personalized referential memory QA over long time horizons (~4 years) with evidence-grounded retrieval and answering.
Paper: According to Me: Long-Term Personalized Referential Memory QA
Overview
Existing long-term memory benchmarks focus primarily on dialogue history, failing to capture realistic personalized references grounded in lived experience.… See the full description on the dataset page: https://huggingface.co/datasets/Jingbiao/ATM-Bench.jee-neet-benchmark
JEE/NEET LLM Benchmark Dataset
🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost.
A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.UniDoc-Bench
UNIDOC-BENCH Dataset
A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG).
Dataset Description
UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.RPC-Bench
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
🌐 Project Page •
💻 GitHub •
📖 Paper
RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations.
Data Structure
RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/RPC-Bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.BenchCAD
BenchCAD
Three-config dataset for CAD evaluation:
edit-bench — held-out CAD edit benchmark.
code_gen — 17,900 synthetic CadQuery samples (compact 12-column variant)
covering 106 mechanical part families. Each row contains the GT CadQuery code
plus 5 normalized renders.
QA — CAD question-answering benchmark.
code_gen schema (12 columns)
Column
Type
Description
stem
string
unique sample identifier
family
string
mechanical part family (106 distinct)… See the full description on the dataset page: https://huggingface.co/datasets/BenchCAD/BenchCAD.LLaVA-NeXT-Interleave-Bench
LLaVA-Interleave Bench Dataset Card
Dataset details
Dataset type:
LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API.
It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs.
Dataset date:
LLaVA-Interleave Bench was collected in April 2024, and released in June 2024.
Paper or resources for more information:
Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.RefSpatial-Bench
🎉 RefSpatial-Expand-Bench is officially released!
The new version not only extends indoor scenes (e.g., factories, stores), but also introduces brand-new outdoor scenarios (e.g., streets, parking lots) — enabling more comprehensive evaluation of spatial referring tasks.
👉 Try it now: RefSpatial-Expand-Bench
🏆 The paper associated with this benchmark, RoboRefer, has been accepted to NeurIPS 2025!
Thank you all for your attention and support! 🙌… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/RefSpatial-Bench.SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation.
CogSense-Bench
CogSense-Bench
Project Page | Paper | GitHub
CogSense-Bench is a comprehensive visual question answering (VQA) benchmark designed to evaluate the cognitive capabilities of Multimodal Large Language Models (MLLMs). It was introduced in the paper "Toward Cognitive Supersensing in Multimodal Large Language Model".
The benchmark assesses MLLMs across five cognitive dimensions:
Fluid intelligence
Crystallized intelligence
Visuospatial cognition
Mental simulation
Visual routines… See the full description on the dataset page: https://huggingface.co/datasets/PediaMedAI/CogSense-Bench.Terminal-Bench-Hard
Terminal-Bench Hard
Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks.
The tasks cover software engineering, debugging, data processing, system
administration, security, scientific computing, and related command-line
workflows.
Contents
tasks/: runnable tasks in Harbor format.
metadata/tasks.parquet: searchable task metadata and instructions.
Each task directory contains task.toml, instruction.md, an
environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.TIR-Bench
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Introduction:
TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.MMSI-Bench
MMSI-Bench
This repo contains evaluation code for the paper "MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence"
🌐 Homepage | 🤗 Dataset | 📑 Paper | 💻 Code | 📖 arXiv
🔔News
🔥[2025-10-23]: We added the normalized human response time for each MMSI-Bench sample and its difficulty level to our dataset on Hugging Face.
🔥[2025-06-18]: MMSI-Bench has been supported in the LMMs-Eval repository.
✨[2025-06-11]: MMSI-Bench was used for evaluation in the… See the full description on the dataset page: https://huggingface.co/datasets/RunsenXu/MMSI-Bench.olmOCR_bench
Dataset Card for olmocr-bench
This is a FiftyOne dataset with 7019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/olmOCR_bench")
# Launch the App
session = fo.launch_app(dataset)
Here is the completed dataset card, filled in… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/olmOCR_bench.R-Bench
R-Bench
Introduction
R-Bench is a graduate-level multi-disciplinary benchmark for evaluating the complex reasoning capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). R stands for Reasoning.
According to statistics on R-Bench, the benchmark spans 19 departments, including mathematics, physics, biology, computer science, and chemistry, covering over 100 subjects such as Inorganic Chemistry, Chemical Reaction Kinetics, and… See the full description on the dataset page: https://huggingface.co/datasets/R-Bench/R-Bench.Act2Cap_benchmarkCollected data from GUI-Action-Narrator
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.Q-Spatial-Bench
Dataset Card for Q-Spatial Bench
Q-Spatial Bench is a benchmark designed to measure the quantitative spatial reasoning 📏 in large vision-language models.
🔥The paper associated with Q-Spatial Bench is accepted by EMNLP 2024 main track!
Our paper: Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models [arXiv link]
Project website: [link]
Dataset Details
Q-Spatial Bench is a benchmark designed to measure the… See the full description on the dataset page: https://huggingface.co/datasets/andrewliao11/Q-Spatial-Bench.MRAG-Bench
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models
🌐 Homepage | 📖 Paper | 💻 Evaluation
Intro
MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios, providing a robust and systematic evaluation of Large Vision Language Model (LVLM)’s vision-centric multimodal retrieval-augmented generation (RAG) abilities.
Results
Evaluated upon 10 open-source and 4 proprietary… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/MRAG-Bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.EVADE-Bench
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
🤗 Dataset | Paper | GitHub
E-commerce platforms increasingly rely on Large Language Models and Vision-Language Models to detect illicit or misleading product content. However, these models remain vulnerable to evasive content, which refers to inputs that superficially comply with platform policies while covertly conveying prohibited claims. Unlike traditional adversarial attacks that aim to… See the full description on the dataset page: https://huggingface.co/datasets/koenshen/EVADE-Bench.ppu-bench
PPU-Bench: Real-World Multimodal Benchmark for Personalized Partial Unlearning
PPU-Bench is a real-world multimodal benchmark designed to evaluate personalized partial unlearning in vision-language models. It supports multiple unlearning settings and provides training/evaluation data for different VLM backbones.
O3-Bench
Benchmarking High-Resolution, Multi-Hop Multimodal Reasoning over Digital Maps and Composite Charts
Can your AI agent truly "think with images"?
O3-Bench is an ICLR 2026 multimodal reasoning benchmark that evaluates visual search, fine-grained perception, and multi-hop reasoning over high-resolution digital maps and composite charts.
It tests how well an AI agent can truly "think with images" with interleaved attention to visual details.
O3-Bench is designed… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/O3-Bench.OmniEarth-Bench_MCQ
Dataset Summary
Each example provides:
Field
Type
Description
index
int32
Row ID
query
string
Prompt that embeds both the image context and the instruction template
question
string
Human-readable question without answer options
question_type
string
"Single Choice", "Multiple Choice"
options
list[string]
letter-labelled options
answer
string
Correct letter
image
list[Image]Images for each question, range from 1 to more than 20
L1-task..L4-task
string… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/OmniEarth-Bench_MCQ.MathCanvas-Bench
MathCanvas-Bench
🚀 Data Usage
from datasets import load_dataset
dataset = load_dataset("shiwk24/MathCanvas-Bench")
print(dataset)
📖 Introduction
MathCanvas-Bench is a challenging new benchmark designed to evaluate the intrinsic Visual Chain-of-Thought (VCoT) capabilities of Large Multimodal Models (LMMs). It serves as the primary evaluation testbed for the [MathCanvas] framework.… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Bench.
