datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JailBreakV-28k
⛓💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
🌐 GitHub | 🛎 Project Page | 👉 Download full datasets
If you like our project, please give us a star ⭐ on Hugging Face for the latest update.
📰 News
Date
Event
2024/07/09
🎉 Our paper is accepted by COLM 2024.
2024/06/22
🛠️ We have updated our version to V0.2, which supports users to customize their attack models… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k.StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.global-piqa-nonparallel
Global PIQA Non-Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems.
In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements.
Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.MMAD
MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
💡 This dataset is the full version of MMAD
Content:Containing both questions, images, and captions.
Questions: All questions are presented in a multiple-choice format with manual verification, including options and answers.
Images:Images are collected from the following links:
DS-MVTec
, MVTec-AD
, MVTec-LOCO
, VisA
, GoodsAD.
We retained the mask… See the full description on the dataset page: https://huggingface.co/datasets/jiang-cc/MMAD.global-piqa-parallel
Global PIQA Parallel
Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world.
The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems.
In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.V
WorldMedQA-V: A Multilingual, Multimodal Medical Examination Dataset
Overview
WorldMedQA-V is a multilingual and multimodal benchmarking dataset designed to evaluate vision-language models (VLMs) in healthcare contexts. The dataset includes medical examination questions from four countries—Brazil, Israel, Japan, and Spain—in both their original languages and English translations. Each multiple-choice question is paired with a corresponding medical image, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WorldMedQA/V.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.SNAP
SNAP Benchmark
Code and annotations: [https://github.com/ykotseruba/SNAP]
SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings.
This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms.
SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.PEACE
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
[Code] [Paper] [Data]
Introduction
We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table.
Property
Description
Source
USGS(English)
CGS(Chinese)
Content
Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.SocialNav-SUB
SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation
This is the accompying dataset for the Social Navigation Scene Understanding Benchmark (SocialNav-SUB) which is a Visual Question Answering (VQA) dataset and benchmark designed to evaluate Vision-Language Models (VLMs) for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across… See the full description on the dataset page: https://huggingface.co/datasets/michaelmunje/SocialNav-SUB.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.M4Benchifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.GLUE3D
GLUE3D: General Language Understanding Evaluation for 3D Point Clouds
Data repository containing all necessary data for the GLUE3D evaluation benchmark.
GLUE3D is a Q&A benchmark for evaluation of 3D-LLMs object understanding capabilities. It is built around 128 richly textured
surfaces spanning creatures, objects, architecture and transport. Each surface is provided as a 50 k-point RGB point cloud,
a 8K-point RGB point cloud, a 512 × 512 RGB rendering, and five RGB-D multiviews.… See the full description on the dataset page: https://huggingface.co/datasets/giorgio-mariani-1/GLUE3D.sop-bench
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
📄 Paper: SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
🏭 Human Expert-Authored SOPs · 🤖 Human-AI Collaborative Framework · 📊 Executable Interfaces · 🔧 Two Agent Architectures · 📈 11 Frontier Models Evaluated
Dataset Summary
SOP-Bench is a comprehensive benchmark for evaluating LLM-based agents on complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial… See the full description on the dataset page: https://huggingface.co/datasets/amazon/sop-bench.MuSP-Bench
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Official benchmark website
Modalities
Each question specifies the minimum source of musical evidence needed to answer it:
S (score): answer from the written score.
P (performance): answer from the performance recording.
S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.UiPad
UiPad - UI Parsing and Accessibility Dataset
📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned.
Curated by: MacPaw Way Ltd.
Language(s): Mostly EN, UA
License: MIT
Overview
UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.MarineLife-16K
Dataset Card for MarineLife-16K
We introduce MarineLife-16K, a marine-domain video benchmark designed to evaluate the video understanding capabilities of Vision-Language Models (VLMs). MarineLife-16K contains 2,000 video-text pairs and 16,080 video-question-answer pairs across a collection of 2,000 marine videos, including 12,080 multiple-choice questions and 4,000 open-ended questions. The benchmark emphasizes specialized marine knowledge, visual reasoning, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MarineLife-16K/MarineLife-16K.AraDiCE
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic.
As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.JMedQA
JMedQA: Benchmarking Large Language Models and Vision-Language Models on the Japanese Medical Licensing Examination
JMedQA is a Japanese medical question-answering benchmark derived from Japan's National Medical Examination materials publicly released by the Ministry of Health, Labour and Welfare (MHLW).
The dataset supports both text-only large language model (LLM) evaluation and vision-language model (VLM) evaluation using associated examination images.
Its image-dependency… See the full description on the dataset page: https://huggingface.co/datasets/SIP-med-LLM/JMedQA.StreamingBench-Slice
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.MuSP-Bench
MuSP-Bench
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Contents
data/questions.csv: all 490 questions, accepted answers, and the
response contract for each.
inputs/pdf/without_context/: one context-removed PDF per piece.
inputs/images/: rendered score-page images for every piece.
inputs/abc/: one ABC score per piece.
inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.OmniBrainBench
OmniBrainBench
🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper
This repository is the official implementation of the paper [OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks].
🚀 News
[02/2026] Our OmniBrainBench is accepted by CVPR2026!
[12/2025] We have released the evaluation code and dataset for OmniBrainBench.
[11/2025] The manuscript can be found on arXiv.
🚀Overview
we introduce… See the full description on the dataset page: https://huggingface.co/datasets/FrankPN/OmniBrainBench.MolPuzzle_data
MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation
Dataset Description:
The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.EndoBench
EndoBench
🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper
This repository is the official implementation of the paper EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis.
🚀 News
[03/2026] We release the EndoVQA-Instruct Dataset at here.
[21/10/2025] We release a new open-set challenging VQA benchmark EndoBench-extended.
[19/09/2025] 🎉🎉Our EndoBench was accepted by NeurIPS'25 D&B Track!!!
☀️ Tutorial
EndoBench is… See the full description on the dataset page: https://huggingface.co/datasets/Saint-lsy/EndoBench.
