CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01futurehouse /lab-bench LAB-Bench The Language Agent Biology Benchmark, or LAB-Bench, is an evaluation dataset for AI systems intended to benchmark capabilities foundational to scientific research in biology. The dataset currently consists of 8 broad categories, comprising 30 narrower subtasks, including extracting information from the scientific literature (LitQA2), retrieving information from databases (DbQA) and supplementary information (SuppQA), reasoning about scientific figures (FigQA) and tables… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/lab-bench.imagequestion-answering1K<n<10K51 likes33k downloads1y agoHugging Face02ShaofantuoshuzhengzhiSha /GUIGuard-Bench GUIGuard-Bench (Public Ladder) GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents. This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots. For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F. Dataset Summary GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.imagequestion-answering1K<n<10K0 likes10k downloads5mo agoHugging Face03bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face04initiacms /XLRS-Bench-lite 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite.imagevisual-question-answering1K<n<10K4 likes5.9k downloads11mo agoHugging Face05Jingbiao /ATM-Bench ATM-Bench: Long-Term Personalized Referential Memory QA ATM-Bench is the first benchmark for multimodal, multi-source personalized referential memory QA over long time horizons (~4 years) with evidence-grounded retrieval and answering. Paper: According to Me: Long-Term Personalized Referential Memory QA Overview Existing long-term memory benchmarks focus primarily on dialogue history, failing to capture realistic personalized references grounded in lived experience.… See the full description on the dataset page: https://huggingface.co/datasets/Jingbiao/ATM-Bench.imagequestion-answering1K<n<10K7 likes4k downloads6mo agoHugging Face06Reja1 /jee-neet-benchmark JEE/NEET LLM Benchmark Dataset 🏆 View the live leaderboard → — interactive results across JEE Advanced, JEE Main & NEET, with open/closed-weight badges, contamination flags, and per-run cost. A benchmark for evaluating vision-capable LLMs on Indian competitive exam questions (JEE Advanced & NEET). Each question is the original exam image; models answer via the OpenRouter API and are scored with authentic, exam-specific marking schemes — including partial credit for JEE… See the full description on the dataset page: https://huggingface.co/datasets/Reja1/jee-neet-benchmark.imagevisual-question-answeringn<1K16 likes3.6k downloads3mo agoHugging Face07Salesforce /UniDoc-Bench UNIDOC-BENCH Dataset A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG). Dataset Description UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.imagequestion-answering1K<n<10K15 likes3.3k downloads10mo agoHugging Face08zai-org /RPC-Bench RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension 🌐 Project Page • 💻 GitHub • 📖 Paper RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations. Data Structure RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/RPC-Bench.imagequestion-answering100K<n<1M2 likes3.3k downloads4mo agoHugging Face09sylvainHellin /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.documentquestion-answering1K<n<10K20 likes3k downloads3d agoHugging Face10BenchCAD /BenchCAD BenchCAD Three-config dataset for CAD evaluation: edit-bench — held-out CAD edit benchmark. code_gen — 17,900 synthetic CadQuery samples (compact 12-column variant) covering 106 mechanical part families. Each row contains the GT CadQuery code plus 5 normalized renders. QA — CAD question-answering benchmark. code_gen schema (12 columns) Column Type Description stem string unique sample identifier family string mechanical part family (106 distinct)… See the full description on the dataset page: https://huggingface.co/datasets/BenchCAD/BenchCAD.imageimage-to-text10K<n<100K19 likes2.6k downloads3mo agoHugging Face11lmms-lab-encoder /LLaVA-NeXT-Interleave-Bench LLaVA-Interleave Bench Dataset Card Dataset details Dataset type: LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API. It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs. Dataset date: LLaVA-Interleave Bench was collected in April 2024, and released in June 2024. Paper or resources for more information: Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.imagevisual-question-answering10K<n<100K15 likes2.6k downloads2y agoHugging Face12BAAI /RefSpatial-Bench 🎉 RefSpatial-Expand-Bench is officially released! The new version not only extends indoor scenes (e.g., factories, stores), but also introduces brand-new outdoor scenarios (e.g., streets, parking lots) — enabling more comprehensive evaluation of spatial referring tasks. 👉 Try it now: RefSpatial-Expand-Bench 🏆 The paper associated with this benchmark, RoboRefer, has been accepted to NeurIPS 2025! Thank you all for your attention and support! 🙌… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/RefSpatial-Bench.imagequestion-answeringn<1K20 likes2.4k downloads2mo agoHugging Face13franky-veteran /SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation. imagequestion-answering1K<n<10K3 likes2.3k downloads7mo agoHugging Face14PediaMedAI /CogSense-Bench CogSense-Bench Project Page | Paper | GitHub CogSense-Bench is a comprehensive visual question answering (VQA) benchmark designed to evaluate the cognitive capabilities of Multimodal Large Language Models (MLLMs). It was introduced in the paper "Toward Cognitive Supersensing in Multimodal Large Language Model". The benchmark assesses MLLMs across five cognitive dimensions: Fluid intelligence Crystallized intelligence Visuospatial cognition Mental simulation Visual routines… See the full description on the dataset page: https://huggingface.co/datasets/PediaMedAI/CogSense-Bench.imageimage-text-to-text1K<n<10K0 likes2k downloads8mo agoHugging Face15Zhongzhi1228 /Terminal-Bench-Hard Terminal-Bench Hard Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks. The tasks cover software engineering, debugging, data processing, system administration, security, scientific computing, and related command-line workflows. Contents tasks/: runnable tasks in Harbor format. metadata/tasks.parquet: searchable task metadata and instructions. Each task directory contains task.toml, instruction.md, an environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.imagequestion-answeringn<1K0 likes2k downloads2mo agoHugging Face16Agents-X /TIR-Bench TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Introduction: TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.imagequestion-answering1K<n<10K3 likes1.7k downloads9mo agoHugging Face17RunsenXu /MMSI-Bench MMSI-Bench This repo contains evaluation code for the paper "MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence" 🌐 Homepage | 🤗 Dataset | 📑 Paper | 💻 Code | 📖 arXiv 🔔News 🔥[2025-10-23]: We added the normalized human response time for each MMSI-Bench sample and its difficulty level to our dataset on Hugging Face. 🔥[2025-06-18]: MMSI-Bench has been supported in the LMMs-Eval repository. ✨[2025-06-11]: MMSI-Bench was used for evaluation in the… See the full description on the dataset page: https://huggingface.co/datasets/RunsenXu/MMSI-Bench.imagequestion-answering1K<n<10K17 likes1.5k downloads11mo agoHugging Face18Voxel51 /olmOCR_bench Dataset Card for olmocr-bench This is a FiftyOne dataset with 7019 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/olmOCR_bench") # Launch the App session = fo.launch_app(dataset) Here is the completed dataset card, filled in… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/olmOCR_bench.imagequestion-answering1K<n<10K0 likes1.3k downloads7mo agoHugging Face19R-Bench /R-Bench R-Bench Introduction R-Bench is a graduate-level multi-disciplinary benchmark for evaluating the complex reasoning capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). R stands for Reasoning. According to statistics on R-Bench, the benchmark spans 19 departments, including mathematics, physics, biology, computer science, and chemistry, covering over 100 subjects such as Inorganic Chemistry, Chemical Reaction Kinetics, and… See the full description on the dataset page: https://huggingface.co/datasets/R-Bench/R-Bench.imagequestion-answering1K<n<10K22 likes994 downloads1y agoHugging Face20FRank62Wu /Act2Cap_benchmarkCollected data from GUI-Action-Narrator imagequestion-answeringn<1K0 likes890 downloads1y agoHugging Face21bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes737 downloads2y agoHugging Face22andrewliao11 /Q-Spatial-Bench Dataset Card for Q-Spatial Bench Q-Spatial Bench is a benchmark designed to measure the quantitative spatial reasoning 📏 in large vision-language models. 🔥The paper associated with Q-Spatial Bench is accepted by EMNLP 2024 main track! Our paper: Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models [arXiv link] Project website: [link] Dataset Details Q-Spatial Bench is a benchmark designed to measure the… See the full description on the dataset page: https://huggingface.co/datasets/andrewliao11/Q-Spatial-Bench.imagequestion-answeringn<1K7 likes689 downloads2y agoHugging Face23uclanlp /MRAG-Bench MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models 🌐 Homepage | 📖 Paper | 💻 Evaluation Intro MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios, providing a robust and systematic evaluation of Large Vision Language Model (LVLM)’s vision-centric multimodal retrieval-augmented generation (RAG) abilities. Results Evaluated upon 10 open-source and 4 proprietary… See the full description on the dataset page: https://huggingface.co/datasets/uclanlp/MRAG-Bench.imagequestion-answering1K<n<10K13 likes639 downloads2y agoHugging Face24orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes612 downloads2y agoHugging Face25SiloLink /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.documentquestion-answering1K<n<10K0 likes606 downloads24d agoHugging Face26koenshen /EVADE-Bench EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection 🤗 Dataset | Paper | GitHub E-commerce platforms increasingly rely on Large Language Models and Vision-Language Models to detect illicit or misleading product content. However, these models remain vulnerable to evasive content, which refers to inputs that superficially comply with platform policies while covertly conveying prohibited claims. Unlike traditional adversarial attacks that aim to… See the full description on the dataset page: https://huggingface.co/datasets/koenshen/EVADE-Bench.imagetext-classification10K<n<100K2 likes571 downloads8mo agoHugging Face27closerG /ppu-bench PPU-Bench: Real-World Multimodal Benchmark for Personalized Partial Unlearning PPU-Bench is a real-world multimodal benchmark designed to evaluate personalized partial unlearning in vision-language models. It supports multiple unlearning settings and provides training/evaluation data for different VLM backbones. imagequestion-answering100K<n<1M0 likes568 downloads5mo agoHugging Face28m-Just /O3-Bench Benchmarking High-Resolution, Multi-Hop Multimodal Reasoning over Digital Maps and Composite Charts Can your AI agent truly "think with images"? O3-Bench is an ICLR 2026 multimodal reasoning benchmark that evaluates visual search, fine-grained perception, and multi-hop reasoning over high-resolution digital maps and composite charts. It tests how well an AI agent can truly "think with images" with interleaved attention to visual details. O3-Bench is designed… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/O3-Bench.imagequestion-answeringn<1K17 likes526 downloads2mo agoHugging Face29initiacms /OmniEarth-Bench_MCQ Dataset Summary Each example provides: Field Type Description index int32 Row ID query string Prompt that embeds both the image context and the instruction template question string Human-readable question without answer options question_type string "Single Choice", "Multiple Choice" options list[string] letter-labelled options answer string Correct letter image list[Image]Images for each question, range from 1 to more than 20 L1-task..L4-task string… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/OmniEarth-Bench_MCQ.imagequestion-answering1K<n<10K0 likes510 downloads1y agoHugging Face30shiwk24 /MathCanvas-Bench MathCanvas-Bench                   🚀 Data Usage from datasets import load_dataset dataset = load_dataset("shiwk24/MathCanvas-Bench") print(dataset) 📖 Introduction MathCanvas-Bench is a challenging new benchmark designed to evaluate the intrinsic Visual Chain-of-Thought (VCoT) capabilities of Large Multimodal Models (LMMs). It serves as the primary evaluation testbed for the [MathCanvas] framework.… See the full description on the dataset page: https://huggingface.co/datasets/shiwk24/MathCanvas-Bench.imageimage-text-to-text1K<n<10K0 likes494 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.