CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JailbreakV-28K /JailBreakV-28k ⛓‍💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks 🌐 GitHub | 🛎 Project Page | 👉 Download full datasets If you like our project, please give us a star ⭐ on Hugging Face for the latest update. 📰 News Date Event 2024/07/09 🎉 Our paper is accepted by COLM 2024. 2024/06/22 🛠️ We have updated our version to V0.2, which supports users to customize their attack models… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k.imagetext-generation10K<n<100K70 likes24k downloads2y agoHugging Face02mjuicem /StreamingBench StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.imagequestion-answering1K<n<10K13 likes12k downloads1y agoHugging Face03bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9k downloads2y agoHugging Face04mrlbenchmarks /global-piqa-nonparallel Global PIQA Non-Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The non-parallel split covers 136 language varieties, covering five continents, 18 language families, and 24 writing systems. In this non-parallel split, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. Details are in our preprint:… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-nonparallel.imagequestion-answering10K<n<100K40 likes6.4k downloads4mo agoHugging Face05jiang-cc /MMAD MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection 💡 This dataset is the full version of MMAD Content:Containing both questions, images, and captions. Questions: All questions are presented in a multiple-choice format with manual verification, including options and answers. Images:Images are collected from the following links: DS-MVTec , MVTec-AD , MVTec-LOCO , VisA , GoodsAD. We retained the mask… See the full description on the dataset page: https://huggingface.co/datasets/jiang-cc/MMAD.imagequestion-answering10K<n<100K17 likes4.2k downloads1y agoHugging Face06mrlbenchmarks /global-piqa-parallel Global PIQA Parallel Global PIQA is a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The parallel split is a multi-parallel dataset for 131 language varieties, covering five continents, 16 language families, and 23 writing systems. In this parallel split, each example was machine-translated from English, then manually corrected by a native speaker of the target language.… See the full description on the dataset page: https://huggingface.co/datasets/mrlbenchmarks/global-piqa-parallel.imagequestion-answering10K<n<100K10 likes4k downloads4mo agoHugging Face07sylvainHellin /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.documentquestion-answering1K<n<10K20 likes3k downloads13d agoHugging Face08bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes752 downloads2y agoHugging Face09WorldMedQA /V WorldMedQA-V: A Multilingual, Multimodal Medical Examination Dataset Overview WorldMedQA-V is a multilingual and multimodal benchmarking dataset designed to evaluate vision-language models (VLMs) in healthcare contexts. The dataset includes medical examination questions from four countries—Brazil, Israel, Japan, and Spain—in both their original languages and English translations. Each multiple-choice question is paired with a corresponding medical image, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WorldMedQA/V.imagequestion-answering1K<n<10K19 likes743 downloads2y agoHugging Face10orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes596 downloads2y agoHugging Face11SiloLink /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.documentquestion-answering1K<n<10K0 likes487 downloads20d agoHugging Face12ykotseruba /SNAP SNAP Benchmark Code and annotations: [https://github.com/ykotseruba/SNAP] SNAP (stands for Shutter speed, ISO seNsitivity, and APerture) is a new benchmark consisting of images of objects taken under controlled lighting conditions and with densely sampled camera settings. This benchmark allows testing the effects of capture bias, which includes camera settings and illumination, on performance of vision algorithms. SNAP contains 37,558 images of 100 scenes (10 scenes per 10 object… See the full description on the dataset page: https://huggingface.co/datasets/ykotseruba/SNAP.imageimage-classification10K<n<100K0 likes373 downloads1y agoHugging Face13jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes371 downloads5mo agoHugging Face14microsoft /PEACE PEACE: Empowering Geologic Map Holistic Understanding with MLLMs [Code] [Paper] [Data] Introduction We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table. Property Description Source USGS(English) CGS(Chinese) Content Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.imagequestion-answering1K<n<10K22 likes349 downloads2y agoHugging Face15michaelmunje /SocialNav-SUB SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation This is the accompying dataset for the Social Navigation Scene Understanding Benchmark (SocialNav-SUB) which is a Visual Question Answering (VQA) dataset and benchmark designed to evaluate Vision-Language Models (VLMs) for scene understanding in real-world social robot navigation scenarios. SocialNav-SUB provides a unified framework for evaluating VLMs against human and rule-based baselines across… See the full description on the dataset page: https://huggingface.co/datasets/michaelmunje/SocialNav-SUB.imagequestion-answeringn<1K0 likes345 downloads1y agoHugging Face16bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes340 downloads2y agoHugging Face17Anonymous8976 /M4Benchimagequestion-answering1K<n<10K3 likes319 downloads1y agoHugging Face18quenfly /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.documentquestion-answering1K<n<10K1 likes289 downloads2mo agoHugging Face19giorgio-mariani-1 /GLUE3D GLUE3D: General Language Understanding Evaluation for 3D Point Clouds Data repository containing all necessary data for the GLUE3D evaluation benchmark. GLUE3D is a Q&A benchmark for evaluation of 3D-LLMs object understanding capabilities. It is built around 128 richly textured surfaces spanning creatures, objects, architecture and transport. Each surface is provided as a 50 k-point RGB point cloud, a 8K-point RGB point cloud, a 512 × 512 RGB rendering, and five RGB-D multiviews.… See the full description on the dataset page: https://huggingface.co/datasets/giorgio-mariani-1/GLUE3D.imagetext-generation1K<n<10K1 likes271 downloads11mo agoHugging Face20amazon /sop-bench SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents 📄 Paper: SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents 🏭 Human Expert-Authored SOPs · 🤖 Human-AI Collaborative Framework · 📊 Executable Interfaces · 🔧 Two Agent Architectures · 📈 11 Frontier Models Evaluated Dataset Summary SOP-Bench is a comprehensive benchmark for evaluating LLM-based agents on complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial… See the full description on the dataset page: https://huggingface.co/datasets/amazon/sop-bench.imagetext-classification1K<n<10K1 likes266 downloads4mo agoHugging Face21bryel-labs /MuSP-Bench MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance MuSP-Bench is a 490-question benchmark for musical score understanding, performance listening, and combined score-performance reasoning. Official benchmark website Modalities Each question specifies the minimum source of musical evidence needed to answer it: S (score): answer from the written score. P (performance): answer from the performance recording. S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.imagequestion-answeringn<1K1 likes262 downloads25d agoHugging Face22macpaw-research /UiPad UiPad - UI Parsing and Accessibility Dataset 📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned. Curated by: MacPaw Way Ltd. Language(s): Mostly EN, UA License: MIT Overview UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.imagequestion-answering1K<n<10K16 likes223 downloads1mo agoHugging Face23MarineLife-16K /MarineLife-16K Dataset Card for MarineLife-16K We introduce MarineLife-16K, a marine-domain video benchmark designed to evaluate the video understanding capabilities of Vision-Language Models (VLMs). MarineLife-16K contains 2,000 video-text pairs and 16,080 video-question-answer pairs across a collection of 2,000 marine videos, including 12,080 multiple-choice questions and 4,000 open-ended questions. The benchmark emphasizes specialized marine knowledge, visual reasoning, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MarineLife-16K/MarineLife-16K.imagevisual-question-answeringn<1K0 likes217 downloads3mo agoHugging Face24QCRI /AraDiCE AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.imagetext-classificationn<1K3 likes194 downloads2y agoHugging Face25SIP-med-LLM /JMedQA JMedQA: Benchmarking Large Language Models and Vision-Language Models on the Japanese Medical Licensing Examination JMedQA is a Japanese medical question-answering benchmark derived from Japan's National Medical Examination materials publicly released by the Ministry of Health, Labour and Welfare (MHLW). The dataset supports both text-only large language model (LLM) evaluation and vision-language model (VLM) evaluation using associated examination images. Its image-dependency… See the full description on the dataset page: https://huggingface.co/datasets/SIP-med-LLM/JMedQA.imagequestion-answering1K<n<10K0 likes194 downloads27d agoHugging Face26jayzhu486 /StreamingBench-Slice StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding 🏠 Project Page | 📄 arXiv Paper | 📦 Dataset | 🏅Leaderboard StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟 [NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output. [NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.imagequestion-answering1K<n<10K1 likes190 downloads6mo agoHugging Face27milan477 /MuSP-Bench MuSP-Bench MuSP-Bench is a 490-question benchmark for musical score understanding, performance listening, and combined score-performance reasoning. Contents data/questions.csv: all 490 questions, accepted answers, and the response contract for each. inputs/pdf/without_context/: one context-removed PDF per piece. inputs/images/: rendered score-page images for every piece. inputs/abc/: one ABC score per piece. inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.imagequestion-answeringn<1K1 likes178 downloads1mo agoHugging Face28FrankPN /OmniBrainBench OmniBrainBench 🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper This repository is the official implementation of the paper [OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks]. 🚀 News [02/2026] Our OmniBrainBench is accepted by CVPR2026! [12/2025] We have released the evaluation code and dataset for OmniBrainBench. [11/2025] The manuscript can be found on arXiv. 🚀Overview we introduce… See the full description on the dataset page: https://huggingface.co/datasets/FrankPN/OmniBrainBench.imagequestion-answering1K<n<10K3 likes163 downloads6mo agoHugging Face29kguo2 /MolPuzzle_data MolPuzzle: A Multimodal Benchmark for Molecular Structure Elucidation Dataset Description: The MolPuzzle dataset is a newly developed resource designed to challenge Large Language Models with multi-modal, multi-step reasoning tasks (molecular structure elucidation). This dataset consists of 217 diverse and intricate structure elucidation challenges that require LLMs to demonstrate advanced reasoning capabilities, integrating multimodal data and deep chemical understanding… See the full description on the dataset page: https://huggingface.co/datasets/kguo2/MolPuzzle_data.imagequestion-answering10K<n<100K0 likes152 downloads2y agoHugging Face30Saint-lsy /EndoBench EndoBench 🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper This repository is the official implementation of the paper EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis. 🚀 News [03/2026] We release the EndoVQA-Instruct Dataset at here. [21/10/2025] We release a new open-set challenging VQA benchmark EndoBench-extended. [19/09/2025] 🎉🎉Our EndoBench was accepted by NeurIPS'25 D&B Track!!! ☀️ Tutorial EndoBench is… See the full description on the dataset page: https://huggingface.co/datasets/Saint-lsy/EndoBench.imagequestion-answering1K<n<10K8 likes145 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.