CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ShaofantuoshuzhengzhiSha /GUIGuard-Bench GUIGuard-Bench (Public Ladder) GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents. This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots. For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F. Dataset Summary GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.imagequestion-answering1K<n<10K0 likes10k downloads5mo agoHugging Face02johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes6.8k downloads8mo agoHugging Face03mirobody /ESL-Bench ESL-bench ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework. ⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.textquestion-answering1K<n<10K18 likes6.7k downloads19d agoHugging Face04mirobody /MedHall-Bench MedHall-Bench MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework. ⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.textquestion-answeringn<1K6 likes4.4k downloads1mo agoHugging Face05mirobody /MedHarm-Bench MedHarm-Bench MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework. ⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.textquestion-answeringn<1K2 likes4.2k downloads1mo agoHugging Face06hashmortar /spreadsheet-bench-v2-modified SpreadsheetBench V2 Modified: Multi-Document QA 1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents. An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.documentquestion-answering1K<n<10K2 likes3.4k downloads18d agoHugging Face07lhpku20010120 /Data-Prep-Bench Data-Prep-Bench This repository contains the data presented in DataPrep-Bench: Benchmarking LLMs as Training Data Preparators. Code: https://github.com/OpenDCAI/Data-Preparation-Bench Dataset Overview This dataset is a comprehensive resource built for Supervised Fine-Tuning (SFT) and evaluation of Large Language Models (LLMs), covering six domains: Finance, Medicine, Law, Mathematics, Science, and General. A key feature of this dataset is that we employed 12… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Data-Prep-Bench.texttext-generation1M<n<10M1 likes3k downloads2mo agoHugging Face08mib-bench /copycolors_mcqaThis dataset consists of formatted n-way multiple choice questions, where n is in [2,10]. The task itself is simply to copy the prototypical color from the context and produce the corresponding color's answer choice letter. The "prototypical colors" dataset instances themselves come from Memory Colors (Norland et al. 2021) and corypaik/coda (instances whose object_group is 0, indicating participants agreed on a prototypical color of that object). tabularquestion-answering1K<n<10K0 likes2.4k downloads2y agoHugging Face09franky-veteran /SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation. imagequestion-answering1K<n<10K3 likes2.3k downloads7mo agoHugging Face10PediaMedAI /CogSense-Bench CogSense-Bench Project Page | Paper | GitHub CogSense-Bench is a comprehensive visual question answering (VQA) benchmark designed to evaluate the cognitive capabilities of Multimodal Large Language Models (MLLMs). It was introduced in the paper "Toward Cognitive Supersensing in Multimodal Large Language Model". The benchmark assesses MLLMs across five cognitive dimensions: Fluid intelligence Crystallized intelligence Visuospatial cognition Mental simulation Visual routines… See the full description on the dataset page: https://huggingface.co/datasets/PediaMedAI/CogSense-Bench.imageimage-text-to-text1K<n<10K0 likes2k downloads8mo agoHugging Face11sorry-bench /sorry-bench-202503gated Dataset Card for SORRY-Bench Dataset (2025/03) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM 🪧UPDATE: In this iteration, we removed the category "Impersonation" due to its ambiguous definition, and that most models more or less fulfill such requests.This dataset contains 9.2K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation. Particularly, our base dataset consists of 440 unsafe… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202503.texttext-generation1K<n<10K23 likes1.7k downloads2y agoHugging Face12Agents-X /TIR-Bench TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning Introduction: TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.imagequestion-answering1K<n<10K3 likes1.7k downloads9mo agoHugging Face13michael7ma /ogd4all-benchmark OGD4All Benchmark This is a 199-question benchmark that was used to evaluate the overall performance of OGD4All and different configurations (LLM, orchestration, ...). OGD4All is an LLM-based prototype system enabling an easy-to-use, transparent interaction with Geospatial Open Government Data through natural language. Each question requires GIS, SQL and/or topological operations on zero, one, or multiple datasets in GPKG or CSV formats to be answered. Tasks The… See the full description on the dataset page: https://huggingface.co/datasets/michael7ma/ogd4all-benchmark.geospatialquestion-answeringn<1K1 likes1.3k downloads7mo agoHugging Face14a8cheng /SR-3D-Bench Spatial Region 3D (SR-3D) Aware Benchmark Paper: https://arxiv.org/abs/2509.13317Project page: https://www.anjiecheng.me/sr3dCode: https://github.com/AnjieCheng/SR-3D [!IMPORTANT] [Feb. 18, 2026] UPDATE: To improve compatibility with general-purpose VLMs, the benchmark is reformulated into multiple-choice and numerical questions following the VSI-Bench evaluation protocol. Videos are annotated with set-of-marks to explicitly indicate regions. The benchmark will be compatible… See the full description on the dataset page: https://huggingface.co/datasets/a8cheng/SR-3D-Bench.textquestion-answering1K<n<10K1 likes1.1k downloads7mo agoHugging Face15swe-qa /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.textquestion-answering1K<n<10K5 likes1.1k downloads2mo agoHugging Face16LLaMAX /BenchMAX_Science Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Science is a dataset of BenchMAX, sourcing from GPQA, which evaluates the natural science reasoning capability in multilingual scenarios. We extend the original English dataset to 16 non-English languages. The data is first translated by Google… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Science.textquestion-answering1K<n<10K2 likes986 downloads2y agoHugging Face17RogoAI /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.textquestion-answeringn<1K12 likes943 downloads2mo agoHugging Face18FRank62Wu /Act2Cap_benchmarkCollected data from GUI-Action-Narrator imagequestion-answeringn<1K0 likes890 downloads1y agoHugging Face19sorry-bench /sorry-bench-202406gated Dataset Card for SORRY-Bench Dataset (2024/06) 🏠Website 📑Paper 📚Dataset 💻Github 🧑‍⚖️Human Judgment Dataset 🤖Judge LLM This dataset contains 9.5K potentially unsafe instructions, intended to be used for LLM safety refusal evaluation. Particularly, our base dataset consists of 450 unsafe instructions in total, spanning across 45 finegrained safety categories (10 data points per category). The dataset we present here equally captures risks from… See the full description on the dataset page: https://huggingface.co/datasets/sorry-bench/sorry-bench-202406.texttext-generation1K<n<10K22 likes866 downloads2y agoHugging Face20paperinstruments /diligence-bench DiligenceBench A 150-item benchmark of analytical tasks across large-accelerated US equities spanning energy, banking, biotech, insurance, technology, REITs, restaurants, industrials, and utilities. Task distribution span: Cash-flow quality. Gap between GAAP operating cash flow and economic cash generation when non-cash items distort the headline — interest credited to policyholder deposits, insurance-liability growth, working-capital releases, stock-based… See the full description on the dataset page: https://huggingface.co/datasets/paperinstruments/diligence-bench.textquestion-answeringn<1K8 likes827 downloads2mo agoHugging Face21bofenghuang /mt-bench-french MT-Bench-French This is a French version of MT-Bench, created to evaluate the multi-turn conversation and instruction-following capabilities of LLMs. Similar to its original version, MT-Bench-French comprises 80 high-quality, multi-turn questions spanning eight main categories. All questions have undergone translation into French and thorough human review to guarantee the use of suitable and authentic wording, meaningful content for assessing LLMs' capabilities in the French… See the full description on the dataset page: https://huggingface.co/datasets/bofenghuang/mt-bench-french.textquestion-answeringn<1K8 likes773 downloads2y agoHugging Face22isaacus /legal-rag-bench Legal RAG Bench ‍⚖️ Legal RAG Bench by Isaacus is a reasoning-intensive benchmark for assessing the end-to-end, real-world performance of production-grade legal RAG systems. Legal RAG Bench is composed of 4,876 passages sampled from the Judicial College of Victoria’s Criminal Charge Book alongside 100 complex, handwritten questions demanding expert-level knowledge of Victorian criminal law and procedure to be answered correctly. Legal RAG Bench is the first open dataset for the… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/legal-rag-bench.texttext-retrieval1K<n<10K25 likes723 downloads7mo agoHugging Face23initiacms /XLRS-Bench-lite_VLM 🐙GitHub Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench 📜Dataset License Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from: DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply). ITCVDLicensed under CC-BY-NC-SA-4.0. MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.textvisual-question-answering1K<n<10K0 likes720 downloads11mo agoHugging Face24TIGER-Lab /SWE-QA-Pro-Bench SWE-QA-Pro Bench (A Repository-level QA Benchmark Built from Diverse Long-tail Repositories) 💻 GitHub | 📖 Paper | 🤗 SWE-QA-Pro 📢 News 🚀 [2026-5-19] The evaluation code is released on GitHub. 🔥 [2026-3-23] SWE-QA-Pro Bench is publicly released! The model and code will be released soon. Introduction SWE-QA-Pro Bench is a repository-level question answering dataset designed to evaluate whether models can perform grounded, agentic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/SWE-QA-Pro-Bench.textquestion-answeringn<1K5 likes697 downloads4mo agoHugging Face25zhiyuan218 /Think-Bench THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models Official repository for "THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models". For more details, please refer to the project page with dataset exploration and visualization tools. [Paper] [Github] [ModelScope Dataset] [Visualization] 👀 About Think-Bench Reasoning models have made remarkable progress in complex tasks… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuan218/Think-Bench.textquestion-answering1K<n<10K2 likes678 downloads1y agoHugging Face26Snowflake /dare-bench DARE-Bench [ICLR 2026] DARE-Bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science Fan Shu1, Yite Wang2, Ruofan Wu1, Boyi Liu2, Zhewei Yao2, Yuxiong He2, Feng Yan1 1University of Houston   2Snowflake AI Research 🔎 Overview DARE-Bench (ICLR 2026) is a benchmark for evaluating LLM agents on data science tasks, focusing on modeling and instruction fidelity. This Hugging Face repository provides a selected subset of the full benchmark for public release.… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/dare-bench.texttext-generation1K<n<10K6 likes604 downloads7mo agoHugging Face27ASLP-lab /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.audioaudio-classification1K<n<10K1 likes541 downloads3mo agoHugging Face28drive-bench /arenatextquestion-answering1K<n<10K14 likes533 downloads2y agoHugging Face29AMA-bench /AMA-bench AMA-Bench: Open-Ended QA for Long-Horizon Memory Evaluation This repository contains the Open-Ended QA dataset for AMA-Bench, introduced in the paper AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications. 🔗 Quick Links Website - Official AMA-Bench website GitHub Repository - Complete AMA-Bench project Dataset - Full dataset on Hugging Face Leaderboard - Model rankings and results Paper - Research paper Overview AMA-Bench Open-Ended… See the full description on the dataset page: https://huggingface.co/datasets/AMA-bench/AMA-bench.tabularquestion-answeringn<1K9 likes527 downloads4mo agoHugging Face30Psychotherapy-LLM /CBT-Bench CBT-Bench Dataset Overview CBT-Bench is a benchmark dataset designed to evaluate the proficiency of Large Language Models (LLMs) in assisting cognitive behavior therapy (CBT). The dataset is organized into three levels, each focusing on different key aspects of CBT, including basic knowledge recitation, cognitive model understanding, and therapeutic response generation. The goal is to assess how well LLMs can support various stages of professional mental health care… See the full description on the dataset page: https://huggingface.co/datasets/Psychotherapy-LLM/CBT-Bench.textquestion-answering1K<n<10K43 likes506 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.