CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K266 likes9.9k downloads2y agoHugging Face02bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face03sylvainHellin /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.documentquestion-answering1K<n<10K20 likes3k downloads1d agoHugging Face04AI4Sec /cti-bench Dataset Card for CTIBench A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks. Dataset Details Dataset Description CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI. Components: CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.textzero-shot-classification1K<n<10K21 likes1.9k downloads2y agoHugging Face05ncbi /MedCalc-Bench [!Note] Please visit MedCalc-Bench Verified at this url: https://github.com/nikhilk7153/MedCalc-Bench-Verified for the latest changes. Here is the HuggingFace link: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified. The first version of MedCalc-Bench Verified is an update from v1.2 on this repository. MedCalc-Bench is the first medical calculation dataset used to benchmark LLMs ability to serve as clinical calculators. Each instance in the dataset consists of a patient note, a… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench.textquestion-answering10K<n<100K2 likes1.7k downloads9mo agoHugging Face06ncbi /MedCalc-Bench-v1.2 [!note] Please visit MedCalc-Bench Verified at this url: https://github.com/nikhilk7153/MedCalc-Bench-Verified for the latest changes. Here is the HuggingFace link: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified. The first version of MedCalc-Bench Verified is an update from v1.2 on this repository. We recommend using v1.0-v1.2 for reproducibility purposes only. You should specify which version you are using when using this dataset. MedCalc-Bench is the first medical… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.2.textquestion-answering10K<n<100K3 likes1.1k downloads9mo agoHugging Face07DolphinAI /u2-bench U2-BENCH: Ultrasound Understanding Benchmark U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding. It provides a diverse, multi-task dataset curated from 40 licensed sources, covering 15 anatomical regions and 8 clinically inspired tasks across classification, detection, regression, and text generation. Check the 🌟Leaderboard🌟here: https://dolphin-sound.github.io/u2-bench/… See the full description on the dataset page: https://huggingface.co/datasets/DolphinAI/u2-bench.textquestion-answering1K<n<10K10 likes994 downloads1y agoHugging Face08ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes800 downloads2y agoHugging Face09bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes742 downloads2y agoHugging Face10orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes615 downloads2y agoHugging Face11SiloLink /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.documentquestion-answering1K<n<10K0 likes591 downloads23d agoHugging Face12bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes357 downloads2y agoHugging Face13deepscholar-bench /DeepScholarBench DeepScholarBench Dataset A comprehensive dataset of academic papers with extracted related works sections and recovered citations, designed for training and evaluating research generation systems. 📊 Dataset Overview This dataset contains 63 academic papers from ArXiv with their related works sections and 1630 recovered citations, providing a rich resource for research generation and citation analysis tasks. 🎯 Use Cases Research Generation: Train models… See the full description on the dataset page: https://huggingface.co/datasets/deepscholar-bench/DeepScholarBench.tabularfeature-extraction1K<n<10K2 likes318 downloads1y agoHugging Face14amazon /sop-bench SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents 📄 Paper: SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents 🏭 Human Expert-Authored SOPs · 🤖 Human-AI Collaborative Framework · 📊 Executable Interfaces · 🔧 Two Agent Architectures · 📈 11 Frontier Models Evaluated Dataset Summary SOP-Bench is a comprehensive benchmark for evaluating LLM-based agents on complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial… See the full description on the dataset page: https://huggingface.co/datasets/amazon/sop-bench.imagetext-classification1K<n<10K1 likes301 downloads4mo agoHugging Face15quenfly /ifc-bench IFC-Bench A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations. Dataset snapshot: question ground_truth ifc_model project category 0 What modelling program and IFC standard were used to create this model? The model was created using... arc 4351 1 1 What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.documentquestion-answering1K<n<10K1 likes297 downloads2mo agoHugging Face16latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K8 likes270 downloads2mo agoHugging Face17bryel-labs /MuSP-Bench MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance MuSP-Bench is a 490-question benchmark for musical score understanding, performance listening, and combined score-performance reasoning. Official benchmark website Modalities Each question specifies the minimum source of musical evidence needed to answer it: S (score): answer from the written score. P (performance): answer from the performance recording. S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.imagequestion-answeringn<1K1 likes259 downloads27d agoHugging Face18plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes212 downloads24d agoHugging Face19ncbi /MedCalc-Bench-v1.0gated Important: This dataset is kept for reproducibility purposes only. Please use the most up-to-date version 1.2, for the most revised and corrected dataset available. We have fixed 12 calculator implementations, ensured the Relevant Entities section best matches with what was specified by a patient note, and have also replaced notes which are better fits for a given calculator to make the dataset more applicable for real-life sitatuons. Because of the number of changes, we find this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.0.tabularquestion-answering10K<n<100K2 likes196 downloads10mo agoHugging Face20milan477 /MuSP-Bench MuSP-Bench MuSP-Bench is a 490-question benchmark for musical score understanding, performance listening, and combined score-performance reasoning. Contents data/questions.csv: all 490 questions, accepted answers, and the response contract for each. inputs/pdf/without_context/: one context-removed PDF per piece. inputs/images/: rendered score-page images for every piece. inputs/abc/: one ABC score per piece. inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.imagequestion-answeringn<1K1 likes167 downloads1mo agoHugging Face21lianghsun /tw-legal-benchmark-v1 Taiwan Legal Benchmark v1 A multiple-choice benchmark for evaluating large language models on Taiwan law in Traditional Chinese (繁體中文). It covers six legal domains with 209 questions drawn from Taiwan bar exam and certification-style questions. Overview Property Value Language Traditional Chinese (zh-TW) Questions 209 Format 4-choice multiple choice (A / B / C / D) Domain Taiwan law License Apache 2.0 Legal Domains Covered… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v1.textquestion-answeringn<1K7 likes159 downloads6mo agoHugging Face22el7982 /aware-bench EvalDetectBench Companion dataset for EvalDetectBench, a benchmark that measures whether LLMs can detect that they are in evaluation or deployment contexts. Three folders, each with its own README.md and croissant.json (Croissant 1.1): collected_trajectories/ raw trajectory pool (per-model JSON) measure_logs/ measure-stage outputs (.eval logs + CSV export) paper_replication/ CSV inputs that feed the paper figures and tables Folder map… See the full description on the dataset page: https://huggingface.co/datasets/el7982/aware-bench.tabulartabular-classification100K<n<1M0 likes157 downloads3mo agoHugging Face23JonathanSu /singapore-legal-ai-benchmark Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.tabulartext-generationn<1K0 likes154 downloads6d agoHugging Face24ncbi /MedCalc-Bench-v1.1gatedWe have revised and improved upon v1.1 based on the updates in 1.2: https://github.com/ncbi-nlp/MedCalc-Bench/releases/tag/version-1.2 Please use the most up-to-date version here: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.2 tabularquestion-answering10K<n<100K1 likes147 downloads10mo agoHugging Face25atroposhealth /precision-evidence-bench Precision Evidence Bench Precision Evidence Bench is a Precision Medicine Benchmark from Atropos Health, the world's largest creator of real-world evidence (RWE) for clinical decision support. It evaluates how well large language models (LLMs) answer clinical questions that are grounded in patient context and inclusive of patient history: not "which treatment is better in general", but "which treatment is better for this patient", with a specific comorbidity, age, prior therapy… See the full description on the dataset page: https://huggingface.co/datasets/atroposhealth/precision-evidence-bench.textquestion-answeringn<1K5 likes140 downloads7d agoHugging Face26Joysouo /hse-bench HSE-Bench HSE-Bench is a graduate-level benchmark designed to evaluate the legal and safety reasoning capabilities of Large Language Models (LLMs) in high-stakes, regulation-intensive domains concerning Health, Safety, and the Environment (HSE). This benchmark focuses on scenario-based, single-choice questions crafted under the IRAC (Issue, Rule, Application, Conclusion) reasoning framework, with all content presented in English. Overview HSE-Bench comprises 1,020… See the full description on the dataset page: https://huggingface.co/datasets/Joysouo/hse-bench.textquestion-answering1K<n<10K2 likes135 downloads1y agoHugging Face27lianghsun /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K3 likes127 downloads1mo agoHugging Face28jerogo /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/jerogo/or-bench.imagetext-generation10K<n<100K0 likes123 downloads2mo agoHugging Face29JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes120 downloads6mo agoHugging Face30nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes115 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.