datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
frames-benchmark
FRAMES: Factuality, Retrieval, And reasoning MEasurement Set
FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning.
Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941.
Dataset Overview
824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles
Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/sylvainHellin/ifc-bench.cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.MedCalc-Bench
[!Note]
Please visit MedCalc-Bench Verified at this url: https://github.com/nikhilk7153/MedCalc-Bench-Verified for the latest changes. Here is the HuggingFace link: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified.
The first version of MedCalc-Bench Verified is an update from v1.2 on this repository.
MedCalc-Bench is the first medical calculation dataset used to benchmark LLMs ability to serve as clinical calculators. Each instance in the dataset consists of a patient note, a… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench.MedCalc-Bench-v1.2
[!note]
Please visit MedCalc-Bench Verified at this url: https://github.com/nikhilk7153/MedCalc-Bench-Verified for the latest changes. Here is the HuggingFace link: https://huggingface.co/datasets/nsk7153/MedCalc-Bench-Verified.
The first version of MedCalc-Bench Verified is an update from v1.2 on this repository. We recommend using v1.0-v1.2 for reproducibility purposes only. You should specify which version you are using when using this dataset.
MedCalc-Bench is the first medical… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.2.u2-bench
U2-BENCH: Ultrasound Understanding Benchmark
U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding. It provides a diverse, multi-task dataset curated from 40 licensed sources, covering 15 anatomical regions and 8 clinically inspired tasks across classification, detection, regression, and text generation.
Check the 🌟Leaderboard🌟here: https://dolphin-sound.github.io/u2-bench/… See the full description on the dataset page: https://huggingface.co/datasets/DolphinAI/u2-bench.Wikipedia_contradict_benchmark
Wikipedia contradict benchmark
Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/SiloLink/ifc-bench.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.DeepScholarBench
DeepScholarBench Dataset
A comprehensive dataset of academic papers with extracted related works sections and recovered citations, designed for training and evaluating research generation systems.
📊 Dataset Overview
This dataset contains 63 academic papers from ArXiv with their related works sections and 1630 recovered citations, providing a rich resource for research generation and citation analysis tasks.
🎯 Use Cases
Research Generation: Train models… See the full description on the dataset page: https://huggingface.co/datasets/deepscholar-bench/DeepScholarBench.sop-bench
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
📄 Paper: SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
🏭 Human Expert-Authored SOPs · 🤖 Human-AI Collaborative Framework · 📊 Executable Interfaces · 🔧 Two Agent Architectures · 📈 11 Frontier Models Evaluated
Dataset Summary
SOP-Bench is a comprehensive benchmark for evaluating LLM-based agents on complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial… See the full description on the dataset page: https://huggingface.co/datasets/amazon/sop-bench.ifc-bench
IFC-Bench
A benchmark dataset for evaluating BIM (Building Information Modeling) comprehension and reasoning capabilities in AI systems. Provides curated IFC models with question-answer pairs across 4 complexity categories for testing BIM-related AI implementations.
Dataset snapshot:
question
ground_truth
ifc_model
project
category
0
What modelling program and IFC standard were used to create this model?
The model was created using...
arc
4351
1
1
What are the… See the full description on the dataset page: https://huggingface.co/datasets/quenfly/ifc-bench.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.MuSP-Bench
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Official benchmark website
Modalities
Each question specifies the minimum source of musical evidence needed to answer it:
S (score): answer from the written score.
P (performance): answer from the performance recording.
S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.MedCalc-Bench-v1.0
Important: This dataset is kept for reproducibility purposes only. Please use the most up-to-date version 1.2, for the most
revised and corrected dataset available.
We have fixed 12 calculator implementations, ensured the Relevant Entities
section best matches with what was specified by a patient note, and have also replaced notes which are better fits
for a given calculator to make the dataset more applicable for real-life sitatuons.
Because of the number of changes, we find this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.0.MuSP-Bench
MuSP-Bench
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Contents
data/questions.csv: all 490 questions, accepted answers, and the
response contract for each.
inputs/pdf/without_context/: one context-removed PDF per piece.
inputs/images/: rendered score-page images for every piece.
inputs/abc/: one ABC score per piece.
inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.tw-legal-benchmark-v1
Taiwan Legal Benchmark v1
A multiple-choice benchmark for evaluating large language models on Taiwan law in Traditional Chinese (繁體中文). It covers six legal domains with 209 questions drawn from Taiwan bar exam and certification-style questions.
Overview
Property
Value
Language
Traditional Chinese (zh-TW)
Questions
209
Format
4-choice multiple choice (A / B / C / D)
Domain
Taiwan law
License
Apache 2.0
Legal Domains Covered… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v1.aware-bench
EvalDetectBench
Companion dataset for EvalDetectBench, a benchmark that measures whether LLMs can detect that they are in evaluation or deployment contexts.
Three folders, each with its own README.md and croissant.json
(Croissant 1.1):
collected_trajectories/ raw trajectory pool (per-model JSON)
measure_logs/ measure-stage outputs (.eval logs + CSV export)
paper_replication/ CSV inputs that feed the paper figures and tables
Folder map… See the full description on the dataset page: https://huggingface.co/datasets/el7982/aware-bench.singapore-legal-ai-benchmark
Singapore Legal AI Benchmark
Public research release of 102 Singapore legal research questions, model
responses from 6 systems, and overlapping grades on five dimensions.
Headline metrics are overlapping binary flags, not a ranking and not a
partition of 100%.
Interactive explorer
Open the explorer →
— comparison table, category heatmap, per-question comparison, and every answer
with its sources and grades.
(Space page)
Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.MedCalc-Bench-v1.1We have revised and improved upon v1.1 based on the updates in 1.2: https://github.com/ncbi-nlp/MedCalc-Bench/releases/tag/version-1.2
Please use the most up-to-date version here: https://huggingface.co/datasets/ncbi/MedCalc-Bench-v1.2
precision-evidence-bench
Precision Evidence Bench
Precision Evidence Bench is a Precision Medicine Benchmark from
Atropos Health, the world's largest creator of
real-world evidence (RWE) for clinical decision support. It evaluates how well
large language models (LLMs) answer clinical questions that are grounded in
patient context and inclusive of patient history: not "which treatment is
better in general", but "which treatment is better for this patient", with a
specific comorbidity, age, prior therapy… See the full description on the dataset page: https://huggingface.co/datasets/atroposhealth/precision-evidence-bench.hse-bench
HSE-Bench
HSE-Bench is a graduate-level benchmark designed to evaluate the legal and safety reasoning capabilities of Large Language Models (LLMs) in high-stakes, regulation-intensive domains concerning Health, Safety, and the Environment (HSE). This benchmark focuses on scenario-based, single-choice questions crafted under the IRAC (Issue, Rule, Application, Conclusion) reasoning framework, with all content presented in English.
Overview
HSE-Bench comprises 1,020… See the full description on the dataset page: https://huggingface.co/datasets/Joysouo/hse-bench.tw-legal-benchmark-v2
Taiwan Legal Benchmark v2
A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese,
built from 15 years (2012–2026) of national examinations published by the
Ministry of Examination (考選部).
Supersedes tw-legal-benchmark-v1
(209 questions) with 17,002 deduplicated questions across 15 legal domains.
Overview
Property
Value
Questions
17,002 (deduplicated)
Years
2012–2026
Source papers
1,040 official exam papers
Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/jerogo/or-bench.DeepResearch-Bench-Multilingual
DeepResearch Bench Multilingual Prompts
This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset.
The translations cover eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
This repository focuses on the benchmark prompts only.
On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.nasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.
