datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ogd4all-benchmark
OGD4All Benchmark
This is a 199-question benchmark that was used to evaluate the overall performance of OGD4All and different configurations (LLM, orchestration, ...).
OGD4All is an LLM-based prototype system enabling an easy-to-use, transparent interaction with Geospatial Open Government Data through natural language.
Each question requires GIS, SQL and/or topological operations on zero, one, or multiple datasets in GPKG or CSV formats to be answered.
Tasks
The… See the full description on the dataset page: https://huggingface.co/datasets/michael7ma/ogd4all-benchmark.SWE-QA-Benchmark
SWE-QA Benchmark
A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories.
Dataset Summary
Total Questions: 720
Repositories: 15
Format: JSONL (JSON Lines)
Fields: question, answer
Repository Coverage
Each repository contains 48 questions:
astropy
conan
django
flask
matplotlib
pylint
pytest
reflex
requests
scikit-learn
sphinx
sqlfluff
streamlink
sympy
xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.Act2Cap_benchmarkCollected data from GUI-Action-Narrator
MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.legal-llm-benchmark
Legal LLM Benchmark Dataset
Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models
Quick Start
from datasets import load_dataset
# Core datasets
questions = load_dataset("marvintong/legal-llm-benchmark", "questions")
phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations")
phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations")
# Additional helpful datasets
contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tiny_qa_benchmark_pp
Tiny QA Benchmark++ (TQB++)
Tiny QA Benchmark++ (TQB++) is an ultra-lightweight evaluation suite designed to expose critical failures in Large Language Model (LLM) systems within seconds. It serves as the LLM analogue of software unit tests, ideal for rapid CI/CD checks, prompt engineering, and continuous quality assurance in modern LLMOps.
This Hugging Face dataset repository hosts the core English dataset and various synthetically generated multilingual and topical dataset packs… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark_pp.Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/big-finance-benchmark.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.k-verify-benchmark-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
K-Verify Benchmark v1
Author: Yachay / SZL Holdings · Version: 1.0.0 · Items: 100
K-Verify measures whether an AI's claimed factual answer is verifiable via a
receipt chain — not just whether it is correct. It is the first benchmark we
know of that scores provenance and honest refusal as… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/k-verify-benchmark-v1.Yue-Benchmark
How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models
Homepage: https://github.com/jiangjyjy/Yue-Benchmark
Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark
Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models.
Introduction
The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.multilingual-queries-2026
MAST Multilingual Queries 2026
This dataset contains the multilingual query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/multilingual-queries-2026.indic-queries-2026
MAST Indic Queries 2026
This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages.
MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.benchmark-radar
Benchmark Radar Dataset
Overview
Benchmark Radar is a living registry, search engine, and discovery pipeline
for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the
Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and
Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).
As described in the paper's Two Input Paths framework, Benchmark Radar
combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.Japanese-RAG-Generator-Benchmark
Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク
Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。
実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。
J-RAGBenchの評価カテゴリ
Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く
Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する
Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く
Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.blockchain-benchmark
Dataset Card for LLM Blockchain Benchmark
Dataset Summary
The Blockchain Benchmark Dataset is a comprehensive collection of data specifically curated for benchmarking Language Models (LMs) in the domain of blockchain technology. This dataset is designed to facilitate research and development in natural language understanding within the blockchain domain.
A complete list of tasks: ['general-reasoning', 'code', 'math']
Supported Tasks and Leaderboards
Model… See the full description on the dataset page: https://huggingface.co/datasets/revflask/blockchain-benchmark.MEME
MEME: Multi-Entity and Evolving Memory Evaluation
A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.turkish_cyber_security_controls_benchmark
Turkish Cyber Security Controls Benchmark
Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için
hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir.
v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5,
Release 5.2.0 kontrol kataloğunu hedefler.
Kapsam
100 Türkçe senaryo
NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru
64 kontrol seçimi sorusu
17 denetim kanıtı sorusu
19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.ragu_benchmarks
RAGU_Benchmarks
MultiQ Dataset
Dataset Description
Dataset Summary
MultiQ is a small but rich dataset designed for question answering (QA) and multi-document information retrieval tasks. It contains 169 Russian-language questions, each accompanied by a correct answer and a set of relevant Wikipedia articles serving as context for locating the answer. This dataset is suitable for evaluating models’ ability to identify precise answers based on multiple… See the full description on the dataset page: https://huggingface.co/datasets/RaguTeam/ragu_benchmarks.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/MaitriVasa/big-finance-benchmark.llmsql-benchmark
LLMSQL Benchmark
⚠️ A newer version of this dataset is available:👉 https://huggingface.co/datasets/llmsql-bench/llmsql-2.0
This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see https://github.com/LLMSQL/llmsql-benchmark.
Arxiv Article: https://arxiv.org/abs/2510.02350
Files
tables.jsonl — Database table metadata
questions.jsonl — All available questions
train_questions.jsonl, val_questions.jsonl, test_questions.jsonl — Data… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-benchmark.LogiTraj-Benchmark
LogiTraj Benchmark
Dual license. Dataset material is licensed under CC BY 4.0;
software in evaluation/code/ is licensed under Apache-2.0. See LICENSE,
LICENSE-DATA, and LICENSE-CODE. Commercial-model raw outputs are not
included.
Synthetic Chinese enterprise logistics sandboxes, tasks, documents, versioned evaluators, verdicts, and Core/Silver/Audit quality views.
Task-family coverage is source-faithful rather than imputed: the 20260628_v45, 20260628_v46, and 20260628_v50 SFT… See the full description on the dataset page: https://huggingface.co/datasets/Linmumu009/LogiTraj-Benchmark.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/big-finance-benchmark.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.UltraRAG_Benchmark
UltraRAG 2.0: Accelerating RAG for Scientific Research
UltraRAG 2.0 (UR-2.0) is jointly released by THUNLP, NEUIR, OpenBMB, and AI9Stars. It is the first lightweight RAG system construction framework built on the Model Context Protocol (MCP) architecture, designed to provide efficient modeling support for scientific research and exploration. The framework offers a full suite of teaching examples from beginner to advanced levels, integrates 17 mainstream benchmark tasks and a wide… See the full description on the dataset page: https://huggingface.co/datasets/UltraRAG/UltraRAG_Benchmark.marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
This dataset is associated with the paper "ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark".
Abstract
Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning… See the full description on the dataset page: https://huggingface.co/datasets/Shalyt/ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark.SEAR
SEAR: Spoofing Evidence-Grounded Audio Reasoning
SEAR is an audio question-answering benchmark for testing whether audio language models
can identify and quantify signal-level acoustic anomalies and use them as evidence for
audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR
separates deepfake detection, forgery-cue identification, acoustic measurement, and
forensic rationale generation.
SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.HALT_Benchmark_0.1_v1
HALT Benchmark Dataset v1.0
HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse
Overview
HALT is a benchmark for evaluating bounded agentic decision-making under partial observability,
constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity
workflows, where acting too early, failing to escalate, or over-escalating can all be costly.
The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.
