CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01michael7ma /ogd4all-benchmark OGD4All Benchmark This is a 199-question benchmark that was used to evaluate the overall performance of OGD4All and different configurations (LLM, orchestration, ...). OGD4All is an LLM-based prototype system enabling an easy-to-use, transparent interaction with Geospatial Open Government Data through natural language. Each question requires GIS, SQL and/or topological operations on zero, one, or multiple datasets in GPKG or CSV formats to be answered. Tasks The… See the full description on the dataset page: https://huggingface.co/datasets/michael7ma/ogd4all-benchmark.geospatialquestion-answeringn<1K1 likes1.3k downloads7mo agoHugging Face02swe-qa /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/swe-qa/SWE-QA-Benchmark.textquestion-answering1K<n<10K5 likes1.1k downloads2mo agoHugging Face03RogoAI /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.textquestion-answeringn<1K12 likes908 downloads2mo agoHugging Face04FRank62Wu /Act2Cap_benchmarkCollected data from GUI-Action-Narrator imagequestion-answeringn<1K0 likes863 downloads1y agoHugging Face05ASLP-lab /MSU-Benchmark MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.audioaudio-classification1K<n<10K1 likes567 downloads3mo agoHugging Face06marvintong /legal-llm-benchmark Legal LLM Benchmark Dataset Safety-Utility Trade-offs in Legal AI: An LLM Evaluation Across 12 Models Quick Start from datasets import load_dataset # Core datasets questions = load_dataset("marvintong/legal-llm-benchmark", "questions") phase1_evals = load_dataset("marvintong/legal-llm-benchmark", "phase1_evaluations") phase3_evals = load_dataset("marvintong/legal-llm-benchmark", "phase3_evaluations") # Additional helpful datasets contracts =… See the full description on the dataset page: https://huggingface.co/datasets/marvintong/legal-llm-benchmark.tabulartext-generation1K<n<10K1 likes533 downloads11mo agoHugging Face07vincentkoc /tiny_qa_benchmark_pp Tiny QA Benchmark++ (TQB++) Tiny QA Benchmark++ (TQB++) is an ultra-lightweight evaluation suite designed to expose critical failures in Large Language Model (LLM) systems within seconds. It serves as the LLM analogue of software unit tests, ideal for rapid CI/CD checks, prompt engineering, and continuous quality assurance in modern LLMOps. This Hugging Face dataset repository hosts the core English dataset and various synthetically generated multilingual and topical dataset packs… See the full description on the dataset page: https://huggingface.co/datasets/vincentkoc/tiny_qa_benchmark_pp.textquestion-answeringn<1K3 likes490 downloads1y agoHugging Face08TreeAILab /Multi-turn_Long-context_Benchmark_for_LLMs LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues Arxiv: https://www.arxiv.org/abs/2507.13681 Huggingface: https://huggingface.co/papers/2507.13681 Introduction LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios. Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.textquestion-answering1K<n<10K0 likes487 downloads1y agoHugging Face09idleengine /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/big-finance-benchmark.textquestion-answeringn<1K0 likes396 downloads1mo agoHugging Face10EigenformAI /groundtruth-dynamic-benchmarking Groundtruth Dynamic Benchmarking — Geology Question sets and grading rubrics for evaluating LLMs on real-world geological reasoning. Every question is authored from a real source corpus, and every claim in the grading key carries an evidence locator back to that corpus — nothing is synthetic. Licensing/redistribution status varies by corpus — see License. This dataset holds the questions, grading rubrics, and source corpora. Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.textquestion-answeringn<1K0 likes387 downloads28d agoHugging Face11SZLHOLDINGS /k-verify-benchmark-v1 Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance. K-Verify Benchmark v1 Author: Yachay / SZL Holdings · Version: 1.0.0 · Items: 100 K-Verify measures whether an AI's claimed factual answer is verifiable via a receipt chain — not just whether it is correct. It is the first benchmark we know of that scores provenance and honest refusal as… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/k-verify-benchmark-v1.textquestion-answeringn<1K0 likes332 downloads24d agoHugging Face12BillBao /Yue-Benchmark How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models Homepage: https://github.com/jiangjyjy/Yue-Benchmark Repository: https://huggingface.co/datasets/BillBao/Yue-Benchmark Paper: How Far Can Cantonese NLP Go? Benchmarking Cantonese Capabilities of Large Language Models. Introduction The rapid evolution of large language models (LLMs), such as GPT-X and Llama-X, has driven significant advancements in NLP, yet much of this… See the full description on the dataset page: https://huggingface.co/datasets/BillBao/Yue-Benchmark.textmultiple-choice1K<n<10K8 likes313 downloads2y agoHugging Face13mast-benchmark /multilingual-queries-2026 MAST Multilingual Queries 2026 This dataset contains the multilingual query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages. MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/multilingual-queries-2026.textquestion-answeringn<1K1 likes295 downloads1mo agoHugging Face14mast-benchmark /indic-queries-2026 MAST Indic Queries 2026 This dataset contains the Indic query set for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages. MAST builds on BrowseComp-Plus (ACL 2026), a reproducible and verifiable extension of BrowseComp with challenging English queries, a verified English corpus of roughly 100K web-sourced documents, and human judgments. In the 2026 MAST… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/indic-queries-2026.textquestion-answeringn<1K1 likes279 downloads2mo agoHugging Face15ktwu01 /benchmark-radar Benchmark Radar Dataset Overview Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115). As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:… See the full description on the dataset page: https://huggingface.co/datasets/ktwu01/benchmark-radar.tabulartext-generation10K<n<100K1 likes222 downloads23h agoHugging Face16neoai-inc /Japanese-RAG-Generator-Benchmark Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。 実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。 J-RAGBenchの評価カテゴリ Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.textquestion-answeringn<1K4 likes184 downloads10mo agoHugging Face17revflask /blockchain-benchmark Dataset Card for LLM Blockchain Benchmark Dataset Summary The Blockchain Benchmark Dataset is a comprehensive collection of data specifically curated for benchmarking Language Models (LMs) in the domain of blockchain technology. This dataset is designed to facilitate research and development in natural language understanding within the blockchain domain. A complete list of tasks: ['general-reasoning', 'code', 'math'] Supported Tasks and Leaderboards Model… See the full description on the dataset page: https://huggingface.co/datasets/revflask/blockchain-benchmark.textquestion-answeringn<1K2 likes179 downloads2y agoHugging Face18meme-benchmark /MEME MEME: Multi-Entity and Evolving Memory Evaluation A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.tabularquestion-answeringn<1K4 likes169 downloads5mo agoHugging Face19logicBombExe /turkish_cyber_security_controls_benchmark Turkish Cyber Security Controls Benchmark Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir. v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5, Release 5.2.0 kontrol kataloğunu hedefler. Kapsam 100 Türkçe senaryo NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru 64 kontrol seçimi sorusu 17 denetim kanıtı sorusu 19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.textquestion-answeringn<1K4 likes167 downloads2mo agoHugging Face20RaguTeam /ragu_benchmarks RAGU_Benchmarks MultiQ Dataset Dataset Description Dataset Summary MultiQ is a small but rich dataset designed for question answering (QA) and multi-document information retrieval tasks. It contains 169 Russian-language questions, each accompanied by a correct answer and a set of relevant Wikipedia articles serving as context for locating the answer. This dataset is suitable for evaluating models’ ability to identify precise answers based on multiple… See the full description on the dataset page: https://huggingface.co/datasets/RaguTeam/ragu_benchmarks.textquestion-answering1K<n<10K1 likes162 downloads9mo agoHugging Face21MaitriVasa /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/MaitriVasa/big-finance-benchmark.textquestion-answeringn<1K0 likes153 downloads2mo agoHugging Face22llmsql-bench /llmsql-benchmark LLMSQL Benchmark ⚠️ A newer version of this dataset is available:👉 https://huggingface.co/datasets/llmsql-bench/llmsql-2.0 This benchmark is designed to evaluate text-to-SQL models. For usage of this benchmark see https://github.com/LLMSQL/llmsql-benchmark. Arxiv Article: https://arxiv.org/abs/2510.02350 Files tables.jsonl — Database table metadata questions.jsonl — All available questions train_questions.jsonl, val_questions.jsonl, test_questions.jsonl — Data… See the full description on the dataset page: https://huggingface.co/datasets/llmsql-bench/llmsql-benchmark.textquestion-answering10K<n<100K2 likes149 downloads7mo agoHugging Face23Linmumu009 /LogiTraj-Benchmark LogiTraj Benchmark Dual license. Dataset material is licensed under CC BY 4.0; software in evaluation/code/ is licensed under Apache-2.0. See LICENSE, LICENSE-DATA, and LICENSE-CODE. Commercial-model raw outputs are not included. Synthetic Chinese enterprise logistics sandboxes, tasks, documents, versioned evaluators, verdicts, and Core/Silver/Audit quality views. Task-family coverage is source-faithful rather than imputed: the 20260628_v45, 20260628_v46, and 20260628_v50 SFT… See the full description on the dataset page: https://huggingface.co/datasets/Linmumu009/LogiTraj-Benchmark.textquestion-answering10K<n<100K0 likes138 downloads2mo agoHugging Face24Koplos /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/big-finance-benchmark.textquestion-answeringn<1K0 likes134 downloads2mo agoHugging Face25oliversayshi /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.textquestion-answeringn<1K0 likes133 downloads2mo agoHugging Face26UltraRAG /UltraRAG_Benchmark UltraRAG 2.0: Accelerating RAG for Scientific Research UltraRAG 2.0 (UR-2.0) is jointly released by THUNLP, NEUIR, OpenBMB, and AI9Stars. It is the first lightweight RAG system construction framework built on the Model Context Protocol (MCP) architecture, designed to provide efficient modeling support for scientific research and exploration. The framework offers a full suite of teaching examples from beginner to advanced levels, integrates 17 mainstream benchmark tasks and a wide… See the full description on the dataset page: https://huggingface.co/datasets/UltraRAG/UltraRAG_Benchmark.textquestion-answering100K<n<1M6 likes123 downloads11mo agoHugging Face27adsgpt /marketing-benchmark-of-more-than-10-ai-models Marketing Benchmark of 10+ AI Models A 5,000-question benchmark for evaluating LLMs across six dimensions of modern marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas. Every question is independently authored by the AdsGPT Marketing Bench team. Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.textquestion-answering1K<n<10K5 likes122 downloads4mo agoHugging Face28Shalyt /ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark This dataset is associated with the paper "ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark". Abstract Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization with genuine reasoning. To address this gap, we present ASyMOB, a high-resolution dataset of 35,368 validated symbolic math problems spanning… See the full description on the dataset page: https://huggingface.co/datasets/Shalyt/ASyMOB-Algebraic_Symbolic_Mathematical_Operations_Benchmark.textquestion-answering10K<n<100K3 likes117 downloads4mo agoHugging Face29SEAR-benchmark /SEAR SEAR: Spoofing Evidence-Grounded Audio Reasoning SEAR is an audio question-answering benchmark for testing whether audio language models can identify and quantify signal-level acoustic anomalies and use them as evidence for audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR separates deepfake detection, forgery-cue identification, acoustic measurement, and forensic rationale generation. SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.textaudio-classification100K<n<1M0 likes115 downloads8d agoHugging Face30supreme-lab /HALT_Benchmark_0.1_v1 HALT Benchmark Dataset v1.0 HALT: Benchmarking When Language Agents Should Stop, Investigate, Escalate, or Refuse Overview HALT is a benchmark for evaluating bounded agentic decision-making under partial observability, constrained tools, and explicit escalation options. It is grounded in defensive cybersecurity workflows, where acting too early, failing to escalate, or over-escalating can all be costly. The benchmark contains 1,248 instances across four decision regimes… See the full description on the dataset page: https://huggingface.co/datasets/supreme-lab/HALT_Benchmark_0.1_v1.texttext-classification1K<n<10K1 likes113 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.