CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K266 likes9.9k downloads2y agoHugging Face02ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes800 downloads2y agoHugging Face03latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K8 likes270 downloads2mo agoHugging Face04plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes212 downloads24d agoHugging Face05lianghsun /tw-legal-benchmark-v1 Taiwan Legal Benchmark v1 A multiple-choice benchmark for evaluating large language models on Taiwan law in Traditional Chinese (繁體中文). It covers six legal domains with 209 questions drawn from Taiwan bar exam and certification-style questions. Overview Property Value Language Traditional Chinese (zh-TW) Questions 209 Format 4-choice multiple choice (A / B / C / D) Domain Taiwan law License Apache 2.0 Legal Domains Covered… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v1.textquestion-answeringn<1K7 likes159 downloads6mo agoHugging Face06JonathanSu /singapore-legal-ai-benchmark Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.tabulartext-generationn<1K0 likes154 downloads6d agoHugging Face07lianghsun /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K3 likes127 downloads1mo agoHugging Face08nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes115 downloads8mo agoHugging Face09roskosmos19 /agentic-reasoning-benchmark Agentic & Reasoning Benchmark (ARB) – Expanded Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning. Überblick Eigenschaft Wert Anzahl Beispiele 2.550 Kategorien 8 Schwierigkeitsgrade easy / medium / hard Formate CSV + JSON Reproduzierbarkeit Generator-Skript (seed=42) enthalten Lizenz CC-BY-4.0 Kategorien Kategorie Anzahl Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.textquestion-answering1K<n<10K1 likes102 downloads20d agoHugging Face10emre /TARA_Turkish_LLM_Benchmark TARA: Turkish Advanced Reasoning Assessment Veri Seti *Img Credit: Open AI ChatGPT **English version is given below.** Evaluation Notebook / Değerlendirme Not Defteri Dataset Summary TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.textquestion-answeringn<1K28 likes90 downloads1y agoHugging Face11boczkakaroly /trilingual-cultural-bias-redteaming-benchmark Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU) Overview This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU). The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language. The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.texttext-generationn<1K0 likes70 downloads9mo agoHugging Face12TPelc /Current_Trivia_Knowledge-benchmark Current Trivia Knowledge RAG Benchmark Short Summary: A 140-QA pair (70 train, 70 test) dataset for real-world RAG evaluation. It features current knowledge questions unavailable to LLMs trained before 2024 (e.g., GPT-4o) across diverse domains, and includes human feedback for the training set, enabling robust assessment of contextual information's critical impact on LLM accuracy. Introduction & Motivation: This dataset addresses the critical need for a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Current_Trivia_Knowledge-benchmark.textquestion-answeringn<1K0 likes68 downloads1y agoHugging Face13Jel1f1sh /tw-legal-benchmark-v2 Taiwan Legal Benchmark v2 A multiple-choice benchmark for evaluating LLMs on Taiwan law in Traditional Chinese, built from 15 years (2012–2026) of national examinations published by the Ministry of Examination (考選部). Supersedes tw-legal-benchmark-v1 (209 questions) with 17,002 deduplicated questions across 15 legal domains. Overview Property Value Questions 17,002 (deduplicated) Years 2012–2026 Source papers 1,040 official exam papers Format… See the full description on the dataset page: https://huggingface.co/datasets/Jel1f1sh/tw-legal-benchmark-v2.tabularquestion-answering10K<n<100K0 likes60 downloads28d agoHugging Face14vkshdev /rag-hallucination-benchmark RAG Hallucination Benchmark Context Retrieval-Augmented Generation (RAG) is the industry standard for reducing LLM hallucinations, but detecting when a RAG system fails is a massive challenge. Most existing benchmarks focus only on massive Deep Learning models and lack tabular features. This dataset provides a clean, engineered setup to train models (from XGBoost to RoBERTa) to detect hallucinations, predict context faithfulness, and measure answer relevance.… See the full description on the dataset page: https://huggingface.co/datasets/vkshdev/rag-hallucination-benchmark.tabulartext-classification10K<n<100K0 likes59 downloads22d agoHugging Face15aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes58 downloads6mo agoHugging Face16OpenStellarTeam /BeyongSafeAnswer_Benchmark 🌐 Website • 📃 Paper • 📊 Leader Board Overview Beyond Safe Answers is a novel benchmark meticulously designed to evaluate the true risk awareness of Large Reasoning Models (LRMs), particularly focusing on their internal reasoning processes rather than just superficial outputs. This benchmark addresses a critical issue termed Superficial Safety Alignment (SSA), where LRMs generate superficially safe responses but fail in genuine internal risk assessment, leading to… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/BeyongSafeAnswer_Benchmark.textquestion-answering1K<n<10K0 likes46 downloads1y agoHugging Face17cynosural /semantic-montecarlo-benchmark Semantic Monte Carlo Benchmark A synthetic benchmark of numeric research and forecasting questions for evaluating the semantic-montecarlo pipeline. This release contains only benchmark inputs. Cached experiments, individual run artifacts, and aggregate results are intentionally excluded. At a glance Questions Language Splits License 300 English Validation and test CC0 1.0 Dataset structure The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.tabularquestion-answeringn<1K1 likes38 downloads2mo agoHugging Face18boczkakaroly /hungarian-riddles-benchmark Hungarian Riddles Benchmark Overview This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles. The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency. Dataset structure Each row contains one riddle with reference material for evaluation. Fields ID – unique identifier topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.imagequestion-answeringn<1K0 likes34 downloads9mo agoHugging Face19BCCard /bc-finance-llm-benchmark BC Card Finance LLM Benchmark 한국 금융 도메인 특화 LLM 성능 평가를 위한 벤치마크 데이터셋입니다.BC카드-연세대 DSL 산학협력(S2026 LLMOps 프로젝트)의 산출물입니다. 데이터셋 개요 항목 내용 총 문항 수 800 언어 한국어 도메인 금융 (BC카드 FAQ, 금융 일반) 형식 Question / Ground Truth 컬럼 설명 컬럼 설명 no 순번 id 문항 ID 대분류 데이터 출처 대분류 (BC카드FAQ, general 등) 금융토픽 세부 금융 토픽 문제유형 문제 유형 (단일추론, 다중추론 등) question 평가 질문 ground_truth 정답 (참조 답변) source_tag 원본 출처 문서 태그 활용 방법 LLM-as-Judge… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/bc-finance-llm-benchmark.textquestion-answeringn<1K2 likes32 downloads3mo agoHugging Face20emre /El-TARA_Spanish_LLM_Benchmark El-Tara: Evaluación de Razonamiento Avanzado en Español Dataset Summary El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset. Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.textquestion-answeringn<1K1 likes31 downloads1y agoHugging Face21bilalabic /math-toolcall-tr-benchmark math-toolcall-tr-benchmark bilalabic/gemma_4_math-toolcall-tr_lora LoRA adaptörünü temel Gemma-4 E4B modeliyle karşılaştıran benchmark sonuçları. Bu depo yalnızca değerlendirme çıktılarını içerir. Eğitim veri seti ayrı olarak bilalabic/math-toolcall-tr adresinde yayımlanmaktadır. Benchmark'lar Benchmark Örnek Ölçülen davranış Türkçe MMLU 250 Genel bilgi doğruluğu ve eğitim sonrası bilgi kaybı Matematik Tool-Call 150 Araç seçimi, çekimserlik ve çıktı… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/math-toolcall-tr-benchmark.tabulartext-generation1K<n<10K0 likes29 downloads2mo agoHugging Face22qwqw3535 /econcausal-benchmark📊 EconCausal: A Context-Aware Causal Reasoning Benchmark for LLMs Donggyu Lee, Hyeok Yun, Meeyoung Cha, Sungwon Park, Sangyoon Park, Jihee Kim 🌍 Overview Socio-economic causal effects depend heavily on their specific institutional and environmental context. A single intervention can produce opposite results depending on regulatory or market factors. EconCausal is a large-scale benchmark comprising 10,490 context-annotated causal triplets extracted from 2,595… See the full description on the dataset page: https://huggingface.co/datasets/qwqw3535/econcausal-benchmark.tabulartext-classification10K<n<100K0 likes24 downloads7mo agoHugging Face23Firmansyah-Ibrahim /IndoBloom-AQG-Benchmark-Corpus 📚 Indo-Bloom AQG Benchmark Corpus (All Models) ⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1) This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator Next Stage: Expert Annotation (Stage 2) → Gold Standard 🔒 FROZEN — Benchmark v1.0 This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.tabulartext-generation10K<n<100K0 likes23 downloads6mo agoHugging Face24Piyushdash94 /odia_reasoning_benchmark Odia Reasoning Benchmark This benchmark contains reasoning questions in Odia (logical, Mathematical,Arithmetic, Deductive, Critical Thinking) with answers and optional explanations. Ideal for evaluating Odia QA and reasoning models. Dataset structure Column Description Question Reasoning question in Odia Answer Correct answer (text or number) Explanation Optional step-by-step explanation (some blank) Type Of Question Category (e.g., Math, Deductive)… See the full description on the dataset page: https://huggingface.co/datasets/Piyushdash94/odia_reasoning_benchmark.textquestion-answeringn<1K1 likes19 downloads1y agoHugging Face25cyan12343 /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse… See the full description on the dataset page: https://huggingface.co/datasets/cyan12343/frames-benchmark.texttext-classificationn<1K0 likes18 downloads3mo agoHugging Face26YusufSimsek /turkce-atasozleri-coktan-secmeli-benchmark Türkçe Atasözleri Çoktan Seçmeli Benchmark Türkçe atasözlerinin anlamını ölçmek amacıyla hazırlanmış 100 soruluk çoktan seçmeli bir benchmark veri setidir. Her soruda dört seçenek bulunur. Doğru cevaplar dengeli dağıtılmıştır: A: 25 B: 25 C: 25 D: 25 Rastgele tahmin seviyesi: %25 Sorular, YusufSimsek/turkce-atasozleri-dataset veri setindeki 100 benzersiz atasözünden türetilmiştir. Veri alanları Alan Açıklama id Sorunun benzersiz kimliği proverb… See the full description on the dataset page: https://huggingface.co/datasets/YusufSimsek/turkce-atasozleri-coktan-secmeli-benchmark.textquestion-answeringn<1K0 likes18 downloads2mo agoHugging Face27Kubermatic /Benchmark-Questions Q&A Dataset for Benchmarking DeepCNCF This is a question-and-answer dataset using multiple-choice questions created for benchmarking our DeepCNCF LLM. Since there is no reliable LLM benchmark specified for CNCF projects we decided to use it to measure the performance of our model based on its performance on these questions. This dataset was gathered from openly available online courses about CNCF projects. So they are created by humans to measure students' understanding from these… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/Benchmark-Questions.textquestion-answeringn<1K0 likes17 downloads2y agoHugging Face28RoscommonSystems /MMLU-Phrasing-Benchmark MMLU Phrasing Benchmark This dataset is a phrasing variant of cais/mmlu, put together by Roscommon Systems to see whether the way a question is worded affects how accurately language models answer it. Each of the 2,650 questions appears four ways: the original text from MMLU, a polite version, a formal academic version, and an angry/demanding version. The answer choices and correct answers are identical to the source dataset in all cases. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/RoscommonSystems/MMLU-Phrasing-Benchmark.textquestion-answering100K<n<1M1 likes16 downloads2mo agoHugging Face29TPelc /Fictional_Persona_Dialogs_Anonymized-benchmark Fictional_Persona_Dialogs_Anonymized Benchmark Dataset Short Summary: A 68-pair synthetic Question-Answering (QA) dataset derived from anonymized fictional dialogues, specifically designed for rigorous Retrieval-Augmented Generation (RAG) system evaluation. It isolates and demonstrates the critical impact of contextual information on LLM accuracy. Introduction & Motivation: This dataset addresses the need for a clean, bias-minimized benchmark to accurately… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Fictional_Persona_Dialogs_Anonymized-benchmark.textquestion-answeringn<1K0 likes15 downloads1y agoHugging Face30groundlens /human-confabulation-benchmark Human-Confabulated Hallucination Benchmark Read this before you report a number Known confound: authorship. In this dataset every grounded_response was written by a language model from a source, and every fabricated_response was written by a human from memory. Authorship is therefore perfectly correlated with the label. A detector can score very highly here by recognising who wrote the text rather than whether it is grounded. This is not a hypothetical. In… See the full description on the dataset page: https://huggingface.co/datasets/groundlens/human-confabulation-benchmark.texttext-classificationn<1K0 likes12 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.