CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mtec-TUB /GPT-4o-evaluation-biases A database to support the evaluation of gender biases in GPT-4o output The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025). Introduction This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.question-answering10K<n<100K0 likes1.6k downloads2y agoHugging Face02ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes1k downloads2y agoHugging Face03Lux0926 /ASPRM-BON-Evaluation-Dataset-MathThis repository contains the datasets for the paper AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence. question-answering0 likes347 downloads2y agoHugging Face04FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes346 downloads2y agoHugging Face05plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes211 downloads23d agoHugging Face06openbmb /RLPR-Evaluation Dataset Card for RLPR-Evaluation GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.textquestion-answeringn<1K3 likes195 downloads1y agoHugging Face07ProlificAI /humaine-evaluation-dataset HUMAINE: Human-AI Interaction Evaluation Dataset Dataset Description Dataset Summary The HUMAINE dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts. This dataset powers the HUMAINE Leaderboard, providing insights into how different AI models perform across various user populations and use cases. The dataset consists of two main components: Feedback Comparisons: Pairwise model comparisons… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/humaine-evaluation-dataset.tabularquestion-answering100K<n<1M6 likes169 downloads5mo agoHugging Face08jizej /Competence-Based-Evaluation Competence-Based Evaluation (Invariance Benchmark) A benchmark for testing whether language models give the same answer to semantically equivalent reformulations of a logical-ordering question. Given a set of pairwise constraints (e.g. Alice is in front of Bob), a model should answer transitive-closure queries (Is Carol in front of Dave?) consistently whether the constraints are stated using a relation or its inverse. Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.textquestion-answering100K<n<1M0 likes128 downloads5mo agoHugging Face09joonnam /lifeos-jaram-gaia-evaluation LifeOS Jaram GAIA Level 1 Evaluation Metadata from LifeOS Jaram 4.0 agent's evaluation on the GAIA benchmark Level 1 (2023 test set). 🏆 Results Metric Value Total Tasks 93 Tasks Solved 90 Accuracy ~96.8% HAL Leaderboard World #1 82.1% 🤖 Agent Overview Jaram 4.0 is the core problem-solving agent of the LifeOS multi-agent council. Base Models: Claude 3.5 Sonnet, Gemini 2.0 Flash Architecture: Multi-agent council (Koram + Boran) Live… See the full description on the dataset page: https://huggingface.co/datasets/joonnam/lifeos-jaram-gaia-evaluation.question-answering0 likes123 downloads7mo agoHugging Face10Naholav /claude_4_math_evaluation_500 Claude 4 Mathematical Evaluation Dataset 🎯 Dataset Overview This dataset contains 500 original mathematical evaluation problems specifically designed to test Claude 4 Sonnet's ability to assess mathematical answers as correct or incorrect. The problems were created from a single prototype and expanded to ensure the model had never encountered them during training, eliminating memorization bias. Why This Dataset Matters Exposes a critical flaw in LLM… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/claude_4_math_evaluation_500.documenttext-classificationn<1K1 likes92 downloads1y agoHugging Face11dougdotcon /douvras-ptbr-enterprise-ai-evaluation Douvras PT-BR Enterprise AI Evaluation Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada: responder somente a partir de um documento fornecido; reconhecer quando a informação não está disponível; resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.textquestion-answering0 likes85 downloads10d agoHugging Face12CEM888AI /cem888-independent-runtime-evaluation What Happens When the Model Is Disposable? A Technical Evaluation of the CEM888 Agent Runtime Independent evaluation performed by an AI engineering assistant running on Hugging Face Jobs infrastructure (ephemeral CPU sandboxes), September 18, 2026. Not an official Hugging Face evaluation or endorsement. The evaluator has no affiliation with the CEM888 project; the project maintainer supplied the installer command, an install code, and the provider API key used for the live model… See the full description on the dataset page: https://huggingface.co/datasets/CEM888AI/cem888-independent-runtime-evaluation.question-answering1 likes72 downloads5d agoHugging Face13aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes57 downloads6mo agoHugging Face14AnjanSB /NQ-RAG-DPO-Evaluation Dataset Card Dataset Summary This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO). The system is organized into three interconnected pipelines: 1️. RAG Pipeline The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark. For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.texttext-generation1K<n<10K1 likes49 downloads7mo agoHugging Face15Sr523 /big-red-bark-chat-evaluation Big Red Bark Chat Q&A Dataset Dataset Description This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.textquestion-answering10K<n<100K0 likes46 downloads3mo agoHugging Face16mlahmy /seal-rag-evaluation-data SEAL-RAG Evaluation Data This repository contains the evaluation slices used in the paper Replace, Don't Expand: Mitigating Context Dilution in Multi-Hop RAG. Files hotpotqa_1000_v1.csv: The 1,000-sample slice from HotpotQA used for the primary evaluation. 2WikiMultihopQA_200_v1.csv: The 200-sample slice from 2WikiMultihopQA used for additional testing. Citation If you use this data, please cite the original paper. question-answering0 likes44 downloads7mo agoHugging Face17SahmBenchmark /fatwa-qa-evaluation Fatwa QA Evaluation Dataset Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers. Dataset Statistics Total Samples: 2,000 Average Question Length: 243.9 characters Average Answer Length: 492.3 characters Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.tabularquestion-answering1K<n<10K0 likes43 downloads10mo agoHugging Face18stucksam /BaldursGate3-Evaluation-Datasettextquestion-answeringn<1K0 likes42 downloads3y agoHugging Face19SahmBenchmark /fatwa-mcq-evaluation_standardized Fatwa MCQ Evaluation Dataset (Standardized) Standardized multiple-choice question dataset for evaluating Islamic jurisprudence knowledge. Dataset Description This dataset contains MCQ versions of Islamic fatwa Q&A pairs, standardized for evaluation purposes. Dataset Summary Language: Arabic Domain: Islamic Finance, Jurisprudence (Fiqh) Format: Multiple choice questions (4 options) Task: Islamic jurisprudence knowledge evaluation… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-mcq-evaluation_standardized.textmultiple-choice1K<n<10K0 likes37 downloads9mo agoHugging Face20jayzou3773 /less-is-moe-gpqa-diamond-evaluation Less-is-MoE GPQA-Diamond evaluation set This private dataset stores the 198-question GPQA-Diamond evaluation file used by the MoE-Honing evaluation format. Upstream source: Idavidrein/gpqa, config gpqa_diamond Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd Split: test Rows: 198 SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3 Fields: problem, solution, domain The problem field contains the formatted four-choice prompt, solution stores the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.textquestion-answeringn<1K0 likes35 downloads3d agoHugging Face21Heng666 /Traditional_Chinese-aya_evaluation_suite 資料集描述 繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集 概述 繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。 此資料集結合了來自 CohereForAI/aya_evaluation_suite,過濾掉除繁體中文、簡體中文內容之外的所有內容。 目標 繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。 資料集來源與資訊 資料來源: 從 CohereForAI/aya_evaluation_suite 3 個子集而來。 語言: 繁體中文、簡體中文('zho') 應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。 論文連結: 2402.06619 維護人: Heng666 License:… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_evaluation_suite.textquestion-answeringn<1K3 likes30 downloads3y agoHugging Face22open-biosciences /biosciences-evaluation-inputs Biosciences RAG Evaluation Inputs Dataset Description This dataset contains RAG inference outputs from 4 retrieval strategies evaluated on 12 biosciences research questions. Each retriever was tested on the same golden testset, producing 48 total records with retrieved contexts and LLM-generated answers ready for RAGAS evaluation. Dataset Summary Total Examples: 48 records (12 questions x 4 retrievers) Retrievers Compared: Naive — Dense vector similarity… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-evaluation-inputs.textquestion-answeringn<1K0 likes29 downloads7mo agoHugging Face23stindardlogic /llm-evaluation-sft-100k LLM Evaluation SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering LLM evaluation methodology, benchmarking, safety assessment, and prompt optimization. Designed to train AI assistants that can help ML engineers and researchers rigorously evaluate and improve language models. Dataset Description This dataset covers the full spectrum of LLM evaluation practice across 7 specialized categories. Each record follows the… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/llm-evaluation-sft-100k.text-generation100K<n<1M0 likes27 downloads2mo agoHugging Face24zli12321 /pedants_qa_evaluation_bench pedants_qa_evaluation This dataset evaluates candidate answers for various question-answering (QA) tasks across multiple datasets such as Jeopardy!, hotpotQA, nq-open, narrativeQA, and BIOMRC, etc. See details in paper. It contains questions, reference answers (ground truth), model-generated candidate answers, and human judgments indicating whether the candidate answers are correct. Dataset Details Column Type Description question string The question asked… See the full description on the dataset page: https://huggingface.co/datasets/zli12321/pedants_qa_evaluation_bench.textquestion-answering10K<n<100K1 likes26 downloads2y agoHugging Face25dwb2023 /gdelt-rag-evaluation-metrics GDELT RAG Detailed Evaluation Results Dataset Description This dataset contains detailed RAGAS evaluation results with per-question metric scores for 5 different retrieval strategies tested on the GDELT RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores. Dataset Summary Total Examples: ~1,400+ evaluation records with metric scores Retrievers Evaluated: Baseline, Naive, BM25, Ensemble, Cohere… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-metrics.tabularquestion-answeringn<1K0 likes22 downloads11mo agoHugging Face26dwb2023 /gdelt-rag-evaluation-inputs GDELT RAG Evaluation Datasets Dataset Description This dataset contains consolidated RAGAS evaluation input datasets from 5 different retrieval strategies tested on the GDELT (Global Database of Events, Language, and Tone) RAG system. Each strategy was evaluated on the same golden testset of 12 questions, providing a direct comparison of retrieval performance. Dataset Summary Total Examples: ~1,400+ evaluation records across 5 retrievers Retrievers Compared:… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-inputs.textquestion-answeringn<1K0 likes20 downloads11mo agoHugging Face27dwb2023 /gdelt-rag-evaluation-metrics-v3 GDELT RAG Detailed Evaluation Results Dataset Description This dataset contains detailed RAGAS evaluation results with per-question metric scores for 4 different retrieval strategies tested on the GDELT RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores. Dataset Summary Total Examples: 48 evaluation records with metric scores (12 questions × 4 retrievers) Retrievers Evaluated: Naive (baseline)… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-metrics-v3.tabularquestion-answeringn<1K0 likes18 downloads11mo agoHugging Face28open-biosciences /biosciences-evaluation-metrics Biosciences RAG Evaluation Metrics Dataset Description This dataset contains detailed RAGAS evaluation results with per-question metric scores for 4 retrieval strategies tested on the biosciences RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores. Dataset Summary Total Examples: 48 records (12 questions x 4 retrievers) Retrievers Evaluated: Naive, BM25, Ensemble, Cohere Rerank Metrics Per… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-evaluation-metrics.tabularquestion-answeringn<1K0 likes15 downloads7mo agoHugging Face29sujitpandey /mobile_sft_evaluation Mobile Sft Evaluation Dataset Description Mobile QA evaluation dataset with 200 randomly sampled questions and rewritten answers for SFT model evaluation. Answers maintain semantic equivalence with different phrasing for robust evaluation. Dataset Summary Total Examples: 200 Task: Question Answering Language: English Format: JSONL (one JSON object per line) Dataset Structure Example Entry { "question": "How can mobile technology expand… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/mobile_sft_evaluation.textquestion-answeringn<1K0 likes14 downloads11mo agoHugging Face30abirT /combined_ophthalmology_evaluation_datasettexttext-generation10K<n<100K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.