datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-4o-evaluation-biases
A database to support the evaluation of gender biases in GPT-4o output
The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025).
Introduction
This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.ASPRM-BON-Evaluation-Dataset-MathThis repository contains the datasets for the paper AdaptiveStep: Automatically Dividing Reasoning Step through Model Confidence.
Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.RLPR-Evaluation
Dataset Card for RLPR-Evaluation
GitHub | Paper
News:
[2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here!
Dataset Summary
We include the following seven benchmarks for evaluation of RLPR:
Mathematical Reasoning Benchmarks:
MATH-500 (Cobbe et al., 2021)
Minerva (Lewkowycz et al., 2022)
AIME24
General Domain Reasoning Benchmarks:
MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.humaine-evaluation-dataset
HUMAINE: Human-AI Interaction Evaluation Dataset
Dataset Description
Dataset Summary
The HUMAINE dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts. This dataset powers the HUMAINE Leaderboard, providing insights into how different AI models perform across various user populations and use cases.
The dataset consists of two main components:
Feedback Comparisons: Pairwise model comparisons… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/humaine-evaluation-dataset.Competence-Based-Evaluation
Competence-Based Evaluation (Invariance Benchmark)
A benchmark for testing whether language models give the same answer to
semantically equivalent reformulations of a logical-ordering question. Given a
set of pairwise constraints (e.g. Alice is in front of Bob), a model should
answer transitive-closure queries (Is Carol in front of Dave?) consistently
whether the constraints are stated using a relation or its inverse.
Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.lifeos-jaram-gaia-evaluation
LifeOS Jaram GAIA Level 1 Evaluation
Metadata from LifeOS Jaram 4.0 agent's evaluation on the GAIA benchmark Level 1 (2023 test set).
🏆 Results
Metric
Value
Total Tasks
93
Tasks Solved
90
Accuracy
~96.8%
HAL Leaderboard World #1
82.1%
🤖 Agent Overview
Jaram 4.0 is the core problem-solving agent of the LifeOS multi-agent council.
Base Models: Claude 3.5 Sonnet, Gemini 2.0 Flash
Architecture: Multi-agent council (Koram + Boran)
Live… See the full description on the dataset page: https://huggingface.co/datasets/joonnam/lifeos-jaram-gaia-evaluation.claude_4_math_evaluation_500
Claude 4 Mathematical Evaluation Dataset
🎯 Dataset Overview
This dataset contains 500 original mathematical evaluation problems specifically designed to test Claude 4 Sonnet's ability to assess mathematical answers as correct or incorrect. The problems were created from a single prototype and expanded to ensure the model had never encountered them during training, eliminating memorization bias.
Why This Dataset Matters
Exposes a critical flaw in LLM… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/claude_4_math_evaluation_500.douvras-ptbr-enterprise-ai-evaluation
Douvras PT-BR Enterprise AI Evaluation
Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada:
responder somente a partir de um documento fornecido;
reconhecer quando a informação não está disponível;
resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.cem888-independent-runtime-evaluation
What Happens When the Model Is Disposable? A Technical Evaluation of the CEM888 Agent Runtime
Independent evaluation performed by an AI engineering assistant running on Hugging Face Jobs infrastructure (ephemeral CPU sandboxes), September 18, 2026. Not an official Hugging Face evaluation or endorsement. The evaluator has no affiliation with the CEM888 project; the project maintainer supplied the installer command, an install code, and the provider API key used for the live model… See the full description on the dataset page: https://huggingface.co/datasets/CEM888AI/cem888-independent-runtime-evaluation.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.NQ-RAG-DPO-Evaluation
Dataset Card
Dataset Summary
This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO).
The system is organized into three interconnected pipelines:
1️. RAG Pipeline
The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark.
For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.big-red-bark-chat-evaluation
Big Red Bark Chat Q&A Dataset
Dataset Description
This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.seal-rag-evaluation-data
SEAL-RAG Evaluation Data
This repository contains the evaluation slices used in the paper Replace, Don't Expand: Mitigating Context Dilution in Multi-Hop RAG.
Files
hotpotqa_1000_v1.csv: The 1,000-sample slice from HotpotQA used for the primary evaluation.
2WikiMultihopQA_200_v1.csv: The 200-sample slice from 2WikiMultihopQA used for additional testing.
Citation
If you use this data, please cite the original paper.
fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.BaldursGate3-Evaluation-Datasetfatwa-mcq-evaluation_standardized
Fatwa MCQ Evaluation Dataset (Standardized)
Standardized multiple-choice question dataset for evaluating Islamic jurisprudence knowledge.
Dataset Description
This dataset contains MCQ versions of Islamic fatwa Q&A pairs, standardized for evaluation purposes.
Dataset Summary
Language: Arabic
Domain: Islamic Finance, Jurisprudence (Fiqh)
Format: Multiple choice questions (4 options)
Task: Islamic jurisprudence knowledge evaluation… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-mcq-evaluation_standardized.less-is-moe-gpqa-diamond-evaluation
Less-is-MoE GPQA-Diamond evaluation set
This private dataset stores the 198-question GPQA-Diamond evaluation file used
by the MoE-Honing evaluation format.
Upstream source: Idavidrein/gpqa, config gpqa_diamond
Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd
Split: test
Rows: 198
SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3
Fields: problem, solution, domain
The problem field contains the formatted four-choice prompt, solution stores
the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.Traditional_Chinese-aya_evaluation_suite
資料集描述
繁體中文 Aya (Traditional Chinese Aya Chinese;TCA):專注於繁體中文處理的 Aya 集合的精選子集
概述
繁體中文 Aya 是一個精心策劃的資料集,源自 CohereForAI 的綜合 Aya 集合,特別關注繁體中文文本資料。
此資料集結合了來自 CohereForAI/aya_evaluation_suite,過濾掉除繁體中文、簡體中文內容之外的所有內容。
目標
繁體中文 Aya 的目標是為研究人員、技術專家和語言學家提供即用型繁體中文文本資源,顯著減少專注於繁體中文的 NLP 和 AI 專案中數據預處理所需的時間和精力。
資料集來源與資訊
資料來源: 從 CohereForAI/aya_evaluation_suite 3 個子集而來。
語言: 繁體中文、簡體中文('zho')
應用: 非常適合語言建模、文本分類、情感分析、和機器翻譯等任務。
論文連結: 2402.06619
維護人: Heng666
License:… See the full description on the dataset page: https://huggingface.co/datasets/Heng666/Traditional_Chinese-aya_evaluation_suite.biosciences-evaluation-inputs
Biosciences RAG Evaluation Inputs
Dataset Description
This dataset contains RAG inference outputs from 4 retrieval strategies evaluated on 12 biosciences research questions. Each retriever was tested on the same golden testset, producing 48 total records with retrieved contexts and LLM-generated answers ready for RAGAS evaluation.
Dataset Summary
Total Examples: 48 records (12 questions x 4 retrievers)
Retrievers Compared:
Naive — Dense vector similarity… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-evaluation-inputs.llm-evaluation-sft-100k
LLM Evaluation SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering LLM evaluation methodology, benchmarking, safety assessment, and prompt optimization. Designed to train AI assistants that can help ML engineers and researchers rigorously evaluate and improve language models.
Dataset Description
This dataset covers the full spectrum of LLM evaluation practice across 7 specialized categories. Each record follows the… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/llm-evaluation-sft-100k.pedants_qa_evaluation_bench
pedants_qa_evaluation
This dataset evaluates candidate answers for various question-answering (QA) tasks across multiple datasets such as Jeopardy!, hotpotQA, nq-open, narrativeQA, and BIOMRC, etc. See details in paper. It contains questions, reference answers (ground truth), model-generated candidate answers, and human judgments indicating whether the candidate answers are correct.
Dataset Details
Column
Type
Description
question
string
The question asked… See the full description on the dataset page: https://huggingface.co/datasets/zli12321/pedants_qa_evaluation_bench.gdelt-rag-evaluation-metrics
GDELT RAG Detailed Evaluation Results
Dataset Description
This dataset contains detailed RAGAS evaluation results with per-question metric scores for 5 different retrieval strategies tested on the GDELT RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores.
Dataset Summary
Total Examples: ~1,400+ evaluation records with metric scores
Retrievers Evaluated: Baseline, Naive, BM25, Ensemble, Cohere… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-metrics.gdelt-rag-evaluation-inputs
GDELT RAG Evaluation Datasets
Dataset Description
This dataset contains consolidated RAGAS evaluation input datasets from 5 different retrieval strategies tested on the GDELT (Global Database of Events, Language, and Tone) RAG system. Each strategy was evaluated on the same golden testset of 12 questions, providing a direct comparison of retrieval performance.
Dataset Summary
Total Examples: ~1,400+ evaluation records across 5 retrievers
Retrievers Compared:… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-inputs.gdelt-rag-evaluation-metrics-v3
GDELT RAG Detailed Evaluation Results
Dataset Description
This dataset contains detailed RAGAS evaluation results with per-question metric scores for 4 different retrieval strategies tested on the GDELT RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores.
Dataset Summary
Total Examples: 48 evaluation records with metric scores (12 questions × 4 retrievers)
Retrievers Evaluated: Naive (baseline)… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-evaluation-metrics-v3.biosciences-evaluation-metrics
Biosciences RAG Evaluation Metrics
Dataset Description
This dataset contains detailed RAGAS evaluation results with per-question metric scores for 4 retrieval strategies tested on the biosciences RAG system. Each record includes the full evaluation context (question, contexts, response) plus 4 RAGAS metric scores.
Dataset Summary
Total Examples: 48 records (12 questions x 4 retrievers)
Retrievers Evaluated: Naive, BM25, Ensemble, Cohere Rerank
Metrics Per… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-evaluation-metrics.mobile_sft_evaluation
Mobile Sft Evaluation
Dataset Description
Mobile QA evaluation dataset with 200 randomly sampled questions and rewritten answers for SFT model evaluation. Answers maintain semantic equivalence with different phrasing for robust evaluation.
Dataset Summary
Total Examples: 200
Task: Question Answering
Language: English
Format: JSONL (one JSON object per line)
Dataset Structure
Example Entry
{
"question": "How can mobile technology expand… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/mobile_sft_evaluation.combined_ophthalmology_evaluation_dataset
