datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.Chat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.Cultural-Evaluation-Kalahi
Kalahi
Kalahi evaluates the ability of LLMs to generate responses relevant to Filipino culture in terms of shared knowledge and ethics. This dataset contains a MCQ-compatible version of the Kalahi dataset that is used in SEA-HELM.
Supported Tasks and Leaderboards
Kalahi is designed for evaluating Filipino cultural representations in instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Tagalog (tl)… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Cultural-Evaluation-Kalahi.VeriLoop-E2-Evaluation-Evidence
VeriLoop E2 Evaluation Evidence
Public evaluation evidence for VeriLoop E2 across nine code, agentic, mathematical, and scientific reasoning benchmarks.
This dataset repository is the canonical public evidence layer for the reported benchmark results of VeriLoop E2, a post-trained model based on Qwen 3.8-27B. It is designed to separate headline benchmark reporting from the underlying auditable artifacts required to inspect, reproduce, and verify those results.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence.patient-evaluations
Patient Evaluations Dataset
This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data.
Dataset Description
The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions.
Dataset Structure
The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.sdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.RLPR-Evaluation
Dataset Card for RLPR-Evaluation
GitHub | Paper
News:
[2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at here!
Dataset Summary
We include the following seven benchmarks for evaluation of RLPR:
Mathematical Reasoning Benchmarks:
MATH-500 (Cobbe et al., 2021)
Minerva (Lewkowycz et al., 2022)
AIME24
General Domain Reasoning Benchmarks:
MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Evaluation.ZEDA-Evaluation
ZEDA Dataset
This repository contains the training data for ZEDA (Zero-Expert Self-Distillation Adaptation), a framework introduced in the paper Post-Trained MoE Can Skip Half Experts via Self-Distillation.
ZEDA is a low-cost framework that transforms post-trained static MoE models into efficient dynamic ones by injecting zero experts and using self-distillation.
Paper: Post-Trained MoE Can Skip Half Experts via Self-Distillation
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/TsinghuaC3I/ZEDA-Evaluation.Crab-role-playing-evaluation-benchmark
📄 Paper
|
📄 Github
💬 Role-playing Model
|
💬 Role-palying Evaluation Model
💬 Training Dataset
|
💬 Evaluation Benchmark
|
💬 Annotated Role-playing Evaluation Dataset
|
💬 Human-preference Dataset
1. Introduction
This is the dataset used for evalauating a role‑playing LLM.
More details can be seen at GitHub and Crab… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/Crab-role-playing-evaluation-benchmark.Competence-Based-Evaluation
Competence-Based Evaluation (Invariance Benchmark)
A benchmark for testing whether language models give the same answer to
semantically equivalent reformulations of a logical-ordering question. Given a
set of pairwise constraints (e.g. Alice is in front of Bob), a model should
answer transitive-closure queries (Is Carol in front of Dave?) consistently
whether the constraints are stated using a relation or its inverse.
Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1338_peixian_equity_evaluation_corpus_sentiment_classifier
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1338_peixian_equity_evaluation_corpus_sentiment_classifier.turkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.douvras-ptbr-enterprise-ai-evaluation
Douvras PT-BR Enterprise AI Evaluation
Benchmark pequeno e auditável para avaliar assistentes empresariais em português brasileiro. A versão 0.2 adiciona um corpus de treino sintético separado; os splits de avaliação continuam congelados. O benchmark mede quatro capacidades que aparecem em projetos reais de IA aplicada:
responder somente a partir de um documento fornecido;
reconhecer quando a informação não está disponível;
resistir a instruções maliciosas inseridas no conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-ptbr-enterprise-ai-evaluation.llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/BushNate/llm-refusal-evaluation.llm-output-evaluation
LLM_OUTPUT_EVALUATION
A preference dataset for LLM_OUTPUT_EVALUATION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/llm-output-evaluation.mHallucination_Evaluation
Multilingual Hallucination Evaluation in the wild
The dataset was as part of the paper: How Much Do LLMs Hallucinate across Languages? On Multilingual Estimation of LLM Hallucination in the Wild
Below is the figure summarizing the multilingual hallucination evaluation dataset creation (and multilingual hallucination detection dataset):
Dataset Details
The dataset is a high quality synthetic query/prompt and wikipedia reference pair for estimating hallucinations in the… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/mHallucination_Evaluation.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.advent_of_code_evaluations
Advent of Code Evaluation
This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.japanese-humor-evaluation-v2
Japanese Multimodal Humor Evaluation Dataset (v2)
画像/テキストのお題に対する面白い回答のデータセット。bokete(画像→テキスト)とkeitai(テキスト→テキスト)を統合。
使い方
from datasets import load_dataset
dataset = load_dataset("iammytoo/japanese-humor-evaluation-v2")
データ構造
odai_type: 'image' or 'text'
image: 画像お題(textタイプではNone)
odai: テキストお題(imageタイプではNone)
response: 回答テキスト
score: 0-4の正規化スコア
ソース
YANS-official/ogiri-bokete
YANS-official/ogiri-keitai
NQ-RAG-DPO-Evaluation
Dataset Card
Dataset Summary
This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO).
The system is organized into three interconnected pipelines:
1️. RAG Pipeline
The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark.
For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.sdf_evaluation_traits_15M
Models That Know How Evaluations Are Designed Score Safer
This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.less-is-moe-gpqa-diamond-evaluation
Less-is-MoE GPQA-Diamond evaluation set
This private dataset stores the 198-question GPQA-Diamond evaluation file used
by the MoE-Honing evaluation format.
Upstream source: Idavidrein/gpqa, config gpqa_diamond
Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd
Split: test
Rows: 198
SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3
Fields: problem, solution, domain
The problem field contains the formatted four-choice prompt, solution stores
the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.Contextual_Response_Evaluation_for_ESL_and_ASD_Support
Dataset Card for "Contextual Response Evaluation for ESL and ASD Support💜💬🌐""
Dataset Description 📖
Dataset Summary 📝
Curated by Eric Soderquist, this dataset is a collection of English prompts and responses generated by the Phi-2 model, designed to evaluate and improve NLP models for supporting ESL (English as a Second Language) and ASD (Autism Spectrum Disorder) user bases. Each prompt is paired with multiple AI-generated responses and evaluated using a… See the full description on the dataset page: https://huggingface.co/datasets/yunjaeys/Contextual_Response_Evaluation_for_ESL_and_ASD_Support.keyphrase_homogeneity_evaluation
license: cc-by-nc-4.0
language:
- en
size_categories:
- n<1K
Data pairs used in the evaluation of the paper "[Evaluating the Homogeneity of Keyphrase Prediction Models]"(https://arxiv.org/abs/2602.12989), Maël Houbre, Florian Boudin and Béatrice Daille, LREC 2026
Molding-Generation-Evaluation
Molding-Generation-Evaluation Dataset
Molding 도메인(사출 성형, 금형)에 대한 LLM의 생성 품질을 평가하기 위한 데이터셋입니다.
tokenized-IELTS-writing-task-2-evaluation-DialoGPT-medium
