datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.chess-position-evaluations
Dataset Card for the Lichess Evaluations dataset
Dataset Description
394,669,566 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 957,860,115 rows.
This dataset is updated monthly, and was last updated on July 8th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.chess-evaluations
Chess Evaluations Dataset
This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations:
tactics: Includes chess positions, their evaluations, and the best move in the position.
randoms: Contains random chess positions and their evaluations.
chess_data: General chess positions with evaluations.
This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.SWE-MERA
SWE-MERA
Continuously updated SWE-MERA dataset
SWE-MERA splits:
dev: for testing (10 samples)
lite: presented at the leaderboard here (750 samples)
full: continuously updated to collect more data (2738 samples)
Load dataset
from datasets import load_dataset
ds = load_dataset("MERA-evaluation/SWE-MERA", split='dev')
Evaluation
Description
The main tool to validate tasks is repotest (available at PyPI or GitHub)
data.jsonl -… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/SWE-MERA.wmt-da-human-evaluation-long-context
Dataset Summary
Long-context / document-level dataset for Quality Estimation of Machine Translation.
It is an augmented variant of the sentence-level WMT DA Human Evaluation dataset.
In addition to individual sentences, it contains augmentations of 2, 4, 8, 16, and 32 sentences, among each language pair lp and domain.
The raw column represents a weighted average of scores of augmented sentences using character lengths of src and mt as weights.
The code used to apply the augmentation… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/wmt-da-human-evaluation-long-context.HIP-training-and-evaluation-data
HIP Training and Evaluation Data
This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation.
Configs
training
data/train.parquet contains 10,581 supervised HIP training pairs with seven columns:
dataset: upstream dataset family, either raid or mage.
source: selected source domain or subcorpus.
text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.CoderForge-Preview-32B-SWE-Bench-Verified-Evaluation-trajectoriesEvaluation-Multilingual-VC
Evaluation-Multilingual-VC
We use dataset https://huggingface.co/datasets/sarulab-speech/commonvoice22_sidon,
Filter languages that support by Whisper Large V3 to evaluate WER automatically,
Only take test set, sort by up votes.
Because VC required to source text, source audio, target text, we make sure the target text is not same as source text, target text we take from other rows.
Only build first 500 rows for each language
Github issue at… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Evaluation-Multilingual-VC.RAG-Evaluation-Dataset-KO
Dataset Card for Reconstructed RAG Evaluation Dataset (KO)
Dataset Summary
본 데이터셋은 allganize/RAG-Evaluation-Dataset-KO를 기반으로 PDF 파일을 포함하도록 재구성한 한국어 평가 데이터셋입니다. 원본 데이터셋에서는 PDF 파일의 경로만 제공되어 수동으로 파일을 다운로드해야 하는 불편함이 있었고, 일부 PDF 파일의 경로가 유효하지 않은 문제를 보완하기 위해 PDF 파일을 포함한 데이터셋을 재구성하였습니다.
Supported Tasks and Leaderboards
RAG Evaluation: 본 데이터는 한국어 RAG 파이프라인에 대한 E2E Evaluation이 가능합니다.
Languages
The dataset is in Korean (ko).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/datalama/RAG-Evaluation-Dataset-KO.humaine-evaluation-dataset
HUMAINE: Human-AI Interaction Evaluation Dataset
Dataset Description
Dataset Summary
The HUMAINE dataset contains human evaluations of AI model interactions across diverse demographic groups and conversation contexts. This dataset powers the HUMAINE Leaderboard, providing insights into how different AI models perform across various user populations and use cases.
The dataset consists of two main components:
Feedback Comparisons: Pairwise model comparisons… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/humaine-evaluation-dataset.evaluationturkish-brand-bias-evaluations
Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri
Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü
incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir.
An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in
Turkish product and service recommendations.
Veri seti özeti
300 tamamlanmış ve judge edilmiş yanıt
Domainler: VPN (150) ve kozmetik (150)
Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.InferBench-evaluation-resultskg-gen-MINE-evaluation-datasetMMDocIR_Evaluation_Dataset
Evaluation Datasets
Evaluation Set Overview
MMDocIR evaluation set includes 313 long documents averaging 65.1 pages, categorized into ten main domains: research reports, administration&industry, tutorials&workshops, academic papers, brochures, financial reports, guidebooks, government documents, laws, and news articles. Different domains feature distinct distributions of multi-modal information. Overall, the modality distribution is: Text (60.4%), Image (18.8%)… See the full description on the dataset page: https://huggingface.co/datasets/yyyyyyyyy111/MMDocIR_Evaluation_Dataset.evaluation-results-task1evalap-legalbenchrag-evaluation-v1-111
LegalBenchRAG Evaluation v1 (ID: 111)
A extensive RAG evaluation on the LegalBenchRAG dataset. See [complete me]
Overview
This dataset contains 36 experiments
from the EvalAP evaluation platform.
Datasets: LegalBenchRAG
Models evaluated: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Metrics: judge_precision, output_length
Scores
LegalBenchRAG
model
judge_precision
output_length
model_semantic_20_qwen3_lbrv5
0.82 ± 0.38
163.13 ± 157.94… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-legalbenchrag-evaluation-v1-111.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.fatwa-qa-evaluation
Fatwa QA Evaluation Dataset
Dataset Description
This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers.
Dataset Statistics
Total Samples: 2,000
Average Question Length: 243.9 characters
Average Answer Length: 492.3 characters
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.cardio_evaluationsrelease-datasetBeaverTails-IT-Evaluation
BeaverTails-IT-Evaluation
This dataset is an Italian machine translated version of https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.
This dataset is created by automatically translating the original examples from Beavertails-Evaluation into Italian.
Multiple state-of-the-art translation models were employed to generate alternative translations.
You can find more information in our Paper.
Translation Models
The dataset includes translations… See the full description on the dataset page: https://huggingface.co/datasets/MIND-Lab/BeaverTails-IT-Evaluation.Retrival_evaluation_HR2onepane-llm-evaluation-geminievalap-mediatechs-legi-chunking-evaluation-v1-116
MediaTech's LEGI Chunking Evaluation V1 (ID: 116)
Evaluation of severals chunking strategies for MediaTech's LEGI dataset.
Overview
This dataset contains 51 experiments
from the EvalAP evaluation platform.
Datasets: LEGI Synthetic QA Dataset
Metrics: contextual_precision, contextual_recall, contextual_relevancy, faithfulness, judge_precision
Scores
LEGI Synthetic QA Dataset
model
contextual_precision
contextual_recall
contextual_relevancy… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-mediatechs-legi-chunking-evaluation-v1-116.speaker_evaluation_multi_test_v0
Seamless Interaction Pairs
This dataset contains paired query and document audio clips for interaction-based
speaker evaluation. Each row describes a query clip and a related document clip,
with segment metadata and durations for analysis.
Data structure
The dataset uses a single split stored in data.parquet.
Audio files are stored under audio/ and referenced by relative paths in the
parquet file.
Columns
pair_id (string): Pair identifier.
interaction… See the full description on the dataset page: https://huggingface.co/datasets/humanify/speaker_evaluation_multi_test_v0.Human-Evaluation
Human Evaluation Dataset
The dataset includes human evaluation for General and Health domains. It was created as part of my two papers:
“Domain-Specific Text Generation for Machine Translation” (Moslem et al., 2022)
"Adaptive Machine Translation with Large Language Models" (Moslem et al., 2023)
The evaluators were asked to assess the acceptability of each translation
using a scale ranging from 1 to 4, where 4 is ideal and 1 is unacceptable translation.
For the paper Moslem et al.… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Human-Evaluation.llama3-ultrafeedback-armo-test-evaluation-rewards-logprobsDeepEval-Question-Answer-Dataset-for-RAG-Evaluation-A2A-And-ACP-PDFeval_my_smolvla_evaluationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 14000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hangVLA/eval_my_smolvla_evaluation.
