datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MERA
MERA (Multimodal Evaluation for Russian-language Architectures)
Summary
MERA (Multimodal Evaluation for Russian-language Architectures) is a new open independent benchmark for the evaluation of SOTA models for the Russian language.
The MERA benchmark unites industry and academic partners in one place to research the capabilities of fundamental models, draw attention to AI-related issues, foster collaboration within the Russian Federation and in the international arena… See the full description on the dataset page: https://huggingface.co/datasets/MERA-evaluation/MERA.Chat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.BeaverTails-Evaluation
Dataset Card for BeaverTails-Evaluation
BeaverTails is an AI safety-focused collection comprising a series of datasets.
This repository contains test prompts specifically designed for evaluating language model safety.
It is important to note that although each prompt can be connected to multiple categories, only one category is labeled for each prompt.
The 14 harm categories are defined as follows:
Animal Abuse: This involves any form of cruelty or harm inflicted on animals… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/BeaverTails-Evaluation.java_evaluation_benchmarkswebvoyager_evaluation_dataMedical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.CodeTruthAgent-V3-Module1-Evaluation
Source Code
https://github.com/Zeeshan78699/CodeTruthAgent — tag v3.0.0-module1
CodeTruth Agent V3 — Module 1 Evaluation
Validation results for Module 1: Repository Cognition Engine —
a deterministic, rule-based engine that scans a software repository
and determines its application type, primary framework, technology
stack, and file inventory.
What's in this dataset
FULL_DOMAIN_SUMMARY.md — summary table of all 69 validated repositories… See the full description on the dataset page: https://huggingface.co/datasets/ZeeshanSaud/CodeTruthAgent-V3-Module1-Evaluation.UGround-Offline-Evaluationsdf_evaluation_traits
Models That Know How Evaluations Are Designed Score Safer
This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer.
Project Page | GitHub Repository
Dataset Description
These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations.
Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.Bias-Evaluation-TurkishTranslation of bias evaluation framework of May et al. (2019) from this repository and this paper into Turkish. There is a total of 37 tests including tests addressing gender-bias as well as tests designed to evaluate the ethnic bias toward Kurdish people in Türkiye context.
Abstract of the paper:
While the growing size of pre-trained language models has led to large improvements in a variety of natural language processing tasks, the success of these models comes with a price: They are trained… See the full description on the dataset page: https://huggingface.co/datasets/orhunc/Bias-Evaluation-Turkish.VC-startup-evaluation-for-investmentThis data set includes the completion pairs for evaluating startups before investing in them.
This data set iincludes completion examples for Chain of Thought reasoning to perform financial calculations.
This data set includes completion examples for evaluating risk profile, growth propspects, cost, ratios, market size, asset, liability, debt, equity and other ratios.
This data set includes comparison of different startups.
normistral-11b-thinking-evaluationairep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.Competence-Based-Evaluation
Competence-Based Evaluation (Invariance Benchmark)
A benchmark for testing whether language models give the same answer to
semantically equivalent reformulations of a logical-ordering question. Given a
set of pairwise constraints (e.g. Alice is in front of Bob), a model should
answer transitive-closure queries (Is Carol in front of Dave?) consistently
whether the constraints are stated using a relation or its inverse.
Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.Evaluation-Dataset-of-AI-Agent-Security-Guardrails
DKnownAI Agent Security Evaluation Dataset
Data Fields
Field
Type
Description
text
string
The adversarial input (prompt) to be evaluated by a security guardrail
action
string
Human-annotated label: blocked or allowed
Citation
@misc{li2026comparativeevaluationaiagent,
title={A Comparative Evaluation of AI Agent Security Guardrails},
author={Qi Li and Jiu Li and Pingtao Wei and Jianjun Xu and Xueyi Wei and Jiwei Shi and Xuan… See the full description on the dataset page: https://huggingface.co/datasets/CaiZhiTech/Evaluation-Dataset-of-AI-Agent-Security-Guardrails.stockfish-evaluation-SAN
Dataset Card for the Stockfish Evaluations
A dataset of chess positions evaluated with various flavours of Stockfish running within user browsers. Produced by, and for, the Lichess analysis board. Evaluations are formatted as JSON; one position per line.
The schema of a position looks like this:
{
"fen": "8/8/2B2k2/p4p2/5P1p/Pb6/1P3KP1/8 w - -",
"depth": 42,
"evaluation": 5.64,
"best_move": "Kg1",
"best_line": "Kg1 Ke6 Kh2 Kd6 Be8 Kc5 Kh3 Kd6 Bb5 Ke7"
}
fen: string, the… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/stockfish-evaluation-SAN.Comparative-Idea-Evaluation
🔬 Comparative Idea Evaluation
Dataset accompanying our Findings of ACL 2026 paper:
Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation.
Srujan P Mule · Aniketh Garikaparthi · Manasi Patwardhan
📄 ACL Anthology · arXiv
Research overview from Figure 1 of the paper. This release contains the comparison datasets; reasoning-training variants illustrated in the figure are not included.
🧠 What is this dataset for?
Given a research… See the full description on the dataset page: https://huggingface.co/datasets/anikethh/Comparative-Idea-Evaluation.data_for_evaluationrag-evaluation-lab-20260829-dataset
RAG Evaluation Lab Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for RAG systems often ship without a stable regression set or failure taxonomy.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260829-dataset.benchmark-evaluation-resultsrepro-accurate-evaluation-of-quickest-changepoint-detectors-via-non-parametric-survival-analysis
Accurate Evaluation of Quickest Changepoint Detectors via Non-parametric Survival Analysis
This is a reproduction logbook for ICML 2026.
OpenReview ID: LhGxRnGmGJ
Paper Abstract
This logbook reproduces KM-ARL and KM-ADD estimators for changepoint detection.
See logbook.json for full claim verification details.
telugu-indicf5-evaluationrivers-evaluation-results
Rivers Evaluation Results - Comprehensive LLM Benchmarking
All results from the paper's five experimental conditions: baseline LLMs, fine-tuned models, RAG, and Graph-RAG with Licensing Oracle. This repository contains baseline evaluations for Claude Sonnet 4.5, Gemini 2.5 Flash Lite, and Gemma 3-4B, along with fine-tuning results for both factual recall and abstention behavior. It also includes outputs from the embedding-based RAG system and the Graph-RAG with Licensing Oracle… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-evaluation-results.rag-evaluation-lab-20260908-dataset
RAG Evaluation Lab Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for RAG systems often ship without a stable regression set or failure taxonomy.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier
variant: generation pattern
synthetic: always true… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/rag-evaluation-lab-20260908-dataset.certainty-robustness-llm-evaluation
Certainty Robustness Benchmark
This repository accompanies the paper:
Certainty robustness: Evaluating LLM stability under self-challenging promptsMohammadreza Saadat, Steve NemzerarXiv:2603.03330, 2026https://arxiv.org/abs/2603.03330
Overview
The Certainty Robustness Benchmark evaluates how large language models (LLMs) behave when their initial answers are challenged by follow-up prompts such as:
“Are you sure?”
“You are wrong!”
confidence elicitation prompts
Rather… See the full description on the dataset page: https://huggingface.co/datasets/Reza-Telus/certainty-robustness-llm-evaluation.Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset
📌 Dataset Contents
Each sample includes:
category: The evaluation domain
prompt: The question given to the LLM
temperature: Environmental temperature input
humidity: Environmental humidity input
context: A scenario label (e.g., cool_humid, hot_dry, average_day)
reference: Expert-crafted expected output
All data is provided in a single JSON file.
🧪 Intended Use
This dataset supports research on:
LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.NQ-RAG-DPO-Evaluation
Dataset Card
Dataset Summary
This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO).
The system is organized into three interconnected pipelines:
1️. RAG Pipeline
The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark.
For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.soas-english-uzbek-rag-evaluation
SOAS English-Uzbek Retrieval Pilot
Dataset Summary
This folder documents a bilingual English-Uzbek retrieval evaluation benchmark for culturally grounded RAG systems. The 400-row public pilot release is retrieval-only: it contains questions and source-document targets, but it intentionally excludes answer, context, excerpt, and source-text fields.
This is a pilot benchmark with documented quality flags, template-generated examples, and domain mismatches. The rows… See the full description on the dataset page: https://huggingface.co/datasets/Rajan2026/soas-english-uzbek-rag-evaluation.
