datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
slovenian-llm-eval
Slovenian LLM Evaluation Dataset
This dataset is designed for evaluating Slovenian language models and builds upon the work of gordicaleksa/slovenian-llm-eval-v0 which translated some of the popular English benchmarks into Slovenian by using Google Translate. We have further improved the quality of the Slovenian translations.
The dataset contains the following benchmarks:
ARC Challenge
ARC Easy
BoolQ
GSM8K
HellaSwag
NQ Open
OpenBookQA
PIQA
TriviaQA
TruthfulQA
Winogrande… See the full description on the dataset page: https://huggingface.co/datasets/cjvt/slovenian-llm-eval.llm-fol-reasoning-eval
LLM FOL Reasoning Eval
This dataset is derived from ProverQA, a First-Order Logic reasoning benchmark designed to test the ability of large language models (LLMs) to perform structured logical reasoning.It restructures and normalizes the ProverQA development and training data into a unified, clean format suitable for evaluating chain-of-thought (CoT) and symbolic reasoning capabilities in LLMs.
Source
Original dataset: ProverQA: A First-Order Logic Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/llm-fol-reasoning-eval.spa-bench-eval-rollouts-groot-n1-7-frozen-llm-full
Spa-Bench Evaluation Rollouts — GR00T-N1.7 Frozen LLM (Full)
Open this dataset in the LeRobot visualizer
This repository contains physical Spa-Bench rollout trajectories with synchronized middle and wrist RGB video, robot state, action, timestamps, episode indices, and task indices. The native Hugging Face Data Studio viewer is enabled through the Parquet files declared above.
Dataset summary
Coverage: 900 scored recordings.
Format: LeRobot v3 at 30 FPS.
Task… See the full description on the dataset page: https://huggingface.co/datasets/Spa-Bench/spa-bench-eval-rollouts-groot-n1-7-frozen-llm-full.llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.ptb-llmevalmedpde-llm-eval-code-perturbation-dataset
pde-llm-eval-code-perturbation-dataset
Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.llm-eval-requestssummeval-annotated-latest
SummEval-LLMEval Dataset
Overview
The original SummEval dataset (Fabbri et al., 2021) consists of 1,600 summaries annotated by human expert evaluators using a 5-point Likert scale across 4 criteria: coherence, consistency, fluency, and relevance. These 1,600 summaries are based on 100 source articles from the CNN/DailyMail dataset (Hermann et al., 2015). For each source article, SummEval collects 16 summaries generated by 16 different automatic summarization systems. Each… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/summeval-annotated-latest.Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset
📌 Dataset Contents
Each sample includes:
category: The evaluation domain
prompt: The question given to the LLM
temperature: Environmental temperature input
humidity: Environmental humidity input
context: A scenario label (e.g., cool_humid, hot_dry, average_day)
reference: Expert-crafted expected output
All data is provided in a single JSON file.
🧪 Intended Use
This dataset supports research on:
LLM evaluation methods (semantic similarity, contextual… See the full description on the dataset page: https://huggingface.co/datasets/wayne-redemption/Sensor_Driven_Environmental_Monitoring_LLM_Evaluation_Dataset.repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
llm-commit-message-evaluation
Dataset Card for LLM Commit Message Evaluation
The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest).
For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.LLM-Eval-Filteredonepane-llm-evaluation-geminimtbench-annotated-latest
MT-Bench-Select Dataset
Introduction
The MT-Bench-Select dataset is a refined subset of the original MT-Bench dataset introduced by Zheng et al. (2023). The original MT-Bench dataset comprises 80 questions with answers generated by six models. Each question and each pair of models form an evaluation task, resulting in 1,200 tasks.
For this dataset, we used a curated subset of the original MT-Bench dataset, as prepared by the authors of the LLMBar paper (Zeng et al.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/mtbench-annotated-latest.llm-metric-mm-eval-pairwisestep2-evaluated-dataset-test2
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-test2
Total Samples: 92
Successfully Evaluated (Rubric): 92
Failed Evaluations (Rubric): 0
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence: 3.51… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-test2.LLM_Eval_Small_Examplellmbar-annotated-latest
LLMBar-Select Dataset
Introduction
The LLMBar-Select dataset is a curated subset of the original LLMBar dataset introduced by Zeng et al. (2024). The LLMBar dataset consists of 419 instances, each containing an instruction paired with two outputs: one that faithfully follows the instruction and another that deviates while presenting superficially appealing qualities. It is designed to evaluate LLM-based evaluators more rigorously and objectively than previous benchmarks.… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmbar-annotated-latest.dv-llm-eval-resultsinumulaisk__eval_model-details
Dataset Card for Evaluation run of inumulaisk/eval_model
Dataset automatically created during the evaluation run of model inumulaisk/eval_model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/inumulaisk__eval_model-details.step2-evaluated-dataset-Qwen3-14B-cp32
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32
Total Samples: 60
Successfully Evaluated (Rubric): 53
Failed Evaluations (Rubric): 7
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp32.jp-llm-evaluator-training
Japanese LLM Evaluator Training Dataset
It realeased on NLP2025 Constructing Open-source Large Language Model Evaluator for Japanese
Overview
Japanese LLM Evaluator Training Dataset is a dataset using for training Japanese LLM evaluator, which is focus on evaluate Japanese LLM from mutiple perspectives and meeting diverse evaluation requirements.
Content
The dataset includes 1000 diveser score rubrics. For every score rubrics, we generate 20 different… See the full description on the dataset page: https://huggingface.co/datasets/ku-nlp/jp-llm-evaluator-training.step2-evaluated-dataset-Qwen3-14B-cp40
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40
Total Samples: 58
Successfully Evaluated (Rubric): 53
Failed Evaluations (Rubric): 5
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B-cp40.memefact-llm-evaluations
MemeFact LLM Evaluations Dataset
This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments.
Dataset Description
Overview
The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.energy_D_eval_llm_as_judge_granular_v6step2-evaluated-dataset-Qwen3-14B
Complete Evaluation Dataset (Rubric + LogP)
This dataset contains chain-of-thought explanations evaluated using both comprehensive rubric assessment and LogP evaluation.
Overview
Source Dataset: llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B
Total Samples: 156
Successfully Evaluated (Rubric): 135
Failed Evaluations (Rubric): 21
Evaluation Model: Qwen/Qwen3-32B
Rubric Evaluation Results
Average Rubric Scores (0-4 scale)
logical_coherence:… See the full description on the dataset page: https://huggingface.co/datasets/llm-compe-2025-kato/step2-evaluated-dataset-Qwen3-14B.dataset_hanna_96_prompts_llm_evalllmeval2-annotated-latestllmeval2-annotated-latest
LLMEval²-Select Dataset
Introduction
The LLMEval²-Select dataset is a curated subset of the original LLMEval² dataset introduced by Zhang et al. (2023). The original LLMEval² dataset comprises 2,553 question-answering instances, each annotated with human preferences. Each instance consists of a question paired with two answers.
To construct LLMEval²-Select, Zeng et al. (2024) followed these steps:
Labelled each instance with the human-preferred answer.
Removed all… See the full description on the dataset page: https://huggingface.co/datasets/bay-calibration-llm-evaluators/llmeval2-annotated-latest.hugging-face-LLM-evaluation-results
Dataset Summary
Data comes from hugging face evaluation results using the https://github.com/EleutherAI/lm-evaluation-harness . See https://huggingface.co/datasets/open-llm-leaderboard/results for full results.
