datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers-pr
Transformers PR Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from huggingface/transformers.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
reviews.parquet
pr_files.parquet
pr_diffs.parquet
review_comments.parquet
links.parquet
events.parquet
new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/transformers-pr.diffusers-pr
Diffusers PR Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from huggingface/diffusers.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
reviews.parquet
pr_files.parquet
pr_diffs.parquet
review_comments.parquet
links.parquet
events.parquet
new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/diffusers-pr.evals-speech-recognition-cy-en
Welsh ASR Model Evaluation Transcription Dataset
This resource compiles the output transcriptions from multiple Welsh Automatic Speech Recognition (ASR) models across several test sets.
The data is structured hierarchically:
Splits delineate the individual test sets.
Configs within each split detail the performance (transcriptions) of a specific ASR model and its version on that set.
Metrics Results
model
test
task
wer
cer… See the full description on the dataset page: https://huggingface.co/datasets/techiaith/evals-speech-recognition-cy-en.fishnet-evals
Lichess database positions with Stockfish evaluations
Positions with evaluations extracted from https://database.lichess.org/.
Evaluations were computed using Lichess's distributed Stockfish analysis fishnet, using various Stockfish versions.
column
type
description
fen
string
FEN of the evaluated position. En passant square included only if a fully legal en passant move exists.
cp / mate
int or null
Signed evaluation of the position in centipawns or moves to mate, as… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/fishnet-evals.SpreadSheetBenchmcp-clients
MCP Clients Dataset
MCP client identity and capability observations from huggingface.co/mcp.
The pipeline incrementally processes completed daily source partitions. Each
release records its source revision and watermark in state/mcp-clients-v1.json
and a sanitized 60-day dashboard snapshot in reports/dashboard-v1.json.
Its window is the requested UTC date range. Omitted client traffic dates are
unavailable data, not zero traffic.
Protocol traffic before 2026-07-27 is a one-time… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/mcp-clients.pythia-memorized-evals
Pythia Memorized Evals
This dataset contains the results of memorization evaluations for all Pythia models. For each model, the dataset lists every training sequence that the fully trained model has memorized.
A training sequence is considered memorized if, when prompted with the first 32 tokens of the sequence, the model's greedy continuation exactly matches the next 32 tokens. This is evaluated over all ~146M training sequences in the Pile.
This dataset was generated for the paper… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pythia-memorized-evals.openclaw-pr
Openclaw PR Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from openclaw/openclaw.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
reviews.parquet
pr_files.parquet
pr_diffs.parquet
review_comments.parquet
links.parquet
events.parquet
new_contributors.parquet
new-contributors-report.json… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/openclaw-pr.glm-simple-evals-dataset
glm-simple-evals-dataset
This repository is dedicated to storing various evaluation data required for the glm-simple-evals evaluation project, to enable industry researchers and developers to reproduce the performance of the GLM-4.5 series models on reported benchmarks.
Currently, this repository covers the data required for the following evaluation tasks:
AIME
GPQA
HLE
LiveCodeBench
MATH 500
SciCode
MMLU Pro
Usage Instructions
To use these evaluation datasets… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/glm-simple-evals-dataset.eval_smolvla-so101-4tasks-aug-v2_stack_30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 25,
"total_frames": 76615,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hjkso1406/eval_smolvla-so101-4tasks-aug-v2_stack_30.OSWorld-Verifiedeval_splatsim_approach_lever_benchmark_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lerobot_splatsim",
"total_episodes": 1000,
"total_frames": 187307,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JennyWWW/eval_splatsim_approach_lever_benchmark_1000.supabase-evals-traces
supabase-evals-traces
Traces generated on DGX Sparks while evaluating various models from supabase/evals — a
leaderboard of agent tasks for building, deploying, investigating, and fixing
Supabase apps/databases. Scoring uses the task-specific EVAL.ts checks as upstream
pnpm eval (named PASS/FAIL rubrics; binary reward only when every check
passes) where certain tasks are also judged via judge-llm (DeepSeek V4 Flash 0731).
This dataset was collected with OpenCode as the agent… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/supabase-evals-traces.Llama-3.1-8B-Instruct-evals
Dataset Card for Llama-3.1-8B-Instruct Evaluation Result Details
This dataset contains the Meta evaluation result details for Llama-3.1-8B-Instruct. The dataset has been created from 30 evaluation tasks. These tasks are human_eval, gorilla_api_bench__huggingface, mmlu_pro, infinite_bench, api_bank, human_eval_plus, ifeval__loose, mmlu__0_shot__cot, nih__multi_needle, multilingual_mmlu_de, mmlu, gsm8k, mgsm, multilingual_mmlu_fr, multilingual_mmlu_pt, math_hard… See the full description on the dataset page: https://huggingface.co/datasets/meta-llama/Llama-3.1-8B-Instruct-evals.eval_smolvla_policy_omx_gelsight_env1_multitask_horizontal_vertical_line_nogel_20260820_192704This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 10,
"total_frames": 7847,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/WoojongKim/eval_smolvla_policy_omx_gelsight_env1_multitask_horizontal_vertical_line_nogel_20260820_192704.malt-public
MALT: Manually-Reviewed Agentic Labeled Transcripts
MALT-public is our collection of agent transcripts. Our public variant only includes data on non-internal tasks, which includes 30 task families and 169 tasks, across ~19 different models (some might be different releases of the same model, from different providers, or internal naming changes).
Here's a summary table:
has_chain_of_thought
labels
model
manually_reviewed
run_source
count
False
bypass_constraints… See the full description on the dataset page: https://huggingface.co/datasets/metr-evals/malt-public.SheetBench-50
⚠️ Migrated to HUD v6
SheetBench-50 now runs on the HUD v6 SDK. The canonical version lives in
hud-evals/hud-remote-browser
(tasks.py, slugs sheetbench-*) and as a taskset on the HUD platform.
The rows in this dataset are the legacy v5 format (mcp_config pointing at
mcp.hud.ai/v3) and are kept for reproducibility of published results. They
will not work with hud >= 0.6.
Task Categories
1. Data Preparation and Hygiene (29 tasks)
De-duplication, type… See the full description on the dataset page: https://huggingface.co/datasets/hud-evals/SheetBench-50.math-evals
math-evals
Uniform {question, answer} math evaluation splits for a single source of
truth across benchmarks. Every split exposes exactly two columns:
question and answer.
split
source
source split
rows
clean_gsm8k_aug
cs-giung/clean-gsm8k-aug @60f9c039
test
1319
clean_gsm8k_aug_val
cs-giung/clean-gsm8k-aug @60f9c039
validation
500
gsm_hard
reasoning-machines/gsm-hard @960448f7
train
1319
gsm1k
ScaleAI/gsm1k @bc09569d
test
1205
gsm8k
openai/gsm8k @740312ad
test… See the full description on the dataset page: https://huggingface.co/datasets/cs-giung/math-evals.eval_showcase_vanilla_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 12564,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_showcase_vanilla_dataset.eval_svla_15b_reflectiveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 21764,
"total_tasks": 5,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/shuohsuan/eval_svla_15b_reflective.OSWorld-Goldeval_so101_car_pick_and_place-96_episodes_v0-real_v0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 157,
"total_frames": 119719,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:157"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jonathm126/eval_so101_car_pick_and_place-96_episodes_v0-real_v0.real01b-marker-d2-r0-blind-teleop-uniform100-sobol100-evalsobol50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-marker-d2-r0-blind-teleop-uniform100-sobol100-evalsobol50.eval_showcase_vanilla_dataset01This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 13786,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_showcase_vanilla_dataset01.lingsoft-pemt-evals
Lingsoft PEMT translation preference evals
Evaluation sets for translation quality and fluency, built from the
Lingsoft-EU-Summaries-PEMT
corpus: professional post-edited machine translations of "Summaries of EU
Legislation" (2016–2025). Each item pairs a raw machine translation
(Raw_MT) with its professional human post-edit (Target_PEMT) for the same
English source sentence. Only pairs where post-editing changed the text are
included. Intended for use with the
LumiOpen… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/lingsoft-pemt-evals.eval_so101-5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 75075,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/miketangmk/eval_so101-5.eval_showcase_vanillaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 16,
"total_frames": 14062,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tersooawai/eval_showcase_vanilla.eval_smovla_tactileThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_tactile_follower",
"total_episodes": 130,
"total_frames": 27426,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:130"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Tna001/eval_smovla_tactile.SpreadSheetBench-200eval_SmolVLA_stackredcube_2026_06_19This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 80,
"total_frames": 66071,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:80"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Alkatt/eval_SmolVLA_stackredcube_2026_06_19.
