datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}mbppplusVideo-MMEmmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.EEE_datastore
Every Eval Ever Datastore
A community database of AI evaluation results, all in one schema. Scores scraped from
leaderboards, pulled out of papers, and produced by local evaluation runs are stored in a
single record format, so results from different sources can be compared, joined, and reused
instead of re-scraped. This dataset is the data itself: one JSON record per
model per evaluation run — which may carry several scored results — with optional
per-sample companion files.… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/EEE_datastore.humanevalplusreinforce-ada-raw-eval
Reinforce-Ada Raw Eval
Raw evaluation artifacts organized by experiment / dataset / step.
Included files when present:
merged_data.jsonl
pass_at_k.json
record.txt
Experiments: grpo_n8, grpo_n16, grpo_n32, reinforce_ada_n8, reinforce_ada_n8_normstdtrue
Datasets: math500, minerva_math, olympiadbench, aime_hmmt_brumo_cmimc_amc23
voicehub-arena-seed-tts-eval
VoiceHub Arena — native TTS evaluations
Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA
speaker SIM and UTMOS22 measurements. The full campaign is still running.
Each generation method is evaluated separately using its publisher's native API.
Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target
diagnostic pilots are stored separately and must not be treated as full scores.
Interactive demo ·
Source code
Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.cmevs-erp-eval
CM-EVS: A Coverage-Curated Panoramic RGB-D Dataset for Indoor Scene Understanding
CM-EVS is a curated panoramic RGB-D dataset built under a single principle: maximize the geometric coverage of a 3D scene with the fewest equirectangular (ERP) frames possible. The release is structured as one redistributable Blender indoor data archive plus four license-aware adapter packages that regenerate matched frames locally from upstream sources whose terms forbid redistribution.
v1.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-cmevs-2026/cmevs-erp-eval.observation-masking-eval-logs
Eval Logs
Paper | Code
This repository contains model evaluation logs for four deep-research / web-agent benchmarks. Each run directory contains evaluated.jsonl judge results and node_0_shard_*.jsonl trajectory logs. Plot files and local bookkeeping files are intentionally excluded.
CM denotes the observation mask context management setting used in the paired run.
Data Access
You can download all released evaluation data, including tasks and… See the full description on the dataset page: https://huggingface.co/datasets/i-DeepSearch/observation-masking-eval-logs.tweet_eval
Dataset Card for tweet_eval
Dataset Summary
TweetEval consists of seven heterogenous tasks in Twitter, all framed as multi-class tweet classification. The tasks include - irony, hate, offensive, stance, emoji, emotion, and sentiment. All tasks have been unified into the same benchmark, with each dataset presented in the same format and with fixed training, validation and test splits.
Supported Tasks and Leaderboards
text_classification: The dataset can be… See the full description on the dataset page: https://huggingface.co/datasets/cardiffnlp/tweet_eval.RAG_Evalcot-eval-traces-2.0auto_evalalpaca_evalData for alpaca_eval, which aims to help automatic evaluation of instruction-following modelstransformers-pr
Transformers PR Dataset
Normalized snapshots of issues, pull requests, comments, reviews, and linkage data from huggingface/transformers.
Files:
issues.parquet
pull_requests.parquet
comments.parquet
issue_comments.parquet (derived view of issue discussion comments)
pr_comments.parquet (derived view of pull request discussion comments)
reviews.parquet
pr_files.parquet
pr_diffs.parquet
review_comments.parquet
links.parquet
events.parquet
new_contributors.parquet… See the full description on the dataset page: https://huggingface.co/datasets/evalstate/transformers-pr.WSC-Evaleval-resultsKling-Audio-Eval
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
🌐 Website | 📖 arXiv
📋 Dataset Structure
The dataset structure is as follows:
Kling-Audio-Eval
├── Folder (first-level label)
│ ├── Folder (second-level label)
│ │ ├── video
│ │ │ └── *.mp4
│ │ ├── audio
│ │ │ └── *.wav
│ │ └── caption.csv # Header: video, audio, audio_tag, video_caption… See the full description on the dataset page: https://huggingface.co/datasets/klingfoley/Kling-Audio-Eval.NexusRaven_API_evaluation
NexusRaven API Evaluation dataset
Please see blog post or NexusRaven Github repo for more information.
License
The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.GDP-Val-Evaluation-Submission
GDPval Submission Dataset
This dataset contains model outputs for GDP-Val evaluation.
Dataset Structure
data/: Contains the main dataset in Parquet format
train-00000-of-00001.parquet: Submission data with model outputs
deliverable_files/: Contains generated files for tasks that produce file deliverables
Organized by task_id
dataset_info.json: Metadata about the dataset
Columns
task_id: Unique identifier for each task
sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.function-calling-eval-dataset-v0The hf dataset contains 2 evaluation datasets
single_turn - The converstaion length for this evaluation dataset is 2. It consists of a user ask followed by a function call by assistant.
multi_turn - The conversation length is variable here but contains a combination of user messages, assistant function calls, assistant messages & tool responses.
Information about the columns
tools - List of functions/tools with specs in JSON format. This is the list of functions the model has to choose from… See the full description on the dataset page: https://huggingface.co/datasets/fireworks-ai/function-calling-eval-dataset-v0.Medical-Eval-HumanityLastExamda-code-evaluation-resultsgrobid-evaluation
GROBID End-to-End Evaluation Dataset
Reference corpora used for GROBID end-to-end
benchmarking of scientific-article structuring.
Documentation: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/
Latest benchmarking scores: https://grobid.readthedocs.io/en/latest/Benchmarking/
Official archive (Zenodo): https://zenodo.org/record/7708580
Dataset summary
These are the datasets used for GROBID end-to-end benchmarking, covering:
metadata extraction… See the full description on the dataset page: https://huggingface.co/datasets/sciencialab/grobid-evaluation.cascade-eval-pool
cascade eval pool — lagged public reveal (exact bytes)
Retired snapshots of the held-out evaluation pool used by the
cascade subnet. Each folder is a
byte-identical mirror of the pool/snapshots/block-<N>.tar that validators
scored — downloaded from the private pool bucket, sha256-verified against the
publisher index, and republished unmodified. A snapshot is revealed only after
a newer snapshot has superseded it, so no revealed pool can be selected by a
current or future round.… See the full description on the dataset page: https://huggingface.co/datasets/Tensor-Link/cascade-eval-pool.data-agent-harbor-eval
🧪 Data Agent — Harbor (eval)
A small, difficulty-balanced validation split — 144 tasks — perfect for quick checkpoints
while you train. Same idea as the rest of the family: your agent gets a real dataset and a
question, explores and answers, and everything is graded deterministically, no LLM judge.
Packaged in Harbor format.
Where it comes from
Built from the jupyter-agent dataset
(real notebooks over Kaggle datasets). Every task was verified — a strong agent… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-harbor-eval.seed-tts-eval-arrowcard_backend
Eval Cards Backend Dataset
Pre-computed evaluation data powering the Eval Cards frontend.
Generated by the eval-cards backend pipeline.
Last generated: 2026-05-05T11:30:42.961096Z
Quick Stats
Stat
Value
Models
5,678
Evaluations (benchmarks)
798
Metric-level evaluations
1321
Source configs processed
52
Benchmark metadata cards
240
File Structure
.
├── README.md # This file
├── manifest.json… See the full description on the dataset page: https://huggingface.co/datasets/evaleval/card_backend.global-piqa-evals
