datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IfEvalCode-testsetpreference-test-sets
Preference Test Sets
Very few preference datasets have heldout test sets for validation of reward model accuracy results.
In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation.
Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples
Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation.
Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.gec-test-setTest data for Icelandic spell and grammar checking, created as part of the Icelandic Language Technology Programme.
The test data is divided into three different formats, type 1, 2 and 3. For every original file corrected, three files are included in the test data when possible: _original, _corrected and _metadata. The original and metadata files are always .txt files, but the format of the corrected file differs between types.
Texts corrected are from the News2 subcorpus of the Icelandic… See the full description on the dataset page: https://huggingface.co/datasets/mideind/gec-test-set.SpatialLM-Testset
SpatialLM Testset
Project page | Paper | Code
We provide a test set of 107 preprocessed point clouds and their corresponding GT layouts, point clouds are reconstructed from RGB videos using MASt3R-SLAM. SpatialLM-Testset is quite challenging compared to prior clean RGBD scan datasets due to the noises and occlusions in the point clouds reconstructed from monocular RGB videos.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialLM-Testset.Easy-Turn-Testset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Testset.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.vindr-cxr-testsetSolarWM-Data_test-set-v1
SolarWM Standalone Test Set v1
This repository contains the complete, self-contained SolarWM test set without
the training shards. It includes 1,300 clips from 13 source views (100 per
view), packaged as 60 uncompressed WebDataset tar files totaling approximately
77.4 GB.
The test identities are the current accepted SolarWM standalone evaluation
split, excluding MIND. The release contains 757 xhigh and 543 high samples. All selected
samples have non-empty captions and finite… See the full description on the dataset page: https://huggingface.co/datasets/junchaoh-cs/SolarWM-Data_test-set-v1.emu_edit_test_set
Dataset Card for the Emu Edit Test Set
Dataset Summary
To create a benchmark for image editing we first define seven different categories of potential image editing operations: background alteration (background), comprehensive image changes (global), style alteration (style), object removal (remove), object addition (add), localized modifications (local), and color/texture alterations (texture).
Then, we utilize the diverse set of input images from the MagicBrush… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set.video-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pyvideo-SALMONN_2_testset
video-SALMONN 2 Benchmark
Github Link
Paper Link
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.py
test_import_dataset_from_hub_using_settings_with_recordsFalse
Dataset Card for test_import_dataset_from_hub_using_settings_with_recordsFalse
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import… See the full description on the dataset page: https://huggingface.co/datasets/argilla-internal-testing/test_import_dataset_from_hub_using_settings_with_recordsFalse.StableSR-TestSets
StableSR TestSets Card
These test sets are used associated with the StableSR, available here.
Data Details
Developed by: Jianyi Wang
Data type: Synthetic and real-world test sets for image super-resolution
License: S-Lab License 1.0
Data Description: The test sets are used to reproduce the metric results shown in Paper.
Resources for more information: GitHub Repository.
Cite as:
@InProceedings{wang2023exploiting,
author = {Wang, Jianyi and Yue, Zongsheng and… See the full description on the dataset page: https://huggingface.co/datasets/Iceclear/StableSR-TestSets.SpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialGen-Testset.indexed-open-image-v4-test-set
Dataset Card for "indexed-open-image-v4-test-set"
More Information needed
coco_test_set_pybboxes
COCO Test Set
This is a coco test set which is used for unit testing in pybboxes library.
inference-test-settestset_popqatestset_piqaMCABSA_testsetMindGuard-testset
MindGuard-testset: Expert-Annotated Evaluation Data for Mental Health AI Safety
MindGuard-testset is a clinically grounded benchmark dataset for evaluating safety classifiers in mental health AI systems. This dataset was developed by Sword Health in collaboration with licensed clinical psychologists to address the critical need for contextually appropriate safety measures in therapeutic AI applications.
Overview
MindGuard-testset contains 1,134 annotated user turns… See the full description on the dataset page: https://huggingface.co/datasets/swordhealth/MindGuard-testset.unitary_compilation_testset_3to5qubit
Testset: Compile discrete-continuous quantum circuits 3 to 5 qubits
Paper: "Synthesis of discrete-continuous quantum circuits with multimodal diffusion models".
Key Features and limitations
Unitary compilation from 3 to 5 qubits
Quantum circuits up to 32 gates
Dataset details in the [paper-arxiv]
Usage
The dataset can be loaded with genQC. First install or upgrade genQC using
pip install -U genQC
A guide on how to use this dataset can be found in the… See the full description on the dataset page: https://huggingface.co/datasets/Floki00/unitary_compilation_testset_3to5qubit.WorldRenderer-Testsettestset_mmluFakeVV_testset_videotestset_hellaswagtestset_winogrande-infillseedtts_testset
SeedTTS Evaluation Dataset
This dataset contains evaluation data for SeedTTS text-to-speech model testing in multiple languages.
Original repo from: https://github.com/BytedanceSpeech/seed-tts-eval
Languages
English (en): Contains test_wer and test_sim splits
Chinese (zh): Contains test_wer, test_sim, and test_wer_hardcase splits
Usage
# makesure: pip install datasets==3.5.1
import os
from datasets import load_dataset
repo_dir = "hhqx/seedtts_testset"… See the full description on the dataset page: https://huggingface.co/datasets/hhqx/seedtts_testset.emu_edit_test_set_generations
Dataset Card for the Emu Edit Generations on Emu Edit Test Set
Dataset Summary
This dataset contains Emu Edit's generations on the Emu Edit test set. For more information please read our paper or visit our homepage.
Licensing Information
Licensed with CC-BY-NC 4.0 License available here.
Citation Information
@inproceedings{Sheynin2023EmuEP,
title={Emu Edit: Precise Image Editing via Recognition and Generation Tasks},
author={Shelly Sheynin and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set_generations.TEST_DCAgent_dev_set_71_tasks_Qwen_Qwen3-8B_thinking_false_20260215_073417
