datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.MCABSA_testsetMMEdit-TestSet
MMEdit Test Set
A paired audio editing test set for text-guided audio manipulation evaluation, released with MMEdit.
Overview
This dataset contains 3,317 aligned triplets:
Component
Description
raw/
Source audio before editing
target/
Target audio after editing
content.jsonl
Editing instruction (caption) keyed by audio_id
Each sample is linked by a shared audio_id. For example, sample add_017221 corresponds to:
raw/add_017221.wav — original… See the full description on the dataset page: https://huggingface.co/datasets/CocoBro/MMEdit-TestSet.test-data-set-Arabic-lettervi-en-ast-testsetsortformer-diarization-test-set
Sortformer Diarization Test Set
100 real speech samples extracted from LibriSpeech test-clean for speaker diarization testing and benchmarking with NVIDIA Sortformer 4spk-v2 ONNX models.
Usage with Sortformer ONNX
from huggingface_hub import snapshot_download
import soundfile as sf
# Download the test set
dataset_path = snapshot_download("DimQ1/sortformer-diarization-test-set")
# Load audio
audio, sr = sf.read(f"{dataset_path}/audio/ls_real_000.wav")… See the full description on the dataset page: https://huggingface.co/datasets/DimQ1/sortformer-diarization-test-set.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/YoonSeon/TTS-Multilingual-Test-Set.en-vi-ast-testsetX-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.kinyarwanda_cleaned_testset_verified_20HRSMSA_test_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1.
TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/jeshica/TTS-Multilingual-Test-Set.Audio-Understanding-Test-Set
Audio Understanding Test Set
A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite.
Overview
Property
Value
Total prompts
137
Implemented (with prompt text)
49
Suggested (description only)
88
Completed outputs
49
Categories
22
Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.cv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.kinyarwanda_cleaned_testset_verified_200HRScorpus-siarad-test-seteval_framework_testsettest_setkinyarwanda_cleaned_testset_verifiedaudio_testsetpersian-solo-setar_testamharic_cleaned_testset_verifiedkinyarwanda_cleaned_testset_verified_100HRSamharic_cleaned_testset_fleurs_currentTestSet_2librispeech_pc_testsetlibri_augmented_test_set
Dataset Card for "libri_augmented_test_set"
More Information needed
Debug-Test-Setcv10-uk-testset-clean-punctuatedThe same as https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean but with restored punctuations and capitalizations by https://huggingface.co/dchaplinsky/punctuation_uk_bert model.
SageLM_testset_audio
