CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MiniMaxAI /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K44 likes1.1k downloads1y agoHugging Face02Eureka-Leo /MCABSA_testsetaudion<1K2 likes324 downloads1y agoHugging Face03CocoBro /MMEdit-TestSet MMEdit Test Set A paired audio editing test set for text-guided audio manipulation evaluation, released with MMEdit. Overview This dataset contains 3,317 aligned triplets: Component Description raw/ Source audio before editing target/ Target audio after editing content.jsonl Editing instruction (caption) keyed by audio_id Each sample is linked by a shared audio_id. For example, sample add_017221 corresponds to: raw/add_017221.wav — original… See the full description on the dataset page: https://huggingface.co/datasets/CocoBro/MMEdit-TestSet.audioaudio-to-audio1K<n<10K0 likes196 downloads4mo agoHugging Face04masumtechnonext /test-data-set-Arabic-letteraudio10K<n<100K0 likes105 downloads2mo agoHugging Face05hoangducanh1865 /vi-en-ast-testsetaudio1K<n<10K0 likes105 downloads19d agoHugging Face06DimQ1 /sortformer-diarization-test-set Sortformer Diarization Test Set 100 real speech samples extracted from LibriSpeech test-clean for speaker diarization testing and benchmarking with NVIDIA Sortformer 4spk-v2 ONNX models. Usage with Sortformer ONNX from huggingface_hub import snapshot_download import soundfile as sf # Download the test set dataset_path = snapshot_download("DimQ1/sortformer-diarization-test-set") # Load audio audio, sr = sf.read(f"{dataset_path}/audio/ls_real_000.wav")… See the full description on the dataset page: https://huggingface.co/datasets/DimQ1/sortformer-diarization-test-set.audioaudio-classificationn<1K0 likes99 downloads2mo agoHugging Face07YoonSeon /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/YoonSeon/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K0 likes84 downloads7mo agoHugging Face08hoangducanh1865 /en-vi-ast-testsetaudio1K<n<10K0 likes69 downloads15d agoHugging Face09XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes65 downloads5mo agoHugging Face10KYAGABA /kinyarwanda_cleaned_testset_verified_20HRSaudio10K<n<100K0 likes49 downloads2y agoHugging Face11otozz /MSA_test_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1. audio10K<n<100K0 likes49 downloads2y agoHugging Face12jeshica /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/jeshica/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K0 likes45 downloads6mo agoHugging Face13danielrosehill /Audio-Understanding-Test-Set Audio Understanding Test Set A structured dataset for evaluating audio understanding capabilities of multimodal AI models. Contains 137 test prompts across 22 categories, paired with a 20-minute voice sample and 49 completed model outputs from Gemini 3.1 Flash Lite. Overview Property Value Total prompts 137 Implemented (with prompt text) 49 Suggested (description only) 88 Completed outputs 49 Categories 22 Model under test… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Audio-Understanding-Test-Set.audioaudio-classificationn<1K0 likes45 downloads6mo agoHugging Face14Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes43 downloads2y agoHugging Face15KYAGABA /kinyarwanda_cleaned_testset_verified_200HRSaudio100K<n<1M0 likes40 downloads2y agoHugging Face16wanasash /corpus-siarad-test-setaudio1K<n<10K0 likes31 downloads2y agoHugging Face17prvInSpace /eval_framework_testsetaudio1K<n<10K0 likes28 downloads1y agoHugging Face18Sammau /test_setaudio1K<n<10K0 likes25 downloads1y agoHugging Face19KYAGABA /kinyarwanda_cleaned_testset_verifiedaudio100K<n<1M0 likes20 downloads2y agoHugging Face20jykim310 /audio_testsetaudion<1K0 likes19 downloads2y agoHugging Face21Razavipour /persian-solo-setar_testaudion<1K0 likes18 downloads1y agoHugging Face22KYAGABA /amharic_cleaned_testset_verifiedaudio10K<n<100K1 likes17 downloads2y agoHugging Face23KYAGABA /kinyarwanda_cleaned_testset_verified_100HRSaudio10K<n<100K0 likes17 downloads2y agoHugging Face24KYAGABA /amharic_cleaned_testset_fleurs_currentaudio1K<n<10K0 likes17 downloads2y agoHugging Face25tasyaworkspace /TestSet_2audion<1K0 likes14 downloads6mo agoHugging Face26yuekai /librispeech_pc_testsetaudio1K<n<10K0 likes12 downloads10mo agoHugging Face27DTU54DL /libri_augmented_test_set Dataset Card for "libri_augmented_test_set" More Information needed audio1K<n<10K0 likes11 downloads4y agoHugging Face28o0dimplz0o /Debug-Test-Setaudio1K<n<10K0 likes11 downloads2y agoHugging Face29Yehor /cv10-uk-testset-clean-punctuatedThe same as https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean but with restored punctuations and capitalizations by https://huggingface.co/dchaplinsky/punctuation_uk_bert model. audio1K<n<10K1 likes10 downloads2y agoHugging Face30LGB666 /SageLM_testset_audioaudion<1K0 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.