datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vi-en-ast-testseten-vi-ast-testsetkinyarwanda_cleaned_testset_verified_20HRScv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.corpus-siarad-test-setContinuo-Testset
Continuo-Testset
Continuo-Testset is a benchmark for long-form, multi-speaker zero-shot speech generation. Each case asks a system to synthesize a complete speaker-attributed script as one recording, given a reference prompt for every speaker. The output is scored against a human-verified target recording.
This dataset accompanies an anonymous ICLR 2027 submission.
Statistics
Chinese (zh)
English (en)
Total
Cases
59
59
118
Single-speaker long-form
18… See the full description on the dataset page: https://huggingface.co/datasets/AnonyData/Continuo-Testset.kinyarwanda_cleaned_testset_verified_200HRSMSA_test_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1.
eval_framework_testsetamharic_cleaned_testset_verifiedtest_setamharic_cleaned_testset_fleurs_currentkinyarwanda_cleaned_testset_verifiedpersian-solo-setar_testkinyarwanda_cleaned_testset_verified_100HRSlibrispeech_pc_testsetTestSet_2Debug-Test-SetNEW_TEST_SET_AMHARIC_FINALlibri_augmented_test_set
Dataset Card for "libri_augmented_test_set"
More Information needed
test_set_100_de_en_de
Dataset Card for "test_set_100_de_en_de"
More Information needed
kinyarwanda_cleaned_testset_verified_10HRSDecodis_Test_Settest_set_100_de_en
Dataset Card for "test_set_100_de_en"
More Information needed
cv10-uk-testset-clean-punctuatedThe same as https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean but with restored punctuations and capitalizations by https://huggingface.co/dchaplinsky/punctuation_uk_bert model.
Darija_Audio_Test_Settest_data_set_2test_set_100_cs_en
Dataset Card for "test_set_100_cs_en"
More Information needed
test_set_100_cs_en_cs
Dataset Card for "test_set_100_cs_en_cs"
More Information needed
test_set_100_en_de
Dataset Card for "test_set_100_en_de"
More Information needed
