datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
THCHS-30-tests
THCHS-30 Test Set
THCHS-30 test split for Mandarin Chinese speech recognition benchmarking.
Dataset Info
Language: Mandarin Chinese (zh-CN)
Samples: 2,495
Speakers: 10
Sample Rate: 16 kHz
License: Apache 2.0
Usage
from datasets import load_dataset
# After uploading to HuggingFace
dataset = load_dataset("your-username/thchs30-test")
# Example
print(dataset['train'][0])
# {
# 'audio': {'array': [...], 'sampling_rate': 16000, 'path': 'audio/D11_750.wav'},
#… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/THCHS-30-tests.MCABSA_testset2024.09.14_AGH_manual_Temperature_tests_20-55_Degrees
Temperature tests (20-55 Degrees)
Dataset author: Oğuzhan Berke Özdil
What was measured?
Methods
Type of needle: Quincke 22G 9cm
Needle Was attached with 3d printed needle attachment directly in front of the sound hole of the microphone
Type of microphone:
How the measurements were taken: manual
Phantoms
thickness of foams : 3 cm
Foam is the same as we used in Magdeburg foam phantom.
Foams were soaked in warm/hot water to create different… See the full description on the dataset page: https://huggingface.co/datasets/VibroNav/2024.09.14_AGH_manual_Temperature_tests_20-55_Degrees.MMEdit-TestSet
MMEdit Test Set
A paired audio editing test set for text-guided audio manipulation evaluation, released with MMEdit.
Overview
This dataset contains 3,317 aligned triplets:
Component
Description
raw/
Source audio before editing
target/
Target audio after editing
content.jsonl
Editing instruction (caption) keyed by audio_id
Each sample is linked by a shared audio_id. For example, sample add_017221 corresponds to:
raw/add_017221.wav — original… See the full description on the dataset page: https://huggingface.co/datasets/CocoBro/MMEdit-TestSet.vi-en-ast-testseten-vi-ast-testsetX-Voice-TestsetX-Voice Multilingual Test Set
High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages
This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model.
Dataset Summary
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.kinyarwanda_cleaned_testset_verified_20HRScv10-uk-testset-clean
The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦
Overview
This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios.
All audios have been checked by a human to be sure that they are correct.
This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk
Community
Discord: https://bit.ly/discord-uds
Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.kinyarwanda_cleaned_testset_verified_200HRSswahili_small_testSwahilidata_77swahili_small_testSwahilidata_88test_stats_erroreval_framework_testsetswahili_small_testSwahilidata_22swahili_small_testSwahilidata_66test_speechswahili_small_testSwahilidata_55test_setswahili_small_testSwahilidata_11swahili_small_testSwahilidata_33whisperkit_testsAll files are from: earnings22
Rencoded to 24kbps MP3 using:
ffmpeg -i 4446796.wav -vn -map_metadata -1 -ac 1 -c:a libmp3lame -b:a 24k -application voip -y 4446796.mp3
hf7192-ocr-testsuite-504791swahili_small_testSwahilidata_44TestSynthkinyarwanda_cleaned_testset_verifiedaudio_testsetamharic_cleaned_testset_verifiedkinyarwanda_cleaned_testset_verified_100HRSamharic_cleaned_testset_fleurs_current
