CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01D4nt3 /esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset. def add_duration(sample): y, sr = sample['audio']["array"], sample['audio']["sampling_rate"] sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000 return sample tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True) # compute duration to filter tedlium = tedlium.map(add_duration) tedlium = tedlium.select(range(512)) # Whisper max supported duration tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.audion<1K0 likes7.9k downloads2y agoHugging Face02AdoCleanCode /korea_speech_mfa_aligned_validationaudio100K<n<1M0 likes455 downloads8mo agoHugging Face03Suchae /Korea-AIHub-middlesenior-dialect-speech-validation-part2audio10K<n<100K0 likes295 downloads2y agoHugging Face04etechgrid /ttm-validation-datasetaudio1K<n<10K0 likes157 downloads2y agoHugging Face05Suchae /Korea-AIHub-middlesenior-dialect-speech-validation-part1audio10K<n<100K0 likes125 downloads2y agoHugging Face06mosama /sada-validation-preprocessed Details This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo. In addtition, the following filters were applied to this data: All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.audioautomatic-speech-recognition1K<n<10K0 likes121 downloads1y agoHugging Face07ALLA1N /In-the-wild_validation-dataset DFE-Val: In-the-Wild Audio Deepfake Proxy Validation Set DFE-Val is a small curated collection of 102 audio clips (51 real, 51 fake) gathered from public social media platforms to approximate the distributional characteristics of the :contentReference[oaicite:0]{index=0} benchmark. This dataset was created as part of the :contentReference[oaicite:1]{index=1} research project and is released as an open-source contribution for the audio deepfake detection community. Why… See the full description on the dataset page: https://huggingface.co/datasets/ALLA1N/In-the-wild_validation-dataset.audioaudio-classificationn<1K0 likes88 downloads4mo agoHugging Face08Leon299 /validation_GTaudion<1K0 likes84 downloads5mo agoHugging Face09midoiv /common_voice_17_validation_cleanedaudio10K<n<100K0 likes78 downloads1y agoHugging Face10EYEDOL /swahili_MEDIUM_validationSwahilidata_22audio1K<n<10K0 likes76 downloads1y agoHugging Face11DTU54DL /librispeech-augmentated-validation-prepared Dataset Card for "librispeech-augmentated-validation-prepared" More Information needed audio1K<n<10K0 likes74 downloads4y agoHugging Face12cmu-mlsp /encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features" More Information needed audion<1K1 likes65 downloads3y agoHugging Face13mosama /sada-validation-wav2vec2-xls-r-300m-ar-preprocessedaudio1K<n<10K0 likes56 downloads1y agoHugging Face14EYEDOL /swahili_small_validationSwahilidata_66audio1K<n<10K0 likes44 downloads1y agoHugging Face15CristianaLazar /librispeech_augm_validation-tiny Dataset Card for "librispeech_augm_validation-tiny" More Information needed audio1K<n<10K0 likes40 downloads4y agoHugging Face16arda-argmax /common_voice_17_0_en_test_validation_pseudo_labelledaudio1K<n<10K0 likes40 downloads1y agoHugging Face17EYEDOL /swahili_small_validationSwahilidata_33audio1K<n<10K0 likes39 downloads1y agoHugging Face18AhmedBadawy11 /UAE_validation_WAVaudion<1K1 likes36 downloads2y agoHugging Face19EYEDOL /swahili_small_validationSwahilidata_55audio1K<n<10K0 likes33 downloads1y agoHugging Face20CristianaLazar /librispeech_validation Dataset Card for "librispeech_validation" More Information needed audio1K<n<10K0 likes29 downloads4y agoHugging Face21etechgrid /music-validation-datasetaudio1K<n<10K1 likes28 downloads2y agoHugging Face22sivakgp /hindi-books-audio-train-validationaudion<1K0 likes26 downloads2y agoHugging Face23EYEDOL /swahili_small_validationSwahilidata_22audio1K<n<10K0 likes26 downloads1y agoHugging Face24EYEDOL /swahili_small_validationSwahilidata_88audio1K<n<10K0 likes25 downloads1y agoHugging Face25lichenda /urgent26_track1_leaderboard_validationaudio1K<n<10K0 likes25 downloads1y agoHugging Face26EYEDOL /swahili_small_validationSwahilidata_11audio1K<n<10K0 likes22 downloads1y agoHugging Face27cmu-mlsp /encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features" More Information needed audio1K<n<10K0 likes21 downloads3y agoHugging Face28rishi70612 /validation_nepali_asr Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/validation_nepali_asr.audioautomatic-speech-recognitionn<1K0 likes18 downloads1y agoHugging Face29Jenny038 /nepali-validation_dataaudion<1K0 likes18 downloads1y agoHugging Face30AhmedBadawy11 /UAE_validation_mp3audion<1K0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.