datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.korea_speech_mfa_aligned_validationKorea-AIHub-middlesenior-dialect-speech-validation-part2ttm-validation-datasetKorea-AIHub-middlesenior-dialect-speech-validation-part1sada-validation-preprocessed
Details
This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo.
In addtition, the following filters were applied to this data:
All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.In-the-wild_validation-dataset
DFE-Val: In-the-Wild Audio Deepfake Proxy Validation Set
DFE-Val is a small curated collection of 102 audio clips (51 real, 51 fake) gathered from public social media platforms to approximate the distributional characteristics of the :contentReference[oaicite:0]{index=0} benchmark.
This dataset was created as part of the :contentReference[oaicite:1]{index=1} research project and is released as an open-source contribution for the audio deepfake detection community.
Why… See the full description on the dataset page: https://huggingface.co/datasets/ALLA1N/In-the-wild_validation-dataset.validation_GTcommon_voice_17_validation_cleanedswahili_MEDIUM_validationSwahilidata_22librispeech-augmentated-validation-prepared
Dataset Card for "librispeech-augmentated-validation-prepared"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features"
More Information needed
sada-validation-wav2vec2-xls-r-300m-ar-preprocessedswahili_small_validationSwahilidata_66librispeech_augm_validation-tiny
Dataset Card for "librispeech_augm_validation-tiny"
More Information needed
common_voice_17_0_en_test_validation_pseudo_labelledswahili_small_validationSwahilidata_33UAE_validation_WAVswahili_small_validationSwahilidata_55librispeech_validation
Dataset Card for "librispeech_validation"
More Information needed
music-validation-datasethindi-books-audio-train-validationswahili_small_validationSwahilidata_22swahili_small_validationSwahilidata_88urgent26_track1_leaderboard_validationswahili_small_validationSwahilidata_11encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features"
More Information needed
validation_nepali_asr
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/rishi70612/validation_nepali_asr.nepali-validation_dataUAE_validation_mp3
