datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
igbo_tts_normalizedFleurs_Irish_normalizedArabic-Diacritized-TTS-Normalized
Arabic-Diacritized-TTS Dataset
Overview
The Arabic-Diacritized-TTS dataset contains Arabic audio samples and their corresponding text with full diacritization. This dataset is designed to support research in Arabic speech processing, text-to-speech (TTS) synthesis, automatic diacritization, and other natural language processing (NLP) tasks.
Dataset Contents
Audio Samples: High-quality Arabic speech recordings.
Text Transcriptions: Fully diacritized Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/hana92/Arabic-Diacritized-TTS-Normalized.masc_filtered_normalizedcommon-voice-20-mn-normalized
Common Voice 20.0 Mongolian Dataset
This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release.
Dataset Structure
The dataset contains:
Audio clips in .mp3 format
Transcriptions for each audio clip
Train/test/dev splits
Additional metadata including speaker demographics
Usage
This dataset can be used for:
Speech Recognition
Voice Analysis
Linguistic Research
Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.ASVspoof_2021_DF_Balanced_Normalizednew-twi-tts-aligned_normalised
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
google-latam-spanish-boundary-normalized
Google LATAM Spanish Boundary-Normalized Audio
Female Spanish speech from the following upstream datasets:
Argentina: ylacombe/google-argentinian-spanish
Chile: ylacombe/google-chilean-spanish
Colombia: ylacombe/google-colombian-spanish
Attribution and Thanks
Many thanks to ylacombe for publishing
and maintaining the original Argentinian, Chilean, and Colombian Spanish
datasets. The recordings, transcripts, speaker labels, and original dataset
structure come… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-boundary-normalized.ASVspoof_2021_LA_Balanced_NormalizedMCV_Fleurs_Combined_Irish_normalizedexpresso-concatenated-half-normalatco2_normalized_augmentednormalized_train_ATC_datasetnormalized_test_ATC_datasetmy-voice-normalizedMCV25_Irish_normalizediqra_curated_normalised_1s_20s_finalvocalsound-normalizednormalized_khmer_dataset_14kASVspoof_2021_DF1_Balanced_NormalizedNormalorpheus-synthetic-dataset-normalizedMathSpeech_whisper_transcribed_normalizednormalized_khmer_datasetmalagasy-asr-normalized-v2hinglish-normal-005-fix-speaker-smokeKimi-7B-Audio-Instruct-Normalnormal_audiotrain82normal_audiotest46grandpa-interview-dataset-normalized
