datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
igbo_tts_normalizedFleurs_Irish_normalizedArabic-Diacritized-TTS-Normalized
Arabic-Diacritized-TTS Dataset
Overview
The Arabic-Diacritized-TTS dataset contains Arabic audio samples and their corresponding text with full diacritization. This dataset is designed to support research in Arabic speech processing, text-to-speech (TTS) synthesis, automatic diacritization, and other natural language processing (NLP) tasks.
Dataset Contents
Audio Samples: High-quality Arabic speech recordings.
Text Transcriptions: Fully diacritized Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/hana92/Arabic-Diacritized-TTS-Normalized.ASVspoof_2021_DF_Balanced_Normalizedcommon-voice-20-mn-normalized
Common Voice 20.0 Mongolian Dataset
This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release.
Dataset Structure
The dataset contains:
Audio clips in .mp3 format
Transcriptions for each audio clip
Train/test/dev splits
Additional metadata including speaker demographics
Usage
This dataset can be used for:
Speech Recognition
Voice Analysis
Linguistic Research
Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.masc_filtered_normalizednew-twi-tts-aligned_normalised
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
ASVspoof_2021_LA_Balanced_Normalizedatco2_normalized_augmentedexpresso-concatenated-half-normalnormalized_train_ATC_datasetMCV_Fleurs_Combined_Irish_normalizednormalized_test_ATC_datasetMCV25_Irish_normalizeduzbek-normal-speech-10hiqra_curated_normalised_1s_20s_finalvocalsound-normalizedASVspoof_2021_DF1_Balanced_Normalizedorpheus-synthetic-dataset-normalizedMathSpeech_whisper_transcribed_normalizednormalized_khmer_dataset_14kgrandpa-interview-dataset-normalizednormalized_khmer_datasetmalagasy-asr-normalized-v2normal_audiotrain82normal_audiotest46hinglish-normal-005-fix-speaker-smokecv24_ds_normalizedgoogle_waxal_ds_normalizedtest_data_normalized
Dataset Card for "test_data_normalized"
More Information needed
