datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FOR-normnst-da-norm
Dataset Card for NST-da Normalized
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): da
License: cc0-1.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/JackismyShephard/nst-da-norm.igbo_tts_normalizedopenslr-sinhala-asr-normxh-tts-vixsd-norm
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd-norm.Fleurs_Irish_normalizedfake_or_real_dataset_for_normxh-tts-slr32-norm
isiXhosa TTS — SLR32 prepared for VITS
Multi-speaker isiXhosa speech, resampled and text-normalised for VITS training.
Each audio file is paired with its transcript in metadata.csv.
Attribution (required by the licence)
Derived from OpenSLR SLR32, "High quality TTS data for four South African
languages (af, st, tn, xh)", created by North West University and Google
(2017), released under CC BY-SA 4.0.
Source: https://openslr.org/32/
This derivative is likewise CC… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-slr32-norm.Arabic-Diacritized-TTS-Normalized
Arabic-Diacritized-TTS Dataset
Overview
The Arabic-Diacritized-TTS dataset contains Arabic audio samples and their corresponding text with full diacritization. This dataset is designed to support research in Arabic speech processing, text-to-speech (TTS) synthesis, automatic diacritization, and other natural language processing (NLP) tasks.
Dataset Contents
Audio Samples: High-quality Arabic speech recordings.
Text Transcriptions: Fully diacritized Arabic text… See the full description on the dataset page: https://huggingface.co/datasets/hana92/Arabic-Diacritized-TTS-Normalized.ASVspoof_2021_DF_Balanced_Normalizedcommon-voice-20-mn-normalized
Common Voice 20.0 Mongolian Dataset
This dataset is a subset of Mozilla's Common Voice project, containing Mongolian speech data. It's part of Common Voice 20.0 release.
Dataset Structure
The dataset contains:
Audio clips in .mp3 format
Transcriptions for each audio clip
Train/test/dev splits
Additional metadata including speaker demographics
Usage
This dataset can be used for:
Speech Recognition
Voice Analysis
Linguistic Research
Speech Processing… See the full description on the dataset page: https://huggingface.co/datasets/warmestman/common-voice-20-mn-normalized.masc_filtered_normalizednew-twi-tts-aligned_normalised
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
ASVspoof_2021_LA_Balanced_Normalizedgoogle-latam-spanish-boundary-normalized
Google LATAM Spanish Boundary-Normalized Audio
Female Spanish speech from the following upstream datasets:
Argentina: ylacombe/google-argentinian-spanish
Chile: ylacombe/google-chilean-spanish
Colombia: ylacombe/google-colombian-spanish
Attribution and Thanks
Many thanks to ylacombe for publishing
and maintaining the original Argentinian, Chilean, and Colombian Spanish
datasets. The recordings, transcripts, speaker labels, and original dataset
structure come… See the full description on the dataset page: https://huggingface.co/datasets/groxaxo/google-latam-spanish-boundary-normalized.expresso-concatenated-half-normalatco2_normalized_augmentednormalized_train_ATC_datasetMCV_Fleurs_Combined_Irish_normalizedmy-voice-normalizednormalized_test_ATC_datasetuzbek-normal-speech-10hUrduTTS-normalizedMCV25_Irish_normalizediqra_curated_normalised_1s_20s_finalUrduTTS-normalized-shortvocalsound-normalizedASVspoof_2021_DF1_Balanced_Normalizedorpheus-synthetic-dataset-normalizedopenslr-sinhala-asr-norm-noise-rem
