CoolFace
20 results

multilingual-tts

multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.8k downloads3mo agoHugging Facemalaysia-ai /Multilingual-TTS Multilingual-TTS A large multilingual corpus for pretraining TTS/STT models, gathered and normalized from 230+ public sources. ~191k hours of audio across 150+ languages, organized into 1,544 dataset configs and tokenized to 34.5B NeuCodec speech tokens (50 Hz, single-codebook) over 111.1M clips. Each config is one source dataset normalized to rows of {audio_filename, text, speaker}: audio_filename — clip path inside that config's <config>_audio.zip (mono MP3). text —… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS.24 likes1.3k downloads27d agoHugging FaceMiniMaxAI /TTS-Multilingual-Test-Set Overview To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts. Specifically, the test set for each language includes: 100 distinct test sentences. Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning. Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.audiotext-to-speechn<1K44 likes1.1k downloads1y agoHugging Facemalaysia-ai /Multilingual-TTS-language Multilingual-TTS-language malaysia-ai/Multilingual-TTS with two extra columns: column description audio_filename, text, speaker unchanged from malaysia-ai/Multilingual-TTS language language detected from the text (transcript) column of every row post-normalized text after rule-based punctuation / capitalization normalization All original columns and the file/folder layout are preserved: 1552 subsets / 1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.text-to-speech100M<n<1B0 likes1.1k downloads27d agoHugging FaceScicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes702 downloads6mo agoHugging Faceakuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes482 downloads9d agoHugging Face