datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.hindi-english-bilingual
Hindi/English/Hinglish Bilingual TTS Dataset
Synthetic TTS dataset generated by Rani voice (ai4bharat/indic-parler-tts) for training a lightweight bilingual student TTS model.
Designed for natural-sounding Hindi, English, and Hinglish (code-switched) speech synthesis.
Dataset Summary
Property
Value
Total utterances
23,277
Total audio
~4.7GB (24kHz WAV)
Languages
Hindi (hi), English (en), Hinglish (bi)
Sample rate
24kHz
Voice
Rani —… See the full description on the dataset page: https://huggingface.co/datasets/nameissakthi/hindi-english-bilingual.audio_configs_single_nondefault_namenamed-entities-tts-pt-br-audio-qwen3tts
Qwen3-TTS (checkpoint base) — áudio sintetizado de entidades nomeadas
Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do
projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação
ao reconhecimento de nomes próprios", usando o modelo Qwen3-TTS (Qwen/Qwen3-TTS-12Hz-1.7B-Base) em seu
checkpoint base, sem fine-tuning — etapa de baseline do projeto.
Splits
Split
Nº de exemplos
train
17309… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio-qwen3tts.shoebox_rir_with_room_namesnamed-entities-tts-pt-br-audio
YourTTS (checkpoint base) — áudio sintetizado de entidades nomeadas
Dataset de áudio gerado sinteticamente a partir do glossário de entidades nomeadas do
projeto de IC "Aprimoramento de modelos de reconhecimento automático de fala em relação
ao reconhecimento de nomes próprios", usando o modelo YourTTS (tts_models/multilingual/multi-dataset/your_tts) em seu
checkpoint base, sem fine-tuning — etapa de baseline do projeto.
Splits
Split
Nº de exemplos… See the full description on the dataset page: https://huggingface.co/datasets/RodrigoLimaRFL/named-entities-tts-pt-br-audio.tamil_names_audiotamil_names_audio_v2NamedEntityRecognition_SLUE-VoxPopulichemical-namessingaporean_accent_district_names_dataset_ultravoxREPO_NAME-figure3-dsd-rp-random350dataset_nameNamedEntityLocalization_SLUE-VoxPopulisingaporean_accent_district_names_continuationtest-dataset-namename_synthesizedEnter-Your-hub-name
Dataset Card for "Enter-Your-hub-name"
More Information needed
polymer-namesyour-dataset-nameafri-names
Afri-names: Read Speech Dataset of Numbers and African Named Entities
This work is licensed under aCreative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Overview
Afri-names is a curated African-accented read speech dataset comprising 6,307 single-speaker audio samples, totaling 8.92 hours of speech data. Each sample is densely populated with numbers or African named entities or voice commands (with African named entities), making it ideal… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afri-names.REPO_NAMEcompany_namecompany_name_2test-large-names_speaker_similarityyour-dataset-nametest-base-names_speaker_similarity12embedded_world_2026_rag_tts_INC_names
