CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01palshub /phonemizer-dicts Phonemizer Dicts Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup. Files en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA Provenance Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here. Regeneration See scripts/generate-espeak-dict.py in the tts-rd-team repo. text100K<n<1M0 likes23k downloads5mo agoHugging Face02suchirsalhan /Phonemized-UD Phoneme-UD: A Multilingual Phonemized Universal Dependencies Corpus for 34+ Languages G2P+ Phonemizer We use G2P+ to phonemize Universal Dependencies. Here is an example usage: # Install required packages !apt-get install -y espeak-ng !pip install phonemizer g2p-plus # Set the environment variable from Python import os os.environ["PHONEMIZER_ESPEAK_LIBRARY"] = "/usr/lib/x86_64-linux-gnu/libespeak-ng.so.1" # Now run your transcription from g2p_plus import… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/Phonemized-UD.text1M<n<10M0 likes840 downloads1y agoHugging Face03ghananlpcommunity /ghana-speech-phonemized Ghana Speech (Phonemized) Preprocessed multilingual speech dataset from ghananlpcommunity/ghana-speech, covering 42 Ghanaian and West African languages with IPA phoneme transcriptions. This dataset contains only the metadata and phonemes — no audio. Use it together with the source dataset to load audio on-the-fly. Contents Each language is a separate config (Parquet). Columns: Column Type Description id string Unique clip ID (matches audio IDs in the… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-speech-phonemized.0 likes673 downloads2mo agoHugging Face04mugezhang /eng_spa_owt_phonemized_fulltext10M<n<100M0 likes405 downloads9mo agoHugging Face05ghananlpcommunity /ghana-speech-lfn-phonemizedtabular1M<n<10M0 likes317 downloads2mo agoHugging Face06AymanMansour /tarteel-ai-EA-DI-phonemized-Finalaudio100K<n<1M0 likes262 downloads1y agoHugging Face07Rosany /catalan-dataset-phonemizedtabular10K<n<100K0 likes63 downloads2y agoHugging Face08mugezhang /eng_spa_owt_phonemized_full_romanizedtext10M<n<100M0 likes54 downloads9mo agoHugging Face09h3110Fr13nd /voxpopuli-spanish-espeak-phonemizedtext0 likes48 downloads2y agoHugging Face10Rosany /catalan-dataset-phonemized-splittext10K<n<100K0 likes45 downloads2y agoHugging Face11mugezhang /xlsum_unseen_phonemized_romanizedtext100K<n<1M0 likes32 downloads4mo agoHugging Face12speech-uk /opentts-phonemized-sentencesPhonemized version of https://huggingface.co/datasets/speech-uk/text-to-speech-sentences with some additional fields. tabulartext-to-speech1K<n<10K1 likes11 downloads1y agoHugging Face13AymanMansour /tarteel-ai-EA-DI-phonemizedtext100K<n<1M0 likes8 downloads1y agoHugging Face14CLAUSE-Bielefeld /10M_phonemized_German_datasetgated Small Phonemized German dataset Description This dataset is a smaller and phonemized version of Bastian Bunzeck's dataset (Bunzeck et al., 2025) and contains approximately 10 million words. The phonemization was done using Phonemizer by Bernard & Titeux, 2021. This dataset was used to train the Phoneme-based German babyLlama model. Source Composition The dataset has been compiled from the following sources: Source Description Words… See the full description on the dataset page: https://huggingface.co/datasets/CLAUSE-Bielefeld/10M_phonemized_German_dataset.text1M<n<10M0 likes7 downloads1y agoHugging Face15mimba /pl-bert-phonemized-pltgatedtext1M<n<10M0 likes6 downloads2mo agoHugging Face16therealvul /parlertts_pony_speech_phonemizedtabular10K<n<100K1 likes4 downloads2y agoHugging Face17commotion /cloning_pairs_text_non_phonemizedtext100K<n<1M0 likes4 downloads3mo agoHugging Face18Ramora0 /phonemized_transcriptions0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.