CoolFace
20 results

g2p

fdemelo /vox-communis-parallel-g2p VoxCommunis Parallel G2P dataset This dataset was derived from the VoxCommunis Corpus to provide pairs of utterances along with their corresponding phonemes, side by side, as to ease the training of grapheme-to-phoneme (G2P) models. The original VoxCommunis Corpus features force-aligned TextGrids with phone- and word-level segmentations derived from the Mozilla Common Voice Corpus. The lexicons were developed using Epitran, the XPF Corpus, Charsiu, and some custom dictionaries.… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/vox-communis-parallel-g2p.n<1K3 likes1.1k downloads1y agoHugging Facecstr /g2p-dicts G2P Pronunciation Dictionaries for CrispASR IPA pronunciation dictionaries for text-to-phoneme conversion in TTS backends (piper, kokoro, melotts). Auto-downloaded by CrispASR on first use. Pre-generated IPA Dictionaries (primary — piper-compatible) Pre-generated IPA transcriptions matching the phoneme inventory of piper TTS models. These are factual phonetic data — word-to-IPA mappings generated by processing vocabulary lists through a phonemization engine.… See the full description on the dataset page: https://huggingface.co/datasets/cstr/g2p-dicts.texttext-to-speech1M<n<10M0 likes589 downloads7d agoHugging FaceTigreGotico /mirandese_g2ptextn<1K1 likes564 downloads1y agoHugging FaceMahtaFetrat /HomoRich-G2P-Persian HomoRich: A Persian Homograph Dataset for G2P Conversion Overview HomoRich is the first large-scale, sentence-level Persian homograph dataset designed for grapheme-to-phoneme (G2P) conversion tasks. It addresses the scarcity of balanced, contextually annotated homograph data for low-resource languages. The dataset was created using a semi-automated pipeline combining human expertise and LLM-generated samples, as described in the paper:"Fast, Not Fancy: Rethinking G2P… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/HomoRich-G2P-Persian.texttranslation100K<n<1M10 likes352 downloads1y agoHugging FaceReza2kn /netjets-g2p-benchmark-artifacts NetJets G2P continuation bundle This repository is the reproducibility bundle for the NameCoach NetJets pronunciation-distribution experiments. Model The matching adapter is in Reza2kn/t5gemma-2-4b-netjets-balanced-0p1, based on google/t5gemma-2-4b-4b. The uploaded checkpoint-2472 contains the final LoRA adapter, optimizer state, scheduler state, RNG state, tokenizer, and trainer state so training can resume. This run used QLoRA (4-bit NF4 base weights), 4 GPUs… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/netjets-g2p-benchmark-artifacts.0 likes225 downloads21d agoHugging Facecrane-local-ai /g2p-lexicons Crane G2P Lexicons Word-to-IPA lexicons used by Crane's grapheme-to-phoneme (G2P) module to phonemize text for text-to-speech models. Each <lang>/<lang>.tsv file is a plain word<TAB>ipa list, one pronunciation per line (a word with multiple accepted pronunciations appears on multiple lines). This is the format Lexicon::from_tsv in crane-core/src/models/g2p/lexicon.rs expects. IPA here follows standard, linguistics-reference-style conventions (e.g. stress marks ˈ/ˌ before the… See the full description on the dataset page: https://huggingface.co/datasets/crane-local-ai/g2p-lexicons.0 likes196 downloads28d agoHugging Face