g2p
Datasets
All datasets matching “g2p”vox-communis-parallel-g2p
VoxCommunis Parallel G2P dataset
This dataset was derived from the VoxCommunis Corpus to provide pairs of utterances along with their
corresponding phonemes, side by side, as to ease the training of grapheme-to-phoneme (G2P) models.
The original VoxCommunis Corpus features force-aligned TextGrids with phone- and word-level segmentations derived from the Mozilla Common Voice Corpus.
The lexicons were developed using Epitran, the XPF Corpus, Charsiu, and some custom dictionaries.… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/vox-communis-parallel-g2p.g2p-dicts
G2P Pronunciation Dictionaries for CrispASR
IPA pronunciation dictionaries for text-to-phoneme conversion in TTS
backends (piper, kokoro, melotts). Auto-downloaded by CrispASR on first use.
Pre-generated IPA Dictionaries (primary — piper-compatible)
Pre-generated IPA transcriptions matching the phoneme inventory of
piper TTS models. These are factual phonetic data — word-to-IPA mappings
generated by processing vocabulary lists through a phonemization engine.… See the full description on the dataset page: https://huggingface.co/datasets/cstr/g2p-dicts.mirandese_g2pHomoRich-G2P-Persian
HomoRich: A Persian Homograph Dataset for G2P Conversion
Overview
HomoRich is the first large-scale, sentence-level Persian homograph dataset designed for grapheme-to-phoneme (G2P) conversion tasks. It addresses the scarcity of balanced, contextually annotated homograph data for low-resource languages. The dataset was created using a semi-automated pipeline combining human expertise and LLM-generated samples, as described in the paper:"Fast, Not Fancy: Rethinking G2P… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/HomoRich-G2P-Persian.netjets-g2p-benchmark-artifacts
NetJets G2P continuation bundle
This repository is the reproducibility bundle for the NameCoach NetJets pronunciation-distribution experiments.
Model
The matching adapter is in Reza2kn/t5gemma-2-4b-netjets-balanced-0p1, based on google/t5gemma-2-4b-4b. The uploaded checkpoint-2472 contains the final LoRA adapter, optimizer state, scheduler state, RNG state, tokenizer, and trainer state so training can resume.
This run used QLoRA (4-bit NF4 base weights), 4 GPUs… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/netjets-g2p-benchmark-artifacts.g2p-lexicons
Crane G2P Lexicons
Word-to-IPA lexicons used by Crane's
grapheme-to-phoneme (G2P) module to phonemize text for text-to-speech
models. Each <lang>/<lang>.tsv file is a plain word<TAB>ipa list, one
pronunciation per line (a word with multiple accepted pronunciations
appears on multiple lines). This is the format Lexicon::from_tsv in
crane-core/src/models/g2p/lexicon.rs
expects.
IPA here follows standard, linguistics-reference-style conventions (e.g.
stress marks ˈ/ˌ before the… See the full description on the dataset page: https://huggingface.co/datasets/crane-local-ai/g2p-lexicons.
