datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-g2p-v4-dataset
Thai G2P v4 dataset
Thai G2P v4 dataset is a Thai grapheme-to-phoneme dataset that was built from Thai W2P and Wiktionary th-pron transliterator. We split the dataset by using the first 2 characters for the split group.
Author: Wannaphong Phatthiyaphaibun
GitHub: https://github.com/PyThaiNLP/thai-g2p-v4
sources
Thai W2P: https://huggingface.co/datasets/wannaphong/thai-w2p
Wiktionary th-pron transliterator: https://github.com/PyThaiNLP/pythainlp/pull/1437
arabic-mantoq-synthetic-g2pportuguese-sentences-synthetic-g2p
Dataset Card for 'TigreGotico/portuguese_g2p'
Dataset Description
Dataset Summary
TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants.
It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.galician_g2p
