datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mirandese_g2parabic-mantoq-synthetic-g2pPersian_G2P
🏢 About Neura Company
Neura Company focuses on state-of-the-art speech and language technologies.
🤝 Contact
🌐 Website: neura.info
📧 Email: info@neura.info
📧 Email: zahedi.esmaeil@gmail.com
For questions, issues, or collaboration requests, feel free to open an issue or contact us directly.
portuguese-sentences-synthetic-g2p
Dataset Card for 'TigreGotico/portuguese_g2p'
Dataset Description
Dataset Summary
TigreGotico/portuguese_g2p is a Grapheme-to-Phoneme (G2P) dataset for Portuguese, offering phonetic transcriptions for sentences across ten different regional variants.
It is derived from the portuguese_phonetic_lexicon and is designed to aid in the development of robust Speech Recognition (ASR) and Text-to-Speech (TTS) models that account for dialectal variation in Portuguese.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-sentences-synthetic-g2p.galician_g2pkabyle-g2p-training-data
Kabyle G2P Training Data
Phonetically-annotated Kabyle (Taqbaylit) text corpus for training Grapheme-to-Phoneme (G2P) models. Generated using the orthography2ipa rule-based phonemizer for Kabyle.
Dataset Overview
Property
Value
Language
Kabyle (kab) — Afro-Asiatic, Berber
Total pairs
59,462
Source
boffire/kabyle-piper-22khz
Phonemizer
orthography2ipa (dev branch)
IPA standard
Narrow transcription with Kabyle-specific allophony
License
CC0… See the full description on the dataset page: https://huggingface.co/datasets/boffire/kabyle-g2p-training-data.
