datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-English-dictionary-resources
English and Chinese IPA Lexicons and Phoneme Sets
This repository provides English and Chinese IPA pronunciation lexicons and phoneme inventories collected from the vocabularies of multiple ASR corpora. The resources can be used for ASR, grapheme-to-phoneme conversion, pronunciation modeling, TTS, and related speech research.
Repository Structure
.
├── en/
│ ├── lexicon.txt
│ └── phone_list
└── zh/
├── lexicon.txt
└── phone_list
Source… See the full description on the dataset page: https://huggingface.co/datasets/maxwellziweiwei/Chinese-English-dictionary-resources.urdu-g2p-dictionary
Urdu G2P Phoneme Dictionary
Dataset Description
A comprehensive Grapheme-to-Phoneme (G2P) dictionary for Urdu, containing 478,000+ word-to-IPA mappings. This is the largest publicly available Urdu phoneme dictionary, designed for:
🎙️ Text-to-Speech (TTS) systems
🔊 Automatic Speech Recognition (ASR)
📚 Linguistic research
🧠 NLP applications
Dataset Summary
Metric
Value
Total Words
478,000+
Language
Urdu (ur)
Script
Arabic (Nastaliq)… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu-g2p-dictionary.
