datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
urdu-g2p-dictionary
Urdu G2P Phoneme Dictionary
Dataset Description
A comprehensive Grapheme-to-Phoneme (G2P) dictionary for Urdu, containing 478,000+ word-to-IPA mappings. This is the largest publicly available Urdu phoneme dictionary, designed for:
🎙️ Text-to-Speech (TTS) systems
🔊 Automatic Speech Recognition (ASR)
📚 Linguistic research
🧠 NLP applications
Dataset Summary
Metric
Value
Total Words
478,000+
Language
Urdu (ur)
Script
Arabic (Nastaliq)… See the full description on the dataset page: https://huggingface.co/datasets/humair025/urdu-g2p-dictionary.ja-tts-g2p-bench
ja-tts-g2p-bench — 日本語 TTS の「漢字の読み(g2p)」性能ベンチマーク
日本語の漢字は同じ表記でも文脈によって読みが変わります(例: 市場→しじょう/いちば)。
このデータセットは TTS が漢字を正しく読み分けられるか を評価する 151 問のベンチマークです。
データ構成
data/items_read_bench_v1.jsonl — 151 問のベンチマーク項目(文・対象漢字・期待読み・代替読み)
results_read_bench.jsonl — 4 エンジン × 2 ASR の評価結果
blindspot_wav/{engine}/ — 各 TTS エンジンが生成した WAV ファイル
gemini/ — Google Gemini (gemini-2.5-flash-preview-tts)
openai/ — OpenAI (gpt-4o-mini-tts)
qwen3/ — Qwen3 (cosyvoice2-0.5b via DashScope)
voicevox/ —… See the full description on the dataset page: https://huggingface.co/datasets/kawajiri-tellernovel/ja-tts-g2p-bench.SND_G2P
Sindhi G2P Dataset (SND_G2P)
A Grapheme-to-Phoneme (G2P) dataset for the Sindhi language, mapping written words to their IPA (International Phonetic Alphabet) pronunciations.
Dataset Description
This dataset was compiled by scraping Wiktionary's Sindhi terms with IPA pronunciation category. For each Sindhi word, the corresponding IPA pronunciation was extracted from the word's Wiktionary entry under the Sindhi language section.
Intended use cases:
Grapheme-to-Phoneme… See the full description on the dataset page: https://huggingface.co/datasets/DanishMahdi/SND_G2P.
