datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipa-lexicon-4v0-7M
IPA Phonetic Lexicon (6.9M words)
This is an International Phonetic Language lexicon containing 6.9 million utterances across 350+ languages.
All speech files were converted into IPA using our neurlang/ipa-whisper-medium model.
Postprocessing to join multiword IPA into single-word IPA record for single-word headwords was performed for appropriate languages.
Global Map
Stats
Total words: 6908075
Countries: 211
Languages: 356
Unique Speakers: 81756… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/ipa-lexicon-4v0-7M.ganjoor-ipa-scansion
Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration
A corpus of 124,404 classical Persian poems collected via the Ganjoor API,
enriched with two things every poem now has:
Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem,
including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet).
Phonemic transliteration — Latin and IPA for every hemistich, produced by the
Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.
