phonetic
Datasets
All datasets matching “phonetic”ucla_phonetic_corpus
Dataset Card for "ucla_phonetic_corpus"
More Information needed
quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.phonetic-piper-recording-studio-prompts
Phonetic Piper Studio Recordings Prompts
This dataset is a processed version of an utterance dataset made available for various languages as prompts for the Piper recording studio. Along with the original prompts, we include:
columns ipa_espeak and ipa_epitran containing phonemized versions of the original sentences according to espeak-ng and Epitran phonemizers, respectively
columns lang, espeak_lang_code, epitran_lang_code containing the language codes as reported by piper… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/phonetic-piper-recording-studio-prompts.PhoneticQA-SO762
PhoneticQA-SO762 v0.1
PhoneticQA-SO762 is a small, word-level multiple-choice AudioQA benchmark derived from SpeechOcean762. It was created and released by Haopeng Geng as an early benchmark for comparing human and speech-language-model sensitivity to salient mispronunciations.
Benchmark design
160 full-utterance audio questions: 32 dev and 128 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-SO762.ml-phonetic-lexicon
Malayalam Phonetic Lexicon
This dataset contains words in Malayalam script and their pronunciation in International Phonetic Alphabet (IPA)
The words in the lexicon are sourced from
The most frequest 100 thousand words from Indic NLP corpus
Curated collection of word categories from Mlmorph project
This pronunciations are created using Mlphon python Library.
Applications
Ready to use pronunciation lexicons for ASR and TTS
To train datadriven grapheme to phoneme… See the full description on the dataset page: https://huggingface.co/datasets/smcproject/ml-phonetic-lexicon.PhoneticQA-L2ARCTIC
PhoneticQA-L2ARCTIC v0.1
PhoneticQA-L2ARCTIC is a small, word-level multiple-choice AudioQA probe derived from L2-ARCTIC. It was created and released by Haopeng Geng to study the gap between phoneme-level error annotations and perceptually salient mispronunciations.
Benchmark design
120 full-utterance audio questions: 24 dev and 96 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most clearly… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-L2ARCTIC.
