datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quran-tajweed-phonetics
The complete phonetic layer of the Quran in the riwaya of Hafs 'an
'Asim via tariq al-Shatibiyyah: 6,236 ayat, 522,475 phones, every
phone carrying its tajweed attribution: madd class with its transmitted
length range, ghunna grade, qalqalah class, tafkheem with its rank, sakt,
the seventeen sifat, and the rule that produced it.
Built and maintained by Quran Lab, a waqf building open technology in
the service of the Quran.
How it was built and verified
Indexed from the… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quran-tajweed-phonetics.PhoneticQA-SO762
PhoneticQA-SO762 v0.1
PhoneticQA-SO762 is a small, word-level multiple-choice AudioQA benchmark derived from SpeechOcean762. It was created and released by Haopeng Geng as an early benchmark for comparing human and speech-language-model sensitivity to salient mispronunciations.
Benchmark design
160 full-utterance audio questions: 32 dev and 128 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-SO762.PhoneticQA-L2ARCTIC
PhoneticQA-L2ARCTIC v0.1
PhoneticQA-L2ARCTIC is a small, word-level multiple-choice AudioQA probe derived from L2-ARCTIC. It was created and released by Haopeng Geng to study the gap between phoneme-level error annotations and perceptually salient mispronunciations.
Benchmark design
120 full-utterance audio questions: 24 dev and 96 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most clearly… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-L2ARCTIC.adaption-african-ideophone-phonetics
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-african_ideophone_phonetics
This dataset contains phonetic measurements of ideophones, nouns, and verbs in African tonal languages including Igbo, Yoruba, and Zaar. It compares careful and natural speech realizations, documenting segmental changes, tonal patterns, reduplication types, and fundamental frequency (F0) metrics. Each entry includes orthographic targets, sentence context… See the full description on the dataset page: https://huggingface.co/datasets/ChiamakaNwokolo/adaption-african-ideophone-phonetics.
