datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhoneticQA-SO762
PhoneticQA-SO762 v0.1
PhoneticQA-SO762 is a small, word-level multiple-choice AudioQA benchmark derived from SpeechOcean762. It was created and released by Haopeng Geng as an early benchmark for comparing human and speech-language-model sensitivity to salient mispronunciations.
Benchmark design
160 full-utterance audio questions: 32 dev and 128 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-SO762.PhoneticQA-L2ARCTIC
PhoneticQA-L2ARCTIC v0.1
PhoneticQA-L2ARCTIC is a small, word-level multiple-choice AudioQA probe derived from L2-ARCTIC. It was created and released by Haopeng Geng to study the gap between phoneme-level error annotations and perceptually salient mispronunciations.
Benchmark design
120 full-utterance audio questions: 24 dev and 96 test.
Each item shows the canonical transcript and four candidate words.
The task is to select the word that sounds most clearly… See the full description on the dataset page: https://huggingface.co/datasets/Haopeng/PhoneticQA-L2ARCTIC.
