THU-SPMI/librispeech-phoneme-labels
LibriSpeech IPA Phoneme Labels This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset. All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit: https://pypi.org/project/pinyin-to-ipa The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/THU-SPMI/librispeech-phoneme-labels.
LibriSpeech IPA Phoneme Labels
This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset.
All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit:
- https://pypi.org/project/pinyin-to-ipa
The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments.
Dataset Structure
train-clean-100-phonemetrain-clean-360-phonemetrain-other-500-phonemedev-clean-phonemedev-other-phonemetest-clean-phonemetest-other-phonemelexicon.txtphone_list
