CoolFace
Datasetpublic

THU-SPMI/librispeech-phoneme-labels

LibriSpeech IPA Phoneme Labels This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset. All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit: https://pypi.org/project/pinyin-to-ipa The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/THU-SPMI/librispeech-phoneme-labels.

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
0likes140downloads
Dataset Card

LibriSpeech IPA Phoneme Labels

This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset.

All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit:

  • https://pypi.org/project/pinyin-to-ipa

The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments.

Dataset Structure

  • train-clean-100-phoneme
  • train-clean-360-phoneme
  • train-other-500-phoneme
  • dev-clean-phoneme
  • dev-other-phoneme
  • test-clean-phoneme
  • test-other-phoneme
  • lexicon.txt
  • phone_list