phoneme
wav2vec2-xls-r-300m-timit-phonemegraphemes_to_phonemes_en_usjapanese-hubert-base-phoneme-ctc-v3japanese-hubert-base-phoneme-ctc-v4wav2vec2-large-robust-L2-english-phoneme-recognitionjapanese-hubert-base-phoneme-ctc-v2wav2vec2-large-lv60_phoneme-timit_english_timit-4kwav2vec2-large-xlsr-53-l2-arctic-phoneme
Datasets
All datasets matching “phoneme”phoneme2IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
phonikud-phonemes-dataHebrew text with diacritics and phonemes.
The dataset contains millions lines of text and phonemes in Hebrew.
The format is text<TAB>phonemes
Sample: הַאִם זֶה אֲנַ֫חְנוּ וְֽ|הֵם אוֹ כֻּו֯לָּ֫נוּ בְּֽיַחַד? haʔˈim zˈe ʔanˈaχnu vehˈem ʔˈo kulˈanu bejaχˈad?
See Phonikued
Files
hedc4-phonemes.txt - 2 million lines
knesset_phonemes.txt - 5 million lines
This datasets contain lines of text and phonemes generated with phonikud
It may still include some errors, as the project… See the full description on the dataset page: https://huggingface.co/datasets/Phonikud/phonikud-phonemes-data.multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.librispeech-phoneme-featuresIPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.
