datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phoneme2IPA-BabyLM
Phonemized BabyLM Pre-training Data
This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available.
The scripts used to produce the dataset are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here.
phonikud-phonemes-dataHebrew text with diacritics and phonemes.
The dataset contains millions lines of text and phonemes in Hebrew.
The format is text<TAB>phonemes
Sample: הַאִם זֶה אֲנַ֫חְנוּ וְֽ|הֵם אוֹ כֻּו֯לָּ֫נוּ בְּֽיַחַד? haʔˈim zˈe ʔanˈaχnu vehˈem ʔˈo kulˈanu bejaχˈad?
See Phonikued
Files
hedc4-phonemes.txt - 2 million lines
knesset_phonemes.txt - 5 million lines
This datasets contain lines of text and phonemes generated with phonikud
It may still include some errors, as the project… See the full description on the dataset page: https://huggingface.co/datasets/Phonikud/phonikud-phonemes-data.multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.librispeech-phoneme-featuresIPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.common_voice_13_french_phoneme
Common Voice 13 French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Common Voice 13.0, enriched with a phonetic transcription column (phoneme).
It was created by the Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) to support research in speech processing, specifically for tasks requiring phonetic alignment, phoneme recognition, and robust speech-to-text applications in French.
The dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/common_voice_13_french_phoneme.wiki-phonemeasante-twi-bible-speech-phonemes
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Asante Twi Bible Speech — Phonemes
Phoneme-labelled version of
ghananlpcommunity/asante-twi-bible-speech-text,
built for training a wav2vec2 (CTC) phoneme recogniser for Asante Twi.
Each example adds a phonemes column: a… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-bible-speech-phonemes.libriSpeech_phonemeIPA-BabyLM-evaluation
BabyLM 2024 evaluation data in IPA
A version of the BabyLM 2024 evalution data converted to IPA using G2P+. Scripts for producing this data are available here.
This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes.
tonebooks-mfa-phonemes-only-hard-sIqra_train_processed_whisper-phonemelibrispeech-phoneme-labels
LibriSpeech IPA Phoneme Labels
This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset.
All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit:
https://pypi.org/project/pinyin-to-ipa
The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/THU-SPMI/librispeech-phoneme-labels.multilingual_librispeech_french_phoneme
Multilingual LibriSpeech French Phoneme
Dataset Summary
This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.ljspeech_phonemes
Dataset Card for "ljspeech_phonemes"
More Information needed
phoneme-blimpmultilingual_librispeech_spanish_phoneme
Multilingual LibriSpeech Spanish Phoneme
Dataset Summary
This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.phoneme-dataset
Hebrew Phoneme Dataset
This dataset contains Hebrew text paired with IPA phoneme transcriptions.
The data was produced from ivrit-ai/VoxKnesset. Audio was segmented with Silero VAD, transcribed with ivrit-ai/whisper-large-v3-turbo for Hebrew text, and with renikud/whisper-he-ipa for IPA.
Files
voxknesset-whisper-abjad-he-ipa-raw.tsv
Raw extraction output. Includes source filenames and unnormalized fields.
voxknesset-whisper-abjad-he-ipa-normalized.tsv
Normalized… See the full description on the dataset page: https://huggingface.co/datasets/renikud/phoneme-dataset.PhonemeFakeV2
PhonemeFake - A Phonetic Deepfake Dataset
We introduce PhonemeFake, a DF attack that manipulates critical speech segments using language reasoning, significantly reducing human perception and SoTA
model accuracies.
Dataset Details
We provide example spoof audio in the viewers tab along with the transcription of the bonafide sample, the manipulated transcription and the audio timings for a small set of the data.
The dataset is split into three subsets. The spoof samples… See the full description on the dataset page: https://huggingface.co/datasets/phonemefake/PhonemeFakeV2.kids_phoneme_asr
Dataset Card for "kids_phoneme_asr"
More Information needed
kids_phoneme_md
Dataset Card for "kids_phoneme_md"
More Information needed
phoneme-ctc-spanish-52h-noisyphoneme-babylm-100Msq-wikipedia-phonemesphoneme-ctc-english-60hllasa_stage1_russian-dolphin_phonemes
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_russian-dolphin_phonemes.persian-phoneme-cleanwav2vec2_phoneme_spot_data_allmultilingual_librispeech_italian_phoneme
Multilingual LibriSpeech Italian Phoneme
Dataset Summary
This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme).
The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.
