datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-phonemes-10k-alpha
Multilingual Phonemes 10K Alpha
This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows.
Languages
We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.phonikud-phonemes-dataHebrew text with diacritics and phonemes.
The dataset contains millions lines of text and phonemes in Hebrew.
The format is text<TAB>phonemes
Sample: הַאִם זֶה אֲנַ֫חְנוּ וְֽ|הֵם אוֹ כֻּו֯לָּ֫נוּ בְּֽיַחַד? haʔˈim zˈe ʔanˈaχnu vehˈem ʔˈo kulˈanu bejaχˈad?
See Phonikued
Files
hedc4-phonemes.txt - 2 million lines
knesset_phonemes.txt - 5 million lines
This datasets contain lines of text and phonemes generated with phonikud
It may still include some errors, as the project… See the full description on the dataset page: https://huggingface.co/datasets/Phonikud/phonikud-phonemes-data.asante-twi-bible-speech-phonemes
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Asante Twi Bible Speech — Phonemes
Phoneme-labelled version of
ghananlpcommunity/asante-twi-bible-speech-text,
built for training a wav2vec2 (CTC) phoneme recogniser for Asante Twi.
Each example adds a phonemes column: a… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-bible-speech-phonemes.tonebooks-mfa-phonemes-only-hard-sljspeech_phonemes
Dataset Card for "ljspeech_phonemes"
More Information needed
sq-wikipedia-phonemesBisher_ClArTTS-HF-format-with_BW_phonemesllasa_stage1_russian-dolphin_phonemes
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_russian-dolphin_phonemes.llasa_stage1_turkish_tinystories_phonemes
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_turkish_tinystories_phonemes.ipa-phonemes-word-pairs
license: cc-by-sa 4.0
size: ~275k pairs, ~7mb (~4mb parquet)
generated using: phonemizer/espeak
check out openphonemizer for more details!
llasa_stage1_text_phonemes_v4_tinystories_533_us
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes_v4_tinystories_533_us.500k_vi_phonemesllasa_stage1_text_phonemes_v3_YOUTUBE_us
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes_v3_YOUTUBE_us.PhonemeSegmentCounting_Librispeech-wordsllasa_stage1_text_phonemes_v2_uk
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes_v2_uk.phonemes_viPhonemeSegmentCounting_VoxAngelesALLSSTAR_2_phonemesPhonemeSegmentCounting_Librispeech-wordsllasa_stage1_text_phonemes
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes.librispeech_phonemesl2-arctic-dataset-250_phonemesSAA_phonemescmu-arctic-train_phonemesamharic_new_with_phonemes-v1CommonVoice_accent_stratified_phonemessandi_phonemesllasa_stage1_french_simplified_phonemes
LLASA Stage 1 Training Dataset (Text-Phoneme Conversion)
This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA.
Purpose
Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities.
Columns
phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world")
text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_french_simplified_phonemes.PhonemeSegmentCounting_VoxAngelescam_assess_phonemes
