CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01styletts2-community /multilingual-phonemes-10k-alpha Multilingual Phonemes 10K Alpha This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows. Languages We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.text100K<n<1M39 likes649 downloads3y agoHugging Face02Phonikud /phonikud-phonemes-dataHebrew text with diacritics and phonemes. The dataset contains millions lines of text and phonemes in Hebrew. The format is text<TAB>phonemes Sample: הַאִם זֶה אֲנַ֫חְנוּ וְֽ|הֵם אוֹ כֻּו֯לָּ֫נוּ בְּֽיַחַד? haʔˈim zˈe ʔanˈaχnu vehˈem ʔˈo kulˈanu bejaχˈad? See Phonikued Files hedc4-phonemes.txt - 2 million lines knesset_phonemes.txt - 5 million lines This datasets contain lines of text and phonemes generated with phonikud It may still include some errors, as the project… See the full description on the dataset page: https://huggingface.co/datasets/Phonikud/phonikud-phonemes-data.texttext-to-speech5 likes562 downloads5mo agoHugging Face03ghananlpcommunity /asante-twi-bible-speech-phonemes This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Asante Twi Bible Speech — Phonemes Phoneme-labelled version of ghananlpcommunity/asante-twi-bible-speech-text, built for training a wav2vec2 (CTC) phoneme recogniser for Asante Twi. Each example adds a phonemes column: a… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-bible-speech-phonemes.audioautomatic-speech-recognition10K<n<100K0 likes228 downloads3mo agoHugging Face04qklent /tonebooks-mfa-phonemes-only-hard-saudio10K<n<100K0 likes217 downloads10mo agoHugging Face05bookbot /ljspeech_phonemes Dataset Card for "ljspeech_phonemes" More Information needed audio10K<n<100K11 likes138 downloads4y agoHugging Face06fatlonder /sq-wikipedia-phonemestext10K<n<100K0 likes53 downloads2y agoHugging Face07Bisher /Bisher_ClArTTS-HF-format-with_BW_phonemesaudio1K<n<10K0 likes52 downloads1y agoHugging Face08AdoCleanCode /llasa_stage1_russian-dolphin_phonemes LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_russian-dolphin_phonemes.text100K<n<1M0 likes50 downloads8mo agoHugging Face09AdoCleanCode /llasa_stage1_turkish_tinystories_phonemes LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_turkish_tinystories_phonemes.text100K<n<1M0 likes37 downloads8mo agoHugging Face10mrfakename /ipa-phonemes-word-pairs license: cc-by-sa 4.0 size: ~275k pairs, ~7mb (~4mb parquet) generated using: phonemizer/espeak check out openphonemizer for more details! text100K<n<1M3 likes29 downloads3y agoHugging Face11AdoCleanCode /llasa_stage1_text_phonemes_v4_tinystories_533_us LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes_v4_tinystories_533_us.text100K<n<1M0 likes28 downloads9mo agoHugging Face12namkuner /500k_vi_phonemestext100K<n<1M0 likes22 downloads1y agoHugging Face13AdoCleanCode /llasa_stage1_text_phonemes_v3_YOUTUBE_us LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes_v3_YOUTUBE_us.text100K<n<1M0 likes22 downloads9mo agoHugging Face14speech31 /PhonemeSegmentCounting_Librispeech-wordsaudio1K<n<10K0 likes19 downloads2y agoHugging Face15AdoCleanCode /llasa_stage1_text_phonemes_v2_uk LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes_v2_uk.text100K<n<1M0 likes19 downloads9mo agoHugging Face16namkuner /phonemes_vitext1M<n<10M0 likes16 downloads2y agoHugging Face17speech31 /PhonemeSegmentCounting_VoxAngelesaudio1K<n<10K0 likes15 downloads2y agoHugging Face18PranavBhalerao /ALLSSTAR_2_phonemesaudio100K<n<1M0 likes13 downloads1y agoHugging Face19DynamicSuperb /PhonemeSegmentCounting_Librispeech-wordsaudio1K<n<10K1 likes10 downloads2y agoHugging Face20AdoCleanCode /llasa_stage1_text_phonemes LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_text_phonemes.text1K<n<10K0 likes10 downloads9mo agoHugging Face21jaeeewon /librispeech_phonemestext100K<n<1M0 likes9 downloads11mo agoHugging Face22cheikh1499 /libriSpeech_phonemes0 likes6 downloads1y agoHugging Face23PranavBhalerao /l2-arctic-dataset-250_phonemesaudio100K<n<1M0 likes6 downloads1y agoHugging Face24PranavBhalerao /SAA_phonemesaudio100K<n<1M0 likes6 downloads1y agoHugging Face25PranavBhalerao /cmu-arctic-train_phonemesaudio100K<n<1M0 likes6 downloads1y agoHugging Face26surafelabebe /amharic_new_with_phonemes-v1tabular10K<n<100K0 likes5 downloads2y agoHugging Face27PranavBhalerao /CommonVoice_accent_stratified_phonemesaudio100K<n<1M0 likes5 downloads1y agoHugging Face28PranavBhalerao /sandi_phonemesaudio100K<n<1M0 likes5 downloads1y agoHugging Face29AdoCleanCode /llasa_stage1_french_simplified_phonemes LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_french_simplified_phonemes.text100K<n<1M0 likes5 downloads8mo agoHugging Face30DynamicSuperb /PhonemeSegmentCounting_VoxAngelesaudio1K<n<10K1 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.