CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecti7 /phoneme2audio100K<n<1M0 likes970 downloads6mo agoHugging Face02phonemetransformers /IPA-BabyLM Phonemized BabyLM Pre-training Data This corpus contains the strict and strict-small portions of the BabyLM 2024 pre-training data converted to phonemes using G2P+. The original orthographic data is also available. The scripts used to produce the dataset are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. See the G2P+ paper here. text10M<n<100M2 likes846 downloads1y agoHugging Face03Phonikud /phonikud-phonemes-dataHebrew text with diacritics and phonemes. The dataset contains millions lines of text and phonemes in Hebrew. The format is text<TAB>phonemes Sample: הַאִם זֶה אֲנַ֫חְנוּ וְֽ|הֵם אוֹ כֻּו֯לָּ֫נוּ בְּֽיַחַד? haʔˈim zˈe ʔanˈaχnu vehˈem ʔˈo kulˈanu bejaχˈad? See Phonikued Files hedc4-phonemes.txt - 2 million lines knesset_phonemes.txt - 5 million lines This datasets contain lines of text and phonemes generated with phonikud It may still include some errors, as the project… See the full description on the dataset page: https://huggingface.co/datasets/Phonikud/phonikud-phonemes-data.texttext-to-speech5 likes645 downloads5mo agoHugging Face04styletts2-community /multilingual-phonemes-10k-alpha Multilingual Phonemes 10K Alpha This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows. Languages We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.text100K<n<1M39 likes632 downloads3y agoHugging Face05Peacockery /librispeech-phoneme-featurestabular100K<n<1M0 likes548 downloads7mo agoHugging Face06phonemetransformers /IPA-CHILDES IPA-CHILDES Dataset This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here. Description Key Columns The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.tabular10M<n<100M7 likes286 downloads1y agoHugging Face07Cnam-LMSSC /common_voice_13_french_phoneme Common Voice 13 French Phoneme Dataset Summary This dataset is a curated version of the French subset of Common Voice 13.0, enriched with a phonetic transcription column (phoneme). It was created by the Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) to support research in speech processing, specifically for tasks requiring phonetic alignment, phoneme recognition, and robust speech-to-text applications in French. The dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/common_voice_13_french_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes285 downloads8mo agoHugging Face08Evan-Lin /wiki-phonemetext1M<n<10M0 likes263 downloads2y agoHugging Face09ghananlpcommunity /asante-twi-bible-speech-phonemes This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Asante Twi Bible Speech — Phonemes Phoneme-labelled version of ghananlpcommunity/asante-twi-bible-speech-text, built for training a wav2vec2 (CTC) phoneme recogniser for Asante Twi. Each example adds a phonemes column: a… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-bible-speech-phonemes.audioautomatic-speech-recognition10K<n<100K0 likes225 downloads3mo agoHugging Face10cheikh1499 /libriSpeech_phonemeaudio10K<n<100K0 likes216 downloads8mo agoHugging Face11phonemetransformers /IPA-BabyLM-evaluation BabyLM 2024 evaluation data in IPA A version of the BabyLM 2024 evalution data converted to IPA using G2P+. Scripts for producing this data are available here. This data was used in From Babble to Words: Pre-Training Language Models on Continuous Streams of Phonemes. 0 likes204 downloads1y agoHugging Face12qklent /tonebooks-mfa-phonemes-only-hard-saudio10K<n<100K0 likes195 downloads10mo agoHugging Face13Bisher /Iqra_train_processed_whisper-phoneme10K<n<100K0 likes169 downloads1y agoHugging Face14THU-SPMI /librispeech-phoneme-labels LibriSpeech IPA Phoneme Labels This repository provides IPA-based phoneme annotations and lexicon for the LibriSpeech dataset. All phoneme labels are converted from CMU Pronouncing Dictionary (CMU-Dict) phonemes into IPA symbols using deterministic rules, with the help of the following toolkit: https://pypi.org/project/pinyin-to-ipa The data is intended for phoneme-based ASR, P2G/G2P research, phoneme CTC / AED models, and cross-lingual phoneme experiments. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/THU-SPMI/librispeech-phoneme-labels.automatic-speech-recognition10M<n<100M0 likes140 downloads9mo agoHugging Face15Cnam-LMSSC /multilingual_librispeech_french_phoneme Multilingual LibriSpeech French Phoneme Dataset Summary This dataset is a curated version of the French subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into French acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_french_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes139 downloads8mo agoHugging Face16bookbot /ljspeech_phonemes Dataset Card for "ljspeech_phonemes" More Information needed audio10K<n<100K11 likes131 downloads4y agoHugging Face17bbunzeck /phoneme-blimptext10K<n<100K0 likes116 downloads2y agoHugging Face18Cnam-LMSSC /multilingual_librispeech_spanish_phoneme Multilingual LibriSpeech Spanish Phoneme Dataset Summary This dataset is a curated version of the Spanish subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Spanish acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_spanish_phoneme.audioautomatic-speech-recognition100K<n<1M1 likes110 downloads7mo agoHugging Face19renikud /phoneme-dataset Hebrew Phoneme Dataset This dataset contains Hebrew text paired with IPA phoneme transcriptions. The data was produced from ivrit-ai/VoxKnesset. Audio was segmented with Silero VAD, transcribed with ivrit-ai/whisper-large-v3-turbo for Hebrew text, and with renikud/whisper-he-ipa for IPA. Files voxknesset-whisper-abjad-he-ipa-raw.tsv Raw extraction output. Includes source filenames and unnormalized fields. voxknesset-whisper-abjad-he-ipa-normalized.tsv Normalized… See the full description on the dataset page: https://huggingface.co/datasets/renikud/phoneme-dataset.automatic-speech-recognition0 likes65 downloads4mo agoHugging Face20phonemefake /PhonemeFakeV2 PhonemeFake - A Phonetic Deepfake Dataset We introduce PhonemeFake, a DF attack that manipulates critical speech segments using language reasoning, significantly reducing human perception and SoTA model accuracies. Dataset Details We provide example spoof audio in the viewers tab along with the transcription of the bonafide sample, the manipulated transcription and the audio timings for a small set of the data. The dataset is split into three subsets. The spoof samples… See the full description on the dataset page: https://huggingface.co/datasets/phonemefake/PhonemeFakeV2.audion<1K0 likes62 downloads1y agoHugging Face21mirfan899 /kids_phoneme_asr Dataset Card for "kids_phoneme_asr" More Information needed audio1K<n<10K1 likes60 downloads3y agoHugging Face22mirfan899 /kids_phoneme_md Dataset Card for "kids_phoneme_md" More Information needed audio1K<n<10K1 likes57 downloads3y agoHugging Face23bobboyms /phoneme-ctc-spanish-52h-noisyaudio10K<n<100K0 likes54 downloads9mo agoHugging Face24bbunzeck /phoneme-babylm-100Mtext10M<n<100M0 likes53 downloads2y agoHugging Face25fatlonder /sq-wikipedia-phonemestext10K<n<100K0 likes49 downloads2y agoHugging Face26bobboyms /phoneme-ctc-english-60haudio10K<n<100K0 likes46 downloads9mo agoHugging Face27AdoCleanCode /llasa_stage1_russian-dolphin_phonemes LLASA Stage 1 Training Dataset (Text-Phoneme Conversion) This dataset contains filtered text-phoneme pairs for Stage 1 curriculum learning in LLASA. Purpose Stage 1 training teaches the model text ↔ phoneme conversion using LLAMA's existing text understanding capabilities. Columns phoneme_text_sequence: Phoneme tokens → Text conversion sequence ("<|ph_0068|>... Convert into text: Hello world") text_phoneme_sequence: Text → Phoneme tokens conversion sequence… See the full description on the dataset page: https://huggingface.co/datasets/AdoCleanCode/llasa_stage1_russian-dolphin_phonemes.text100K<n<1M0 likes46 downloads8mo agoHugging Face28Reza2kn /persian-phoneme-cleantext100K<n<1M0 likes45 downloads2mo agoHugging Face29Chijioke-Mgbahurike /wav2vec2_phoneme_spot_data_allaudio1K<n<10K0 likes44 downloads2y agoHugging Face30Cnam-LMSSC /multilingual_librispeech_italian_phoneme Multilingual LibriSpeech Italian Phoneme Dataset Summary This dataset is a curated version of the Italian subset of Multilingual LibriSpeech (MLS), enriched with a phonetic transcription column (phoneme). The Laboratoire de Mécanique des Structures et des Systèmes Couplés (Cnam-LMSSC) created this version to facilitate research into Italian acoustic modeling, phoneme recognition, and speech synthesis. It builds upon the high-quality audio derived from LibriVox audiobooks… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/multilingual_librispeech_italian_phoneme.audioautomatic-speech-recognition10K<n<100K1 likes44 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.