CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes1.5k downloads2mo agoHugging Face02maliced /librispeech_mfcc_alignedtextautomatic-speech-recognition100K<n<1M0 likes448 downloads1y agoHugging Face03ghananlpcommunity /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes410 downloads2mo agoHugging Face04manassehzw /shona-bible-bdsc-aligned Shona Bible Speech Alignment Dataset Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC source audio made available by Biblica, Inc. through Open.Bible. This release contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech segments covering approximately 75.55 hours. Dataset summary Language: Shona (sna) Speaker: narrator 1 Speaker sex: male Books: 66 Clips: 31,284 Audio: approximately 75.55 hours Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.audioautomatic-speech-recognition10K<n<100K3 likes259 downloads16d agoHugging Face05duplexio /emilia-yodas-en-aligned Emilia-YODAS EN Word-Aligned Word-level forced-alignment timestamps for the English subset of amphion/Emilia-Dataset (Emilia-YODAS split), produced with Qwen/Qwen3-ForcedAligner-0.6B. No audio is redistributed — this dataset contains only metadata (IDs, transcripts already present in Emilia-YODAS, and per-word [start, end] timestamps). To use it, join on id with the original Emilia-YODAS audio. Stats Metric Value Utterances 4,516,833 Total audio 11,572.7… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-aligned.textautomatic-speech-recognition1M<n<10M0 likes215 downloads5mo agoHugging Face06MagicLuke /CHILDES-Alignedgated [!IMPORTANT] How to access this dataset: the official public release is hosted by TalkBank at https://talkbank.org/childes/access/Derived/CHILDES-Aligned.html (audio archives + CSV/JSONL metadata, CC BY-NC-SA 4.0). Please obtain the dataset there. This Hugging Face copy is retained gated, for internal use; access requests are approved manually and general requests may be declined — use the TalkBank release instead. CHILDES-Aligned: Curated Child-Speech Dataset (BEACON) English… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/CHILDES-Aligned.audioautomatic-speech-recognition100K<n<1M4 likes11 downloads2mo agoHugging Face07instinct-org /cv_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/cv_chunked Aligned dataset: instinct-org/cv_chunked_nfa_aligned Rows: 71097 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.textautomatic-speech-recognition10K<n<100K0 likes9 downloads1mo agoHugging Face08RakhatM /medical_asr_aligned_27_04gated Medical ASR Aligned Dataset Aligned Kazakh medical speech dataset from the «ТЕЛЕДӘРІГЕР» (TeleDoctor) TV program on Qazaqstan National Channel. Dataset Description Audio-transcript aligned segments of Kazakh-language medical TV broadcasts. Each segment contains the original audio chunk, ASR transcription, human reference transcription, and Character Error Rate (CER). Only segments with CER < 25% are included. Stats Split Segments Avg CER Avg Duration… See the full description on the dataset page: https://huggingface.co/datasets/RakhatM/medical_asr_aligned_27_04.audioautomatic-speech-recognition10K<n<100K0 likes6 downloads5mo agoHugging Face09instinct-org /audio_youtube_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/audio_youtube_chunked Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned Rows: 559484 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes4 downloads4mo agoHugging Face10instinct-org /default_voices_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/default_voices_chunked Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned Rows: 134236 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes4 downloads4mo agoHugging Face11instinct-org /miscellaneous_yt_chunked_speech_restorised_nfa_alignedgated miscellaneous_yt_chunked_speech_restorised_nfa_aligned Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised. Contents Parquet shards: 130 Rows: 528,187 Approx hours: 863.88 Audio column: audio with embedded FLAC bytes Transcript column: transcription Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes4 downloads4mo agoHugging Face12instinct-org /audiobook_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/audiobook_chunked Aligned dataset: instinct-org/audiobook_chunked_nfa_aligned Rows: 1291838 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_nfa_aligned.tabularautomatic-speech-recognition1M<n<10M0 likes3 downloads4mo agoHugging Face13instinct-org /zy_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/zy_chunked Aligned dataset: instinct-org/zy_chunked_nfa_aligned Rows: 534816 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes3 downloads4mo agoHugging Face14instinct-org /tbp_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/tbp_chunked Aligned dataset: instinct-org/tbp_chunked_nfa_aligned Rows: 548483 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_nfa_aligned.textautomatic-speech-recognition100K<n<1M0 likes3 downloads4mo agoHugging Face15instinct-org /espeech_podcasts_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/espeech_podcasts_chunked_speech_restorised Aligned dataset: instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned Rows: 2467471 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition1M<n<10M0 likes3 downloads4mo agoHugging Face16instinct-org /yt_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt_chunked_speech_restorised Aligned dataset: instinct-org/yt_chunked_speech_restorised_nfa_aligned Rows: 416380 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes3 downloads4mo agoHugging Face17instinct-org /omni_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/omni_chunked Aligned dataset: instinct-org/omni_chunked_nfa_aligned Rows: 2699 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_nfa_aligned.textautomatic-speech-recognition1K<n<10K0 likes2 downloads4mo agoHugging Face18instinct-org /espeech_podcasts_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/espeech_podcasts_chunked Aligned dataset: instinct-org/espeech_podcasts_chunked_nfa_aligned Rows: 231121 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face19instinct-org /yt1_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt1_chunked_speech_restorised Aligned dataset: instinct-org/yt1_chunked_speech_restorised_nfa_aligned Rows: 261565 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face20instinct-org /yt2_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt2_chunked_speech_restorised Aligned dataset: instinct-org/yt2_chunked_speech_restorised_nfa_aligned Rows: 809612 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face21instinct-org /yt3_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt3_chunked_speech_restorised Aligned dataset: instinct-org/yt3_chunked_speech_restorised_nfa_aligned Rows: 506444 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face22instinct-org /yt4_chunked_speech_restorised_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/yt4_chunked_speech_restorised Aligned dataset: instinct-org/yt4_chunked_speech_restorised_nfa_aligned Rows: 965192 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes2 downloads4mo agoHugging Face23instinct-org /education_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/education_chunked Aligned dataset: instinct-org/education_chunked_nfa_aligned Rows: 386961 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_nfa_aligned.tabularautomatic-speech-recognition100K<n<1M0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.