CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01humair025 /Urdu-ONYX-WAV-kanade-Annotated Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,217 Total Duration: 42.77 hours Average Duration: 5.87 seconds Duration Range: 0.65s - 122.23s Average Phonemes: 18.5 per sample Average Kanade Tokens: 151.1 per sample Global Embedding Dimension: 128 New Columns This dataset adds the following columns: duration (float): Audio duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Urdu-ONYX-WAV-kanade-Annotated.tabulartext-to-speech100K<n<1M0 likes1.3k downloads8mo agoHugging Face02manassehzw /sna-dataset-annotated manassehzw/sna-dataset-annotated An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline. This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments. Why this annotated release exists The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.audioautomatic-speech-recognition10K<n<100K1 likes437 downloads2mo agoHugging Face03manassehzw /sna-waxal-annotated-unlabeled Shona WAXAL annotated-unlabeled checkpoint This is a self-contained operational checkpoint for pseudo-labeling Shona ASR data. It contains 90,253 conservatively segmented FLAC clips (441.585 hours), but intentionally contains no transcripts. Fields transcription is intentionally empty. speaker_id is an approximate source-blind EOM cluster or unknown; speaker_clip_count is zero for unknown assignments. gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.audioautomatic-speech-recognition10K<n<100K0 likes374 downloads2mo agoHugging Face04Congo-digital-service /audios-lingala-annotatees Annotated Lingala Dataset – Full Version Description This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models. It includes: the original audio files (viewable directly in the Hugging Face viewer) text transcriptions Mel spectrograms tokenized labels Overall statistics Metric Value Total volume 5 h 0 min 18 s Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.audioautomatic-speech-recognition10K<n<100K0 likes318 downloads17d agoHugging Face05humair025 /Rasa-Annotated-25kHz Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 530.7 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.tabulartext-to-speech10K<n<100K0 likes208 downloads8mo agoHugging Face06humair025 /Rasa-Annotated-V1 Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 264.5 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.tabulartext-to-speech10K<n<100K0 likes163 downloads8mo agoHugging Face07Congo-digital-service /audios-lingala-annotatees-v2 Annotated Lingala Audio — canonical corpus Annotated Lingala speech for open automatic speech recognition research and for fine-tuning speech models. This release is a full reconstruction of the corpus from its source recordings and annotations. It supersedes Congo-digital-service/audios-lingala-annotatees, which is deprecated — see Relationship to the previous release below. What this dataset contains Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.audioautomatic-speech-recognition10K<n<100K0 likes163 downloads13d agoHugging Face08ebellob /annotated_catalan_common_voice_v17_cleaned_enhanced Processed Annotated Catalan Common Voice v17 (CleanUNet + FlashSR) Dataset Summary This dataset is a processed and enhanced version of: projecte-aina/annotated_catalan_common_voice_v17. Furthermore, as this is a personal project, we give no guarantees that the audio is completely clean from any artifacts or noise the CleanUNet model could not remove. However, we have personally tested the corpus via the fine-tuning of some SOTA speech models and the results have been… See the full description on the dataset page: https://huggingface.co/datasets/ebellob/annotated_catalan_common_voice_v17_cleaned_enhanced.audiotext-to-speech100K<n<1M2 likes50 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.