CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face02laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face03mitermix /audiosnippetsaudio1M<n<10M6 likes4.4k downloads2y agoHugging Face04mitermix /audiosnippets_small_with_detailed_annotationaudio100K<n<1M1 likes2.7k downloads2y agoHugging Face05mitermix /audiosnippets_small_with_detailed_annotation2audio1M<n<10M1 likes2.7k downloads2y agoHugging Face06mitermix /audiosnippets_long_2_5Maudio1M<n<10M3 likes1.4k downloads2y agoHugging Face07mitermix /audiosnippets_long_1Maudio100K<n<1M0 likes823 downloads2y agoHugging Face08mkrausio /audiosnippets-cleaned Dataset Summary This dataset is a processed version of mitermix/audiosnippets. The dataset contains audio snippets that have been cleaned and resampled, making it suitable for tasks like audio captioning, audio classification, or other audio-based machine learning applications. Processing Details Transcriptions and broken characters were removed. All MP3 audio files were resampled to 16kHz for consistency. The accompanying JSON metadata was made consistent. Entries with… See the full description on the dataset page: https://huggingface.co/datasets/mkrausio/audiosnippets-cleaned.audio1M<n<10M3 likes754 downloads2y agoHugging Face09sleeping-ai /MemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata. We share these files as-part of research initiative. audio100K<n<1M0 likes680 downloads1y agoHugging Face10cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes568 downloads2y agoHugging Face11Zhaowc /AudioCapsaudio100K<n<1M2 likes442 downloads1y agoHugging Face12ssz1111 /SpokenWOZ-Train-Audioaudio1K<n<10K0 likes322 downloads9mo agoHugging Face13laion /talent_plus_rl_groups_of_50_with_audiobox_scoresaudio1M<n<10M0 likes293 downloads10mo agoHugging Face14freococo /voa_myanmar_asr_audio_1 📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology. Overview This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.audioautomatic-speech-recognition1M<n<10M1 likes214 downloads1y agoHugging Face15seungheondoh /cmd-audio-dumpaudio10K<n<100K1 likes194 downloads1y agoHugging Face16laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes190 downloads9mo agoHugging Face17shenyunhang /AudioQA-1Maudio1M<n<10M2 likes177 downloads11mo agoHugging Face18overfitprolabse /subjective_audio_quality Balanced Perceptual Audio Quality Dataset Dataset Summary This is a large-scale, balanced dataset designed for training models for perceptual audio quality assessment. It consists of 612,020 examples, each containing a pair of 1-second audio clips: a high-quality original and a degraded version processed by various audio codecs. Each pair is accompanied by a perceptual quality score (ranging from 0.0 to 1.0) generated by visqol-like algorithms. The key feature of this… See the full description on the dataset page: https://huggingface.co/datasets/overfitprolabse/subjective_audio_quality.audio100K<n<1M1 likes170 downloads11mo agoHugging Face19mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes139 downloads1y agoHugging Face20mitermix /audiosnippets_long_50kaudio10K<n<100K0 likes137 downloads2y agoHugging Face21sheng22213 /speech_text-tts_audioaudio10K<n<100K0 likes136 downloads1y agoHugging Face22acul3 /Audiobook_Noice_V2audio10K<n<100K0 likes128 downloads2y agoHugging Face23krishnakalyan3 /laion-audio-preview-splitaudio1M<n<10M2 likes89 downloads2y agoHugging Face24cahya /laion-audio-smallaudio10K<n<100K0 likes77 downloads2y agoHugging Face25laion /audioset-with-captionsaudio1M<n<10M2 likes75 downloads11mo agoHugging Face26freococo /rohingya_asr_audioThis is the first public Rohingya language ASR dataset in AI history. Overview This dataset contains broadcast audio recordings from the Voice of America (VOA) Rohingya Service. Each file represents a daily news segment, typically 30 minutes in length, automatically segmented into chunks of 5–15 seconds for use in self-supervised ASR, pretraining, language identification, and more. The content was aired publicly as part of VOA’s Rohingya-language radio program and is therefore… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rohingya_asr_audio.audioautomatic-speech-recognition100K<n<1M2 likes61 downloads1y agoHugging Face27cubbk /audio_swedish_2_dataset_cleanedaudio1K<n<10K0 likes56 downloads1y agoHugging Face28freococo /mon_language_asr_audio RFA Mon Language Voices This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. This dataset was created by freococo. The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.audioautomatic-speech-recognition1K<n<10K0 likes43 downloads1y agoHugging Face29PuristanLabs1 /urdu-turn-detection-audio-v2 🗣️ Urdu Turn Detection (Audio Dataset V2) This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech. It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection. 🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.audioaudio-classification10K<n<100K0 likes42 downloads9mo agoHugging Face30freococo /karenni_language_asr_audio RFA Karenni (Kayah) Language Voices This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks. This dataset was created by freococo. The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.audioautomatic-speech-recognition1K<n<10K0 likes37 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.