CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01psdn-ai /bangla-10kgated Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh Bangla-10K is a 10,816-hour Bengali speech corpus with 624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus (567,323 recordings) and a separately collected 745.1-hour evaluation set (57,628 recordings). It combines scripted single-speaker read speech with natural multi-speaker conversations for Bengali automatic speech recognition (ASR). The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.audioautomatic-speech-recognition100K<n<1M0 likes342 downloads2d agoHugging Face02Suprio85 /Bangla_Speech_Corpus 🎙️ Bengali-Loop: A Long-Form Bangla Speech Corpus Dataset Summary Bengali-Loop is a large-vocabulary, long-form Bangla (Bengali) speech corpus designed to push the boundaries of Automatic Speech Recognition (ASR) in low-to-mid resource settings. It comprises 155 hours of naturally occurring Bangla speech sourced from 249 YouTube videos spanning drama serials, audiobooks, and entertainment channels — making it one of the most diverse publicly available Bangla ASR datasets… See the full description on the dataset page: https://huggingface.co/datasets/Suprio85/Bangla_Speech_Corpus.audioautomatic-speech-recognition2 likes246 downloads7mo agoHugging Face03Suprio85 /Bangla_speech_corpus-321 🎙️ BanglaSpeechCorpus-321: Large-Scale Long-Form Bangla Speech Corpus Dataset Summary BanglaSpeechCorpus-321 is an extended, large-scale Bangla (Bengali) speech corpus for Automatic Speech Recognition (ASR), featuring 321.2 hours of naturally occurring Bangla speech across 401 recordings. This is the expanded successor to Bangla_Speech_Corpus, covering a broader set of YouTube channels including drama serials, audiobooks, and entertainment content. With over 303,000… See the full description on the dataset page: https://huggingface.co/datasets/Suprio85/Bangla_speech_corpus-321.audioautomatic-speech-recognition1K<n<10K1 likes235 downloads7mo agoHugging Face04akhikhan123 /BanglaEnglishMixedAsrDatasetautomatic-speech-recognition100K<n<1M1 likes220 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.