CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rishiraj /open-large-bengali-asr-data Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice audioautomatic-speech-recognition100K<n<1M0 likes498 downloads3mo agoHugging Face02SKNahin /open-large-bengali-asr-data Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice openslr madasr shrutilipi flerus kathbath indictts ucla gali audioautomatic-speech-recognition1M<n<10M13 likes479 downloads3y agoHugging Face03nymtheescobar /bengali-talkshow-audio Bengali Talkshow Audio Dataset A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs. Dataset Description This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audioaudio-classification1K<n<10K0 likes294 downloads8mo agoHugging Face04kawshikbuet17 /bengali-telecom-customer-care-speech-v2 Bengali Telecom Customer Care Synthetic Speech Dataset v2 Dataset Description This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts. The dataset is intended for experiments with: Bengali ASR/STT Bengali TTS Speech-to-text preprocessing Telecom/customer-care domain adaptation Synthetic speech research This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.audiotext-to-speech1K<n<10K0 likes71 downloads3mo agoHugging Face05bengaliAI /ben10-asr-results Ben-10 Regional ASR — public results Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard. Field Meaning model_id Hub id or slug model_url Link to weights / paper wer Corpus Word Error Rate on private ben-10-test (lower better) wer_by_region JSON map region → WER backend Decode stack used by maintainers scorer_commit / decode_commit Git SHAs in BengaliAI/reg-speech-aacl evaluated_at ISO date requested_by Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.tabularautomatic-speech-recognitionn<1K0 likes38 downloads2mo agoHugging Face06sha1779 /bengali_regional_dataset_refineThis is the dataset of ভাষা-বিচিত্রা: ASR for Regional Dialects competition. here i preprocessed and make train and eval split. this dataset consist of 10 dialact named 'barishal', 'chittagong', 'habiganj', 'kishoreganj', 'narail', 'narsingdi', 'rangpur', 'sandwip', 'sylhet', 'tangail'. barishal district has 796 samples chittagong district has 1406 samples habiganj district has 940 samples kishoreganj district has 1638 samples narail district has 1488 samples narsingdi district has… See the full description on the dataset page: https://huggingface.co/datasets/sha1779/bengali_regional_dataset_refine.textautomatic-speech-recognition10K<n<100K0 likes32 downloads2y agoHugging Face07InfoBayAI /Bengali_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes30 downloads9d agoHugging Face08parambharat /bengali_asr_corpusThe corpus contains roughly 500 hours of audio and transcripts in Bangla language. The transcripts have beed de-duplicated using exact match deduplication and audio has be converted to 16000 samplesautomatic-speech-recognition100K<n<1M0 likes29 downloads3y agoHugging Face09InfoBayAI /Bengali-Call-Center-Audio-Dataset-Single-ChannelgatedDataset Description: This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali-Call-Center-Audio-Dataset-Single-Channel.audioautomatic-speech-recognitionn<1K0 likes26 downloads9d agoHugging Face10IntisarUddin /Bengali_Long_form_ASR Bengali Long-Form ASR Dataset Dataset Summary The Bengali Long-Form ASR Dataset is a large-scale collection of long-duration Bangla speech recordings paired with verified transcripts. The dataset is designed specifically for long-form Automatic Speech Recognition (ASR) research. Key Statistics Total duration: 310.06 hours Number of recordings: 382 Average duration per recording: ~48.7 minutes Language: Bengali (bn) Audio format: WAV Sampling rate: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/IntisarUddin/Bengali_Long_form_ASR.audioautomatic-speech-recognitionn<1K0 likes24 downloads6mo agoHugging Face11kawshikbuet17 /bengali-telecom-customer-care-speech Bengali Telecom Customer Care Synthetic Speech Dataset Dataset Description This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts. The dataset is intended for experiments with: Bengali ASR/STT Bengali TTS Speech-to-text preprocessing Telecom/customer-care domain adaptation Synthetic speech research Important Disclosure This is a synthetic speech dataset generated using the OmniVoice TTS system in… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech.audiotext-to-speech10K<n<100K0 likes22 downloads4mo agoHugging Face12InfoBayAI /Bengali_Podcast_Audio_Dataset_Dual_Channelgated Dataset Description This dataset is a large-scale collection of 7,798 hours of processed Bengali dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali_Podcast_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes20 downloads9d agoHugging Face13psdn-ai /bengali-multi-speaker-speech-samplesgated Bengali Speech: Multi-Speaker Samples This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery. What This Shows Multi-speaker Bengali speech with transcript alignment Conversation-style audio rather than isolated prompt reading Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.audioautomatic-speech-recognitionn<1K0 likes17 downloads3mo agoHugging Face14smam /bengali-tts-folderized-parquet-stage1gated Bengali TTS Folderized Parquet Stage 1 This is the intermediate organized parquet layer before the final fully row-wise Bengali TTS dataset. Generated metadata refresh: 2026-06-18T21:44:49Z Layout <speaker>/part-00000.parquet <speaker>/part-00001.parquet <speaker>/metadata/stats.json Columns speaker video_id chunk_file audio_file duration transcription uuid audio as Hugging Face Audio feature backed by parquet struct<bytes,path>… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-folderized-parquet-stage1.audioautomatic-speech-recognition100K<n<1M1 likes13 downloads3mo agoHugging Face15dipit099 /bengali-tts-folderized-parquet-stage2gated Bengali TTS Folderized Parquet Stage 2 Final filtered (keep=True) Bengali TTS dataset, with merged/combined chunks. Layout <speaker>.parquet (single shard, audio <= ~1GB) <speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size) Audio sourcing convention COMBINED == False -> sourced from extracted_audio/<speaker>/<video_id>/<chunk_file> COMBINED == True -> sourced from… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-folderized-parquet-stage2.audioautomatic-speech-recognition100K<n<1M0 likes13 downloads3mo agoHugging Face16FrancophonIA /17-minute-world-languages_bengali [!NOTE] Dataset origin: https://www.17-minute-world-languages.com/fr/bengali/ Site à scrapper translation0 likes12 downloads1y agoHugging Face17Speech-data /bengali-speech-dataset 🎧 Bengali Speech Dataset The Bengali Speech Dataset is a high-quality speech audio dataset designed to support advanced AI and machine learning systems with reliable audio data. It provides structured voice data for building and evaluating modern speech technologies, including conversational AI and multilingual models. The dataset contains 156 hours of audio data across 643 files, delivered in MP3 and WAV formats, with a total size of 285 MB, making it a scalable resource for… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/bengali-speech-dataset.audioautomatic-speech-recognitionn<1K0 likes8 downloads6mo agoHugging Face18psdn-ai /bengali-speech-samplesgated Bengali Speech Samples This sample shows Bengali read and conversational speech with paired transcripts. It is meant to help buyers review spoken content, transcript alignment, and audio consistency before scoping a larger delivery. What This Shows Bengali speech across read and conversational styles Clip-level transcript alignment Audio metadata that supports format and quality review Dataset Specifications Field Value Modality Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-speech-samples.audioautomatic-speech-recognitionn<1K0 likes7 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.