datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
openslr
madasr
shrutilipi
flerus
kathbath
indictts
ucla
gali
bengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.bengali_regional_dataset_refineThis is the dataset of ভাষা-বিচিত্রা: ASR for Regional Dialects competition.
here i preprocessed and make train and eval split.
this dataset consist of 10 dialact named 'barishal', 'chittagong', 'habiganj', 'kishoreganj', 'narail',
'narsingdi', 'rangpur', 'sandwip', 'sylhet', 'tangail'.
barishal district has 796 samples
chittagong district has 1406 samples
habiganj district has 940 samples
kishoreganj district has 1638 samples
narail district has 1488 samples
narsingdi district has… See the full description on the dataset page: https://huggingface.co/datasets/sha1779/bengali_regional_dataset_refine.Bengali_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali_Call_Center_Audio_Dataset_Dual_Channel.bengali_asr_corpusThe corpus contains roughly 500 hours of audio and transcripts in Bangla language.
The transcripts have beed de-duplicated using exact match deduplication and audio has be converted to 16000 samplesBengali-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 377,909 hours of processed Bengali (BN) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali-Call-Center-Audio-Dataset-Single-Channel.Bengali_Long_form_ASR
Bengali Long-Form ASR Dataset
Dataset Summary
The Bengali Long-Form ASR Dataset is a large-scale collection of long-duration Bangla speech recordings paired with verified transcripts. The dataset is designed specifically for long-form Automatic Speech Recognition (ASR) research.
Key Statistics
Total duration: 310.06 hours
Number of recordings: 382
Average duration per recording: ~48.7 minutes
Language: Bengali (bn)
Audio format: WAV
Sampling rate: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/IntisarUddin/Bengali_Long_form_ASR.bengali-telecom-customer-care-speech
Bengali Telecom Customer Care Synthetic Speech Dataset
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
Important Disclosure
This is a synthetic speech dataset generated using the OmniVoice TTS system in… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech.Bengali_Podcast_Audio_Dataset_Dual_Channel
Dataset Description
This dataset is a large-scale collection of 7,798 hours of processed Bengali dual-channel podcast audio recordings, containing 57,569 hours of processed podcast audio recordings across 12 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It captures real-world podcast conversations across diverse topics and formats. The dataset is organized in a dual-channel format, where corresponding speaker… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Bengali_Podcast_Audio_Dataset_Dual_Channel.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.bengali-tts-folderized-parquet-stage1
Bengali TTS Folderized Parquet Stage 1
This is the intermediate organized parquet layer before the final fully row-wise Bengali TTS dataset.
Generated metadata refresh: 2026-06-18T21:44:49Z
Layout
<speaker>/part-00000.parquet
<speaker>/part-00001.parquet
<speaker>/metadata/stats.json
Columns
speaker
video_id
chunk_file
audio_file
duration
transcription
uuid
audio as Hugging Face Audio feature backed by parquet struct<bytes,path>… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-folderized-parquet-stage1.bengali-tts-folderized-parquet-stage2
Bengali TTS Folderized Parquet Stage 2
Final filtered (keep=True) Bengali TTS dataset, with merged/combined chunks.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
COMBINED == False -> sourced from extracted_audio/<speaker>/<video_id>/<chunk_file>
COMBINED == True -> sourced from… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-folderized-parquet-stage2.17-minute-world-languages_bengali
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/bengali/
Site à scrapper
bengali-speech-dataset
🎧 Bengali Speech Dataset
The Bengali Speech Dataset is a high-quality speech audio dataset designed to support advanced AI and machine learning systems with reliable audio data. It provides structured voice data for building and evaluating modern speech technologies, including conversational AI and multilingual models. The dataset contains 156 hours of audio data across 643 files, delivered in MP3 and WAV formats, with a total size of 285 MB, making it a scalable resource for… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/bengali-speech-dataset.bengali-speech-samples
Bengali Speech Samples
This sample shows Bengali read and conversational speech with paired transcripts. It is meant to help buyers review spoken content, transcript alignment, and audio consistency before scoping a larger delivery.
What This Shows
Bengali speech across read and conversational styles
Clip-level transcript alignment
Audio metadata that supports format and quality review
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-speech-samples.
