datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
openslr
madasr
shrutilipi
flerus
kathbath
indictts
ucla
gali
bengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.bengali_regional_dataset_refineThis is the dataset of ভাষা-বিচিত্রা: ASR for Regional Dialects competition.
here i preprocessed and make train and eval split.
this dataset consist of 10 dialact named 'barishal', 'chittagong', 'habiganj', 'kishoreganj', 'narail',
'narsingdi', 'rangpur', 'sandwip', 'sylhet', 'tangail'.
barishal district has 796 samples
chittagong district has 1406 samples
habiganj district has 940 samples
kishoreganj district has 1638 samples
narail district has 1488 samples
narsingdi district has… See the full description on the dataset page: https://huggingface.co/datasets/sha1779/bengali_regional_dataset_refine.bengali-telecom-customer-care-speech
Bengali Telecom Customer Care Synthetic Speech Dataset
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
Important Disclosure
This is a synthetic speech dataset generated using the OmniVoice TTS system in… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.bengali-tts-folderized-parquet-stage1
Bengali TTS Folderized Parquet Stage 1
This is the intermediate organized parquet layer before the final fully row-wise Bengali TTS dataset.
Generated metadata refresh: 2026-06-18T21:44:49Z
Layout
<speaker>/part-00000.parquet
<speaker>/part-00001.parquet
<speaker>/metadata/stats.json
Columns
speaker
video_id
chunk_file
audio_file
duration
transcription
uuid
audio as Hugging Face Audio feature backed by parquet struct<bytes,path>… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-folderized-parquet-stage1.bengali-tts-folderized-parquet-stage2
Bengali TTS Folderized Parquet Stage 2
Final filtered (keep=True) Bengali TTS dataset, with merged/combined chunks.
Layout
<speaker>.parquet (single shard, audio <= ~1GB)
<speaker>_001.parquet, _002.parquet, ... (multiple shards, split by audio byte size)
Audio sourcing convention
COMBINED == False -> sourced from extracted_audio/<speaker>/<video_id>/<chunk_file>
COMBINED == True -> sourced from… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-folderized-parquet-stage2.bengali-speech-samples
Bengali Speech Samples
This sample shows Bengali read and conversational speech with paired transcripts. It is meant to help buyers review spoken content, transcript alignment, and audio consistency before scoping a larger delivery.
What This Shows
Bengali speech across read and conversational styles
Clip-level transcript alignment
Audio metadata that supports format and quality review
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-speech-samples.
