datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shrutilipi_bengali
Dataset Card for "shrutilipi_bengali"
More Information needed
IndicTTS_Bengali
Bengali Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Bengali
Total Duration: ~15.06 hours (Male: 10.05 hours, Female: 5.01 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Bengali.open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
openslr
madasr
shrutilipi
flerus
kathbath
indictts
ucla
gali
SPRING_INX_Bengali_R1SPRING_INX_Bengali_R2bengali_audio_files
Dataset Card for "bengali_audio_files"
More Information needed
Bengali_Competition_Datasetbengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.Ben-10bengali-tts-dataset-196Kbengali-ai-asr-80kbengali-diarization-synthetic-v3bengali-speech-datasetIndicTTS_BengaliOLD
Bengali Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Bengali
Total Duration: ~15.06 hours (Male: 10.05 hours, Female: 5.01 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/Abdullah500/IndicTTS_BengaliOLD.IndicVoices-R_Bengalibengali-tts-filtered-v3-backup
Bengali TTS Filtered Dataset (v3)
Quality-filtered subset of rwd51/bengali-tts-combined.
Filtering Criteria
Rows with has_english=True -> auto-KEEP (no WER/CER check)
Rows with has_english=False -> KEEP if WER <= 0.5 AND CER <= 0.3
~293,097 rows expected
Columns
Column
Description
uuid
Unique identifier
speaker
Speaker name
video_id
Video ID
chunk_file
Chunk filename
audio_file
Audio filename
duration
Duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-filtered-v3-backup.bengali-asr-dataoriginal_data_bengali_ttsbengali-ai-asr-datasetIndicTTS-Bengali
Bengali Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Bengali
Audio Format: WAV
Sampling Rate: 48000Hz
Speakers: 4 (2 male, 2 female native Bengali speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abdullah500/IndicTTS-Bengali.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.bengaliai-speech-datasetbengali-asr-chunkedbengali-tts-missing-v1
Bengali TTS — Missing Rows
This dataset contains the rows from rwd51/bengali-tts-combined that are not present in the filtered repository dipit099/bengali-tts-filtered-v3.
In other words: source samples that were never scored by Whisper/Gemini in the original transcription pipeline and therefore never made it into the quality-filtered output.
Source
rwd51/bengali-tts-combined (397,437 rows, 780 parquet shards)
dipit099/bengali-tts-filtered-v3 (326,647 KEEP rows)… See the full description on the dataset page: https://huggingface.co/datasets/smam/bengali-tts-missing-v1.indictts_bengali
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/avipri/indictts_bengali.ben10-asr-results
Ben-10 Regional ASR — public results
Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard.
Field
Meaning
model_id
Hub id or slug
model_url
Link to weights / paper
wer
Corpus Word Error Rate on private ben-10-test (lower better)
wer_by_region
JSON map region → WER
backend
Decode stack used by maintainers
scorer_commit / decode_commit
Git SHAs in BengaliAI/reg-speech-aacl
evaluated_at
ISO date
requested_by
Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.Bengali_AI_Speech
Dataset Card for "Bengali_AI_Speech"
More Information needed
indicvoices-bengali
Dataset Description
This dataset has been collected from IndicVoices-R github repository. Only the Bengali language portion is in this dataset.
Train Duration: 393985.0155 seconds
Test Duration: 2719.5667 seconds
bengali-tts-yt-test
Bengali TTS — YouTube Pipeline (test)
YouTube-sourced Bengali speech chunks transcribed with Gemini 2.5 Flash, produced by the YT scraping + VAD chunking + Gemini transcription pipeline.
Columns
Column
Description
uuid
Unique identifier (<video_id>_<chunk#>)
speaker
Speaker / channel name
video_id
YouTube video ID
chunk_file
Chunk filename
audio_file
<speaker>_<video_id>_<chunk_file>
duration
Duration in seconds
transcription
Gemini… See the full description on the dataset page: https://huggingface.co/datasets/dipit099/bengali-tts-yt-test.
