Bengali
Datasets
All datasets matching “Bengali”bengali-ocr-synthetic
Bengali OCR Synthetic Dataset
A high-quality synthetic Bengali OCR dataset for fine-tuning vision-language models like DeepSeek-OCR 2. Generated using 100+ professional Bengali Unicode fonts and 13K+ unique Bengali words with advanced text rendering via FreeType and HarfBuzz.
Dataset Overview
Language: Bengali (বাংলা)
Task: Optical Character Recognition (OCR)
Format: Conversation-based (vision-language)
Total Samples: 30,000
Train: 27,007 samples
Validation: 2,993… See the full description on the dataset page: https://huggingface.co/datasets/rifathridoy/bengali-ocr-synthetic.shrutilipi_bengali
Dataset Card for "shrutilipi_bengali"
More Information needed
winograndeIndicTTS_Bengali
Bengali Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Bengali
Total Duration: ~15.06 hours (Male: 10.05 hours, Female: 5.01 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Bengali.open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
SIQA
