CoolFace
Datasetpublic

rishiraj/open-large-bengali-asr-data

Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes406downloads
Dataset Card

Open Large Bengali ASR Data

This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).

Datasets: