CoolFace
Datasetpublic

SKNahin/open-large-bengali-asr-data

Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice openslr madasr shrutilipi flerus kathbath indictts ucla gali

sourceHugging Faceupdated 3y agoView on Hugging Face
13likes499downloads
settings

This repository belongs to SKNahin on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameopen-large-bengali-asr-data
visibilitypublic
licencenot set
gatedno
ownerSKNahin
Account settings
SKNahin/open-large-bengali-asr-data · CoolFace