CoolFace
Datasetpublic

SKNahin/open-large-bengali-asr-data

Open Large Bengali ASR Data This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps). Datasets: commonvoice openslr madasr shrutilipi flerus kathbath indictts ucla gali

sourceHugging Faceupdated 3y agoView on Hugging Face
13likes507downloads
9 commits on main
b779cb63y ago

Update README.md

SKNahin
eb3d9203y ago

Update README.md

SKNahin
6150ddd3y ago

Update README.md

SKNahin
fd7b6983y ago

Update README.md

SKNahin
fe3dc9b3y ago

Update README.md

SKNahin
83eda1f3y ago

Upload dataset (part 00002-of-00003)

SKNahin
cc0d8f33y ago

Upload dataset (part 00001-of-00003)

SKNahin
30a56a43y ago

Upload dataset (part 00000-of-00003)

SKNahin
4467f4d3y ago

initial commit

SKNahin
SKNahin/open-large-bengali-asr-data · CoolFace