datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iban-speech
Iban Data collected by Sarah Samson Juan and Laurent Besacier
Prepared by Sarah Samson Juan and Laurent Besacier
Created in GETALP, Grenoble, France
INTRODUCTION
This package has iban text and speech corpora used for Automatic Speech Recognition (ASR) experiments. Data is available in the subdirectories of /data. The subdirectories contain:
a. train - train transcript for training ASR system using Kaldi ASR… See the full description on the dataset page: https://huggingface.co/datasets/SaLTUNIMAS/iban-speech.iban-whisper-format
Iban Whisper Format
Originally from https://github.com/sarahjuan/iban, we applied True Case and only selected audio that less than 12 seconds.
Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/speech-to-text/iban
