Racoci/CORAA-v1.1
CORAA-v1.1 CORAA-v1.1 is a publicly available dataset for Automatic Speech Recognition (ASR) in the Brazilian Portuguese language containing 290.77 hours of audios and their respective transcriptions (400k+ segmented audios). The dataset is composed of audios of 5 original projects: ALIP (Gonçalves, 2019) C-ORAL Brazil (Raso and Mello, 2012) NURC-Recife (Oliviera Jr., 2016) SP-2010 (Mendes and Oushiro, 2012) TEDx talks (talks in Portuguese) The audios were either validated by… See the full description on the dataset page: https://huggingface.co/datasets/Racoci/CORAA-v1.1.
Upload LICENSE file and notebook used to upload the dataset
Upload CORAA dataset on Hugging Face format
Upload dataset (part 00002-of-00003)
Upload dataset (part 00001-of-00003)
Upload dataset (part 00000-of-00003)
initial commit
