BSC-LT/CAESAR-TV3
Dataset card for CAESAR-TV3 Dataset Summary This corpus includes 5 hours and 45 minutes of Catalan speech code-switched with Spanish extracted from the original tv3_parla dataset. Supported Tasks and Leaderboards The CAESAR-TV3 dataset is designed for the Automatic Speech Recognition (ASR) task, enabling the transcription of utterances in Catalan, Spanish, and code-switched speech between the two languages. Languages The dataset… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CAESAR-TV3.
157
