datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CAESAR-TV3
Dataset card for CAESAR-TV3
Dataset Summary
This corpus includes 5 hours and 45 minutes of Catalan speech code-switched with Spanish extracted from the original tv3_parla dataset.
Supported Tasks and Leaderboards
The CAESAR-TV3 dataset is designed for the Automatic Speech Recognition (ASR) task, enabling the transcription of utterances in Catalan, Spanish, and code-switched speech between the two languages.
Languages
The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CAESAR-TV3.CommonPhone-SE
CommonPhone-SE
Multilingual, age and gender balanced subset for speech enhancement benchmark.
Dataset Details
Dataset Description
Commonphone-SE is a benchmark dataset derived from Commonphone. It contains audio samples from 7 languages in the age range from 18 to 80. It aims to
provide a speaker diverse dataset to benchmark speech enhancement algorithms in real world conditions.
Curated by: LangTech Lab members from the speech team.
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CommonPhone-SE.CAESAR-TINY
Dataset Card for CAESAR-TINY
Dataset Summary
CAESAR-TINY is a synthetic code-switched dataset generated by combining monolingual samples in Catalan and Spanish.
The process includes trimming silences, normalizing audio volume, and introducing random pauses. It contains 2 hours of speech data, created by concatenating audio from the Common voice 17 Benchmark split and VoxForge Spanish datasets.
Example Usage
To load CAESAR-TINY:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CAESAR-TINY.
