datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.escagleu-64k
Dataset Card for escagleu-64K corpus
Dataset Description
Dataset Summary
This is the second version of escagleu-64k, a parallel corpus containing approximately 64k sentences translated across Spanish, Catalan, Valencian Catalan, Galician, and Basque.
The original sentences are in Spanish and are sourced from the Spanish Common Voice Corpus.
This corpus was prepared with the goal of creating a parallel speech dataset for these languages using the Common Voice… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/escagleu-64k.commonvoice_benchmark_catalan_accentsA new presentation of the corpus Catalan Common Voice v17.0 - metadata annotated version with the splits redefined to benchmark ASR models with various Catalan accentparlament_parlaThis is the ParlamentParla speech corpus for Catalan prepared by Col·lectivaT. The audio segments were extracted from recordings the Catalan Parliament (Parlament de Catalunya) plenary sessions, which took place between 2007/07/11 - 2018/07/17. We aligned the transcriptions with the recordings and extracted the corpus. The content belongs to the Catalan Parliament and the data is released conforming their terms of use.
Preparation of this corpus was partly supported by the Department of Culture of the Catalan autonomous government, and the v2.0 was supported by the Barcelona Supercomputing Center, within the framework of the project AINA of the Departament de Polítiques Digitals.
As of v2.0 the corpus is separated into 211 hours of clean and 400 hours of other quality segments. Furthermore, each speech segment is tagged with its speaker and each speaker with their gender. The statistics are detailed in the readme file.
For more information, go to https://github.com/CollectivaT-dev/ParlamentParla or mail info@collectivat.cat.corts_valencianes_asr_aThis is the first version of CortsValencianes speech corpus for Valencian: a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications.annotated_catalan_common_voice_v17This version of the Catalan sentences of the Common Voice corpus v17
includes metadata (gender and accent) for 263 speakers annotated by a team of experts.cv17_es_other_automatically_verifiedSplit called -other- of the Spanish Common Voice v17.0 that was automatically verified
using various ASR system.
