projecte-aina
CATalog
Dataset Summary
CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words.
Supported Tasks and Leaderboards
Fill-Mask
Text Generation
other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.COPA-ca
Dataset Card for COPA-ca
Dataset Summary
The COPA-ca dataset (Choice of plausible alternatives in Catalan) is a professional translation of the English COPA dataset into Catalan, commissioned by BSC LangTech Unit. The dataset consists of 1000 premises, each given a question and two choices with a label encoding which of the choices is more plausible given the annotator.
The dataset is split into 400 training samples, 100 validation samples, and 500 test samples. It… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/COPA-ca.parlament_parla_v3
Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions
A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.4catac
Dataset Card for 4catac
Dataset Summary
4catac: examples of phonetic transcription in 4 Catalan accents is a dataset of phonetic transcriptions in four Catalan accents: Balearic, Central, North-Western and Valencian.
It consists of 160 sentences transcribed using IPA, following the recommendations of the Institut d'Estudis Catalans.
These sentences are the same for the four accents but may have small morphological adaptations to make them more natural for the accent.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/4catac.festcat_trimmed_denoised
Dataset Card for festcat_trimmed_denoised
This is a post-processed version of the Catalan Festcat speech dataset.
The original data can be found here.
Same license is maintained: Creative Commons Attribution-ShareAlike 3.0 Spain License.
Dataset Details
Dataset Description
We processed the data of the Catalan Festcat with the following recipe:
Trimming: Long silences from the start and the end of clips have been removed.
py-webrtcvad -> Python interface to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/festcat_trimmed_denoised.
