datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
4catac
Dataset Card for 4catac
Dataset Summary
4catac: examples of phonetic transcription in 4 Catalan accents is a dataset of phonetic transcriptions in four Catalan accents: Balearic, Central, North-Western and Valencian.
It consists of 160 sentences transcribed using IPA, following the recommendations of the Institut d'Estudis Catalans.
These sentences are the same for the four accents but may have small morphological adaptations to make them more natural for the accent.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/4catac.GuiaCat
Dataset Card for GuiaCat
Dataset Summary
GuiaCat is a dataset consisting of 5.750 restaurant reviews in Catalan, with 5 associated scores and a label of sentiment. The data was provided by GuiaCat and curated by the BSC.
This work is licensed under a Creative Commons Attribution Non-commercial No-Derivatives 4.0 International License.
Supported Tasks and Leaderboards
This corpus is mainly intended for sentiment analysis.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/GuiaCat.sts-ca
Dataset Card for STS-ca
Dataset Summary
STS-ca corpus is a benchmark for evaluating Semantic Text Similarity in Catalan. This dataset was developed by BSC TeMU as part of Projecte AINA, to enrich the Catalan Language Understanding Benchmark (CLUB).
This work is licensed under a Attribution-ShareAlike 4.0 International License.
Supported Tasks and Leaderboards
This dataset can be used to build and score semantic similarity models in Catalan.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/sts-ca.CA-EN_Parallel_Corpus
Dataset Card for CA-EN Parallel Corpus
Dataset Description
Dataset Summary
The CA-EN Parallel Corpus is a Catalan-English dataset of parallel sentences created to
support Catalan in NLP tasks, specifically Machine Translation.
Supported Tasks and Leaderboards
The dataset can be used to train Bilingual Machine Translation models between English and Catalan in any direction,
as well as Multilingual Machine Translation models.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EN_Parallel_Corpus.
