datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RAG_Multilingual
Dataset Card for RAG_Multilingual
Dataset Summary
RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets.
The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC).
This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.teca
Dataset Card for TE-ca
Dataset Summary
TE-ca is a dataset of textual entailment in Catalan, which contains 21,163 pairs of premises and hypotheses, annotated according to the inference relation they have (implication, contradiction or neutral).
This is the second version of the dataset, released on 29/03/2025, where encoding mistakes from the initial release have been corrected in both the 'hypothesis' and 'premise' column. Check… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/teca.casum
Dataset Card for CaSum
Dataset Summary
CaSum is a summarization dataset. It is extracted from a newswire corpus crawled from the Catalan News Agency (Agència Catalana de Notícies; ACN). The corpus consists of 217,735 instances that are composed by the headline and the body.
Supported Tasks and Leaderboards
The dataset can be used to train a model for abstractive summarization. Success on this task is typically measured by achieving a high Rouge score. The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/casum.MentorES
Dataset Summary
Mentor_ES is an open source dataset of 10,175 instructions in Spanish organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming.
Supported Tasks and Leaderboards
Useful for fine-tuning instructions in large language models for downstream tasks.
Languages
This dataset is in… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorES.viquiquad
ViquiQuAD: An Extractive QA Dataset for Catalan from Wikipedia
Dataset Summary
ViquiQuAD is an extractive Question Answering dataset for Catalan, built from the Catalan Wikipedia (Viquipèdia).
3,111 contexts extracted from 597 high-quality, original (non-translated) articles.
For each context, 1 to 5 questions were created with their corresponding answers.
Total: 15,153 question–answer pairs.
This dataset can be used to fine-tune and evaluate extractive QA models and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/viquiquad.IFEval_ca
Dataset Card for IFEval_ca
IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.tecla
Dataset Card for TeCla
Dataset Summary
TeCla (Text Classification) is a Catalan News corpus for thematic multi-class Text Classification tasks. The present version (2.0) contains 113.376 articles classified under a hierarchical class structure consisting of a coarse-grained and a fine-grained class. Each of the 4 coarse-grained classes accept a subset of fine-grained ones, 53 in total.
The previous version (1.0.1) can still be found at https://zenodo.org/record/4761505… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/tecla.vilaquad
Dataset Card for VilaQuAD
Dataset Summary
VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text.
This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context).
VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence.
This dataset can be used to build extractive-QA and Language Models.
Supported Tasks and Leaderboards
Extractive-QA… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilaquad.vilasum
Dataset Card for VilaSum
Dataset Summary
VilaSum is a summarization dataset for evaluation. It is extracted from a newswire corpus crawled from the Catalan news portal VilaWeb. The corpus consists of 13,843 instances that are composed by the headline and the body.
Supported Tasks and Leaderboards
The dataset can be used to train a model for abstractive summarization. Success on this task is typically measured by achieving a high Rouge score. The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilasum.MentorCA
Dataset Summary
Mentor_CA is an open source dataset of 10,175 instructions in Catalan, machine translated from the original Mentor_ES dataset in Spanish, and organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming.
Supported Tasks and Leaderboards
Useful for fine-tuning instructions in large language… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorCA.InstruCAT
Dataset Card for Instrucat
Dataset Summary
InstruCat is a dataset consisting of 216,826 instructions in Catalan.
It contains data converted to instructions format from the following datasets:
caBreu : The instructions were created in form of summarization tasks. There are 2 types of summarization categories in the dataset: extreme and abstractive. The extreme one summarizes text into one sentence and the abstractive into shorter texts around 3-5 sentences.
CatalanQA :… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/InstruCAT.catalan_government_crawling
Dataset Card for Catalan Government Crawling
Dataset Summary
The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39,117,909 tokens, 1,565,433 sentences and 71,043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_government_crawling.vinclat
Dataset Card for Vinclat
Vinclat is a Catalan-language dataset for multi-step problem solving. It employs a game-based structure to evaluate both the reasoning capabilities and cultural knowledge of large language models.
Language(s): Catalan
Paper: ACL Anthology
Leaderboard: HF Space
Code: GitHub
License: CC-BY-4.0
Funded by: Projecte Aina
Curated by: Barcelona Supercomputing CenterShared by: Barcelona Supercomputing Center
Dataset Description
Vinclat is a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vinclat.CaWikiTC
Dataset Card for CaWikiTC
Dataset Summary
CaWikiTC (Catalan Wikipedia Text Classification) is a text classification dataset authomatically created by scraping Catalan Wikipedia article summaries and their associated thematic category. It contains 21002 texts (19952 and 1050 in the train and dev partitions, respectively) classified under 67 exclusive categories.
For the dataset creation, we selected all the Catalan Wikipedia article summaries from a previously fixed… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CaWikiTC.
