datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
catalanqa
Dataset Card for CatalanQA
Dataset Summary
This dataset can be used to build extractive-QA and Language Models. It is an aggregation and balancing of 2 previous datasets: VilaQuAD and ViquiQuAD.
Splits have been balanced by kind of question, and unlike other datasets like SQuAD, it only contains, per record, one question and one answer for each context, although the contexts can repeat multiple times.
This dataset was developed by BSC TeMU as part of Projecte AINA, to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalanqa.openbookqa_ca
Dataset Card for openbookqa_ca
openbookqa_ca is a question answering dataset in Catalan, professionally translated from the main version of the OpenBookQA dataset in English.
Dataset Details
Dataset Description
openbookqa_ca (Open Book Question Answering - Catalan) is designed to simulate open book exams and assess human-like understanding of a subject. The dataset comprises 500 instances in the validation split and another 500 instances in the test split.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/openbookqa_ca.mgsm_ca
Dataset Card for mgsm_ca
mgsm_ca is a question answering dataset in Catalan that has been professionally translated from the MGSM dataset in English.
Dataset Details
Dataset Description
mgsm_ca (Multilingual Grade School Math - Catalan) is designed to evaluate multi-step mathematical reasoning using grade school math word problems. It includes 8 instances in the train split and another 250
instances in the test split. Each instance contains a math problem… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/mgsm_ca.RAG_Multilingual
Dataset Card for RAG_Multilingual
Dataset Summary
RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets.
The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC).
This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.veritasQA
Dataset Card for VeritasQA
VeritasQA is a context- and time-independent QA benchmark for the evaluation of truthfulness in Language Models.
Dataset Summary
VeritasQA is a context- and time-independent truthfulness benchmark built with multilingual transferability in mind. It is intended to be used to evaluate Large Language Models on truthfulness in a zero-shot setting. VeritasQA comprises 353 question-answer pairs inspired by common misconceptions and falsehoods, not… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/veritasQA.arc_ca
Dataset Card for arc_ca
arc_ca is a question answering dataset in Catalan, professionally translated from the Easy and Challenge versions of the ARC dataset in English.
Dataset Details
Dataset Description
arc_ca (AI2 Reasoning Challenge - Catalan) is based on multiple-choice science questions at elementary school level. The dataset consists of 2950 instances in the Easy version (570 in the test and 2380 instances in the validation split) and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/arc_ca.piqa_ca
Dataset Card for piqa_ca
piqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the PIQA validation set in English.
Dataset Details
Dataset Description
piqa_ca (Physical Interaction Question Answering - Catalan) is designed to evaluate physical commonsense reasoning using question-answer triplets based on everyday situations. It includes 1838 instances in the validation split. Each instance contains… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/piqa_ca.xquad-ca
Dataset Card for XQuAD-Ca
Dataset Summary
Professional translation into Catalan of XQuAD dataset.
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance. The dataset consists of a subset of 240 paragraphs and 1190 question-answer pairs from the development set of SQuAD v1.1 (Rajpurkar, Pranav et al., 2016) together with their professional translations into ten language: Spanish, German, Greek… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xquad-ca.siqa_ca
Dataset Card for siqa_ca
siqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the SIQA
validation set in English.
Dataset Details
Dataset Description
siqa_ca (Social Interaction Question Answering - Catalan) is designed to evaluate social commonsense intelligence using multiple choice question-answer instances based on reasoning about people’s actions and their
social implications. It includes… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/siqa_ca.xstorycloze_ca
Dataset Card for xstorycloze_ca
xstorycloze_ca is a question answering dataset in Catalan, professionally translated from the English StoryCloze dataset (Spring 2016 version), used to create its multilingual version XStoryCloze.
Dataset Details
Dataset Description
xstorycloze_ca (Multilingual Story Cloze Test - Catalan) is based on multiple-choice narrative completions. The dataset consists of 360 instances in the train split and 1510 instances in the test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xstorycloze_ca.oasst1_ca
Dataset Card for oasst1_ca
oasst1_ca is a conversational dataset in Catalan that has been professionally translated from the OASST1 dataset.
Dataset Details
Dataset Description
oasst1_ca (OpenAssistant Conversations Release 1 - Catalan) consists of human-generated, human-annotated assistant-style conversation corpus. It includes 5213 messages in the train split and 273 messages in the validation split. To arrive to this number, we filter the dataset (See… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/oasst1_ca.CoQCat
Dataset Card for CoQCat
Dataset Summary
CoQCat is a dataset for Conversational Question Answering in Catalan. It is based on CoQA dataset.
CoQCat comprises 89,364 question-answer pairs, sourced from conversations related to 6,000 text passages from six different domains.
The questions and responses are designed to maintain a conversational tone.
The answers are presented in a free-form text format, with evidence highlighted from the passage.
For the development and test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CoQCat.MentorES
Dataset Summary
Mentor_ES is an open source dataset of 10,175 instructions in Spanish organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming.
Supported Tasks and Leaderboards
Useful for fine-tuning instructions in large language models for downstream tasks.
Languages
This dataset is in… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorES.viquiquad
ViquiQuAD: An Extractive QA Dataset for Catalan from Wikipedia
Dataset Summary
ViquiQuAD is an extractive Question Answering dataset for Catalan, built from the Catalan Wikipedia (Viquipèdia).
3,111 contexts extracted from 597 high-quality, original (non-translated) articles.
For each context, 1 to 5 questions were created with their corresponding answers.
Total: 15,153 question–answer pairs.
This dataset can be used to fine-tune and evaluate extractive QA models and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/viquiquad.IFEval_ca
Dataset Card for IFEval_ca
IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.vilaquad
Dataset Card for VilaQuAD
Dataset Summary
VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text.
This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context).
VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence.
This dataset can be used to build extractive-QA and Language Models.
Supported Tasks and Leaderboards
Extractive-QA… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilaquad.MentorCA
Dataset Summary
Mentor_CA is an open source dataset of 10,175 instructions in Catalan, machine translated from the original Mentor_ES dataset in Spanish, and organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming.
Supported Tasks and Leaderboards
Useful for fine-tuning instructions in large language… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorCA.hhh_alignment_ca
Dataset Card for hhh_alignment_ca
hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English.
Dataset Details
Dataset Description
hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.vinclat
Dataset Card for Vinclat
Vinclat is a Catalan-language dataset for multi-step problem solving. It employs a game-based structure to evaluate both the reasoning capabilities and cultural knowledge of large language models.
Language(s): Catalan
Paper: ACL Anthology
Leaderboard: HF Space
Code: GitHub
License: CC-BY-4.0
Funded by: Projecte Aina
Curated by: Barcelona Supercomputing CenterShared by: Barcelona Supercomputing Center
Dataset Description
Vinclat is a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vinclat.dolly3k_ca
Dataset Card for dolly3k_ca
dolly3k_ca is a question answering dataset in Catalan, professionally translated from a filtered version of databricks-dolly-15k dataset in English.
Dataset Details
Dataset Description
dolly3k_ca (Dolly 3K instances - Catalan) is based on question-answer instance pairs written by humans. The dataset consists of 3232 instances in the train split. Each instance contains an instruction or question, and one answer. Every instance is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/dolly3k_ca.
