datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CATalog
Dataset Summary
CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words.
Supported Tasks and Leaderboards
Fill-Mask
Text Generation
other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.RAG_Multilingual
Dataset Card for RAG_Multilingual
Dataset Summary
RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets.
The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC).
This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.veritasQA
Dataset Card for VeritasQA
VeritasQA is a context- and time-independent QA benchmark for the evaluation of truthfulness in Language Models.
Dataset Summary
VeritasQA is a context- and time-independent truthfulness benchmark built with multilingual transferability in mind. It is intended to be used to evaluate Large Language Models on truthfulness in a zero-shot setting. VeritasQA comprises 353 question-answer pairs inspired by common misconceptions and falsehoods, not… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/veritasQA.xstorycloze_ca
Dataset Card for xstorycloze_ca
xstorycloze_ca is a question answering dataset in Catalan, professionally translated from the English StoryCloze dataset (Spring 2016 version), used to create its multilingual version XStoryCloze.
Dataset Details
Dataset Description
xstorycloze_ca (Multilingual Story Cloze Test - Catalan) is based on multiple-choice narrative completions. The dataset consists of 360 instances in the train split and 1510 instances in the test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xstorycloze_ca.oasst1_ca
Dataset Card for oasst1_ca
oasst1_ca is a conversational dataset in Catalan that has been professionally translated from the OASST1 dataset.
Dataset Details
Dataset Description
oasst1_ca (OpenAssistant Conversations Release 1 - Catalan) consists of human-generated, human-annotated assistant-style conversation corpus. It includes 5213 messages in the train split and 273 messages in the validation split. To arrive to this number, we filter the dataset (See… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/oasst1_ca.MentorES
Dataset Summary
Mentor_ES is an open source dataset of 10,175 instructions in Spanish organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming.
Supported Tasks and Leaderboards
Useful for fine-tuning instructions in large language models for downstream tasks.
Languages
This dataset is in… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorES.IFEval_ca
Dataset Card for IFEval_ca
IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.MentorCA
Dataset Summary
Mentor_CA is an open source dataset of 10,175 instructions in Catalan, machine translated from the original Mentor_ES dataset in Spanish, and organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming.
Supported Tasks and Leaderboards
Useful for fine-tuning instructions in large language… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorCA.hhh_alignment_ca
Dataset Card for hhh_alignment_ca
hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English.
Dataset Details
Dataset Description
hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.vinclat
Dataset Card for Vinclat
Vinclat is a Catalan-language dataset for multi-step problem solving. It employs a game-based structure to evaluate both the reasoning capabilities and cultural knowledge of large language models.
Language(s): Catalan
Paper: ACL Anthology
Leaderboard: HF Space
Code: GitHub
License: CC-BY-4.0
Funded by: Projecte Aina
Curated by: Barcelona Supercomputing CenterShared by: Barcelona Supercomputing Center
Dataset Description
Vinclat is a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vinclat.
