CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /catalanqa Dataset Card for CatalanQA Dataset Summary This dataset can be used to build extractive-QA and Language Models. It is an aggregation and balancing of 2 previous datasets: VilaQuAD and ViquiQuAD. Splits have been balanced by kind of question, and unlike other datasets like SQuAD, it only contains, per record, one question and one answer for each context, although the contexts can repeat multiple times. This dataset was developed by BSC TeMU as part of Projecte AINA, to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalanqa.textquestion-answering10K<n<100K1 likes284 downloads2y agoHugging Face02projecte-aina /openbookqa_ca Dataset Card for openbookqa_ca openbookqa_ca is a question answering dataset in Catalan, professionally translated from the main version of the OpenBookQA dataset in English. Dataset Details Dataset Description openbookqa_ca (Open Book Question Answering - Catalan) is designed to simulate open book exams and assess human-like understanding of a subject. The dataset comprises 500 instances in the validation split and another 500 instances in the test split.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/openbookqa_ca.textquestion-answering1K<n<10K0 likes250 downloads2y agoHugging Face03projecte-aina /mgsm_ca Dataset Card for mgsm_ca mgsm_ca is a question answering dataset in Catalan that has been professionally translated from the MGSM dataset in English. Dataset Details Dataset Description mgsm_ca (Multilingual Grade School Math - Catalan) is designed to evaluate multi-step mathematical reasoning using grade school math word problems. It includes 8 instances in the train split and another 250 instances in the test split. Each instance contains a math problem… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/mgsm_ca.textquestion-answeringn<1K0 likes237 downloads2y agoHugging Face04projecte-aina /RAG_Multilingual Dataset Card for RAG_Multilingual Dataset Summary RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets. The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC). This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.textquestion-answering10K<n<100K23 likes176 downloads2y agoHugging Face05projecte-aina /veritasQA Dataset Card for VeritasQA VeritasQA is a context- and time-independent QA benchmark for the evaluation of truthfulness in Language Models. Dataset Summary VeritasQA is a context- and time-independent truthfulness benchmark built with multilingual transferability in mind. It is intended to be used to evaluate Large Language Models on truthfulness in a zero-shot setting. VeritasQA comprises 353 question-answer pairs inspired by common misconceptions and falsehoods, not… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/veritasQA.texttext-generation1K<n<10K3 likes163 downloads1y agoHugging Face06projecte-aina /arc_ca Dataset Card for arc_ca arc_ca is a question answering dataset in Catalan, professionally translated from the Easy and Challenge versions of the ARC dataset in English. Dataset Details Dataset Description arc_ca (AI2 Reasoning Challenge - Catalan) is based on multiple-choice science questions at elementary school level. The dataset consists of 2950 instances in the Easy version (570 in the test and 2380 instances in the validation split) and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/arc_ca.textquestion-answering1K<n<10K0 likes150 downloads4mo agoHugging Face07projecte-aina /piqa_ca Dataset Card for piqa_ca piqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the PIQA validation set in English. Dataset Details Dataset Description piqa_ca (Physical Interaction Question Answering - Catalan) is designed to evaluate physical commonsense reasoning using question-answer triplets based on everyday situations. It includes 1838 instances in the validation split. Each instance contains… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/piqa_ca.textquestion-answering1K<n<10K0 likes120 downloads2y agoHugging Face08projecte-aina /xquad-ca Dataset Card for XQuAD-Ca Dataset Summary Professional translation into Catalan of XQuAD dataset. XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance. The dataset consists of a subset of 240 paragraphs and 1190 question-answer pairs from the development set of SQuAD v1.1 (Rajpurkar, Pranav et al., 2016) together with their professional translations into ten language: Spanish, German, Greek… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xquad-ca.textquestion-answering1K<n<10K0 likes119 downloads2y agoHugging Face09projecte-aina /siqa_ca Dataset Card for siqa_ca siqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the SIQA validation set in English. Dataset Details Dataset Description siqa_ca (Social Interaction Question Answering - Catalan) is designed to evaluate social commonsense intelligence using multiple choice question-answer instances based on reasoning about people’s actions and their social implications. It includes… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/siqa_ca.textquestion-answering1K<n<10K0 likes113 downloads5mo agoHugging Face10projecte-aina /xstorycloze_ca Dataset Card for xstorycloze_ca xstorycloze_ca is a question answering dataset in Catalan, professionally translated from the English StoryCloze dataset (Spring 2016 version), used to create its multilingual version XStoryCloze. Dataset Details Dataset Description xstorycloze_ca (Multilingual Story Cloze Test - Catalan) is based on multiple-choice narrative completions. The dataset consists of 360 instances in the train split and 1510 instances in the test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xstorycloze_ca.textquestion-answering1K<n<10K0 likes108 downloads2y agoHugging Face11projecte-aina /oasst1_ca Dataset Card for oasst1_ca oasst1_ca is a conversational dataset in Catalan that has been professionally translated from the OASST1 dataset. Dataset Details Dataset Description oasst1_ca (OpenAssistant Conversations Release 1 - Catalan) consists of human-generated, human-annotated assistant-style conversation corpus. It includes 5213 messages in the train split and 273 messages in the validation split. To arrive to this number, we filter the dataset (See… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/oasst1_ca.tabulartext-generation1K<n<10K0 likes107 downloads2y agoHugging Face12projecte-aina /CoQCat Dataset Card for CoQCat Dataset Summary CoQCat is a dataset for Conversational Question Answering in Catalan. It is based on CoQA dataset. CoQCat comprises 89,364 question-answer pairs, sourced from conversations related to 6,000 text passages from six different domains. The questions and responses are designed to maintain a conversational tone. The answers are presented in a free-form text format, with evidence highlighted from the passage. For the development and test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CoQCat.textquestion-answering1K<n<10K2 likes100 downloads2y agoHugging Face13projecte-aina /MentorES Dataset Summary Mentor_ES is an open source dataset of 10,175 instructions in Spanish organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming. Supported Tasks and Leaderboards Useful for fine-tuning instructions in large language models for downstream tasks. Languages This dataset is in… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorES.textquestion-answering10K<n<100K3 likes96 downloads2mo agoHugging Face14projecte-aina /viquiquad ViquiQuAD: An Extractive QA Dataset for Catalan from Wikipedia Dataset Summary ViquiQuAD is an extractive Question Answering dataset for Catalan, built from the Catalan Wikipedia (Viquipèdia). 3,111 contexts extracted from 597 high-quality, original (non-translated) articles. For each context, 1 to 5 questions were created with their corresponding answers. Total: 15,153 question–answer pairs. This dataset can be used to fine-tune and evaluate extractive QA models and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/viquiquad.textquestion-answeringn<1K0 likes89 downloads1y agoHugging Face15projecte-aina /IFEval_ca Dataset Card for IFEval_ca IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.textquestion-answeringn<1K0 likes83 downloads10mo agoHugging Face16projecte-aina /vilaquad Dataset Card for VilaQuAD Dataset Summary VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text. This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context). VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence. This dataset can be used to build extractive-QA and Language Models. Supported Tasks and Leaderboards Extractive-QA… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilaquad.textquestion-answering1K<n<10K0 likes75 downloads2y agoHugging Face17projecte-aina /MentorCA Dataset Summary Mentor_CA is an open source dataset of 10,175 instructions in Catalan, machine translated from the original Mentor_ES dataset in Spanish, and organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming. Supported Tasks and Leaderboards Useful for fine-tuning instructions in large language… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorCA.textquestion-answering10K<n<100K2 likes71 downloads2mo agoHugging Face18projecte-aina /hhh_alignment_ca Dataset Card for hhh_alignment_ca hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English. Dataset Details Dataset Description hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.textquestion-answeringn<1K0 likes55 downloads2y agoHugging Face19projecte-aina /vinclat Dataset Card for Vinclat Vinclat is a Catalan-language dataset for multi-step problem solving. It employs a game-based structure to evaluate both the reasoning capabilities and cultural knowledge of large language models. Language(s): Catalan Paper: ACL Anthology Leaderboard: HF Space Code: GitHub License: CC-BY-4.0 Funded by: Projecte Aina Curated by: Barcelona Supercomputing CenterShared by: Barcelona Supercomputing Center Dataset Description Vinclat is a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vinclat.texttext-generation1K<n<10K1 likes47 downloads6mo agoHugging Face20projecte-aina /dolly3k_ca Dataset Card for dolly3k_ca dolly3k_ca is a question answering dataset in Catalan, professionally translated from a filtered version of databricks-dolly-15k dataset in English. Dataset Details Dataset Description dolly3k_ca (Dolly 3K instances - Catalan) is based on question-answer instance pairs written by humans. The dataset consists of 3232 instances in the train split. Each instance contains an instruction or question, and one answer. Every instance is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/dolly3k_ca.textquestion-answering1K<n<10K0 likes29 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.