CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /CATalog Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.textfill-mask10M<n<100M8 likes3.9k downloads1y agoHugging Face02projecte-aina /parlament_parlaThis is the ParlamentParla speech corpus for Catalan prepared by Col·lectivaT. The audio segments were extracted from recordings the Catalan Parliament (Parlament de Catalunya) plenary sessions, which took place between 2007/07/11 - 2018/07/17. We aligned the transcriptions with the recordings and extracted the corpus. The content belongs to the Catalan Parliament and the data is released conforming their terms of use. Preparation of this corpus was partly supported by the Department of Culture of the Catalan autonomous government, and the v2.0 was supported by the Barcelona Supercomputing Center, within the framework of the project AINA of the Departament de Polítiques Digitals. As of v2.0 the corpus is separated into 211 hours of clean and 400 hours of other quality segments. Furthermore, each speech segment is tagged with its speaker and each speaker with their gender. The statistics are detailed in the readme file. For more information, go to https://github.com/CollectivaT-dev/ParlamentParla or mail info@collectivat.cat.automatic-speech-recognition100K<n<1M4 likes228 downloads2y agoHugging Face03projecte-aina /RAG_Multilingual Dataset Card for RAG_Multilingual Dataset Summary RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets. The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC). This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.textquestion-answering10K<n<100K23 likes176 downloads2y agoHugging Face04projecte-aina /veritasQA Dataset Card for VeritasQA VeritasQA is a context- and time-independent QA benchmark for the evaluation of truthfulness in Language Models. Dataset Summary VeritasQA is a context- and time-independent truthfulness benchmark built with multilingual transferability in mind. It is intended to be used to evaluate Large Language Models on truthfulness in a zero-shot setting. VeritasQA comprises 353 question-answer pairs inspired by common misconceptions and falsehoods, not… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/veritasQA.texttext-generation1K<n<10K3 likes163 downloads1y agoHugging Face05projecte-aina /xstorycloze_ca Dataset Card for xstorycloze_ca xstorycloze_ca is a question answering dataset in Catalan, professionally translated from the English StoryCloze dataset (Spring 2016 version), used to create its multilingual version XStoryCloze. Dataset Details Dataset Description xstorycloze_ca (Multilingual Story Cloze Test - Catalan) is based on multiple-choice narrative completions. The dataset consists of 360 instances in the train split and 1510 instances in the test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xstorycloze_ca.textquestion-answering1K<n<10K0 likes108 downloads2y agoHugging Face06projecte-aina /oasst1_ca Dataset Card for oasst1_ca oasst1_ca is a conversational dataset in Catalan that has been professionally translated from the OASST1 dataset. Dataset Details Dataset Description oasst1_ca (OpenAssistant Conversations Release 1 - Catalan) consists of human-generated, human-annotated assistant-style conversation corpus. It includes 5213 messages in the train split and 273 messages in the validation split. To arrive to this number, we filter the dataset (See… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/oasst1_ca.tabulartext-generation1K<n<10K0 likes107 downloads2y agoHugging Face07projecte-aina /MentorES Dataset Summary Mentor_ES is an open source dataset of 10,175 instructions in Spanish organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming. Supported Tasks and Leaderboards Useful for fine-tuning instructions in large language models for downstream tasks. Languages This dataset is in… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorES.textquestion-answering10K<n<100K3 likes96 downloads2mo agoHugging Face08projecte-aina /IFEval_ca Dataset Card for IFEval_ca IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.textquestion-answeringn<1K0 likes83 downloads10mo agoHugging Face09projecte-aina /MentorCA Dataset Summary Mentor_CA is an open source dataset of 10,175 instructions in Catalan, machine translated from the original Mentor_ES dataset in Spanish, and organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming. Supported Tasks and Leaderboards Useful for fine-tuning instructions in large language… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorCA.textquestion-answering10K<n<100K2 likes71 downloads2mo agoHugging Face10projecte-aina /hhh_alignment_ca Dataset Card for hhh_alignment_ca hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English. Dataset Details Dataset Description hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.textquestion-answeringn<1K0 likes55 downloads2y agoHugging Face11projecte-aina /vinclat Dataset Card for Vinclat Vinclat is a Catalan-language dataset for multi-step problem solving. It employs a game-based structure to evaluate both the reasoning capabilities and cultural knowledge of large language models. Language(s): Catalan Paper: ACL Anthology Leaderboard: HF Space Code: GitHub License: CC-BY-4.0 Funded by: Projecte Aina Curated by: Barcelona Supercomputing CenterShared by: Barcelona Supercomputing Center Dataset Description Vinclat is a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vinclat.texttext-generation1K<n<10K1 likes47 downloads6mo agoHugging Face12projecte-aina /NLUCatNLUCat - Natural Language Understanding in Catalantext-classification10M<n<100M0 likes29 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.