CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /RAG_Multilingual Dataset Card for RAG_Multilingual Dataset Summary RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets. The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC). This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.textquestion-answering10K<n<100K23 likes179 downloads2y agoHugging Face02projecte-aina /teca Dataset Card for TE-ca Dataset Summary TE-ca is a dataset of textual entailment in Catalan, which contains 21,163 pairs of premises and hypotheses, annotated according to the inference relation they have (implication, contradiction or neutral). This is the second version of the dataset, released on 29/03/2025, where encoding mistakes from the initial release have been corrected in both the 'hypothesis' and 'premise' column. Check… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/teca.texttext-classification10K<n<100K1 likes132 downloads2y agoHugging Face03projecte-aina /casum Dataset Card for CaSum Dataset Summary CaSum is a summarization dataset. It is extracted from a newswire corpus crawled from the Catalan News Agency (Agència Catalana de Notícies; ACN). The corpus consists of 217,735 instances that are composed by the headline and the body. Supported Tasks and Leaderboards The dataset can be used to train a model for abstractive summarization. Success on this task is typically measured by achieving a high Rouge score. The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/casum.textsummarization100K<n<1M0 likes103 downloads2y agoHugging Face04projecte-aina /MentorES Dataset Summary Mentor_ES is an open source dataset of 10,175 instructions in Spanish organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming. Supported Tasks and Leaderboards Useful for fine-tuning instructions in large language models for downstream tasks. Languages This dataset is in… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorES.textquestion-answering10K<n<100K3 likes98 downloads2mo agoHugging Face05projecte-aina /viquiquad ViquiQuAD: An Extractive QA Dataset for Catalan from Wikipedia Dataset Summary ViquiQuAD is an extractive Question Answering dataset for Catalan, built from the Catalan Wikipedia (Viquipèdia). 3,111 contexts extracted from 597 high-quality, original (non-translated) articles. For each context, 1 to 5 questions were created with their corresponding answers. Total: 15,153 question–answer pairs. This dataset can be used to fine-tune and evaluate extractive QA models and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/viquiquad.textquestion-answeringn<1K0 likes92 downloads1y agoHugging Face06projecte-aina /IFEval_ca Dataset Card for IFEval_ca IFEval_ca is a prompt dataset in Catalan, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_ca (Instruction-Following Eval benchmark - Catalan) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/IFEval_ca.textquestion-answeringn<1K0 likes80 downloads10mo agoHugging Face07projecte-aina /tecla Dataset Card for TeCla Dataset Summary TeCla (Text Classification) is a Catalan News corpus for thematic multi-class Text Classification tasks. The present version (2.0) contains 113.376 articles classified under a hierarchical class structure consisting of a coarse-grained and a fine-grained class. Each of the 4 coarse-grained classes accept a subset of fine-grained ones, 53 in total. The previous version (1.0.1) can still be found at https://zenodo.org/record/4761505… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/tecla.texttext-classification100K<n<1M0 likes77 downloads2y agoHugging Face08projecte-aina /vilaquad Dataset Card for VilaQuAD Dataset Summary VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text. This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context). VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence. This dataset can be used to build extractive-QA and Language Models. Supported Tasks and Leaderboards Extractive-QA… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilaquad.textquestion-answering1K<n<10K0 likes73 downloads2y agoHugging Face09projecte-aina /vilasum Dataset Card for VilaSum Dataset Summary VilaSum is a summarization dataset for evaluation. It is extracted from a newswire corpus crawled from the Catalan news portal VilaWeb. The corpus consists of 13,843 instances that are composed by the headline and the body. Supported Tasks and Leaderboards The dataset can be used to train a model for abstractive summarization. Success on this task is typically measured by achieving a high Rouge score. The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilasum.textsummarization10K<n<100K0 likes70 downloads2y agoHugging Face10projecte-aina /MentorCA Dataset Summary Mentor_CA is an open source dataset of 10,175 instructions in Catalan, machine translated from the original Mentor_ES dataset in Spanish, and organized in several of the behavioral categories outlined in the InstructGPT paper, including closed QA, open QA, general QA, classification, information extraction, summarization, creative writing and brainstorming. Supported Tasks and Leaderboards Useful for fine-tuning instructions in large language… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/MentorCA.textquestion-answering10K<n<100K2 likes65 downloads2mo agoHugging Face11projecte-aina /InstruCAT Dataset Card for Instrucat Dataset Summary InstruCat is a dataset consisting of 216,826 instructions in Catalan. It contains data converted to instructions format from the following datasets: caBreu : The instructions were created in form of summarization tasks. There are 2 types of summarization categories in the dataset: extreme and abstractive. The extreme one summarizes text into one sentence and the abstractive into shorter texts around 3-5 sentences. CatalanQA :… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/InstruCAT.text100K<n<1M2 likes60 downloads2y agoHugging Face12projecte-aina /catalan_government_crawling Dataset Card for Catalan Government Crawling Dataset Summary The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39,117,909 tokens, 1,565,433 sentences and 71,043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_government_crawling.textfill-mask10K<n<100K1 likes56 downloads2y agoHugging Face13projecte-aina /vinclat Dataset Card for Vinclat Vinclat is a Catalan-language dataset for multi-step problem solving. It employs a game-based structure to evaluate both the reasoning capabilities and cultural knowledge of large language models. Language(s): Catalan Paper: ACL Anthology Leaderboard: HF Space Code: GitHub License: CC-BY-4.0 Funded by: Projecte Aina Curated by: Barcelona Supercomputing CenterShared by: Barcelona Supercomputing Center Dataset Description Vinclat is a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vinclat.texttext-generation1K<n<10K1 likes47 downloads6mo agoHugging Face14projecte-aina /CaWikiTC Dataset Card for CaWikiTC Dataset Summary CaWikiTC (Catalan Wikipedia Text Classification) is a text classification dataset authomatically created by scraping Catalan Wikipedia article summaries and their associated thematic category. It contains 21002 texts (19952 and 1050 in the train and dev partitions, respectively) classified under 67 exclusive categories. For the dataset creation, we selected all the Catalan Wikipedia article summaries from a previously fixed… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CaWikiTC.texttext-classification10K<n<100K0 likes37 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.