CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /CATalog Dataset Summary CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words. Supported Tasks and Leaderboards Fill-Mask Text Generation other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.textfill-mask10M<n<100M8 likes3.9k downloads1y agoHugging Face02projecte-aina /synthetic_dem Dataset Card for synthetic_dem Dataset Summary The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC). It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.audioautomatic-speech-recognition100K<n<1M2 likes1.5k downloads1y agoHugging Face03projecte-aina /COPA-ca Dataset Card for COPA-ca Dataset Summary The COPA-ca dataset (Choice of plausible alternatives in Catalan) is a professional translation of the English COPA dataset into Catalan, commissioned by BSC LangTech Unit. The dataset consists of 1000 premises, each given a question and two choices with a label encoding which of the choices is more plausible given the annotator. The dataset is split into 400 training samples, 100 validation samples, and 500 test samples. It… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/COPA-ca.tabular1K<n<10K0 likes468 downloads2y agoHugging Face04projecte-aina /parlament_parla_v3 Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.audioautomatic-speech-recognition100K<n<1M1 likes438 downloads2y agoHugging Face05projecte-aina /4catac Dataset Card for 4catac Dataset Summary 4catac: examples of phonetic transcription in 4 Catalan accents is a dataset of phonetic transcriptions in four Catalan accents: Balearic, Central, North-Western and Valencian. It consists of 160 sentences transcribed using IPA, following the recommendations of the Institut d'Estudis Catalans. These sentences are the same for the four accents but may have small morphological adaptations to make them more natural for the accent.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/4catac.texttext-to-speechn<1K1 likes431 downloads3y agoHugging Face06projecte-aina /festcat_trimmed_denoised Dataset Card for festcat_trimmed_denoised This is a post-processed version of the Catalan Festcat speech dataset. The original data can be found here. Same license is maintained: Creative Commons Attribution-ShareAlike 3.0 Spain License. Dataset Details Dataset Description We processed the data of the Catalan Festcat with the following recipe: Trimming: Long silences from the start and the end of clips have been removed. py-webrtcvad -> Python interface to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/festcat_trimmed_denoised.audiotext-to-speech10K<n<100K0 likes366 downloads1y agoHugging Face07projecte-aina /escagleu-64k Dataset Card for escagleu-64K corpus Dataset Description Dataset Summary This is the second version of escagleu-64k, a parallel corpus containing approximately 64k sentences translated across Spanish, Catalan, Valencian Catalan, Galician, and Basque. The original sentences are in Spanish and are sourced from the Spanish Common Voice Corpus. This corpus was prepared with the goal of creating a parallel speech dataset for these languages using the Common Voice… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/escagleu-64k.translation10K<n<100K0 likes300 downloads6mo agoHugging Face08projecte-aina /catalanqa Dataset Card for CatalanQA Dataset Summary This dataset can be used to build extractive-QA and Language Models. It is an aggregation and balancing of 2 previous datasets: VilaQuAD and ViquiQuAD. Splits have been balanced by kind of question, and unlike other datasets like SQuAD, it only contains, per record, one question and one answer for each context, although the contexts can repeat multiple times. This dataset was developed by BSC TeMU as part of Projecte AINA, to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalanqa.textquestion-answering10K<n<100K1 likes284 downloads2y agoHugging Face09projecte-aina /commonvoice_benchmark_catalan_accentsA new presentation of the corpus Catalan Common Voice v17.0 - metadata annotated version with the splits redefined to benchmark ASR models with various Catalan accentautomatic-speech-recognition1M<n<10M3 likes278 downloads2y agoHugging Face10projecte-aina /openbookqa_ca Dataset Card for openbookqa_ca openbookqa_ca is a question answering dataset in Catalan, professionally translated from the main version of the OpenBookQA dataset in English. Dataset Details Dataset Description openbookqa_ca (Open Book Question Answering - Catalan) is designed to simulate open book exams and assess human-like understanding of a subject. The dataset comprises 500 instances in the validation split and another 500 instances in the test split.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/openbookqa_ca.textquestion-answering1K<n<10K0 likes250 downloads2y agoHugging Face11projecte-aina /mgsm_ca Dataset Card for mgsm_ca mgsm_ca is a question answering dataset in Catalan that has been professionally translated from the MGSM dataset in English. Dataset Details Dataset Description mgsm_ca (Multilingual Grade School Math - Catalan) is designed to evaluate multi-step mathematical reasoning using grade school math word problems. It includes 8 instances in the train split and another 250 instances in the test split. Each instance contains a math problem… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/mgsm_ca.textquestion-answeringn<1K0 likes237 downloads2y agoHugging Face12projecte-aina /parlament_parlaThis is the ParlamentParla speech corpus for Catalan prepared by Col·lectivaT. The audio segments were extracted from recordings the Catalan Parliament (Parlament de Catalunya) plenary sessions, which took place between 2007/07/11 - 2018/07/17. We aligned the transcriptions with the recordings and extracted the corpus. The content belongs to the Catalan Parliament and the data is released conforming their terms of use. Preparation of this corpus was partly supported by the Department of Culture of the Catalan autonomous government, and the v2.0 was supported by the Barcelona Supercomputing Center, within the framework of the project AINA of the Departament de Polítiques Digitals. As of v2.0 the corpus is separated into 211 hours of clean and 400 hours of other quality segments. Furthermore, each speech segment is tagged with its speaker and each speaker with their gender. The statistics are detailed in the readme file. For more information, go to https://github.com/CollectivaT-dev/ParlamentParla or mail info@collectivat.cat.automatic-speech-recognition100K<n<1M4 likes228 downloads2y agoHugging Face13projecte-aina /GuiaCat Dataset Card for GuiaCat Dataset Summary GuiaCat is a dataset consisting of 5.750 restaurant reviews in Catalan, with 5 associated scores and a label of sentiment. The data was provided by GuiaCat and curated by the BSC. This work is licensed under a Creative Commons Attribution Non-commercial No-Derivatives 4.0 International License. Supported Tasks and Leaderboards This corpus is mainly intended for sentiment analysis. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/GuiaCat.tabulartext-classification1K<n<10K0 likes212 downloads3y agoHugging Face14projecte-aina /sts-ca Dataset Card for STS-ca Dataset Summary STS-ca corpus is a benchmark for evaluating Semantic Text Similarity in Catalan. This dataset was developed by BSC TeMU as part of Projecte AINA, to enrich the Catalan Language Understanding Benchmark (CLUB). This work is licensed under a Attribution-ShareAlike 4.0 International License. Supported Tasks and Leaderboards This dataset can be used to build and score semantic similarity models in Catalan.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/sts-ca.texttext-classification1K<n<10K1 likes205 downloads2mo agoHugging Face15projecte-aina /ancora-ca-ner Dataset Card for AnCora-Ca-NER Dataset Summary This is a dataset for Named Entity Recognition (NER) in Catalan. It adapts AnCora corpus for Machine Learning and Language Model evaluation purposes. This dataset was developed by BSC TeMU as part of the Projecte AINA, to enrich the Catalan Language Understanding Benchmark (CLUB). Supported Tasks and Leaderboards Named Entities Recognition, Language Model Languages The dataset is in Catalan… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ancora-ca-ner.text10K<n<100K3 likes199 downloads2y agoHugging Face16projecte-aina /raco_forums Dataset Card for Racó Forums Corpus Dataset Summary The Racó Forums Corpus is a 19-million-sentence corpus of Catalan user-generated text built from the forums of Racó Català. Since the existing available corpora in Catalan lacked conversational data, we searched for a major source of such data for Catalan, and we found Racó Català, a popular multitopic online forum. We obtained a database dump and we transformed all the threads so that we obtained documents that… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/raco_forums.textfill-mask1M<n<10M2 likes177 downloads2y agoHugging Face17projecte-aina /RAG_Multilingual Dataset Card for RAG_Multilingual Dataset Summary RAG_Multilingual is an instruction-following synthetic QA dataset created from extractive QA datasets from Catalan, English and Spanish reference sets. The reference datasets were: SQAD (https://huggingface.co/datasets/rajpurkar/squad), Catalanqa (https://huggingface.co/datasets/projecte-aina/catalanqa) and SQAC (https://huggingface.co/datasets/PlanTL-GOB-ES/SQAC). This dataset, of 56.406 instances, was created by… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/RAG_Multilingual.textquestion-answering10K<n<100K23 likes176 downloads2y agoHugging Face18projecte-aina /corts_valencianes_asr_aThis is the first version of CortsValencianes speech corpus for Valencian: a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications.automatic-speech-recognition1K<n<10K2 likes176 downloads1y agoHugging Face19projecte-aina /PAWS-ca Dataset Card for PAWS-ca: Paraphrase Adversaries from Word Scrambling in Catalan Dataset Summary The PAWS-ca dataset (Paraphrase Adversaries from Word Scrambling in Catalan) is a translation of the English PAWS dataset into Catalan, commissioned by BSC LangTech Unit. The dataset contains 4,000 human translated PAWS pairs and 49,000 machine translated pairs. Supported Tasks and Leaderboards Paraphrase Identification, Language Model Languages The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/PAWS-ca.texttext-classification10K<n<100K0 likes173 downloads2y agoHugging Face20projecte-aina /caBreu Dataset Card for caBREU Dataset Summary caBreu is a summarization dataset in Catalan, produced by the BSC LangTech Unit. The dataset consists of 3,000 articles, each averaging about 700 words in length, along with extreme, abstractive and extractive summaries, manually generated by three annotators. The source material for the articles was gathered from various Catalan news sources, including the Catalan News Agency (Agència Catalana de Notícies; ACN), VilaWeb and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/caBreu.textsummarization1K<n<10K0 likes173 downloads2y agoHugging Face21projecte-aina /veritasQA Dataset Card for VeritasQA VeritasQA is a context- and time-independent QA benchmark for the evaluation of truthfulness in Language Models. Dataset Summary VeritasQA is a context- and time-independent truthfulness benchmark built with multilingual transferability in mind. It is intended to be used to evaluate Large Language Models on truthfulness in a zero-shot setting. VeritasQA comprises 353 question-answer pairs inspired by common misconceptions and falsehoods, not… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/veritasQA.texttext-generation1K<n<10K3 likes163 downloads1y agoHugging Face22projecte-aina /wnli-ca WNLI-ca Dataset Summary "A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its resolution. The schema takes its name from Terry Winograd." Source: The Winograd Schema Challenge. The Winograd NLI dataset presents 855 sentence pairs, in which the first sentence contains an ambiguity and the second one a… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/wnli-ca.texttext-classificationn<1K0 likes158 downloads2y agoHugging Face23projecte-aina /arc_ca Dataset Card for arc_ca arc_ca is a question answering dataset in Catalan, professionally translated from the Easy and Challenge versions of the ARC dataset in English. Dataset Details Dataset Description arc_ca (AI2 Reasoning Challenge - Catalan) is based on multiple-choice science questions at elementary school level. The dataset consists of 2950 instances in the Easy version (570 in the test and 2380 instances in the validation split) and… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/arc_ca.textquestion-answering1K<n<10K0 likes150 downloads4mo agoHugging Face24projecte-aina /teca Dataset Card for TE-ca Dataset Summary TE-ca is a dataset of textual entailment in Catalan, which contains 21,163 pairs of premises and hypotheses, annotated according to the inference relation they have (implication, contradiction or neutral). This is the second version of the dataset, released on 29/03/2025, where encoding mistakes from the initial release have been corrected in both the 'hypothesis' and 'premise' column. Check… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/teca.texttext-classification10K<n<100K1 likes140 downloads2y agoHugging Face25projecte-aina /Parafraseja Dataset Card for Parafraseja Dataset Summary Parafraseja is a dataset of 21,984 pairs of sentences with a label that indicates if they are paraphrases or not. The original sentences were collected from TE-ca and STS-ca. For each sentence, an annotator wrote a sentence that was a paraphrase and another that was not. The guidelines of this annotation are available. This work is licensed under a Creative Commons Attribution Non-commercial No-Derivatives 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/Parafraseja.texttext-classification10K<n<100K0 likes138 downloads2y agoHugging Face26projecte-aina /annotated_catalan_common_voice_v17This version of the Catalan sentences of the Common Voice corpus v17 includes metadata (gender and accent) for 263 speakers annotated by a team of experts.automatic-speech-recognition1M<n<10M1 likes128 downloads1y agoHugging Face27projecte-aina /piqa_ca Dataset Card for piqa_ca piqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the PIQA validation set in English. Dataset Details Dataset Description piqa_ca (Physical Interaction Question Answering - Catalan) is designed to evaluate physical commonsense reasoning using question-answer triplets based on everyday situations. It includes 1838 instances in the validation split. Each instance contains… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/piqa_ca.textquestion-answering1K<n<10K0 likes120 downloads2y agoHugging Face28projecte-aina /xquad-ca Dataset Card for XQuAD-Ca Dataset Summary Professional translation into Catalan of XQuAD dataset. XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance. The dataset consists of a subset of 240 paragraphs and 1190 question-answer pairs from the development set of SQuAD v1.1 (Rajpurkar, Pranav et al., 2016) together with their professional translations into ten language: Spanish, German, Greek… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xquad-ca.textquestion-answering1K<n<10K0 likes119 downloads2y agoHugging Face29projecte-aina /siqa_ca Dataset Card for siqa_ca siqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the SIQA validation set in English. Dataset Details Dataset Description siqa_ca (Social Interaction Question Answering - Catalan) is designed to evaluate social commonsense intelligence using multiple choice question-answer instances based on reasoning about people’s actions and their social implications. It includes… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/siqa_ca.textquestion-answering1K<n<10K0 likes113 downloads5mo agoHugging Face30projecte-aina /xstorycloze_ca Dataset Card for xstorycloze_ca xstorycloze_ca is a question answering dataset in Catalan, professionally translated from the English StoryCloze dataset (Spring 2016 version), used to create its multilingual version XStoryCloze. Dataset Details Dataset Description xstorycloze_ca (Multilingual Story Cloze Test - Catalan) is based on multiple-choice narrative completions. The dataset consists of 360 instances in the train split and 1510 instances in the test… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/xstorycloze_ca.textquestion-answering1K<n<10K0 likes108 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.