CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPAI-BSC /CareQA CareQA Dataset Summary CareQA is a healthcare QA dataset with two versions: Closed-Ended Version: A multichoice question answering (MCQA) dataset containing 5,621 QA pairs across six categories. Available in English and Spanish. Open-Ended Version: A free-response dataset derived from the closed version, containing 2,769 QA pairs (English only). The dataset originates from… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CareQA.tabularquestion-answering10K<n<100K18 likes3.9k downloads1y agoHugging Face02HPAI-BSC /medical-specialities Medical Question Classification Dataset Dataset Summary This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly. Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medical-specialities.tabularquestion-answering10K<n<100K7 likes1.4k downloads10mo agoHugging Face03BSC-LT /multi_lmentry Multi-LMentry This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?". Dataset Details Dataset Description Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.textquestion-answering100K<n<1M12 likes876 downloads5mo agoHugging Face04HPAI-BSC /MedQA-Mixtral-CoT Dataset Card for medqa-cot Synthetically enhanced responses to the medqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the question… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedQA-Mixtral-CoT.textmultiple-choice10K<n<100K9 likes786 downloads2y agoHugging Face05HPAI-BSC /headqa-cot-llama31 headqa-cot Synthetically enhanced responses to the HeadQA dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the HeadQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/headqa-cot-llama31.textquestion-answering1K<n<10K2 likes435 downloads1y agoHugging Face06HPAI-BSC /MMLU-medical-cot-llama31 MMLU-medical-cot Synthetically enhanced responses to the medical-related questions of the auxiliary train set of the MMLU dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description First, we use Llama-3.1-70B-Instruct to filter the medical-related questions of the auxiliary train set of the MMLU dataset. Next, we leverage Mixtral-8x7B to… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MMLU-medical-cot-llama31.textquestion-answering1K<n<10K6 likes410 downloads10mo agoHugging Face07BSC-LT /EsBBQ Spanish Bias Benchmark for Question Answering (EsBBQ) The Spanish Bias Benchmark for Question Answering (EsBBQ) is an adaptation of the original BBQ to the Spanish language and the social context of Spain. Dataset Description This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EsBBQ.tabularquestion-answering10K<n<100K0 likes409 downloads1y agoHugging Face08BSC-LT /CaBBQ Catalan Bias Benchmark for Question Answering (CaBBQ) The Catalan Bias Benchmark for Question Answering (CaBBQ) is an adaptation of the original BBQ to the Catalan language and the social context of Spain. Dataset Description This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CaBBQ.tabularquestion-answering10K<n<100K1 likes390 downloads1y agoHugging Face09BSC-LT /BSC_ParaMT_8 Dataset Card for BSC_ParaMT_8 Dataset Summary Large-scale multilingual parallel corpus covering Catalan, Spanish, and English paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish portion of the dataset includes synthetic data generated by translating original English sentences into Spanish. Similarly… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/BSC_ParaMT_8.texttranslation100M<n<1B0 likes389 downloads3mo agoHugging Face10BSC-LT /m-personas mPersonas: Multilingual Persona‑Driven Conversational Dataset Dataset Summary mPersonas is a multilingual open-source dataset with high-quality persona descriptions synthetically generated by DeepSeek-V3–0324. It follows a persona-driven data synthesis methodology, similar to PersonaHub. Instances: 510,000 Total tokens: 173M 28M in personas 145M in conversations (105M in assistant turns) Languages: 15 License: Apache 2.0 Methodology This section… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/m-personas.textquestion-answering100K<n<1M1 likes377 downloads1y agoHugging Face11HPAI-BSC /HEART HEART Dataset Summary The HEART dataset is composed of multiple splits that differ in the type of injected cues. It includes a baseline split with no injected cues and four cue-based splits. Each cue is instantiated in two variants: assistive, where the cue is consistent with the GT, and adversarial, where the cue supports an incorrect option. Cue types, summarized below… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/HEART.imagevisual-question-answering10K<n<100K1 likes309 downloads3mo agoHugging Face12BSC-LT /COPA-es Dataset Card for COPA-es COPA-es is a textual entailment dataset in Spanish, professionally translated from the COPA dataset in English. The dataset consists of 600 premises, each given a question and two choices with a label encoding which of the choices is more plausible given the annotator. Dataset Details Dataset Description COPA-es (Choice of Plausible Alternatives - Spanish) is designed to simulate causal reasoning of text from commonsense subjects.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/COPA-es.texttext-classificationn<1K2 likes287 downloads2y agoHugging Face13HPAI-BSC /MedMCQA-Mixtral-CoT Dataset Card for medmcqa-cot Synthetically enhanced responses to the medmcqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedMCQA-Mixtral-CoT.textquestion-answering100K<n<1M4 likes270 downloads2y agoHugging Face14HPAI-BSC /Egida Dataset Card for Egida Dataset Summary Egida is an expanded collection of unsafe requests gathered from a variety of external sources. This dataset is boosted and extended (1) through a manual fine-grained topic classification, and (2) by applying a variety of jailbreaking attacks to all their samples. Dataset Curation Sources and data collection In total, the dataset is composed of 2,949… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Egida.textquestion-answering100K<n<1M3 likes257 downloads1y agoHugging Face15BSC-LT /openbookqa-es Dataset Card for openbookqa_es openbookqa_es is a question answering dataset in Spanish, professionally translated from the main version of the OpenBookQA dataset in English. Dataset Details Dataset Description openbookqa_es (Open Book Question Answering - Spanish) is designed to simulate open book exams and assess human-like understanding of a subject. The dataset comprises 500 instances in the validation split and another 500 instances in the test split.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/openbookqa-es.textquestion-answering1K<n<10K2 likes228 downloads2y agoHugging Face16HPAI-BSC /Aloe-Beta-Medical-Collection Aloe-Beta-Medical-Collection Collection of curated datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available medical instruction tuning data sources (QA format). Most data samples correspond to single-turn QA pairs, while a small proportion contain multi-turn. All data sources are publicly available for… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-Medical-Collection.textquestion-answering100K<n<1M4 likes220 downloads1y agoHugging Face17HPAI-BSC /medqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedQa dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medqa-cot-llama31.textmultiple-choice10K<n<100K3 likes159 downloads10mo agoHugging Face18BSC-LT /ALIA_mixed_authentic_synthetic_MT Dataset Card for ALIA_mixed_authentic_synthetic_MT Dataset Summary Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.texttranslation100M<n<1B1 likes158 downloads9mo agoHugging Face19HPAI-BSC /MedS-Ins HPAI-BSC MedS-Ins Collection of curated data from the MedS-Ins dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description This is the curated version of the MedS-Ins dataset included in the training set of the Aloe-Beta models. First, we selected 75 out of the 122 existing tasks, excluding the tasks that were already in the training set, and the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MedS-Ins.textquestion-answering100K<n<1M2 likes157 downloads1y agoHugging Face20HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes143 downloads10mo agoHugging Face21HPAI-BSC /medmcqa-cot-llama31 medqa-cot-llama31 Synthetically enhanced responses to the MedMCQA dataset. Used to train Aloe-Beta model. Dataset Details Dataset Description To increase the quality of answers from the training splits of the MedMCQA dataset, we leverage Llama-3.1-70B-Instruct to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/medmcqa-cot-llama31.textmultiple-choice100K<n<1M2 likes138 downloads1y agoHugging Face22BSC-LT /salamandra-guard-dataset Salamandra Guard Dataset Dataset Description The Salamandra Guard dataset is a comprehensive multilingual safety classification corpus designed for training and evaluating content moderation systems in Catalan, Spanish. It consists of 21,335 carefully curated conversational examples annotated across a hierarchical safety taxonomy. This dataset represents a significant advancement in culturally-grounded safety data, with particular emphasis on Catalan—a language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/salamandra-guard-dataset.documenttext-generation10K<n<100K3 likes130 downloads13d agoHugging Face23HPAI-BSC /Aloe-Beta-DPO Aloe-Beta-Medical-Collection Collection of curated DPO datasets used to align Aloe-Beta. Dataset Details Dataset Description The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data: Medical preference data: TsinghuaC3I/UltraMedical-Preference General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.textquestion-answering100K<n<1M2 likes106 downloads1y agoHugging Face24BSC-LT /IFEval_es Dataset Card for IFEval_es IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.textquestion-answeringn<1K1 likes105 downloads10mo agoHugging Face25BSC-LT /EQ-bench_es Dataset Card for EQ Bench Dataset (Spanish Version) This dataset card documents the Spanish adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts. Dataset Details Dataset Description EQ-Bench (Spanish Version) is a translated and linguistically adapted version of the original EQ-Bench dataset. Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_es.textquestion-answeringn<1K0 likes103 downloads1y agoHugging Face26HPAI-BSC /PubmedQA-Mixtral-CoT Dataset Card for pubmedqa-cot Synthetically enhanced responses to the pubmedqa dataset using mixtral. Dataset Details Dataset Description To increase the quality of answers from the training splits of the PubMedQA dataset, we leverage Mixtral-8x7B to generate Chain of Thought(CoT) answers. We create a custom prompt for the dataset, along with a hand-crafted list of few-shot examples. For a multichoice answer, we ask the model to rephrase and explain the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/PubmedQA-Mixtral-CoT.textmultiple-choice100K<n<1M2 likes88 downloads2y agoHugging Face27BSC-LT /EQ-bench_ca Dataset Card for EQ Bench Dataset (Catalan Version) This dataset card documents the Catalan adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts. Dataset Details Dataset Description EQ-Bench (Catalan Version) is a translated and linguistically adapted version of the original EQ-Bench dataset. Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_ca.textquestion-answeringn<1K0 likes84 downloads1y agoHugging Face28BSC-LT /arc_es Dataset Card for ARC (Spanish Version) Dataset summary This dataset provides the Spanish translation and adaptation of the ARC (AI2 Reasoning Challenge) validation set. The original dataset was designed to evaluate scientific reasoning in LLMs by presenting a collection of authentic, grade-school science multiple choice questions split into two sets of varying difficulty: ARC Easy and ARC Challenging. This Spanish adaptation enables evaluation of the ARC task in… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/arc_es.textquestion-answeringn<1K0 likes81 downloads6d agoHugging Face29BSC-LT /Catalan-Aranese_Parallel_Corpus Dataset Card for Catalan-Aranese Parallel Corpus Dataset Summary A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated from… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.texttranslation100K<n<1M2 likes80 downloads8mo agoHugging Face30BSC-LT /LexBOE Dataset Card for LexBOE Dataset summary LexBOE is a Spanish legal text classification dataset built from articles extracted from the Boletín Oficial del Estado (BOE), the official source of legislation and administrative acts in Spain. The articles included in the dataset were published between 2022 and 2024. LexBOE reflects contemporary legal-administrative language and is intended for the training and evaluation of language models on legal text classification tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/LexBOE.texttext-classification10K<n<100K0 likes80 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.