CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BSC-LT /multi_lmentry Multi-LMentry This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?". Dataset Details Dataset Description Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.textquestion-answering100K<n<1M12 likes960 downloads5mo agoHugging Face02BSC-LT /BSC_ParaMT_8 Dataset Card for BSC_ParaMT_8 Dataset Summary Large-scale multilingual parallel corpus covering Catalan, Spanish, and English paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish portion of the dataset includes synthetic data generated by translating original English sentences into Spanish. Similarly… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/BSC_ParaMT_8.texttranslation100M<n<1B0 likes422 downloads4mo agoHugging Face03BSC-LT /m-personas mPersonas: Multilingual Persona‑Driven Conversational Dataset Dataset Summary mPersonas is a multilingual open-source dataset with high-quality persona descriptions synthetically generated by DeepSeek-V3–0324. It follows a persona-driven data synthesis methodology, similar to PersonaHub. Instances: 510,000 Total tokens: 173M 28M in personas 145M in conversations (105M in assistant turns) Languages: 15 License: Apache 2.0 Methodology This section… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/m-personas.textquestion-answering100K<n<1M1 likes360 downloads1y agoHugging Face04BSC-LT /EsBBQ Spanish Bias Benchmark for Question Answering (EsBBQ) The Spanish Bias Benchmark for Question Answering (EsBBQ) is an adaptation of the original BBQ to the Spanish language and the social context of Spain. Dataset Description This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EsBBQ.tabularquestion-answering10K<n<100K0 likes326 downloads1y agoHugging Face05BSC-LT /COPA-es Dataset Card for COPA-es COPA-es is a textual entailment dataset in Spanish, professionally translated from the COPA dataset in English. The dataset consists of 600 premises, each given a question and two choices with a label encoding which of the choices is more plausible given the annotator. Dataset Details Dataset Description COPA-es (Choice of Plausible Alternatives - Spanish) is designed to simulate causal reasoning of text from commonsense subjects.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/COPA-es.texttext-classificationn<1K2 likes307 downloads2y agoHugging Face06BSC-LT /CaBBQ Catalan Bias Benchmark for Question Answering (CaBBQ) The Catalan Bias Benchmark for Question Answering (CaBBQ) is an adaptation of the original BBQ to the Catalan language and the social context of Spain. Dataset Description This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CaBBQ.tabularquestion-answering10K<n<100K1 likes300 downloads1y agoHugging Face07BSC-LT /openbookqa-es Dataset Card for openbookqa_es openbookqa_es is a question answering dataset in Spanish, professionally translated from the main version of the OpenBookQA dataset in English. Dataset Details Dataset Description openbookqa_es (Open Book Question Answering - Spanish) is designed to simulate open book exams and assess human-like understanding of a subject. The dataset comprises 500 instances in the validation split and another 500 instances in the test split.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/openbookqa-es.textquestion-answering1K<n<10K2 likes253 downloads2y agoHugging Face08BSC-LT /ALIA_mixed_authentic_synthetic_MT Dataset Card for ALIA_mixed_authentic_synthetic_MT Dataset Summary Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.texttranslation100M<n<1B1 likes169 downloads9mo agoHugging Face09BSC-LT /salamandra-guard-dataset Salamandra Guard Dataset Dataset Description The Salamandra Guard dataset is a comprehensive multilingual safety classification corpus designed for training and evaluating content moderation systems in Catalan, Spanish. It consists of 21,335 carefully curated conversational examples annotated across a hierarchical safety taxonomy. This dataset represents a significant advancement in culturally-grounded safety data, with particular emphasis on Catalan—a language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/salamandra-guard-dataset.documenttext-generation10K<n<100K3 likes145 downloads16d agoHugging Face10BSC-LT /IFEval_es Dataset Card for IFEval_es IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English. Dataset Details Dataset Description IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.textquestion-answeringn<1K1 likes107 downloads10mo agoHugging Face11BSC-LT /EQ-bench_es Dataset Card for EQ Bench Dataset (Spanish Version) This dataset card documents the Spanish adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts. Dataset Details Dataset Description EQ-Bench (Spanish Version) is a translated and linguistically adapted version of the original EQ-Bench dataset. Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_es.textquestion-answeringn<1K0 likes85 downloads1y agoHugging Face12BSC-LT /LexBOE Dataset Card for LexBOE Dataset summary LexBOE is a Spanish legal text classification dataset built from articles extracted from the Boletín Oficial del Estado (BOE), the official source of legislation and administrative acts in Spain. The articles included in the dataset were published between 2022 and 2024. LexBOE reflects contemporary legal-administrative language and is intended for the training and evaluation of language models on legal text classification tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/LexBOE.texttext-classification10K<n<100K0 likes85 downloads6mo agoHugging Face13BSC-LT /arc_es Dataset Card for ARC (Spanish Version) Dataset summary This dataset provides the Spanish translation and adaptation of the ARC (AI2 Reasoning Challenge) validation set. The original dataset was designed to evaluate scientific reasoning in LLMs by presenting a collection of authentic, grade-school science multiple choice questions split into two sets of varying difficulty: ARC Easy and ARC Challenging. This Spanish adaptation enables evaluation of the ARC task in… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/arc_es.textquestion-answeringn<1K0 likes84 downloads9d agoHugging Face14BSC-LT /Catalan-Aranese_Parallel_Corpus Dataset Card for Catalan-Aranese Parallel Corpus Dataset Summary A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated from… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.texttranslation100K<n<1M2 likes84 downloads8mo agoHugging Face15BSC-LT /ACAData Dataset Card for ACAData Dataset Summary ACAData is a multilingual instruction tuning dataset containing parallel text paragraphs from the academic domain. Supported Tasks and Leaderboards The dataset is meant to be used for fine-tuning and benchmarking general purpose LLM's on Machine Translation tasks. Languages The dataset contains (mainly long) paragraph of scientific texts from the academic domain in many European language pairs. The language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ACAData.texttranslation100K<n<1M2 likes75 downloads7mo agoHugging Face16BSC-LT /EQ-bench_ca Dataset Card for EQ Bench Dataset (Catalan Version) This dataset card documents the Catalan adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts. Dataset Details Dataset Description EQ-Bench (Catalan Version) is a translated and linguistically adapted version of the original EQ-Bench dataset. Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_ca.textquestion-answeringn<1K0 likes70 downloads1y agoHugging Face17BSC-LT /AbSanitas Dataset Card for AbSanitas Dataset summary AbSanitas is a Spanish biomedical information retrieval dataset built from biomedical texts collected from official academic repositories and open-access sources. This dataset is designed to support the training and evaluation of encoder models on biomedical retrieval and semantic matching tasks in Spanish. Curated by: Barcelona Supercomputing Center (BSC) Funded by: ALIA Language(s) (NLP): Spanish (es) License: CC BY-NC-ND 4.0… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/AbSanitas.texttext-retrieval10K<n<100K0 likes69 downloads6mo agoHugging Face18BSC-LT /rpaweb Dataset Card for rpaweb The rpaweb dataset contains multilingual annotations of Robotic Process Automation (RPA) actions in a web environment in English, Spanish and Catalan.It has been designed to train multimodal models capable of autonomously performing browser-based tasks based on natural language user requests. The corpus documents each user request along with the full sequence of web automation actions needed to fulfil it, including structured JSON code, natural language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/rpaweb.image1K<n<10K0 likes64 downloads1y agoHugging Face19BSC-LT /CAESAR-TV3 Dataset card for CAESAR-TV3 Dataset Summary This corpus includes 5 hours and 45 minutes of Catalan speech code-switched with Spanish extracted from the original tv3_parla dataset. Supported Tasks and Leaderboards The CAESAR-TV3 dataset is designed for the Automatic Speech Recognition (ASR) task, enabling the transcription of utterances in Catalan, Spanish, and code-switched speech between the two languages. Languages The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CAESAR-TV3.audioautomatic-speech-recognition1K<n<10K1 likes57 downloads1y agoHugging Face20BSC-LT /ALIA-2606-DPO-safety Dataset Card for BSC Multilingual Synthetic Safety Preferences Dataset Summary This dataset consists of synthetic safety preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician. Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool of safety… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-safety.texttext-generation10K<n<100K0 likes55 downloads2mo agoHugging Face21BSC-LT /geneval_catalan Dataset Card for GenEval Catalan Dataset Summary GenEval Catalan is an English→Catalan extension of MT-GenEval, a benchmark for evaluating gender accuracy in machine translation. It is derived from the original English–Spanish MT-GenEval dataset (Currey et al., 2022), with the Spanish side replaced by Catalan translations produced by professional human translators (with the exception of the single-sentence dev split, which uses automatic translation). The dataset… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/geneval_catalan.texttranslation1K<n<10K0 likes54 downloads6mo agoHugging Face22BSC-LT /AbScientia Dataset Card for AbScientia Dataset summary AbScientia is a Spanish STEM scientific text classification dataset built from scientific abstracts collected from official academic repositories and open-access sources. The dataset focuses on Science, Technology, Engineering, and Mathematics (STEM) disciplines and reflects domain-specific scientific language in Spanish. This dataset is designed to support the training and evaluation of encoder models on STEM scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/AbScientia.texttext-classification10K<n<100K0 likes52 downloads7mo agoHugging Face23BSC-LT /SIQA_es Dataset Card for SIQA (Spanish Version) Dataset summary This dataset provides the Spanish translation and adaptation of the SIQA (Social Interaction Question Answering) validation set. The original dataset was designed to evaluate social commonsense reasoning in LLMs by presenting a collection of questions based on everyday social situations, with the ultimate goal of challenging models to infer motivations, reactions and social implications behind human actions. This… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/SIQA_es.textquestion-answering1K<n<10K0 likes42 downloads5mo agoHugging Face24open-llm-leaderboard /BSC-LT__salamandra-7b-detailsgated Dataset Card for Evaluation run of BSC-LT/salamandra-7b Dataset automatically created during the evaluation run of model BSC-LT/salamandra-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BSC-LT__salamandra-7b-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face25BSC-LT /hhh_alignment_es Dataset Card for hhh_alignment_es hhh_alignment_es is a question answering dataset in Spanish, professionally translated from the main version of the hhh_alignment dataset in English. Dataset Details Dataset Description hhh_alignment_es (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Spanish) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/hhh_alignment_es.texttext-classificationn<1K0 likes39 downloads2y agoHugging Face26BSC-LT /XitXatTools Dataset Card for XitXat Tools XitXat Tools is a dataset comprising simulated Catalan call center conversations. Each conversation is annotated with structured tool calls, making it suitable for training and evaluating language models with function-calling capabilities. Dataset Details Dataset Sources Repository: XitXat Uses The dataset can be utilized for: Training language models to handle function-calling scenarios… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/XitXatTools.texttext-generationn<1K0 likes39 downloads4mo agoHugging Face27BSC-LT /Spanish-Valencian_Catalan_Parallel_Corpus Dataset Card for Spanish-Valencian Catalan Parallel Corpus Dataset Summary A bilingual parallel corpus containing parallel sentences in Spanish and the Valencian variant of Catalan. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus.texttranslation1M<n<10M2 likes39 downloads7mo agoHugging Face28open-llm-leaderboard /BSC-LT__salamandra-7b-instruct-detailsgated Dataset Card for Evaluation run of BSC-LT/salamandra-7b-instruct Dataset automatically created during the evaluation run of model BSC-LT/salamandra-7b-instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BSC-LT__salamandra-7b-instruct-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face29BSC-LT /cabreu_dolly_summarizationtext1K<n<10K0 likes37 downloads3y agoHugging Face30BSC-LT /InstrucatQA Dataset Card for Dataset Name Instructional dataset to finetune models used for RAG applications Dataset Details Dataset Description This dataset is a merge from QA instructions from InstruCAT (ca), SQUAC (es), SQUAD (en), plus generalists CA and ES MENTOR datasets to provide a cognitive background for generating responses. Contains splits of 66139 (train) and 11674 (validation) instructions Curated by: [More Information Needed] Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/InstrucatQA.textquestion-answering10K<n<100K0 likes37 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.