datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_lmentry
Multi-LMentry
This dataset card provides documentation for Multi-LMentry, a multilingual benchmark designed for evaluating large language models (LLMs) on fundamental, elementary-level tasks across nine languages. It is the official dataset release accompanying the EMNLP 2025 paper "Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?".
Dataset Details
Dataset Description
Multi-LMentry is a multilingual extension of LMentry (Efrat et… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/multi_lmentry.open_data_26B_tokens_balanced_es_caThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.BSC_ParaMT_8
Dataset Card for BSC_ParaMT_8
Dataset Summary
Large-scale multilingual parallel corpus covering Catalan, Spanish, and English paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish portion of the dataset includes synthetic data generated by translating original English sentences into Spanish. Similarly… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/BSC_ParaMT_8.EsBBQ
Spanish Bias Benchmark for Question Answering (EsBBQ)
The Spanish Bias Benchmark for Question Answering (EsBBQ) is an adaptation of the original BBQ to the Spanish language and the social context of Spain.
Dataset Description
This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EsBBQ.m-personas
mPersonas: Multilingual Persona‑Driven Conversational Dataset
Dataset Summary
mPersonas is a multilingual open-source dataset with high-quality persona descriptions synthetically generated by DeepSeek-V3–0324. It follows a persona-driven data synthesis methodology, similar to PersonaHub.
Instances: 510,000
Total tokens: 173M
28M in personas
145M in conversations (105M in assistant turns)
Languages: 15
License: Apache 2.0
Methodology
This section… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/m-personas.CaBBQ
Catalan Bias Benchmark for Question Answering (CaBBQ)
The Catalan Bias Benchmark for Question Answering (CaBBQ) is an adaptation of the original BBQ to the Catalan language and the social context of Spain.
Dataset Description
This dataset is used to evaluate social bias in LLMs in a multiple-choice Question Answering (QA) setting and along 10 social categories: Age, Disability Status, Gender, LGBTQIA, Nationality, Physical Appearance, Race/Ethnicity, Religion… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CaBBQ.ALIA-2606-SFT-fc
Dataset Card for ALIA-2606's Function Calling Data
Dataset Summary
This dataset consists of a mixture of publicly available function-calling datasets and a collection of synthetic examples curated in-house for post-training language models with function-calling capabilities.
The dataset was used to train ALIA-40b-fc-2606, by combining these function-calling examples with a subset of 600k instances from the supervised fine-tuning mixture released as ALIA-2606-SFT… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-SFT-fc.COPA-es
Dataset Card for COPA-es
COPA-es is a textual entailment dataset in Spanish, professionally translated from the COPA dataset in English. The dataset consists of 600 premises, each given a question and two choices with a label encoding which of the choices is more plausible given the annotator.
Dataset Details
Dataset Description
COPA-es (Choice of Plausible Alternatives - Spanish) is designed to simulate causal reasoning of text from commonsense subjects.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/COPA-es.openbookqa-es
Dataset Card for openbookqa_es
openbookqa_es is a question answering dataset in Spanish, professionally translated from the main version of the OpenBookQA dataset in English.
Dataset Details
Dataset Description
openbookqa_es (Open Book Question Answering - Spanish) is designed to simulate open book exams and assess human-like understanding of a subject. The dataset comprises 500 instances in the validation split and another 500 instances in the test split.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/openbookqa-es.ALIA-2606-SFT
Dataset Card for BSC Multilingual Synthetic SFT Instructions
Dataset Summary
This dataset consists of 714k conversations mixing human and synthetic instructions generated to post-train language models across five languages: Catalan, Spanish, English, Basque, and Galician.
This dataset has been used in the supervised fine-tuning stage of ALIA-40b-instruct-2606. The training mixture is obtained by combining a selection of (human and synthetic) permissively licensed… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-SFT.ALIA_mixed_authentic_synthetic_MT
Dataset Card for ALIA_mixed_authentic_synthetic_MT
Dataset Summary
Large-scale multilingual parallel corpus covering English and Spanish paired with Arabic, Hindi, Chinese, Japanese, and Korean. Built by aggregating and carefully filtering multiple public sources, it provides sentence-level alignments for training Machine Translation systems. The Spanish–Hindi and Spanish–Chinese portions of the dataset include synthetic Spanish translations generated from English using… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA_mixed_authentic_synthetic_MT.NTEU_Multilingual_Evaluation_Dataset
Dataset Card for NTEU Multilingual Evaluation Dataset
Dataset Summary
This evaluation dataset for Machine Translation was created by the NTEU - Neural Translation for the EU project.
The evaluation dataset includes around 1,000 parallel sentences in the 24 official European languages.
The original NTEU dataset has been cleaned and filtered by removing empty lines and near-duplicates, and it has been augmented with Catalan.
The Catalan version was manually produced by a… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/NTEU_Multilingual_Evaluation_Dataset.salamandra-guard-dataset
Salamandra Guard Dataset
Dataset Description
The Salamandra Guard dataset is a comprehensive multilingual safety classification corpus designed for training and evaluating content moderation systems in Catalan, Spanish. It consists of 21,335 carefully curated conversational examples annotated across a hierarchical safety taxonomy.
This dataset represents a significant advancement in culturally-grounded safety data, with particular emphasis on Catalan—a language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/salamandra-guard-dataset.IFEval_es
Dataset Card for IFEval_es
IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.distilled-yodas-spanishDistilled YODAS Spanish is a high-quality subset of the Spanish portion of the YouTube-Oriented Dataset for Audio and Speech (YODAS). While the full YODAS corpus contains over 37,000 hours of Spanish speech across 43 million files, this dataset provides a distilled version of approximately 8,000 validated hours.distilled-catalan-youtube-speechThe Distilled Catalan YouTube Speech Corpus is the result of the automatic validation of the original
Catalan YouTube Speech Corpus created by SoftCatala and shared trhough this HF repo:
https://huggingface.co/datasets/softcatala/catalan-youtube-speech
The corpus is -distilled- because only recordings with a high certaintity to be correct were taken
and the rest were rejected.EQ-bench_es
Dataset Card for EQ Bench Dataset (Spanish Version)
This dataset card documents the Spanish adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts.
Dataset Details
Dataset Description
EQ-Bench (Spanish Version) is a translated and linguistically adapted version of the original EQ-Bench dataset.
Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_es.LexBOE
Dataset Card for LexBOE
Dataset summary
LexBOE is a Spanish legal text classification dataset built from articles extracted from the Boletín Oficial del Estado (BOE), the official source of legislation and administrative acts in Spain. The articles included in the dataset were published between 2022 and 2024.
LexBOE reflects contemporary legal-administrative language and is intended for the training and evaluation of language models on legal text classification tasks.… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/LexBOE.Catalan-Aranese_Parallel_Corpus
Dataset Card for Catalan-Aranese Parallel Corpus
Dataset Summary
A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated from… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.EQ-bench_ca
Dataset Card for EQ Bench Dataset (Catalan Version)
This dataset card documents the Catalan adaptation of the EQ-Bench benchmark. The original dataset was designed to evaluate emotional reasoning in language models through dialogue-based prompts.
Dataset Details
Dataset Description
EQ-Bench (Catalan Version) is a translated and linguistically adapted version of the original EQ-Bench dataset.
Its design responds to the need to adapt the emotional detection… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/EQ-bench_ca.arc_es
Dataset Card for ARC (Spanish Version)
Dataset summary
This dataset provides the Spanish translation and adaptation of the ARC (AI2 Reasoning Challenge) validation set. The original dataset was designed to evaluate scientific reasoning in LLMs by presenting a collection of authentic, grade-school science multiple choice questions split into two sets of varying difficulty: ARC Easy and ARC Challenging.
This Spanish adaptation enables evaluation of the ARC task in… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/arc_es.ACAData
Dataset Card for ACAData
Dataset Summary
ACAData is a multilingual instruction tuning dataset containing parallel text paragraphs from the academic domain.
Supported Tasks and Leaderboards
The dataset is meant to be used for fine-tuning and benchmarking general purpose LLM's on Machine Translation tasks.
Languages
The dataset contains (mainly long) paragraph of scientific texts from the academic domain in many European language pairs.
The language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ACAData.AbSanitas
Dataset Card for AbSanitas
Dataset summary
AbSanitas is a Spanish biomedical information retrieval dataset built from biomedical texts collected from official academic repositories and open-access sources.
This dataset is designed to support the training and evaluation of encoder models on biomedical retrieval and semantic matching tasks in Spanish.
Curated by: Barcelona Supercomputing Center (BSC)
Funded by: ALIA
Language(s) (NLP): Spanish (es)
License: CC BY-NC-ND 4.0… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/AbSanitas.BSCs_Code_Switching_CA-ES_ASR_TestThe BSC's Code-Switching Catalan-Spanish ASR Test is a speech dataset of 4 hours and 9 minutes. It consists of carefully selected recordings that feature code-switching between Catalan and Spanish. This dataset is designed to be a test set for Catalan ASR systems that need to handle code-switching to Spanish.rpaweb
Dataset Card for rpaweb
The rpaweb dataset contains multilingual annotations of Robotic Process Automation (RPA) actions in a web environment in English, Spanish and Catalan.It has been designed to train multimodal models capable of autonomously performing browser-based tasks based on natural language user requests.
The corpus documents each user request along with the full sequence of web automation actions needed to fulfil it, including structured JSON code, natural language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/rpaweb.Legal_Catalan_Spanish_Parallel_Corpus
Dataset Card for Legal Catalan-Spanish Parallel Corpus
Dataset Summary
The Legal Catalan-Spanish Parallel Corpus is a multilingual dataset of authentic parallel text in Catalan and Spanish from the legal domain, drawn from official Catalan public institution documents. It comprises three sub-corpora organized at different textual granularities: sentence, paragraph, and document level. This multi-level structure enables research and training that goes beyond traditional… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Legal_Catalan_Spanish_Parallel_Corpus.ALIA-2606-DPO-safety
Dataset Card for BSC Multilingual Synthetic Safety Preferences
Dataset Summary
This dataset consists of synthetic safety preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician.
Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool of safety… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-safety.CAESAR-TV3
Dataset card for CAESAR-TV3
Dataset Summary
This corpus includes 5 hours and 45 minutes of Catalan speech code-switched with Spanish extracted from the original tv3_parla dataset.
Supported Tasks and Leaderboards
The CAESAR-TV3 dataset is designed for the Automatic Speech Recognition (ASR) task, enabling the transcription of utterances in Catalan, Spanish, and code-switched speech between the two languages.
Languages
The dataset features… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/CAESAR-TV3.geneval_catalan
Dataset Card for GenEval Catalan
Dataset Summary
GenEval Catalan is an English→Catalan extension of MT-GenEval, a benchmark for evaluating gender accuracy in machine translation. It is derived from the original English–Spanish MT-GenEval dataset (Currey et al., 2022), with the Spanish side replaced by Catalan translations produced by professional human translators (with the exception of the single-sentence dev split, which uses automatic translation).
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/geneval_catalan.AbScientia
Dataset Card for AbScientia
Dataset summary
AbScientia is a Spanish STEM scientific text classification dataset built from scientific abstracts collected from official academic repositories and open-access sources. The dataset focuses on Science, Technology, Engineering, and Mathematics (STEM) disciplines and reflects domain-specific scientific language in Spanish.
This dataset is designed to support the training and evaluation of encoder models on STEM scientific domain… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/AbScientia.
