datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.salamandra-guard-dataset
Salamandra Guard Dataset
Dataset Description
The Salamandra Guard dataset is a comprehensive multilingual safety classification corpus designed for training and evaluating content moderation systems in Catalan, Spanish. It consists of 21,335 carefully curated conversational examples annotated across a hierarchical safety taxonomy.
This dataset represents a significant advancement in culturally-grounded safety data, with particular emphasis on Catalan—a language… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/salamandra-guard-dataset.IFEval_es
Dataset Card for IFEval_es
IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.Aloe-Beta-DPO
Aloe-Beta-Medical-Collection
Collection of curated DPO datasets used to align Aloe-Beta.
Dataset Details
Dataset Description
The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data:
Medical preference data: TsinghuaC3I/UltraMedical-Preference
General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.NotSoTiny-25-12
NotSoTiny: A Large, Living Benchmark for RTL Code Generation
Summary
NotSoTiny is a large, structurally rich, and "living" benchmark designed to assess Large Language Models (LLMs) on the generation of context-aware RTL (Register-Transfer Level) code. Built from hundreds of real hardware designs produced by the Tiny Tapeout community, this benchmark overcomes the limitations of prior static… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-25-12.ALIA-2606-DPO-safety
Dataset Card for BSC Multilingual Synthetic Safety Preferences
Dataset Summary
This dataset consists of synthetic safety preference data generated to align language models across five languages: Catalan, Spanish, English, Basque, and Galician.
Building on the PKU-SafeRLHF and Tulu 3/Ultrafeedback methodologies for creating preference data, this dataset leverages an LLM-as-a-judge approach to automatically score and pair model responses to a massive pool of safety… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/ALIA-2606-DPO-safety.NotSoTiny-26-07NotSoTiny: A Large, Living Benchmark for RTL Code Generation
Summary
NotSoTiny is a large, structurally rich, and "living" benchmark designed to assess Large Language Models (LLMs) on the generation of context-aware RTL (Register-Transfer Level) code. Built from hundreds of real hardware designs produced by the Tiny Tapeout community, this benchmark overcomes the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-26-07.InstrucatQA
Dataset Card for Dataset Name
Instructional dataset to finetune models used for RAG applications
Dataset Details
Dataset Description
This dataset is a merge from QA instructions from InstruCAT (ca), SQUAC (es), SQUAD (en), plus generalists CA and ES MENTOR datasets to provide a cognitive background for generating responses.
Contains splits of 66139 (train) and 11674 (validation) instructions
Curated by: [More Information Needed]
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/InstrucatQA.XitXatTools
Dataset Card for XitXat Tools
XitXat Tools is a dataset comprising simulated Catalan call center conversations. Each conversation is annotated with structured tool calls, making it suitable for training and evaluating language models with function-calling capabilities.
Dataset Details
Dataset Sources
Repository: XitXat
Uses
The dataset can be utilized for:
Training language models to handle function-calling scenarios… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/XitXatTools.
