datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.IFEval_es
Dataset Card for IFEval_es
IFEval_es is a prompt dataset in Spanish, professionally translated from the main version of the IFEval dataset in English.
Dataset Details
Dataset Description
IFEval_es (Instruction-Following Eval benchmark - Spanish) is designed to evaluating chat or instruction fine-tuned language models. The dataset comprises 541 "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times"… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/IFEval_es.Aloe-Beta-DPO
Aloe-Beta-Medical-Collection
Collection of curated DPO datasets used to align Aloe-Beta.
Dataset Details
Dataset Description
The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data:
Medical preference data: TsinghuaC3I/UltraMedical-Preference
General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.NotSoTiny-25-12
NotSoTiny: A Large, Living Benchmark for RTL Code Generation
Summary
NotSoTiny is a large, structurally rich, and "living" benchmark designed to assess Large Language Models (LLMs) on the generation of context-aware RTL (Register-Transfer Level) code. Built from hundreds of real hardware designs produced by the Tiny Tapeout community, this benchmark overcomes the limitations of prior static… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-25-12.NotSoTiny-26-07NotSoTiny: A Large, Living Benchmark for RTL Code Generation
Summary
NotSoTiny is a large, structurally rich, and "living" benchmark designed to assess Large Language Models (LLMs) on the generation of context-aware RTL (Register-Transfer Level) code. Built from hundreds of real hardware designs produced by the Tiny Tapeout community, this benchmark overcomes the… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/NotSoTiny-26-07.InstrucatQA
Dataset Card for Dataset Name
Instructional dataset to finetune models used for RAG applications
Dataset Details
Dataset Description
This dataset is a merge from QA instructions from InstruCAT (ca), SQUAC (es), SQUAD (en), plus generalists CA and ES MENTOR datasets to provide a cognitive background for generating responses.
Contains splits of 66139 (train) and 11674 (validation) instructions
Curated by: [More Information Needed]
Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/InstrucatQA.XitXatTools
Dataset Card for XitXat Tools
XitXat Tools is a dataset comprising simulated Catalan call center conversations. Each conversation is annotated with structured tool calls, making it suitable for training and evaluating language models with function-calling capabilities.
Dataset Details
Dataset Sources
Repository: XitXat
Uses
The dataset can be utilized for:
Training language models to handle function-calling scenarios… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/XitXatTools.
